Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
Posts made with a user token (xoxp-) arrive with app_id and no
client_msg_id, so _event_declares_bot_sender dropped them as app traffic;
the only workaround was allow_bots: all. Adds
platforms.slack.extra.api_human_users (SLACK_API_HUMAN_USERS fallback), a
users-only allowlist consulted inside the predicate.
Salvaged from #100964 (users only: an app-id allowlist would also admit
the app's own xoxb bot posts, which share the user+app_id shape).
Follow-up to the #99641 salvage:
- One module-level _normalize_security() (ssl/tls/implicit -> tls, starttls,
plain/none -> plain; unknown -> WARNING + secure default) replaces the three
copies of the alias set; _connect_imap/_connect_smtp/_standalone_send all
compare against the canonical value. Unknown modes no longer raise.
- _tls_context(verify, host) is module-level and shared by all sites; when
verification is disabled for a non-loopback host it logs a WARNING.
- _esecret_bool: an unset/empty env var now yields the caller's default
(previously is_truthy_value('') returned False, silently disabling TLS
verification whenever EMAIL_*_TLS_VERIFY was unset).
- Documented surface is platforms.email.extra.{imap,smtp}_security and
{imap,smtp}_tls_verify in config.yaml; env vars remain an internal bridge
and are NOT added to plugin.yaml (optional_env feeds hermes setup prompts).
- Docs: Proton Mail Bridge / local relays recipe in user-guide/messaging/email.md.
- Tests: starttls builds IMAP4 then .starttls(); unknown mode falls back to
tls/starttls with verification still on.
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.
- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.
Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
De-risking for the #48820 behaviour change: before this branch, credentials
in the environment force-enabled twelve platforms regardless of an explicit
enabled: false in config.yaml. Now that the explicit disable wins, users who
relied on the old override would see the platform go dark with no trace.
_enable_from_env (and Slack's inline copy) now emit ONE WARNING per platform
per process when the platform is explicitly disabled AND its env credentials
are present, naming the platform, the winning key
(platforms.<x>.enabled: false), the env var(s) being ignored, and the remedy.
A plain disable with no credentials, an enabled platform, and the env-only
(no YAML opinion) path stay silent; repeated config reloads do not repeat it.
_ENV_ENABLE_CREDENTIALS maps every _enable_from_env platform to its
triggering env var(s); a test pins that the map covers every routed branch.
Docs: messaging/index.md gains a 'Disabling a platform whose credentials are
still in .env' section with the exact warning text.
Live repro (real load_gateway_config on a temp HERMES_HOME with
platforms.weixin/telegram.enabled: false + WEIXIN_TOKEN/TELEGRAM_BOT_TOKEN in
env): before — both stayed disabled with zero log output; after — one
WARNING each ('Platform 'weixin' is explicitly disabled by
platforms.weixin.enabled: false ... (WEIXIN_TOKEN, WEIXIN_ACCOUNT_ID) will
NOT start its adapter ...'), none for the enabled homeassistant, none on the
second load.
A non-admin '/sessions all' or '/resume --all' silently downgraded to
chat-scoped listing with zero feedback, which reads as 'my session
vanished' (community Telegram report). Both surfaces now append a notice
that cross-chat listing requires a configured admin. Follows up the
salvaged current-session '(current)' marker (PR #68556, fixes#68547):
sibling tests updated to pin the new contract, new i18n key
gateway.resume.all_requires_admin added to all 17 locales, docs updated.
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.
`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.
Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
Compose the two salvaged approaches (#78065 + #78511):
- Keep #78065's terminal-only scrub-path exemption (first-party prefix
predicate in _make_run_env / _sanitize_subprocess_env, plain env values
never scope-resolved, snapshot exclusion for cross-profile isolation,
every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
an import-time blocklist discard: the blocklist is shared by every
scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
process env (Buzz Desktop buzz-acp harness, #76243) OR the live
session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
session on a host that also runs a Buzz gateway does NOT get the
signing key in its terminal children (maintainer triage note on
#76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
strips) and positive test (buzz session platform enables carve-out);
docs updated accordingly.
Closes#78026, closes#76243.
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.
Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.
Follow-up hardening from review:
- First-party matches use the merged env value directly instead of
_resolve_passthrough_value: under multiplex with no profile secret scope
installed the resolver raised UnscopedSecretError (fail-closed) at call
sites like the webhook-filter script runner, a regression where the
script previously ran without the var. The vars are the process's own
env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
shared login-shell snapshot (_additional_profile_scoped_passthrough_names
override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
exclusion set (env_passthrough refuses blocklisted names), so without
this a multiplexed gateway would dump profile A's key into
hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
would source it — a cross-profile nsec leak. The names are now excluded
from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
like ddgs, computer-use driver, user-script runners) that also receive
first-party platform vars.
Fixes#78026
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Compose/fix-up on top of the cherry-picked cluster commits:
- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
through _apply_yaml_config and honored by send(), send_image(), and
_standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
_progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
so the synthetic-thread fallback and the progress reply anchor are both
suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
_extract_thread_root (marked root > reply > legacy positional e-tag)
instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
edit_message now implemented, accumulate-style progress works, but without
the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).
- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
pagination logic (unit-testable without python-telegram-bot). First
query token filters; the remainder is carried into the sent command as
its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
enables inline mode via BotFather /setinline) + _handle_inline_query
with the same auth path as inline-button callbacks — unauthorized users
get an empty list, so the installed-skill catalog is not leaked to
arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
message starts with /, which reaches the bot even under default privacy
mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
/setinline setup.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
- The direct adapter.send() confirmation path in /approve and /deny is only
needed on native-streaming platforms (WeCom) where the reply stream is
already finalized; other platforms keep the return-text contract (fixes
4 approve/deny regression tests, guarded with 'is not True' against
MagicMock auto-attributes).
- Remove a stray [DEBUG] logger.info left in _deliver_media_from_response.
- website/docs wecom.md: replace the 'does not stream' notes with the native
msgtype:stream behavior and document the stream keepalive extra keys.
When either unfurl key is set, media captions post as a separate
message before the file (the upload API cannot carry unfurl controls)
and native draft streaming falls back to edit-based delivery. Surface
both side effects in the config reference table.
Follow-up to lkz-de's adapter chunking commit: long Signal messages no
longer truncate on ANY delivery path.
- tools/send_message_tool.py: register Signal's 8000-char limit in
_MAX_LENGTHS (imported from the adapter module so the two paths can't
drift) so hermes send / cron standalone / MCP sends split via the
shared truncate_message() pass instead of signal-cli rejecting them.
Standalone-path idea credited to @5L-hermes01 (#67279).
- tests: regression test proving standalone Signal sends chunk at the
adapter limit with no truncation footer (fails on pre-fix main).
- docs: Long Messages section on the Signal page (en + zh-Hans).
Both fixes verified by sabotage A/B (tests fail with the respective
half reverted to origin/main) and a real-import E2E: 27k-char message
with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles
in-range UTF-16, lossless reassembly.
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
A blackholed IPv6 path to api.telegram.org never errors, so
_await_with_thread_deadline never fires and connect hangs at
"attempt 1/8". Known A-record IPs connect over IPv4 immediately.
DoH timeout now fail-opens to the seed IPv4 list instead of the
hostname. Hostname stays last for IPv6-only hosts.
Closes#87015
Background process completions on messaging platforms now default to a
one-line status message (✅/❌ + command + duration; failures append a
short output tail) instead of dumping the raw output buffer into the
chat. New display.background_process_notifications mode 'concise' is
the default; 'all' keeps the old raw-dump behavior for anyone who wants
it. Config migration v35 moves users still on the old implicit default
'all' to 'concise' on their next update; explicit result/error/off
choices are preserved.
Adds platforms.slack.extra.native_task_cards: when enabled, live tool
calls render as Slack-native plan/task cards via chat.startStream /
chat.appendStream (task_display_mode: plan, task_update chunks) instead
of text/edit progress bubbles. ID-bearing tool_start/tool_complete
callbacks correlate concurrent same-name tool calls correctly; any
native API failure falls back to one continuously edited text update.
The stream is stopped exactly once when the turn finalizes.
Salvaged from PR #29496 onto current main (TurnRunner/TurnContext seam);
closes#29483.
Webhook agent runs default to the constrained hermes-webhook toolset
(web/vision/clarify) because payloads can carry untrusted third-party
content. That default is right for public webhooks but wrong for trusted
local pushes (e.g. an OOM monitor daemon that needs the agent to run
ps/free/py-spy): the only workaround was widening platform_toolsets.webhook,
which elevates EVERY webhook route at once.
This adds a 'toolsets' key on individual webhook route configs (static
routes in config.yaml and dynamic subscriptions in
webhook_subscriptions.json) that replaces the platform-level resolution
for that route only:
- BasePlatformAdapter.toolsets_for_source(): per-source override hook,
default None (no behavior change for any other platform).
- WebhookAdapter.toolsets_for_source(): maps the session chat_id
(webhook:{route}:{delivery_id}) back to its route config and returns
the route's toolsets list.
- GatewayRunner._resolve_enabled_toolsets_for_source(): shared resolver
used by both agent-run call sites; validates the override through the
SAME _get_platform_tools path as platform config, so unknown names and
platform-restricted toolsets (e.g. discord_admin) are dropped rather
than trusted.
Deliberately NOT exposed via 'hermes webhook subscribe': granting elevated
tools is a manual config edit only, so an agent-created subscription
cannot self-grant terminal at runtime.
Cache-safe: the toolset list is resolved before agent construction and is
constant for a route, so the per-session agent signature and frozen system
prompt are unaffected mid-conversation.
PrivilegedIntentsRequired is a Developer Portal config error; surface which
intents Hermes requested as a non-retryable fatal and teach setup/docs.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(relay): ambient token endpoint mode for gateway.idp.token_url
When gateway.idp.token_url is configured WITHOUT client_id/client_secret,
treat the URL as a metadata-server-style ambient credential endpoint:
plain GET, response body is the token (raw JWT or {"access_token": ...}
JSON envelope). Covers workload-identity proxies such as Domino's
$DOMINO_API_PROXY/access-token, which mint short-lived user-scoped OIDC
tokens with no client registration.
Previously this configuration was a hard error (client_id/client_secret
missing), so no working deployment changes behaviour: creds present keeps
the OAuth2 client_credentials POST, no token_url keeps Nous Portal. The
misconfig error now self-diagnoses (names the ambient fallback and how to
select the client_credentials grant instead).
* fix(relay): reject short plain-text bodies in ambient token shape gate
Review finding: the shape gate accepted any base64url-alphabet word, so an
IdP answering the ambient GET with a terse error body ('unauthorized',
'error', 'null') had that word returned as a bearer token instead of the
fail-closed misconfiguration error. Tighten the gate to JWT-like dotted
tokens (3+ segments) or long opaque tokens (>= 32 chars); short bare words
now raise the self-diagnosing ambient error.
* fix(relay): partial IdP client credentials keep the loud error, never select ambient GET
The ambient-endpoint dispatch used 'not client_id or not client_secret',
so configuring exactly one credential (a mistyped client_credentials
setup) silently issued a GET at the IdP token endpoint and then raised
'no client_id/client_secret configured' — factually wrong for that
operator, and a stray request the old hard error never made.
Ambient mode now requires NEITHER credential; a partial pair raises
immediately, names the missing key, and issues no HTTP request (tests
assert urlopen is never called). Docstring and relay.md now say
'neither' instead of 'without'.
* fix(relay): ambient JSON envelope requires a string access_token, no coercion
Review finding (P2): the JSON-envelope branch accepted any truthy
access_token via str() coercion — a number became '12345…', a boolean
became 'True', an object became its Python repr — bypassing the fail-
closed contract and deferring the failure to the connector, where it
hides the real endpoint problem.
The envelope value must now be a non-empty string, the same contract the
client_credentials path enforces on its token response. Deliberately NO
shape gate on envelope values: an envelope is an intentional token
response (mode-1 symmetry), and opaque tokens may use the standard-base64
alphabet the raw-body gate rejects. Mutation check: reverting the branch
to str() coercion sends the 3 coercion tests red (3 failed, 15 passed).
---------
Co-authored-by: Ben Barclay <ben@nousresearch.com>
Document the actual transport split: legacy editable drafts by default, optional rich drafts, persistent rich final sends, and in-place rich final edits for edit-based streams.
- matrix/dingtalk: extract deps-only installers (ensure_matrix_deps,
ensure_dingtalk_deps) and register THOSE as ensure_deps_fn — the prior
check_*_requirements combined credential env checks with the install,
so a platform configured via PlatformConfig.extra (which is_connected
accepts) would pass enablement, reach create_adapter(), and have the
'installer' veto on env-var grounds before installing anything —
re-creating the #79812 deadlock for extra-configured setups. The
combined deps+credentials functions remain for setup/status callers.
- matrix/feishu passive probes: use the existing lazy_deps.is_available()
instead of hand-rolling 'not feature_missing(...)' (reuse finding).
- teams: module docstring no longer recommends bare system pip (the
PEP 668 trap purged everywhere else); docs troubleshooting row updated
to match the new hint text.
- wecom_callback: drop dead 'global ET, DEFUSEDXML_AVAILABLE'
(ensure_and_bind mutates the module dict directly; nothing assigns).
- tests: parametrized wiring contract for all 8 lazy-installable
platforms — ensure_deps_fn present and distinct from check_fn
(behavior contract, not identity snapshot, so renames don't churn it).
- gateway/config.py: rewrite the stale enablement-pass header comment that
still described check_fn as 'the single source of truth for are-my-env-
vars-set' / 'lazy-installs it' — both false under the new contract.
- teams: check_requirements docstring wrongly claimed credential checks
(body checks only SDK/aiohttp presence); derive install_hint from the
canonical LAZY_DEPS pins + sys.executable instead of hardcoding
'~/.hermes/hermes-agent/venv/bin/pip' and version pins (wrong under
HERMES_HOME overrides / profile installs; pins go stale on CVE bumps);
connect() fatal-error hints now point at the venv pip instead of bare
system pip (the PEP 668 trap the docs warn about).
- teams docs: drop exact version pins from the two manual-install commands
(LAZY_DEPS is the source of truth; unpinned installs still work and the
text can't go stale).
- hermes_cli/status.py: per-entry exception guard around check_fn so one
raising probe can't abort the listing of all remaining plugin platforms
(aligns with the other three call sites).
- tests: rename test_register_check_fn_is_active_lazy_installer ->
test_register_splits_passive_probe_from_active_installer (name said the
opposite of what it verifies).
The reset keywords have existed in both CLI and gateway handlers since
June but were undocumented — users couldn't find how to cancel a
personality overlay. Adds a 'Resetting to the default' section to the
personality feature page and mentions the reset in the CLI guide,
slash-command reference (both tables), and messaging command table.
- New website/docs/user-guide/messaging/a2a.md: when/where to use A2A
(cross-machine, specialist peers, being callable) vs delegation/kanban
for same-machine multi-agent; enable, outbound tools, inbound surface,
security model, env reference, quick test, troubleshooting. Registered
in sidebars.ts and the messaging index.
- README/DESIGN/plugin.yaml/protocol.py prose updated to name the A2A
v1.0 canonical discovery path /.well-known/agent-card.json (the code
already served both; only the docs lagged).
Builds on the adapter list_channels() hook (cherry-picked from #43545 by
@Guoen0):
- plugins/platforms/simplex: implement list_channels() — enumerates
contacts (/contacts) and groups (/groups) over the live daemon
WebSocket into the channel directory. Returns None when the WS is
down so the directory falls back to session discovery instead of
wiping known targets.
- hermes send --list: merge configured-but-undiscovered platforms into
the listing. Previously a platform configured only via env (e.g. a
fresh SimpleX setup used for outbound sends) was silently omitted,
leaving users guessing at platform names.
- format_directory_for_display(): accept an explicit platforms view and
render empty platforms with a targeting hint instead of hiding them.
- docs: simplex hermes-send section.
Reported by Fedpostoffice on Discord (simplex missing from
hermes send --list; guessed platform names simplex-chat/simplex-relay).
Google API and authentication packages permit vulnerable httplib2 and pyasn1
transitives, while the Workspace and Google Chat runtime installers previously
treated any importable version as sufficient. Existing environments could
therefore remain vulnerable after the project dependency pins were repaired.
Carry the fixed versions through the Google and Vertex extras, lazy feature
requirements, lockfile, and both runtime installers. Route the documented
Google Chat installation path through its maintained secure requirements
instead of an unconstrained direct pip command.
Detect stale distributions, install only unsatisfied requirements, and verify
the result before continuing. Behavioral tests cover those repair invariants
without freezing manifests, lockfiles, or complete package sets.
Related #72108
Extracted from #72840
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
LINE plugin.yaml plus line/wecom-callback/msgraph-webhook/
whatsapp-cloud/teams user-guide pages still documented the old
IPv4-only 0.0.0.0 defaults; update to the dual-stack unset default.
Telegram webhook env docs live in the adapter docstring (updated with
the code change); its plugin.yaml has no webhook host entry.
Consolidates the native-transport half of PR #73636 by @ScaleLeanChris
onto the merged adapter: persistent NIP-42-authenticated Nostr WebSocket
subscription as the default inbound path (transport=auto|websocket|poll),
kind-44100 membership events for live DM discovery, since-timestamp
resume on reconnect with bounded exponential backoff, and automatic
fallback to CLI polling when the WS can't be established. Events route
through the same _handle_event() pipeline as the poll loop, so de-dupe,
mention gating, p-tag DM latching, and allow-lists behave identically on
both transports. Outbound stays on the CLI (one-shot sends never race a
WS auth handshake — his design).
E2E verified against a real in-process websockets relay: NIP-42
challenge -> signed kind-22242 AUTH (event id re-derived server-side) ->
REQ subscription -> EVENT dispatch -> clean disconnect.
Co-authored-by: ScaleLeanChris <chris@scalelean.com>
The sidecar auth token is generated at spawn (secrets.token_hex) and
existed only in the gateway process memory + sidecar child env, so
_standalone_send from cron subprocesses, hermes send, or the dashboard
structurally could not authenticate (#69960).
The adapter now writes <hermes-home>/runtime/photon-sidecar.json
({port, token, pid}, 0600, atomic tempfile+os.replace) once the sidecar
passes its /healthz readiness check, and deletes it in _stop_sidecar,
on every startup-failure path, and at disconnect so a stale record
never outlives a dead sidecar. _standalone_send falls back to the
record when PHOTON_SIDECAR_TOKEN is unset, validating the recorded pid
is alive first; a stale record yields a clear 'gateway appears to be
down' error. Docs note the gateway-must-be-running requirement and the
Photon-side shared-line initiation policy (#51897).