_persist_credential promises "leaving all else intact", but it seeded
its write from _read_config, which returns {} on ANY read failure, and
_atomic_write_config replaces the whole file via os.replace — which
needs only a writable parent, so a present-but-unreadable honcho.json
did not stop the overwrite. One OSError (EACCES after a root-owned
write, EIO, a stalled mount) during an automatic token refresh or a
fresh login therefore replaced the store with a single-host file,
destroying every other host's credentials and honcho.json's root
config. No user action is required to trigger the refresh path.
Same defect class as #75206 (P1, fixed for the core auth store in
Add _read_config_strict for the write paths: a missing file still
bootstraps as {}; an unreadable file raises with the store untouched;
genuine corruption still degrades but preserves a .corrupt copy first,
since a truncated store usually holds the other hosts' tokens verbatim.
_rotate_and_persist now takes its strict read BEFORE the exchange —
rotation is single-use, so an exchange whose result cannot be persisted
loses the grant — and threads the dict through to _persist_credential,
which also closes the re-read race between the locked read and the
persist. install_grant seeds its root-merge from the strict reader for
the same reason. Read paths keep their fail-open contract untouched;
both readers now use utf-8-sig so a BOM'd store is not misclassified
as corruption (the wipe vector needing no filesystem fault at all).
Adds TestPersistReadFailure: six tests, four of which fail against the
previous source; the rotate-ordering test additionally pins that no
exchange is attempted against an unreadable store.
reading the live client's api_key inside _force_reauth races the in-place
rotation: a sibling waiter's apply_token_to_client (or the proactive
refresh entered via the honcho property) can swap the bearer before the
read, so force_refresh_token receives the already-rotated token, disk
matches it, and the adopt branch never fires — reproducing the very
force-refresh burst this PR removes.
_authed_call now snapshots the bearer before invoking the operation and
passes that exact token through to force_refresh_token.
desktop spawns multiple serve processes that share one rotating honcho
refresh token. force_refresh_token treated every 401 as "rotate now" even
when a sibling had already persisted a new grant, which can replay a
single-use refresh token and revoke the whole grant.
re-read honcho.json under the existing file lock and adopt if the token
that 401'd is no longer on disk. on invalid_grant, re-read once more
before marking the grant dead.
- Correct the per-generation reset comment: an in-place updater restart keeps
PTB's update_queue, so old-generation dispatches can briefly exceed
received; the check already treats that as no backlog.
- Make the once-per-stall gate explicit (cap the heartbeat count) instead of
relying on `!=`.
- Drop the dead getattr in _record_updates_received (only reachable after
_record_polling_progress dereferenced the same instance state); keep the
fallbacks in the heartbeat check and group-99 handler, which sibling
watchdogs share because object.__new__ adapter doubles exist in tests.
- Split the single invariant test so a failure names the broken guard:
stall-once, re-arm-on-progress, generation-reset; drop the unused _app mock.
Follow-up to the cherry-picked #102383 commit. The check as written was neither
sensitive nor specific:
- It aged the newest received update, so a wedged PTB dispatcher was never
reported while new updates kept arriving more often than every 300s (probe:
1 update/250s for an hour -> 0 reports).
- `delivered` counted only MessageEvents reaching the gateway handler, while
`received`/`dispatched` counted every Update; a single handled callback_query,
reaction, unauthorized user or unmentioned group message produced a false
ERROR after 300s of quiet.
- `_record_updates_received` skipped the generation/teardown guard
`_record_polling_progress` applies, and the counters never reset across
polling generations, so a late response from a fenced poll or a reconnect
inflated the backlog.
Now `received` and `dispatched` count the same population (every fetched
update reaches the group-99 catch-all) and the report fires when a backlog
persists with no dispatch progress across two 90s heartbeats, once per stall,
re-armed on progress, reset per generation, at WARNING (diagnostic only;
#71240 owns recovery). The delivered counter and the `note_inbound_delivered`
facade method are dropped; the once-per-adapter "no message handler" error on
BasePlatformAdapter.handle_message stays. `_record_polling_progress` returns
whether the round-trip was accepted so the received stamp reuses its gate.
Tests trimmed to two invariants; every guard proven red by mutation.
Refs #102260
Every Telegram health probe measures the transport. A getUpdates round-trip
that returns 200 proves bytes are moving and nothing else: the stall watchdog
(#92991), the pending-update probe (#42909/#55769), the get_me() heartbeat
(#66377) and the polling-progress instrumentation all stay green while updates
arrive and then die downstream. The adapter then publishes "connected", logs
nothing at all, and is indistinguishable from a bot nobody has messaged.
That is #102260: three weeks of telegram.state "connected" plus "polling
confirmed healthy: getUpdates progressing (generation 1)" with zero inbound
reaching the agent, surviving every restart. Two of the issue's three
hypotheses do not hold on this code — _record_polling_progress fires on every
round-trip (not only at start_polling), and _send_path_degraded is cleared on
the first confirmed round-trip — and the reporter's own observation that fresh
messages are received but not processed places the failure downstream of the
transport, in the one stretch with no instrumentation at all.
Add the missing delivered side of the accounting:
- received: updates Telegram handed the process, read from the getUpdates
envelope the adapter already parses (an empty result proves the transport,
not arrival, so only non-empty results count).
- dispatched: updates PTB's dispatcher carried through the whole handler
chain, stamped in the existing group-99 catch-all before its early returns.
- delivered: inbound events that reached the gateway's message handler,
stamped in BasePlatformAdapter.handle_message for every platform.
_check_ingress_delivery_gap runs on the existing heartbeat and, when updates
arrived but nothing was delivered for 300s, names the broken hop: received >
dispatched means the dispatcher is not draining, dispatched > delivered means
Hermes is dropping what arrives. Diagnostic only — a received update
legitimately reaches no gateway turn, and reconnecting a healthy transport
cannot repair a dropped update, so this never drives recovery.
Also make the silent discard on the shared funnel speak: handle_message
returned with no log when no message handler was installed, so a mis-wired
adapter discarded 100% of inbound while connected and able to send.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018pT5hFJBRfLj8KMqFhm3qz
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.
Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).
Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
dingtalk _extra_get, mattermost _extra_or_env and slack _extra_or_env_flag/_channel_set fell
through to env only on None, so `allowed_channels: ""` / `free_response_channels: ""` meant
"no whitelist" rather than "use the env CSV". The shared reader treated blank as unset and
silently widened those to the env value. New `blank_is_unset=False` knob restores the old
semantics at those seven call sites; the default (blank = unset) stays for the readers whose
old body was `extra.get(k) or env`.
WeComAdapter mixed in OwnAccessPolicyMixin without ALLOW_ALL_ENV_PREFIX, so
_allow_all_env_names() read "_ALLOW_ALL_USERS" and every open-DM WeCom deployment
(setup still writes WECOM_ALLOW_ALL_USERS) silently denied all DMs.
- WeComAdapter.ALLOW_ALL_ENV_PREFIX = "WECOM"
- OwnAccessPolicyMixin.__init_subclass__ raises TypeError on an empty prefix so the
omission cannot ship again (every other host already sets one).
- Parity test now runs a third matrix row that sets each host's own
<PREFIX>_ALLOW_ALL_USERS; the GATEWAY-only row is why this slipped through.
Sabotage: prefix removed -> [platform] row fails even with the guard reverted.
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
Fourteen platform plugins hand-rolled the "already configured? Reconfigure? [y/N]"
gate at the top of interactive_setup (env check + info line + prompt_yes_no(..., False)),
with drifting wording ("X: already configured" vs "X is already configured." vs
"already enabled") and, for LINE and SimpleX, raw input() loops with their own
EOF/KeyboardInterrupt handling and no gate at all. Fixes to the gate (non-interactive
handling, wording, default) therefore reached only the core Telegram/BlueBubbles/webhook
wizards.
- hermes_cli/setup_platforms.py: `_declines_reconfigure` becomes the public
`declines_reconfigure(label, question, *env_vars)` (any-of env check, so Matrix's
token-or-password gate fits); `_save_prompted` becomes `save_prompted` alongside it.
No alias kept; the three core callers are updated.
- buzz, dingtalk, discord, feishu, google_chat, irc, matrix, mattermost, raft, slack,
teams, wecom: the hand-rolled gate is replaced by one `declines_reconfigure(...)` call;
post-decline extras (Discord allowlist nudge, Slack manifest refresh, Raft "Keeping"
line) stay local and unchanged.
- line, simplex: the raw input() loops move onto hermes_cli.cli_output.prompt (masked
for secrets, "" on Ctrl-C/EOF) and gain the shared gate on their primary env var.
Behavior change: the gate's info line is now uniformly "<Label>: already configured"
(DingTalk/Feishu/WeCom lose the trailing period + inline ID; Buzz/IRC/Google Chat/Raft/
Teams no longer echo the current value in that line). Feishu and WeCom now gate on the
app/bot ID alone instead of ID AND secret. LINE and SimpleX gain a "Reconfigure?" [y/N]
prompt when already configured; their prompts now honour HERMES_NONINTERACTIVE and print
via the CLI helpers instead of bare print(). Prompt defaults (No) are unchanged everywhere.
Not touched: WhatsApp's gate keys on WHATSAPP_ENABLED being truthy (a "false" value must
not count as configured), which the shared any-set gate cannot express — left hand-rolled.
Test: tests/plugins/platforms/test_interactive_setup_reconfigure_gate.py parametrized over
the 14 wizards — with the primary env var set and the user declining, each wizard must have
called declines_reconfigure with that var and returned without prompting or saving.
Sabotage: reverting mattermost's gate fails that row.
Six adapters defined their own `_cancel_task` and nine more inlined the same
cancel + suppress(CancelledError) + await block; five kept a hand-rolled TTL-dict
`_is_duplicate` next to the existing `helpers.MessageDeduplicator`; three carried a
`_bounded_put`. Each copy fixed the same bugs on its own schedule (self-cancel deadlock,
done-task re-await, eviction under load).
- `helpers.cancel_task`: None/done no-op, never awaits the current task, swallows the
task's own exception at teardown. Replaces qqbot/signal/yuanbao/buzz/photon/simplex
definitions and the inline copies in weixin, discord, email, irc, line, mattermost,
whatsapp and telegram.
- `helpers.MessageDeduplicator` replaces `_is_duplicate` in qqbot, ntfy, photon,
wecom_callback and LINE's `_MessageDeduplicator`; every site keeps its own
max_size/TTL (qqbot and ntfy 1000/300s, photon 4000/48h, wecom_callback 2000/300s,
LINE 1000/no TTL).
- `helpers.bounded_put` replaces photon/wecom/whatsapp_cloud copies; a re-put now
refreshes the key to the newest slot at every site.
- telegram gmail-triage scripts resolve under `get_hermes_home()` instead of a hard
`~/.hermes`, so profiles with HERMES_HOME set find them.
Not changed: `get_chat_info` stays `@abstractmethod` because
tests/gateway/test_relay_capability_surface.py locks the abstract set to exactly
{connect, disconnect, send, get_chat_info} as a cross-repo contract, so the ~17 no-op
overrides remain.
Behavior change: whatsapp_cloud `_bounded_put` was a pure FIFO (no refresh on re-put);
it now refreshes like the other two sites. Task cancellation at the migrated sites
swallows a task's terminal exception where a few copies previously only suppressed
CancelledError (all are shutdown/disconnect paths).
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).
`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.
Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
Weixin, WeCom, QQBot, WhatsApp (common + cloud) and Yuanbao's AccessPolicy each carried
their own `_open_dm_opted_in` / `_is_dm_allowed` / `_is_dm_intake_allowed` /
`_is_group_allowed`, differing only in the platform prefix of the allow-all env var and in
small drifts. The same "allow-all must be scoped / fail closed" fix has landed on this trio
at least four times; a shared rule means it lands once.
gateway/platforms/access_policy_mixin.py::OwnAccessPolicyMixin owns the predicates,
reads every env name through the scoped `get_scoped_secret` and exposes two hooks:
`_entry_matches` (platform allowlist matching) and `_live_dm_allow_from` (env-seeded
lists re-read live). `ALLOW_ALL_ENV_PREFIX` is the only per-adapter datum.
Retained overrides (behavior genuinely differs):
- whatsapp_cloud `_allow_all_env_names` adds WHATSAPP_CLOUD_ALLOW_ALL_USERS;
`_is_dm_allowed` keeps its bare-wa_id normalisation.
- whatsapp_common `_entry_matches` -> phone/LID alias matching; `_live_dm_allow_from`.
- wecom `_is_group_allowed(chat_id, sender_id)` adds the per-group sender allowlist on
top of the shared chat-level rule; `_entry_matches` strips `wecom:user:` prefixes.
- qqbot `_entry_matches` (case-insensitive, `*`).
- yuanbao `AccessPolicy.is_group_allowed`: `open` groups still require the allow-all
opt-in (no runner-side mention gate), so it wraps the shared rule.
Behavior change: Weixin `_is_dm_intake_allowed` now denies a blank/whitespace principal
(the other four already did; safe variant chosen). Weixin's inline group gate now goes
through `_is_group_allowed` with identical verdicts.
Photon forked BasePlatformAdapter._send_with_retry before three fixes landed there: the server's
retry_after is honoured over exponential backoff, a long server penalty (> 60s) returns a typed
failure instead of sleeping inline (#91969), and an exhausted rate-limited send no longer posts the
failure notice inside the flood penalty. The fork got none of them. Photon's two genuine differences
are now hooks on the base — `_send_retry_is_final(result)` (structured auth/target refusals are
returned as-is, no retry, no plain-text resend) and `_send_plain_fallback(...)` (no Markdown banner,
richlink() bypassed) — and the 42-line fork is deleted.
Slack's `_retry_after_from_exc` and Discord's `_extract_discord_retry_after` parsed the header by
hand and only understood the numeric form; both now call agent.retry_utils.parse_retry_after_seconds
(numeric or HTTP-date, either header casing). Discord keeps its `retry_after` attribute path, the
`X-RateLimit-Reset-After` fallback and the 1s floor.
Behavior change: Photon retries now add up to 1s of jitter to the backoff and honour a server
retry_after; Slack/Discord recognise an HTTP-date Retry-After they previously ignored.
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.
`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.
Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
Nine surfaces (feishu, teams, slack, telegram, whatsapp_cloud, qqbot, matrix, discord, relay)
each re-derived the approval choice set — [Allow Once]; session + always unless smart-denied;
[Deny] — and four of them (discord, slack, teams, whatsapp_cloud) never adopted
base._format_exec_approval, so header/reason/smart-deny wording and truncation budgets drifted
per adapter. Three separate commits had to touch 5–9 adapters for one semantic fix.
BasePlatformAdapter.send_exec_approval now builds an ExecApprovalPrompt (shared text via
_format_exec_approval, shared `(label, choice, style)` rows via _exec_approval_actions) and
hands it to the `_send_exec_approval_prompt` hook. Each adapter keeps only its widget mapping
(~10–20 LOC); platform wording stays via the existing `_EA_*` class attrs, and a new
`_exec_approval_cmd_budget` hook lets Slack/Discord budget the command against their hard
message caps (3000-char section / 2000-char message) instead of computing it inline.
`_EA_REASON_BUDGET` covers Slack's 500 / Discord's 300 reason caps.
The runner used to detect button support by `hasattr(type(adapter), "send_exec_approval")`;
that is now true for every adapter, so `_renders_exec_approval_buttons` asks
`supports_exec_approval_buttons()` (hook overridden?) and keeps the duck-typed check for
non-BasePlatformAdapter classes.
Visible text changes (button semantics unchanged everywhere):
- Discord: the smart-deny line now follows the reason (was inside the header before the
fence); the truncation marker is "..." not "\n... [truncated]".
- Slack: smart-deny line follows the reason instead of the header.
- Teams: unchanged (same 2000-char preview, same smart-deny block).
- WhatsApp Cloud: identical text; body still capped at 1024.
- QQBot/relay: unchanged.
Eight adapters (discord, telegram, wecom, matrix, whatsapp, simplex, feishu, weixin) each kept a
copy of the delayed text-batch flush that base._enqueue_text_event schedules. Two correctness
fixes had landed in single copies only: Discord's asyncio.shield around the dispatch (#12444 —
a late chunk cancelling the flush task aborted the in-flight agent turn) and WeCom/Weixin's
synchronous task-identity check before the pop (a superseded task waking late popped the event
and the successor found nothing). The other adapters carried both bugs latent.
The base now owns `_flush_text_batch` with both fixes, plus the `_pending_text_batches` /
`_pending_text_batch_tasks` dicts and `_SPLIT_THRESHOLD` / delay attrs (defaults; adapters set
their own). Platform policy goes through three small hooks instead of a copied body:
`_text_batch_delay_for(pending)` (Telegram's fast/short tiers, WeCom's attachment-only wait),
`_pop_text_batch(key)` (Feishu's side count table) and `_dispatch_text_batch(event)` (Feishu's
per-chat lock). Telegram keeps its `_flush_buffered` body because its contract differs on
purpose — a cancel after the pop must hold-and-re-raise so teardown can stop a flush — and gains
the identity check there. Matrix's `_split_threshold` is renamed to the shared `_SPLIT_THRESHOLD`;
SimpleX exposes its single delay through the shared attr names.
`helpers.TextBatchAggregator` (zero users) sits inside the revert-scheduled PLUGIN-COMPAT block
and is left for that revert.
gateway.status.acquire_scoped_lock returns (acquired, existing_record). The irc, line and
buzz adapters tested `if not acquire_scoped_lock(...)`, and a non-empty tuple is always
truthy, so two profiles could drive one IRC nick / LINE channel / Buzz identity in
parallel. Route the three through BasePlatformAdapter._acquire_platform_lock (the seam
the other 8 adapters use), which unpacks the tuple, names the owning profile + PID in the
fatal error and honours the `--replace` takeover. Release goes through
_release_platform_lock; the private _lock_key bookkeeping is gone.
Error code changes from `lock_conflict` to `{scope}_lock` — both families are already
matched by gateway.restart.is_global_startup_conflict.
The buzz test mocked acquire_scoped_lock as a bare False, which masked the bug; it now
returns the real (False, record) contract, and irc/line gain the same conflict test.
The retained comment still claimed a scope-less multiplex caller could load an
OSS config; identity reads above it raise UnscopedSecretError first. Both the
comment and the regression test now say the same thing: callers are scoped, an
OSS profile whose scope lacks MEM0_API_KEY initializes, and a scope-less caller
is a spawn-site bug that raises.
read_json_or_empty returns {} for malformed JSON; returning that directly from
_load_config made a corrupt profile config.json yield an empty (silently
unconfigured) mapping where the pre-dedup loop kept walking to the legacy
file and then the env branch. Only a non-empty parsed object is now returned.
_finalize_all_traces called _get_langfuse(), which lazily BUILDS a client when the
launch-profile slot is empty. Under a multiplex gateway every turn runs scoped, so
only the per-home slots are populated; at exit (no scope) the credential read
raised UnscopedSecretError and the flush loop was skipped for every profile,
losing all pending traces. The finalizer now iterates the already-settled
clients (launch slot + per-home map) and never initializes. Test drives two
profiles under real home-override + secret scopes, then finalizes unscoped.
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
The 6-line try/except copy of "load_config() or {}" (also in doctor_live and
kanban_decompose, hermes_cli lane) took a full deepcopy per tool call for a
dict that is only read. load_config_readonly already degrades to the
last-known-good / defaults on a broken file, so the outer except was
defense-in-depth around a path that does not raise. Behavior change: none.
Tests already patch tools._load_config by name and continue to pass.
Four bundled plugins wrapped get_hermes_home() in a try/except that fell back
to ~/.hermes on ImportError (a2a/protocol._hermes_home, photon/auth
._auth_json_path, google_chat adapter inline, openviking done in the previous
commit). A bundled plugin cannot lose hermes_constants -- each already imports
gateway.* / agent.* from the same tree -- so the fallback was dead code that
was also wrong on Windows (%LOCALAPPDATA%/hermes) and under a profile override.
telegram's gmail-triage verb path used Path.home()/".hermes" outright, ignoring
profiles. mem0/_oss_providers baked os.path.expanduser("~/.hermes/mem0_qdrant")
into VECTOR_PROVIDERS at import time, so the Qdrant default landed in the
user's ~/.hermes for every profile; the default is now a lazy callable resolved
by vector_default_config(provider_id) at setup time. a2a/security.py only
changes its import (it borrowed protocol._hermes_home).
Behavior change: on Windows and under profiles these paths now follow the
active HERMES_HOME (they were previously anchored to ~/.hermes in the impossible
fallback / at import time); the default-profile POSIX layout is unchanged.
Test: tests/plugins/test_plugin_paths_follow_profile.py asserts each resolver
(a2a conversations, photon auth.json, mem0 qdrant default, openviking log)
lands inside a HERMES_HOME ContextVar override; sabotage red for the mem0
import-time constant and the photon ~/.hermes fallback.
Three shims caught UnscopedSecretError and degraded to "" (mem0._scoped_env)
or to os.environ (langfuse._secret, azure_identity_adapter._scoped_env). Under
multiplex os.environ holds the DEFAULT profile's .env, so the langfuse/azure
fallback could ship another profile's keys, and the mem0 fallback silently
routed a mis-spawned turn's memories into the default profile's account. The
exception exists to surface exactly that spawn-site bug (agent/AGENTS.md:
never add environ fallthrough, never swallow it). All three now call
agent.secret_scope.get_secret directly: with a scope installed a miss returns
the default; single-profile deployments (multiplex off) still read the process
env inside get_secret; a scope-less multiplex caller raises.
Behavior change: a mis-spawned child under gateway.multiplex_profiles now
fails loud with UnscopedSecretError instead of running silently unauthenticated
/ on the default profile's identity. The #99121 contract (OSS mode needs no
MEM0_API_KEY in scope) is unchanged and its test now installs an empty profile
scope, which is the situation the issue described; a genuinely scope-less
caller is asserted to raise in a new test.
Tests: tests/plugins/test_scoped_secret_readers_fail_closed.py (scope wins over
environ; scope-less multiplex raises) for langfuse + azure, sabotage red when
the langfuse fallthrough is restored; tests/plugins/memory/test_mem0_v3.py::
test_load_config_fails_closed_without_scope_even_for_identity_settings,
sabotage red with a swallowing wrapper reinstated.
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.
Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.
holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.
openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.
Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.
hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.
Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
`ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
hermes_state_messages._scrub_surrogates (0 callers) deleted.
The a2a invariant test skipped any synthesized token the canonical
redactor itself let through, so a boundary/shape regression would have
passed silently. Every class scrubs today, so the escape hatch goes and
the soft ">= 50" count becomes the exact registry size.
The gateway body test still asserted a "[REDACTED]" fallback marker that
no longer exists; redact_for_egress masks via _mask_token ("***"), so
assert that alone. honcho oauth.py's `re` import became unused when its
private pattern list moved to the registry.
Six independent line-parsers with three different quoting/comment semantics
read the same .env files: tools/skills_tool.load_env (strip("\"'"), no inline
comments), hermes_cli/managed_scope._parse_env (same, no export, no BOM),
web_server_cron._profile_env_value (plain utf-8, no BOM), profile_cmd
._env_file_has_key, env_loader._env_keys_defined_in_dotenv (utf-8, so a BOM'd
first key stayed "\ufeffKEY" and the dashboard profile scrub missed line 1),
mem0/_setup._prompt_api_key (startswith scan, no quote strip). The boundary
parsers (scrub key set, skill secret capture) therefore disagreed with the
parser that installs the profile scope.
Now every one is a 1-3 line forwarder onto load_env_file, and
hermes_cli.config.load_env is memo over it (public signature unchanged).
_parse_env_value moves next to its only caller in secret_scope.
load_env_file gains the same latin-1 fallback env_loader uses to install
into os.environ, so a mis-encoded file yields the same key set on both sides.
Managed .env keeps its fail-LOUD contract (decode error logs and ignores the
file) instead of load_env_file's fail-soft {}.
Behavior change: managed .env, skills_tool and mem0 setup now honour
`export`, quoted-value escapes and inline comments the way the profile scope
does; web_server_cron and the dashboard scrub tolerate a BOM.
Invariant test: a BOM'd/export/quoted/commented .env yields the same key set
via load_hermes_dotenv (installer), load_env_file (scope) and
_env_keys_defined_in_dotenv (scrub); fails with the old scrub parser.
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.
Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.
Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).
Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.
utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.
Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
The Matrix docs described six agent-exposed matrix_* tools and three MATRIX_TOOLS_ALLOW_* env gates that were never implemented — the tool names and gates appear nowhere in code. Docs now describe actual behavior: no Matrix-specific agent tools; reactions/redactions are internal to approval prompts and pickers; MATRIX_ALLOWED_ROOMS scopes responses. Fixes#100535.
Slack deprecates the Assistant messaging experience (assistant_view) in
February 2027: assistant.threads.setStatus/setTitle are replaced by
agents.sessions.setStatus/rename. slack-sdk 3.44.0 (Aug 27 2026) ships
the typed methods with drop-in-compatible signatures.
- adapter: capability probe on the AsyncWebClient CLASS (never instance —
mock auto-attributes lie), cached; status set/clear + thread title route
through agents.sessions.* when available, legacy otherwise
- pins: slack-sdk 3.43.0 -> 3.44.0 (pyproject messaging+slack extras,
lazy_deps, uv.lock)
- tests: autouse fixture pins the probe to legacy under the mocked SDK;
5 new tests cover both routing paths for typing, clear, and title
- docs: slack.md scope table + status-line notes mention both methods
fal launched H3 Max Turbo on Sep 3 (minimax/h3-max-turbo/{text,image}-to-video):
a throughput-tuned post-train of H3 Max with a 1080P tier Max lacks, at
$0.025/s 480p / $0.04/s 768p / $0.08/s 1080p list ($0.00625-0.02/s promo until
Sep 14). Schema matches Max's shape — required prompt_expansion_mode static
key, int duration 5-15, seed on both endpoints, i2v drops aspect_ratio — plus
the new 1080P resolution enum, so it reuses the existing family capability
flags with a Turbo-specific resolution alias map.
Schema verified against the FAL queue OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403 "Exhausted balance"); portal
allowlist/pricing needed for managed users on the 2 new endpoints.
Port from openclaw/openclaw#140531: users paste the application ID from the
Developer Portal's General Information page instead of the bot token (Bot
page); the gateway then fails at runtime with an opaque 401. A real bot token
is dot-separated base64 and never purely numeric, so the setup wizard now
rejects an all-digit answer with pointed guidance and re-prompts once. A
second consecutive numeric answer is kept (user override), and non-numeric
tokens are saved exactly as before.
Adapted to hermes: the guard lives in the Discord plugin's interactive_setup
(the active setup path per the #9983 review — legacy _setup_standard_platform
is not the dispatch target), mirroring the existing Telegram token-shape
validation in hermes_cli/setup_platforms.py.
The deepseek plugin is bundled and always loads, so the try/except around
the delegation import was defense-in-depth for a path that cannot fail.
Tests reduced to the two contracts that matter: /reasoning none reaches
the wire as thinking.disabled, and DeepSeek ids produce exactly the native
DeepSeek profile's output while non-DeepSeek families stay a no-op.
The bundled-plugin loader pops half-registered modules when a plugin
fails to load, so the lazy 'from plugins.model_providers.deepseek import
deepseek' could raise ImportError on every DeepSeek-routed CommandCode
turn — turning the soft 'thinking uncontrollable' bug into a hard
crash. Catch ImportError, log, and return the pre-fix no-op (review
feedback on #95241).
CommandCode fronts DeepSeek with vendor-prefixed ids
(deepseek/deepseek-v4-flash). DeepSeek V4+ defaults to thinking mode
when the thinking field is omitted, so /reasoning none changed the
Hermes session state but not the actual request -- the turn sat in
reflecting.../brainstorming... for minutes (#95232). Strip the vendor
prefix for DeepSeek-family ids and delegate to the native DeepSeek
profile's build_api_kwargs_extras (extra_body.thinking +
reasoning_effort mapping); other CommandCode model families keep the
base no-op behavior. The prior no-op tests codified the bug and are
rewritten to pin the new contract.
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):
- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
no duration are outside the unified video_generate surface and are dropped)
with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
live limits (nearest by value/height/ratio) and drops generate_audio/seed
for models that lack them (the API 400s otherwise); reference images ride
in input_references; local file inputs are refused (OpenRouter fetches
URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
never to a provider-supplied unsigned_urls host (kept from #103267)
Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.
Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
send_document() accepted reply_to but never passed it down, so attachments
always threaded from the cached per-address context (or not at all) even
when the caller named the message to reply to. The plain-text path
(_send_email) already honored it; the attachment path now does too.
The metadata half of #10131 (send_image rejecting metadata=) was already
fixed on main by the adapter parity pass. Diagnosis from #10131 and the
explicit-reply_to-wins shape from PR #10321 (which targeted the
pre-plugin path).
Fixes#10131
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Text and media batching advance the coalesced MessageEvent.message_id to
the newest message but left source.message_id at the first one, so after
a batch the reply anchor and the event id disagreed. Advance both together
at the two enqueue sites.
Same bug class as the Slack/Feishu picks: buzz, dingtalk, email, google_chat,
line, ntfy, photon, sms, teams, wecom and whatsapp already had the platform
message id on the MessageEvent but built the SessionSource without it, so
source.message_id consumers (reply anchor in run.py, /sethome synthetic-thread
check, relay _event_ids fallback, shutdown notice anchor) saw None. Only sites
where the id variable was already in scope are widened.