Follow-up on the salvaged #110928 (CLI `--parent-chat-id` / `--guild-id`):
- `_claim_for_sub` skipped a thread-shaped row that matched no `profile_routes`
entry at DEBUG on every tick. Legacy rows written before the flags existed can
never match a channel-level route (no `parent_chat_id`), so the notifier now
logs ONE WARNING per row naming the task, the thread and the re-subscribe
command. Still fail-closed: the events stay unclaimed.
- Docs: the kanban user guide explains the anchors and shows the Discord-thread
subscribe command under `profile_routes`.
- Test (red on origin/main): two collects → exactly one WARNING, events unseen.
Follow-up on the salvaged #111034:
- `_start_multiplex` published the enumerated home list before the gate had
filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
#100489) kept the thread alive but ticked every profile UNGATED — racing the
gateway that owns them for the same cron store. The list is now assigned only
after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
chore that respawns a ticker that ended without a stop request and logs the
outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
`executions.db` (no patched recover) no longer kills the ticker; a raising gate
keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
leaves a stopped one alone.
Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
Under `gateway.multiplex_profiles: true` every gateway-hosted session that ended
inside the gateway process logged `Memory provider 'openviking' on_session_end
failed: get_secret('OPENVIKING_API_KEY') called with no profile secret scope
active` and skipped its end-of-session commit. The provider lifecycle hooks
(`flush_pending` -> `on_session_end` -> provider teardown -> `close`) read
credentials and home at call time; the stop-time finalize pass and the
idle-cache sweep run on the main loop outside any adapter handler, so
`_run_in_executor_with_context` copied an EMPTY scope into the worker.
Route `_cleanup_agent_resources_off_loop` and `_finalize_session_off_loop`
through `_run_release_in_profile_scope` (the seam cache eviction already uses):
a scoped caller keeps its scope; an unscoped caller passes the session key and
the owner's profile scope is entered from it. The stop path passes the key for
both active and idle-cached agents. Cron's post-run cleanup thread is the
sibling seam and is fixed by the cherry-picked #110634.
Live repro (fake OpenViking recording the X-API-Key header, alpha profile,
multiplex on): base - WARNING + traceback, no commit; head - no warning,
`POST /api/v1/sessions/<sid>/commit` arrives with alpha's key.
Fixes#110622.
Review follow-up: a sentinel written before the create_time stamp has only
the claim time, but that is epoch seconds too, and the owner was born before
it claimed while a reuser was born after the owner died. One-sided compare
instead of trusting any live PID. Dead monkeypatch line removed from the test.
`detect_unclean_exit` decided "live owner mid-handover" by comparing the
sentinel's `start_time` (the ledger claim, `time.time()` seconds) with
`gateway.status.get_process_start_time(pid)` (proc clock ticks on Linux,
centiseconds elsewhere — its own docstring says it is only comparable with
itself). The two never matched, so every `--replace` takeover whose old owner
was still tearing down read as a crash and was logged/persisted as one.
The sentinel now carries the psutil `create_time` (stamped at claim since
8c3a35b69d), so the guard compares that with the live PID's create time —
same producer, same unit. A pre-stamp sentinel cannot disambiguate PID
reuse; a live PID is taken as the owner, as the psutil-silent case already
was. Test drives all three cases; it fails on the previous ledger.
Found while reviewing the start-attestation follow-ups (#110958).
Review follow-ups on the identity binding:
- The sentinel's `start_time` is `time.time()` at `record_startup`, seconds
after the process was born once imports finish, so comparing it with
psutil's create_time within 2 s would have read every real gateway as
undecidable and silently stopped the #109538 cold-start. `record_startup`
now stamps `create_time` (psutil birth via the existing
`process_identity._process_create_time`), `mark_exited` carries it, and
the attestation compares birth to birth. A sentinel from a gateway older
than the stamp falls back to the PID-only rule.
- A resume token written by pre-generation code and resumed by this code
probes the marker again instead of skipping the spawn.
- Horizon allows a 60 s backwards clock step; the unused `now` parameter is
gone; the create-time tolerance is a named constant; the read-then-unlink
in `_consume_start_attestation` is documented as best-effort.
`_attested_pid_exited_cleanly` matched the lifecycle sentinel by numeric PID
only, so a stale marker for PID 111 flipped from "clean exit" to "crash" once
an unrelated PID 222 lifecycle overwrote the sentinel, and a reused PID's clean
exit could vouch for a different life (#110020 review, gateway_windows.py:937).
`_write_start_attestation` now records `create_times: {pid: create_time}` via
the existing `process_identity._process_create_time`; `mark_exited` carries the
running sentinel's `start_time` onto the exited sentinel; the attested probes
fail closed for a bound PID whenever the sentinel cannot be shown to describe
that incarnation (other PID, start time off by > 2s, or no start time) —
"unknown" never reads as "dead". A missing sentinel still reads as dead, and
markers without `create_times` keep the PID-only rule.
Tests: two attestation tests (stale marker vs. moved-on sentinel → no
authority; own incarnation keeps authority / clean exit / legacy marker) and a
ledger test for the carried `start_time`. Mutation: with HEAD's prod files the
no-authority test and the ledger test fail.
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.
Test proven red on the salvaged head (adapter.sent == [] with `all`).
Review findings (Salt, adversarial pass on the two preceding commits):
- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
counted as an explicit mode because the gate tested key presence, while
the display resolver skips None and inherits. Null resolved to Slack's
tier default `off` and disabled cards, which is the default-off trap the
change exists to avoid. Explicit intent is now a non-None value (or the
env bridge). Tests cover null at each level plus null-over-global-all;
mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
destinations, so the name no longer described the field. Renamed to
`publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
described the opt-in as independent of tool_progress. Rewritten: cards
follow an operator-written off (including /verbose), null inherits, an
un-threaded chat with the card lane active shows no tool progress, other
native failures keep the editable fallback.
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.
Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.
Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.
Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
A /p/<profile>/ cron_job route resolved and fired the job from the gateway's
default home: `_fire_cron_job` ran `execute_job_for_event` outside
`_profile_scope`, so `cron/jobs.json` lookups (`resolve_job_ref`,
`claim_job_for_fire`) and the run's config/secrets came from the wrong
profile — the agent-mode path already scopes its run, this path did not.
`asyncio.to_thread` copies contextvars, so wrapping the call is enough.
Route-level `skills` were also being injected into the rendered prompt on
cron_job routes even though the docs say they are ignored (the job's own
skills apply); skip `_apply_skills` for cron_job routes so the per-run
context stays plain event text.
Test proven red on the pre-fix tree (home resolved to the default profile),
green after.
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.
- gateway/platforms/webhook.py: cron_job route mode — after the same
HMAC auth / rate limit / filters / script / idempotency as agent
routes, the rendered prompt becomes transient per-run context and the
job fires through execute_job_for_event on a worker thread (202
Accepted immediately). Startup validation rejects cron_job +
deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
the shared claimed-run body (_execute_job_now), so event fires share
at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
subscriptions; job ref validated (and canonicalized to the job ID) at
create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
CLI tests in test_webhook_cli.py.
Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).
- hermes_cli/model_selection_guards.py: new context_cache guard +
SelectionContext carrier + selection_context_for_agent() helper;
registry threads live-session facts to guards (6-arg signature with a
TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
Twelve credential-driven env branches were routed through _enable_from_env on
main (867e4158f0, #48820), which honors the loader's `_enabled_explicit`
marker. The flag-driven WhatsApp step was the one survivor: `_whatsapp` still
set `wa_cfg.enabled = True` on WHATSAPP_ENABLED=true regardless of an explicit
YAML disable. The dashboard's disable action writes only
`platforms.whatsapp.enabled: false` and leaves the env flag on disk, so the
Baileys bridge reconnected to real contacts on the next full restart
(reported live on #73289).
Route the truthy branch through _enable_from_env like every other platform;
WHATSAPP_ENABLED=false still forces a disable. Register the flag in
_ENV_ENABLE_CREDENTIALS so the one-time explicit-disable WARNING can name it.
Remaining scope of #96557 (the other ~20 sites) landed on main in 867e4158f0
and the config_env.py extraction; this is the delta. Fix direction from
@CryptoDombili in #73303.
Co-authored-by: Professor Dombili <Cryptodombili@gmail.com>
A message that arrives mid-turn is parked in the adapter's pending slot and drained in-band
by the runner's recursive _run_agent, a path that never touches on_processing_start /
on_processing_complete — so queued, interrupting and steer-demoted messages never got the
read-receipt reaction idle-session messages get, on every adapter implementing the hooks.
Bracket the drain in TurnRunner._run_agent_queued_followup with the hooks, resolving the
adapter from the follow-up's own source (multiplex-safe). Fires only for real inbound
platform events (message_id or raw_message present) and only for adapters that override
on_processing_start, so complete-only adapters (Google Chat, webhook) are not handed an
early completion. Cancels are classified like _process_message_background does.
Re-ported onto current main: the drain moved from gateway/run.py to
gateway/run_turn.py::_run_agent_queued_followup (#102117 decomposition) and the helpers
now live in the topical sibling gateway/run_turn_followup_ack.py.
Original commits (8828a3db37c2, f3c4f7a68c8d) by Mira Solari <268252643+mira-solari@users.noreply.github.com>:
--- 8828a3db37c2 fix(gateway): fire processing hooks for runner-drained queued follow-ups
A message that arrives while a turn is already running never gets the
processing-start acknowledgement — the 👀 read receipt on Slack, and the
equivalent in-progress reaction on Discord, Telegram, Feishu, Matrix,
Signal and Photon. It is not added-then-removed; the hook is never called.
`_run_processing_hook("on_processing_start", …)` has exactly one call site,
inside `BasePlatformAdapter._process_message_background` (base.py:5403).
`handle_message` takes the busy branch at base.py:5165, parks the event and
returns at base.py:5316 — above `_start_session_processing`, which is the
only thing that spawns `_process_message_background`. The parked event is
then drained in-band by the runner (`_dequeue_pending_event`, run.py:23277)
and replayed through a recursive `_run_agent` that touches no adapter hooks.
Because `get_pending_message` pops, the adapter's own drain (base.py:5823)
— the one path that would fire the hook — finds an empty slot. The gap is
structural, not a race, and it is shared by every mode (`queue`,
`interrupt`, steer-demoted-to-queue), by `/queue`, by photo-burst and
text-debounce flushes, and by voice drains.
Fire the existing hook pair around the recursive call. The follow-up now
gets the same lifecycle an idle-session message already gets, on every
platform, through one shared site.
Details that shaped the placement:
- Fired after the depth-cap requeue (run.py:23372) and after every discard
and early return in the block, so no path can strand an in-progress
marker: from that point on, control either reaches the recursion or
raises, and both close the hook.
- Fired before `_refresh_agent_cache_message_count` so that re-baseline
stays adjacent to the recursive call it exists to protect — inserting a
reaction round-trip between them would widen the window in which the
cross-process coherence guard (#45966) can trip on our own writes and
rebuild the agent, destroying the prompt-cache prefix #46237 preserves.
- The hook adapter is resolved from the follow-up's own source, not the
completing turn's: a multiplexed gateway can route it to a different
profile's adapter, and only that instance holds the per-message reaction
state.
- Gated on a truthy `message_id`, which is what every adapter's own hook
already checks. Synthetic drains (`/goal` continuations, wake-ups, CLI
hand-offs) carry no id and stay silent; `interrupt_message` and leftover
`/steer` carry no event at all.
- Cancellation maps to CANCELLED rather than FAILURE, matching
_process_message_background — Telegram clears the marker on CANCELLED and
Signal deliberately leaves it, so the distinction is load-bearing.
Outcome is SUCCESS unless the recursion raises. That mirrors the existing
non-queued contract, where a run returning `failed: True` still delivers a
diagnostic message and reports SUCCESS; making the outcome track agent
failure is a separate change that should apply to both paths at once.
No new config, no new env var, no new hook, no change to message
construction or role alternation.
Fixes#72502
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--- f3c4f7a68c8d fix(gateway): don't hand a completion to complete-only adapters; cover raw-envelope events
Self-review of the previous commit found three defects in it. All three are
about the *guard*, not the mechanism.
1. Complete-only adapters were handed an unpaired completion. Google Chat
(google_chat/adapter.py:2766) and webhook (webhook.py:901) implement
on_processing_complete WITHOUT on_processing_start, and theirs is
end-of-cycle teardown, not a reaction: Google Chat reaps the typing card,
patching it to "(no reply)"/"(interrupted)", and webhook ends the
per-delivery session. The drain fires before the follow-up's reply is
delivered — delivery happens after the whole chain unwinds back into
_process_message_background — so a Google Chat space would get a permanent
"(no reply)" tombstone on every queued follow-up, and the real answer would
then land as a separate message. Now gated on the adapter actually
overriding on_processing_start: we bracket, so both halves must be ours.
2. The message_id gate made the fix a no-op on Signal, which the previous
commit message claimed to fix. SignalAdapter never sets message_id
(signal.py:749-766) — its hook keys off raw_message["sender"] and
["timestamp_ms"] via _extract_reaction_target — and Discord's start hook
reads raw_message too. Gate is now message_id OR raw_message; synthetic
drains still carry neither, so /goal continuations stay silent.
3. Cancellation was classified unconditionally as CANCELLED. base.py:5862-5868
only reports CANCELLED for a task in _expected_cancelled_tasks (/stop, /new,
/reset, adapter cleanup) and downgrades anything else to FAILURE. Signal and
Matrix deliberately LEAVE the marker in place on CANCELLED, so the previous
version would strand exactly the marker this PR exists to clear. Now mirrors
base.py via _followup_cancel_outcome().
Also moves _refresh_agent_cache_message_count inside the try. It awaits DB I/O
and guards it with `except Exception`, which does not catch cancellation, so a
/stop landing there escaped both handlers and stranded the marker. Ordering is
unchanged, so the prompt-cache adjacency the previous commit describes still
holds.
Two new tests, both failing against the previous commit:
test_complete_only_adapter_is_left_alone and
test_raw_envelope_only_followup_is_acknowledged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
resolve_runtime_provider returns the bare billing class 'custom' for every
named providers:/custom_providers: entry; the configured id only survives in
requested_provider. All three fallback resolvers (gateway, TUI/desktop, cron)
persisted runtime['provider'] as the agent identity, so an automatic fallback
labeled the session 'custom' in the UI and billing rows, while a manual
/model switch to the same provider showed the configured name.
New shared helper hermes_cli.fallback_config.effective_runtime_provider()
upgrades the bare class back to the entry's configured identity (ad-hoc
provider: custom entries stay unchanged), applied at all three sites —
same class as the delegate_tool fix.
Under gateway.multiplex_profiles the routed handler runs inside
_profile_runtime_scope, but the adapter's delivery side
(_process_message_background -> _extract_response_content, weixin's own
send()) extracts and validates the reply's MEDIA: / bare-path attachments
after that scope was reset. Docker translation in platforms/base.py
(_docker_sandbox_dir_candidates via get_active_profile_name,
_parse_docker_volume_mounts via the scope-aware TERMINAL_DOCKER_VOLUMES)
therefore resolved a secondary's /output or /root path against the DEFAULT
profile's sandbox and mounts: dropped as "not found on this host", or a
same-named file from the default's mount delivered instead (#109024).
Add GatewayRunner._media_delivery_scope_for_source (home + terminal policy,
no secret hydration: path validation reads no credentials and runs on the
loop) and enter it from BasePlatformAdapter._media_delivery_scope around
the two extraction sites that run outside the turn scope. The streamed path
(_deliver_media_from_response), the background task and cron delivery
already run inside their profile scope.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.
One request now carries a model pick AND its effort on every surface:
- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
<level>` (validated against `parse_reasoning_effort`; unknown level ->
`MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
`ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
high` applies the effort AFTER the agent swap (`switch_model` re-resolves
`reasoning_config` from config.yaml, so an earlier write is clobbered) with the
pick's scope (session; config on `--global`; `--once` snapshots and restores it).
The `/model` picker gains a third stage, "Reasoning effort for <model>", built
from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
`config.set model "X --reasoning high"` applies after the swap; session pin
(`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
one-turn restore carries `reasoning_config`; re-emits `session_info` so the
status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
`<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
`_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
the Copilot-only inline prompt; Copilot keeps its per-model level set via
`github_model_reasoning_efforts`, other routes get the ladder, catalog
`supports_reasoning=False` skips it) plus a "Reasoning effort for the current
model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
with the same step (+ "Provider default"), stored as
`auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
the task list ("openrouter · model · high"), cleared by "Reset all to auto";
tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
it.
Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
"Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
errors "Model names cannot contain spaces"; after switches and `config.get
reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
"reasoning: high", status bar "fable 5.1 high".
`write_runtime_status` re-stamps the previous writer's `gateway_state.json` in place and
only `_record_served_profiles` (multiplex on) ever wrote `served_profiles`, so a
multiplexer's list survived into a later non-multiplex run of the same home. Every
`hermes -p X` surface then kept treating X as served by that live default gateway: exit
78 on start/install, "running via the default-profile multiplexer" on status (review of
#108352, finding D, second half). The secondary-profile phase now writes an empty list
when multiplexing is off; an empty list is the authoritative "serves nobody else" the
readers already honour.
Both default-profile process matchers (`gateway.status._command_line_belongs_to_profile`
and `hermes_cli.gateway._scan_gateway_pids._matches_current_profile`) rejected a named
gateway with a substring test for `--profile ` / ` -p `, which the equals spelling the
CLI pre-parser accepts (`--profile=ops`) slipped past. The default home's identity check
then adopted that gateway's PID, and a default-profile `gateway stop` with no pid file
scanned the process table and could SIGTERM the named gateway (review of #108352,
finding E). Both sites now ask `profile_flag_value()`, the same tokenizer the named
branch already uses.
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.
One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
`main` is a valid profile name (only hermes/default/test/tmp/root/sudo are
reserved), but _session_key_namespace mapped it to `agent:main` — the default
profile's namespace. Both profiles then built byte-identical keys: one routing
entry, one cached agent, and, since 75ae2859b9 pinned default-namespace
keys to the launch store, profiles/main's scoped sessions were written into
the ROOT state.db instead of profiles/main/state.db.
Key the `main` profile as `agent:main~` (`~` is outside the profile-id
alphabet, so the marked form cannot be any other profile's id) and give the
namespace slot one inverse, profile_from_session_key_namespace, used by the
store's key parser, _parse_session_key, the update-marker profile reader and
the profile-delete eviction prefix. Default keys stay byte-identical.
For a multiplexed secondary the middleware wrote a top-level
YUANBAO_HOME_CHANNEL key that load_gateway_config never reads and skipped the
(correctly suppressed) process-env write, so cron and home-channel delivery had
no target in-process and none after a reload either. Persist through the
gateway's persist_home_channel (the profile-aware config path every /sethome
uses) and set the live PlatformConfig.home_channel; the process env is still
untouched under a secondary's scope.
Refs #108440 (ehz0ah inline, gateway/platforms/yuanbao.py)
GatewayAuthorizationMixin._chat_scoped_grant read only the scoped
{PLATFORM}_ALLOW_BOTS env var, so a secondary whose config.yaml said
`allow_bots: all` was admitted by its own adapter and then denied centrally
(Discord, Slack Workflow posts with user=None, Feishu, Telegram). The gate now
resolves the routed adapter's effective policy with the same reader as intake:
scoped env → adapter YAML → none. Mention requirement and loop guard are
unchanged.
Refs #108440 (finding 7)
545e74d0ea correctly stopped writing telegram.proxy_url into TELEGRAM_PROXY
for a multiplexed secondary, but _build_ptb_requests still resolved the proxy
only from that env var, so the secondary silently connected direct (or via the
default's proxy). #100448 had deliberately left this bridge unscoped for that
reason; this finishes the consumer migration instead.
_apply_yaml_config seeds proxy_url into extra and resolve_proxy_url gains a
`configured` rung: scoped TELEGRAM_PROXY → the profile's YAML → HTTPS_PROXY/
HTTP_PROXY/ALL_PROXY (trust_env) → macOS system proxy, with NO_PROXY semantics
unchanged.
Refs #108440 (finding 6)
One reader (gateway.platforms._shared.extra_or_secret) now implements the
precedence every per-profile setting follows for the OWNING profile:
explicit scoped env/.env → that profile's config.yaml (PlatformConfig.extra)
→ the adapter's default. A scoped miss returns the default, never the launch
process's os.environ; single-profile / default-profile installs keep the
documented env-over-YAML contract.
Why: 545e74d0ea (#108705) stopped bridging a secondary's YAML into the
process env and moved readers to config.extra, but the shared reader and the
hand-rolled helpers in Discord/Slack/Matrix/Telegram consulted YAML FIRST and
then fell back to a scoped env read. Two bug classes followed (#108440
post-merge review by andrexibiza, #109032):
- an explicit env value could no longer beat YAML for the owning profile
(DISCORD_ALLOW_MENTION_EVERYONE=false lost to allow_mentions.everyone: true;
TELEGRAM_REACTIONS=true lost to the stock reactions: false);
- a secondary that OMITTED a key inherited the launch profile's bridged env
through the fallback (Matrix process_notices/session_scope, Discord
auto_thread/reactions/mentions, Slack reactions/ignored_channels).
Consumers migrated to the shared reader: Discord _build_allowed_mentions and
_extra_or_env_flag; Slack _slack_allow_bots, _reactions_enabled (the
_extra_or_env_* getters already used it); Matrix _extra_truthy, _extra_csv_set,
session_scope, reactions, require_mention parsers, and — new — the
allowed_users / ignore_user_patterns consumers that never read the seeded YAML
lists; Telegram _extra_bool, _extra_str_set, _reactions_enabled; Feishu
allow_bots; WhatsApp dm_policy/group_policy.
Refs #108440, #109032
Every other WEIXIN_* tunable in this __init__ block (dm_policy,
group_policy, rate_limit_circuit_*, send_chunk_*) already reads
extra-first with a scoped-secret fallback via _extra_or_secret(). This
one field was missed and still fell back to a bare os.getenv(), so a
secondary profile without its own split_multiline_messages setting
silently inherited the default profile's process-env value instead of
the coded default.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 44e3c0d5cf276a4ef78e676f2631fe89b11b5893)
The raw-output watcher modes (all/result/error) and the interim running
update sent the bracketed debug wrapper with the internal process id
(`[Background process proc_… finished with exit code N~ Here's the final
output: …]`) to Telegram/Discord/Slack chats. Reuse the concise one-line
status header for every mode and append the bounded, ANSI-stripped output
tail in a code block; the running update gets the same shape.
Salvaged from #54266 (rebased onto the post-#102117 run_notifications
sibling; the concise mode had landed in between, so the header is shared
rather than reimplemented). Also covers #13122 (ANSI stripping).
`_persisted_billing_route` (idle `/usage` account-limits lookup) was the last
reader of the lifetime-dominant route, so it queried the retired provider's
account after a switch. Point it at `get_recent_session_model_route` and
delete the dominant query, which no longer has a caller.
`GatewayRunner` never had a `get_adapter` method, so every gateway `/save`
(Telegram, Discord, ...) rendered the file and then failed with
"'GatewayRunner' object has no attribute 'get_adapter'". Resolve the adapter
through `_adapter_for_source`, the profile-aware lookup the rest of the runner
uses, so multiplex secondaries deliver through their own bot rather than a
missing key on the default map.
Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Baophan00 <109447498+Baophan00@users.noreply.github.com>
Salvage follow-up to #107955 (Alex Tu) and #109494 (EloquentBrush0x):
- _default_profile_secret_scope: drop the import try/except, the
current_secret_scope() short-circuit and the build-failure fallthrough.
The tick always runs in a fresh Context (no scope can be present) and a
failure to build the launch profile scope must surface, not silently
degrade to an unscoped tick.
- Regression test proven red on origin/main: run the real
auto_decompose_tick through _to_thread_process_service under multiplex
and assert the decomposer reads the launch profile .env value; ports the
contract from #57837 (srojk34) to the post-refactor dispatcher.
- Trim the #109494 test docstring to the invariant.
With gateway.multiplex_profiles on, agent.secret_scope.get_secret() fails closed
whenever no profile secret scope is installed. auto_decompose_tick runs through
_to_thread_process_service in a fresh context, so the decomposer's credential
read raised UnscopedSecretError on every tick before the aux LLM was called,
and every triage card stayed in triage forever (logged at INFO only).
Wrap the tick in _default_profile_secret_scope(): when multiplexing is active
and no scope is installed, build the gateway default profile's scope (the same
home load_gateway_config_for_runner uses) for the duration of the tick. No-op
for single-profile gateways and when a scope is already active.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c47e3ea6f8a5750182b7534160184210b8d5a110)
Same class as the Matrix readers: since 545e74d0 the WhatsApp YAML bridge seeds
`free_response_chats` into extra via the "csv" kind (any non-None value), so a
present-but-blank `free_response_chats: ''` reaches `_whatsapp_free_response_chats`
as '' and its `if raw is None` fallback never reads WHATSAPP_FREE_RESPONSE_CHATS.
Before 545e74d0 the hook returned None and the key never reached extra, so the
env CSV applied. Route it through the shared extra_or_secret reader (blank =
unset; an explicit empty list stays "no chats").
Sibling sweep of every "csv"-kind bridged key: dingtalk, mattermost and slack
readers already go through extra_or_secret and their bridges seeded extra before
0.21.2, so 008caa88 deliberately kept blank-means-clear there; whatsapp allow_from
uses key-presence semantics by design (_select_dm_allowlist); buzz reads env
first. Telegram allowed_chats: '' shadowing the env var is pre-existing (identical
on v2026.9.7, via the shared-key bridge) and left as is.
The adapter fix decoded `'["-100","-200"]'` before comma-splitting, but
the runner's central gate in gateway/authz_mixin.py::_coerce_allow_set
reads the same YAML-bridged env chain (TELEGRAM_GROUP_ALLOWED_CHATS,
TELEGRAM_ALLOWED_USERS via _auth_env) and still produced
{'["1"', '"2"]'}, so a group message admitted by the adapter could
still be rejected upstream.
Move the decoder to gateway/platforms/_shared.py, which both the adapter
and authz_mixin already import from (no plugin -> gateway cycle), and
route _coerce_allow_set through it. One invariant test on the runner
side, red before this change.
Two review follow-ups on request_review's artifact staging.
Staging copies files into attachments/<tid>/ inside the write txn, but
the copy is a filesystem side effect the rollback cannot undo. When a
later step in the same txn raised (anything other than
ArtifactPreservationError, e.g. run bookkeeping), the task correctly
stayed `running` but the copy leaked, so the retry staged `a_1.txt`
beside an orphan `a.txt`. _stage_completion_artifacts now returns the
copies and request_review discards them on any exception around the
txn, reusing the same unlink/rmdir logic the staging helper already
had. complete_task is left alone: its txn has a different shape (the
early-return paths and acceptance recording) and its cleanup runs the
scratch workspace anyway, so it was not the identical one-line change.
The notifier unions payload['artifacts'] with paths parsed from the
summary prose and dedupes by full path only. For `review_requested` the
scratch original still exists (the reviewer's completion is what
deletes it), so a summary naming the original uploaded the file twice:
staged copy and original. Prose-parsed paths whose basename matches a
staged artifact are now skipped; `completed` delivery is unaffected in
practice because there the original is already gone by delivery time.
request_review ignored artifacts entirely, so a review-bound card lost every
file its handoff named: the reviewer's complete_task is what runs
_cleanup_workspace over the managed scratch workspace.
Stage declared (explicit artifacts argument or metadata["artifacts"]) and
prose-referenced files into the task's durable attachments dir at the review
handoff, exactly as complete_task already does, carry the staged paths in the
review_requested event payload, and let the gateway notifier upload them (its
guard widens from completed to review_requested). ArtifactPreservationError
still rolls the whole transition back: the task stays running and retryable
with no attachments and no event.
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.
reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.
test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.