Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
Review of #90448 by @andrexibiza: adding _ensure_reconnect_watcher_running()
to the already-queued branch of a fatal callback is still an event-coupled
check. It needs a later fatal error from some other platform to arrive, and
#81036 makes that less likely rather than more -- it publishes the queue
before disconnect and drops the failed adapter from the live map, so after
the watcher's supervised restart budget is spent there may be no adapter
left to emit the event recovery is waiting on.
That is the state #72366 (salvage of #71867 by @ygd58) restored supervision
to close: queued work exists, the watcher is dead, and nobody owns the
invariant. Supervision being finite is correct; having no owner past the
budget is not.
_spawn_supervised now takes on_give_up, invoked when it abandons a task --
the supervisor is the only thing that knows it has. The reconnect watcher
uses it to hold:
while _running and _failed_platforms is non-empty, either a reconnect
watcher is live or a bounded respawn is scheduled.
Empty queue: leave it down and log; the enqueue path spawns a fresh watcher
the moment something depends on one. Non-empty: a bounded slow tier at
_RECONNECT_WATCHER_SLOW_RETRY_SECS (300s) for _MAX_SLOW_WATCHER_RESPAWNS (6)
attempts, standing down early if the queue drains or a watcher returns on
its own. Exhausted: one loud error naming the platforms left unattended.
The ceiling is (1 + _MAX_SUPERVISED_RESTARTS) x (1 + _MAX_SLOW_WATCHER_RESPAWNS)
spawns -- 42 across at least half an hour -- because each slow attempt hands
the watcher a fresh supervised budget. A test asserts that ceiling so it
cannot quietly become a restart loop.
Deliberately NOT included: requesting a process restart when the slow tier
is also exhausted. Taking down every healthy platform to heal a sick one is
a blast-radius policy decision for a maintainer.
Two things this turned up:
- _spawn_supervised did not thread on_give_up through its own backoff
respawn, so the callback was lost after the first restart and the give-up
branch had no owner at exactly the moment it needed one -- the same defect
the on_spawn docstring warns about, one parameter over.
- Three call sites repeated the (factory, name, on_spawn) triple, whose
on_spawn half is load-bearing. They now go through
_spawn_reconnect_watcher().
_supervised_backoff() names the previously-inline exponential schedule so
the exhaustion tests can collapse it; production behaviour is unchanged.
Refs #90386
_ensure_reconnect_watcher_running() exists for one situation: the reconnect
watcher has exhausted _MAX_SUPERVISED_RESTARTS, so _spawn_supervised has logged
"giving up restarts" and will never bring it back on its own (#70344, and the
supervised-restart half of #71758). It had exactly one call site, inside the
newly-queued branch of _queue_retryable_fatal_platform.
That branch is unreachable for a platform already in _failed_platforms, which
is the only kind of platform the watcher can have been retrying long enough to
burn five rapid restarts on. So the backstop could not fire in the one state it
was written for.
The failure is silent by construction. The early return logs nothing, so there
is no "queued for background reconnection" line. The stranded check in
_handle_adapter_fatal_error_detached deliberately treats a queued platform as
safe, so the gateway does not exit for the service manager either. With another
platform still connected, self.adapters is non-empty and the "gateway staying
alive, watcher will retry in background" branch is skipped too. A retryable
fatal error can therefore produce a single ERROR line and then nothing: the
platform sits in the queue that nobody is draining until someone restarts the
process by hand (#90386 reports 4h17m of that, with cron unaffected throughout).
Call the ensure on the already-queued path as well. It is already idempotent
and already cheap: it returns immediately unless the tracked task is done, and
it routes through the same on_spawn handle tracking, so a live watcher is never
duplicated.
The queue entry itself is deliberately left untouched. Re-enqueueing would
reset attempts and next_retry, restarting the backoff ladder on every fatal
error and hammering a provider that is already refusing the connection.
Follow-up to the salvaged #91986: the per-row clear still left rows the
loop had not reached exposed — a slow send ahead of them could hold the
loop past the inbound-gate timeout and let
_schedule_resume_pending_sessions replay those turns. Clearing every
claimed row up front closes the duplicate window; claiming already
spent the redelivery attempt, so the ledger retry path is unchanged.
Restart notification and obligation redelivery ran before the
startup-restore gate opened, so one hung Telegram send queued inbound
on every platform. Bound those sends with the same timeout the resume
gate already uses, and clear resume_pending before send so a timed-out
redelivery cannot also replay the turn.
Add the foreign-holder guard to gateway/session.py::_rebuild_fts_once(),
the third FTS rebuild path that was not covered by the original fix.
Also add a comment explaining why _fts_runtime_rebuild_attempted is set
before the foreign-holder check: the fail-open path that follows
persists FTS_STALE_KEY so the next startup retries via _recover_stale_fts.
_turn_transcript_messages pre-classified every message with
_is_compressed_summary_message (full content flatten + prefix scan), then
_message_response re-ran the same classifier inside its projection --
2x per non-summary row, 3x per summary row on every run.completed emit.
The outer guard was redundant: _message_response already yields
display_kind hidden for pure handoffs. One projection call per row now.
Surfaced by the post-merge simplify re-review of #91517/#91535.
Artifact receipts live only in memory, so files left behind by a dead
process were unreachable but persisted forever despite the advertised
300s TTL — a retention failure on the surface meant to be ephemeral
(blocker 4 of andrexibiza's #91535 review). A fresh ArtifactStore now
removes every artifact-id-shaped file and stale *.tmp with no index entry
(at construction the index is empty, so all such files are orphans).
Non-artifact-shaped names are untouched. Regression: store -> recreate
store over same root -> orphan+tmp gone, unrelated file kept.
The global broker snapshotted browser.extension_control.developer_mode once
at construction, so flipping it OFF in config did not revoke raw CDP/eval
from already-attached controllers until process restart — a revocation
failure at the highest-privilege browser surface (blocker 3 of
andrexibiza's #91535 review). select() now consults the live config on
every privileged selection (explicit bool still pins for tests); off->on
also unlocks without restart. Regression test drives both directions
against an attached controller. Also drops the dead back-compat
_artifact_store property (zero readers).
The frame handler returns reply dicts (heartbeat/detach acks) that the WS
reader loop sends back; the -> None annotation was the only new ty
diagnostic vs origin/main.
Addresses both merge blockers from @andrexibiza's review of #85351:
1. HTTP-uploaded artifacts could never be consumed by broker dispatch:
artifact_scope_key hashed (principal, session, family), the HTTP routes
store with an EMPTY session (API-key auth has no server session) while
broker validation carries a session-bearing ControllerScope — every
real upload->dispatch journey died with ArtifactScopeMismatch
(reproduced before fixing). Canonical ownership is now
principal/transport-family (documented in the scope-key docstring);
ids stay unguessable server-minted 32-hex and downloads one-shot.
New composition regression: HTTP-shape upload -> registered controller
scope -> broker artifact dispatch, mutation-checked (re-adding session
to the key makes it fail).
2. The 'profile-scoped' artifact store was first-profile-wins process
state: one adapter-level singleton pinned profile B to profile A's
physical root on multiplex listeners (same frozen-handle class as
#88734). Stores are now cached by resolved profile, and the broker
selects the store from the controller scope's profile_id (default-slot
fallback preserves single-profile/test behaviour). New A/B multiplex
regression proves distinct physical roots regardless of touch order.
Also documents the advertised ticket_expires_at as best-effort wall clock
(broker enforces expiry monotonically) per review feedback.
browser_control_enabled()/browser_control_developer_mode() run on every
browser tool call and inside every check_fn evaluation (uncached for bound
sessions). Both are pure reads of nested dicts; load_config()'s defensive
deepcopy (~135us/call) is wasted there. Same pattern as the other read-only
config probes.
Surfaced during review of PR #85351.
- Rename the broker's TicketInvalid to ControllerTicketInvalid: the same
exception name already exists in hermes_cli/dashboard_auth/ws_tickets.py
and BOTH are caught in the same WS auth flow this feature touches — two
unrelated same-named exception types in one blast radius invited a wrong
except clause.
- Import the 'server-internal' sentinel identity from its canonical
definition (ws_tickets.INTERNAL_USER_ID/INTERNAL_PROVIDER) instead of
re-declaring the strings; drift would have silently broken the
internal-peer exclusion in _is_authenticated_identity.
Surfaced during review of PR #85351.
attach/disconnect/detach acquire a per-controller threading.Lock that a
worker-thread dispatch can hold for up to 10s while blocking on the event
loop to transmit its command frame (run_coroutine_threadsafe +
result(timeout=10)). Acquiring that lock synchronously from loop context
(controller WS finally, frame handler, gateway WS teardown) could park the
ENTIRE gateway event loop behind the send bridge — a deterministic
multi-second global stall whenever controller teardown raced an in-flight
command. All loop-context broker calls now go through asyncio.to_thread,
matching the existing offload pattern for _close_sessions_for_transport.
Surfaced during review of PR #85351.
The router treated any server-stamped principal as a bound lane, so with the
flag ON every authenticated dashboard/API session lost the legacy browser
backend even when no extension controller ever registered (scope_for_session
returns None -> ControllerUnavailable, no fallback) — while check_fns still
advertised the tools via the legacy OR-gate.
New broker.lane_registered() distinguishes the two cases:
- lane never registered -> generic callers keep the legacy backend
- lane registered (controller offline/ambiguous) -> fail closed, unchanged —
a control-this-tab session never silently jumps to another browser
Also makes the four non-allowlisted wrapped tools (cdp/console/vision/
get_images) behave correctly for never-registered lanes (legacy backend)
while staying fail-closed for registered lanes.
Surfaced during review of PR #85351.
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.
Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.
Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
The 7-key internal-fields tuple was inlined twice (agent/compaction_display.py
and _project_client_message); a drift between the copies would silently leak
one internal field class through the API projection. Surfaced during review
of PR #85442.
Project client-visible session messages through the canonical compaction classifier. Hide standalone handoffs, unwrap merged carriers to their authentic prior-tail content, strip inherited internal fields, and keep model-facing recovery history unchanged.
The salvaged skip condition keyed on _use_draft_streaming alone, which
also suppressed the explicit REQUIRES_EDIT_FINALIZE pass when a
draft-streaming run had degraded to edit-based delivery (draft failure
fallback sets _message_id). Key the skip on the got_done update being a
fresh persistent send through the native-draft transport (_message_id is
None), which is the only case where the update already carried its own
finalize.
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
Two review nits on the clarify ambiguity fix:
- _approval_send_outcome swallowed the failure detail the old inline
callers logged (scheduling exception text / SendResult error). Both
the approval and clarify lanes now share one warning with the detail,
logged in the classifier itself.
- The clarify tests pinned the disposition helper but nothing proved
the ambiguous branch actually reaches the bounded wait. The
send-then-wait sequence is extracted to _clarify_send_then_wait (the
callback closure now just binds context onto it) and the suite gains
caller-path tests: ambiguous/sent -> wait_for_response with the
generated clarify_id and configured timeout; definitive failure ->
sentinel without waiting; no-response timeout sentinel preserved;
plus caplog assertions that failed sends log their detail.
Relay + gateway sweep 289/289.
The clarify caller treated a send-scheduling timeout as a definitive
failure: clear_session() + '[clarify prompt could not be delivered]'.
Same physics as the approval card fixed earlier in this PR — the card
may well have posted with a late connector ack — so the teardown ran
out from under a rendered clarify card and the user's answer resolved
nothing.
New _clarify_send_disposition() routes the outcome through
_approval_send_outcome: only a DEFINITIVE failure (error result,
non-timeout exception, no future) clears the registration and aborts;
ambiguous logs a warning and falls through to wait_for_response, whose
existing bounded wait already handles the truly-lost-card case. This
makes the boundary rule stated in the ambiguity test docstring hold
for the clarify lane, not just approvals.
Tests: 5 disposition tests mirroring the approval suite, including
clear_session-not-called on timeout. Mutation-verified: folding
ambiguous into the failed branch sends
test_timeout_keeps_registration_armed_and_proceeds_to_wait red.
Relay + gateway sweep 283/283.
Round 2 of the approval-turn stuck-stream hunt. Round 1 (interim-marked
acks) fixed the draft-hijack-by-matching path — live logs confirm the
absorption fallback no longer fires — but the freeze persisted because
of a second, deeper defect on the same codepath:
_consume_prompt_response executes ON the transport read loop (inbound
frame -> _handle_frame -> _inbound handler). The handler awaited
self.send() for its '✅ Approved once' ack — but send() blocks on an
outbound_result future that ONLY the read loop can resolve, and the
read loop is blocked inside this very handler. Guaranteed self-deadlock
for the full outbound timeout (30s) on EVERY button tap. While wedged,
everything on the transport starved: draft appends (the frozen stream
right after approving), sibling approval-card sends (timed out into
'possibly-delivered' — the observed double-approval ambiguity), and the
turn's seal (timed out ambiguous -> plain-send fallback -> duplicate
final). Log signature was the tell: card-send timeout at tap time, no
absorption INFO, no seal-failed WARNING, no suppression line.
Fix: _send_lifecycle_ack() — acks ride a background task with strong
ref retention; the handler returns immediately and the read loop keeps
consuming, so the ack's own result frame resolves normally. Applied to
all six lifecycle sends (approval ack, slash-confirm ack + result text,
clarify acks, expiry notice). Acks are cosmetic by contract; failure
logs at debug and never breaks the reader.
Tests: new deadlock-shape test (gated transport send; handler must
return within 1s and the ack must still egress afterwards — RED on the
awaited version via TimeoutError at the exact deadlock), prior 3 tests
green with a yield for the background task. Targeted sweep 203/203.
Live finding (rc.4 staging, 100% reproducible on approval turns): after
resolving an exec-approval prompt_response, the adapter sends a short
ack ('Approved once'). _prompt_reply_metadata carried only placement
metadata (thread_id) — no per-turn identity, no interim marker — so
send()'s single-open-stream fallback (review B2) matched the approval
turn's OWN live draft and sealed it with the ack text. From there,
silently: every later append died on the post-seal tombstone (built for
millisecond stragglers, deliberately quiet), freezing the visible draft
mid-word; the turn-final found no open draft and fell through to a
plain send — the duplicate 'fallback' message. No suppression line, no
seal-failed warning: the log signature was pure absence.
Fix: _prompt_reply_metadata stamps _interim_send=True, which send()
already honors by bypassing draft matching. One source covers the whole
lifecycle class (approval ack, slash-confirm ack, prompt-expired
notice — all six call sites route through it).
Observability (the quiet parts, out loud):
- single-open-stream absorption now logs at INFO with the absorbed key;
- the FIRST post-seal tombstone swallow per draft key logs at WARNING
(bounded FIFO dedup) — one swallow is the normal straggler race, a
burst means a live stream was sealed mid-flight by someone else.
Tests (RED-first: both ack tests failed on the unfixed adapter at the
'draft still armed' assertion): approval ack leaves the open draft
armed and egresses as a plain send op; expiry notice same; regression
control pins the B2 contract — a real identity-less turn-final still
absorbs into its single open stream.
Three live findings from rc.4 staging, all on the relay-fronted Slack
path, all with the failure observed in live logs before the fix:
1. Approval-send timeout is AMBIGUOUS, not failed (no re-ask).
send_exec_approval through the connector can time out with the card
already rendered — the connector may ack after the deadline (slow
platform API call, transient backpressure, event-loop stall) — and
the timeout-as-failure path re-sent and produced duplicate cards.
The outcome is now tri-state: sent / failed / ambiguous. Ambiguous =
no re-send, no text fallback; the prompt registration stays armed so
a late tap still resolves. Only a definite send error falls back to
text.
2. pending_approval tool results forbid re-issuing the command.
With one card correctly armed, the agent could still mint a SECOND
card by re-running a rephrased variant of the gated command after
reading the pending_approval tool result (observed live: same
command re-issued in a different form, two cards). The tool message
now instructs: do not re-run/rephrase; wait or report pending.
Applied to both the terminal and execute_code arms.
3. Draft interim AND seal frames carry format_hints.
format_hints are stamped on send, edit, and send_for_platform, but
both draft-frame builders (send_draft interim + _seal_open_draft
seal) shipped bare metadata. A streamed final therefore arrived at
the connector hintless and sealed as a plain code block while
non-streamed sends rendered native markdown blocks (observed live:
language-tagged block on send/edit, downgrade on streamed seal).
Both sites now stamp _with_format_hints_for_chat
(destination-resolved, same pattern as the existing lanes).
Verified live after the fix against the platform's stored message
payload: rich_text_preformatted with language field on a streamed
seal.
Tests: tri-state outcome unit tests (5), draft/seal hint stamping + knobs-
off regression control (2, RED-first), existing format-hints suite intact
(14/14). Mutation-verified: reverting the adapter hunk sends
test_draft_interim_and_seal_frames_carry_hints red; restore -> green.
Boundary sweep (text egress lanes crossing the frame contract): send ✓
(pre-existing) edit ✓ (pre-existing) send_for_platform ✓ (pre-existing)
draft-interim ✓ (this PR) draft-seal ✓ (this PR); task_card lane carries
no text content — exempt.
sol-reviewer MINOR: the _enabled_explicit marker is set by config.yaml,
gateway.json, and dashboard PUTs alike, so the WARNING no longer claims
config.yaml specifically. Also names the opt-out env var in the message
so an operator seeing the WARNING knows the escape hatch.
sol-reviewer IMPORTANT: the relay trigger and opt-out read os.getenv()
directly, bypassing the profile secret scope that _apply_env_overrides
uses everywhere else. Under a multiplexed gateway a profile-scoped
GATEWAY_RELAY_URL was invisible, and a process-global one leaked into
every profile, disabling direct platforms in profiles that are not
relay-fronted. Both reads now go through the scope-aware getenv.
Also folds the NIT: opt-out truthiness now uses the shared
is_truthy_value helper instead of a local truthy tuple.
Deployments that intentionally mix connector-fronted and direct ingress
can set GATEWAY_RELAY_ALLOW_DIRECT_PLATFORMS=true to keep directly-
connected messaging adapters enabled beside the relay. Unset, the
GATEWAY_RELAY_URL env stamp keeps its exclusive behavior. Like the
trigger, the opt-out is a deploy-stamp env var, not config.yaml.
A GATEWAY_RELAY_URL set in the process environment marks a
connector-fronted deployment where the connector owns every platform
connection. A directly-connected messaging adapter in the same process
is a second, unmanaged ingress path: it causes duplicate deliveries and
split sessions, and its live socket disarms scale-to-zero.
At the end of _apply_env_overrides, after all enablement passes, the
env stamp now disables every other enabled messaging platform:
- Explicitly-enabled platforms (config.yaml enabled: true) are disabled
with a WARNING that names the platform.
- Credential-auto-enabled platforms are disabled with an INFO line.
- Non-messaging surfaces (local, api_server, webhook) are untouched --
the same exclusion set as the scale-to-zero arm gate.
- gateway.relay_url in config.yaml alone (no env stamp) keeps the old
additive behavior: relay runs beside direct adapters.
With the gateway owning the suspend, the idle predicate covers every
work source (agent turns, cron jobs, API-server runs, background work,
fail-awake on unreadable sources) and the relay drains + flips before
the freeze, so a long timeout no longer buys safety — real work always
blocks the suspend and resume is sub-second. 5 idle minutes just bills
idle RAM. Per-instance override stays config.yaml
gateway.scale_to_zero.idle_timeout_minutes (D2).
New behavior-contract test: invalid config values degrade to the module
default (whatever it is), never zero/negative — asserts the RELATION,
not the literal, per the no-change-detector-tests rule.
Address sol-reviewer findings:
- The shared shutdown-drain counters swallow exceptions to 0 — fine for
a drain, unsafe for a suspend predicate (a transient read failure made
live work look idle, reopening the mid-job freeze). The suspend path
now reads both sources itself and treats an unreadable source as work
(sentinel 1, fail-awake) with a debug log. A MISSING api_server
adapter remains a normal not-work state.
- is_idle()'s parameter renamed running_agent_count -> active_work_count:
it receives the broad aggregate, and the old name invited future
callers to pass only agents again.
- New failure-path tests: unreadable cron source and unreadable API
source each hold the machine awake (both fail against the fail-open
shape); missing adapter stays idle-capable.
Addresses teknium1 sweeper review on PR #67696:
1. Gateway bridge: Skip str(None) bridging when YAML value is Python None
(from or bare ). Previously str(None) → None → unlimited
instead of default 90. Now clears stale env var so resolver applies default.
2. TUI: Route _cfg_max_turns through resolve_turn_limit instead of bare
int(). Old code crashed on none/unlimited and swallowed 0 via
. HERMES_TUI_MAX_TURNS env var also routed through
resolver.
3. Docs: Document unlimited spellings (none/unlimited/infinite/0/-1) in
configuration.md.
4. Tests: Add TestGatewayBridgeNullHandling (4 tests) and TestTUIResolver
(8 tests) covering null handling, string spellings, env var override,
and legacy root-level config.
All 50 tests pass.
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.
This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:
- int/float → int(raw) (floats truncated)
- numeric string ('120') → int(raw)
- 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
whitespace-tolerant) → sys.maxsize sentinel
- YAML None/null → default (90)
- bool/list/dict/garbage → default (with debug log)
All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.
The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.
Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:
- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
spawn-ledger.json self-registration keyed on (pid, create_time) — PID
reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
themselves at startup and attach to the job; Desktop legacy
HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
_ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
backends (purpose reapable + recorded spawner provably dead) in ANY update
context. Ledger-unknown holders fall through to the existing rungs.
22 new tests, sabotage-verified.
One conflict, gateway/relay/adapter.py send_for_platform: main added the
turn-final draft-seal interception (_sfp_metadata with the _interim_send
marker stripped, seal-or-fall-through); this branch added format-hint
stamping on the same frame. COMPOSED: the plain-send frame now stamps
_with_format_hints_for_platform over _sfp_metadata (the stripped copy),
so both the seal fall-through contract and the cron-lane block hints
hold. Note: the seal frame itself (op:draft final) does not stamp hints
— cron sends are never open drafts, so the flagship path is unaffected;
noted as a connector-PR follow-up for streamed interactive finals.
_scale_to_zero_is_idle() consumed _running_agent_count(), but cron jobs
run through a standalone AIAgent on the scheduler's own thread pool and
API-server runs live on the adapter — both outside _running_agents (the
same blind spot the #60432 shutdown-drain fix addressed with
_active_work_count()). The idle predicate therefore read True DURING a
running cron job; a suspend at that moment freezes the job mid-flight.
Observed live on staging 2026-08-20: is_idle held True throughout the
10:45:04-22 cron run — only watcher-tick timing (next tick 9s after
completion) avoided a mid-job freeze.
Use _active_work_count() (agents + cron + API runs). New tests cover a
running cron job and an active API run each blocking idle, plus the
all-quiet True case; both blocking tests fail without the fix.
#84327 excluded _spawn_supervised's permanent watchers (session-expiry,
kanban, reconnect, the scale-to-zero watcher itself, ...) from
_scale_to_zero_has_live_background_work() via a _hermes_supervised_watcher
tag, because counting them made an armed gateway consider itself busy
forever and never go dormant.
Two more permanent, infinite-loop tasks are added to _background_tasks
OUTSIDE _spawn_supervised and were untagged:
- _loop_heartbeat_task (loop_heartbeat_forever, #66892): a `while True`
loop started unconditionally in start() on every gateway boot. Extracted
the inline spawn block into _start_loop_heartbeat_task() so it's
independently testable, matching the existing _start_heartbeat_poller()
pattern.
- _heartbeat_poll_task (_poll_loop in _start_heartbeat_poller): also a
`while True` loop, started the first time a session registers a
heartbeat watch, and then permanent for the rest of the process.
Because _loop_heartbeat_task starts on every boot, it alone made
_scale_to_zero_has_live_background_work() return True forever on every
armed instance, regardless of the #84327 fix -- confirmed empirically
against the real method with the exact untagged-task shape this task has.
Two new regression tests spawn each task through its real production
entry point and assert the busy check returns False; both fail against
the unfixed code (missing method / real assertion failure).
Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
Co-authored-by: Ben Barclay <ben@nousresearch.com>
RelayAdapter.supports_inchannel_continuable is a scalar adopted from the
PRIMARY identity's handshake descriptor, but one RelayAdapter fronts N
platforms and the connector advertises the bit per platform. Reading the
scalar for every logical platform both leaked a Slack-primary True onto
other fronted platforms (activating the flat surface their descriptor
never advertised) and suppressed a non-primary platform's advertised
True (forcing thread mode on capable Slack behind a Discord primary).
Add supports_inchannel_continuable_for_platform(platform): resolves the
platform's own negotiated descriptor via descriptor_for_platform (the
same Phase 1.5 seam max_message_length uses), scalar fallback only when
the per-platform descriptor is unavailable. The scheduler's D6 gate
prefers the query when the adapter provides it; native adapters keep
the class-attribute path byte-identically.
Tests: two-platform descriptor matrix (primary-True no-leak,
non-primary-True honored, unknown-platform scalar fallback).
Two gaps in the block-formatting hint stamping:
1. Wrong descriptor: _format_hints gated on self.descriptor — the PRIMARY
identity's scalar — while one RelayAdapter fronts N platforms. A
Slack-primary adapter stamped Slack hints onto known Discord chats; a
Discord-primary adapter suppressed hints for Slack chats whose own
negotiated descriptor advertised the bit. Resolve per destination:
send/edit use _descriptor_for_chat (the same seam max_message_length
already uses) plus the chat's logical platform for the config
sub-block; the knob lookup is now per-logical-platform
(platforms.relay.extra.<platform>.*) instead of hardwired to slack.
2. Missing lane: send_for_platform — the scheduled/persisted-home lane
(gateway/delivery.py), i.e. the CRON delivery path, the flagship
consumer of the in_channel brief — never stamped hints at all. Stamp
there too, resolving descriptor_for_platform(logical) off the
transport; the scalar descriptor is used only when it belongs to that
exact platform (fail closed).
Tests: Slack-primary/Discord-chat no-leak, Discord-primary/Slack-chat
still-stamps, send_for_platform stamps for capable platform and stays
clean for incapable — all against a two-platform negotiated-descriptor
transport. Existing single-platform suite unchanged and green.