Commit Graph

3331 Commits

Author SHA1 Message Date
fangliquan 91fb175188 fix(gateway): preserve systemd handoff recovery 2026-08-22 15:27:26 +05:30
fangliquanflq 5b024c7ccc fix(gateway): make systemd the sole restart owner 2026-08-22 15:27:26 +05:30
fangliquan f12cd04015 fix(gateway): isolate kanban dispatcher to_thread context
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
2026-08-22 15:25:50 +05:30
fangliquan bf3a0bb99d fix(gateway): isolate supervised watcher contexts 2026-08-22 15:25:50 +05:30
Jack Lau e173720774 fix(gateway): give supervision exhaustion an owner for queued platforms
Review of #90448 by @andrexibiza: adding _ensure_reconnect_watcher_running()
to the already-queued branch of a fatal callback is still an event-coupled
check. It needs a later fatal error from some other platform to arrive, and
#81036 makes that less likely rather than more -- it publishes the queue
before disconnect and drops the failed adapter from the live map, so after
the watcher's supervised restart budget is spent there may be no adapter
left to emit the event recovery is waiting on.

That is the state #72366 (salvage of #71867 by @ygd58) restored supervision
to close: queued work exists, the watcher is dead, and nobody owns the
invariant. Supervision being finite is correct; having no owner past the
budget is not.

_spawn_supervised now takes on_give_up, invoked when it abandons a task --
the supervisor is the only thing that knows it has. The reconnect watcher
uses it to hold:

  while _running and _failed_platforms is non-empty, either a reconnect
  watcher is live or a bounded respawn is scheduled.

Empty queue: leave it down and log; the enqueue path spawns a fresh watcher
the moment something depends on one. Non-empty: a bounded slow tier at
_RECONNECT_WATCHER_SLOW_RETRY_SECS (300s) for _MAX_SLOW_WATCHER_RESPAWNS (6)
attempts, standing down early if the queue drains or a watcher returns on
its own. Exhausted: one loud error naming the platforms left unattended.

The ceiling is (1 + _MAX_SUPERVISED_RESTARTS) x (1 + _MAX_SLOW_WATCHER_RESPAWNS)
spawns -- 42 across at least half an hour -- because each slow attempt hands
the watcher a fresh supervised budget. A test asserts that ceiling so it
cannot quietly become a restart loop.

Deliberately NOT included: requesting a process restart when the slow tier
is also exhausted. Taking down every healthy platform to heal a sick one is
a blast-radius policy decision for a maintainer.

Two things this turned up:

- _spawn_supervised did not thread on_give_up through its own backoff
  respawn, so the callback was lost after the first restart and the give-up
  branch had no owner at exactly the moment it needed one -- the same defect
  the on_spawn docstring warns about, one parameter over.
- Three call sites repeated the (factory, name, on_spawn) triple, whose
  on_spawn half is load-bearing. They now go through
  _spawn_reconnect_watcher().

_supervised_backoff() names the previously-inline exponential schedule so
the exhaustion tests can collapse it; production behaviour is unchanged.

Refs #90386
2026-08-22 15:25:45 +05:30
Jack Lau 92018e76a8 fix(gateway): heal a dead reconnect watcher when the platform is already queued
_ensure_reconnect_watcher_running() exists for one situation: the reconnect
watcher has exhausted _MAX_SUPERVISED_RESTARTS, so _spawn_supervised has logged
"giving up restarts" and will never bring it back on its own (#70344, and the
supervised-restart half of #71758). It had exactly one call site, inside the
newly-queued branch of _queue_retryable_fatal_platform.

That branch is unreachable for a platform already in _failed_platforms, which
is the only kind of platform the watcher can have been retrying long enough to
burn five rapid restarts on. So the backstop could not fire in the one state it
was written for.

The failure is silent by construction. The early return logs nothing, so there
is no "queued for background reconnection" line. The stranded check in
_handle_adapter_fatal_error_detached deliberately treats a queued platform as
safe, so the gateway does not exit for the service manager either. With another
platform still connected, self.adapters is non-empty and the "gateway staying
alive, watcher will retry in background" branch is skipped too. A retryable
fatal error can therefore produce a single ERROR line and then nothing: the
platform sits in the queue that nobody is draining until someone restarts the
process by hand (#90386 reports 4h17m of that, with cron unaffected throughout).

Call the ensure on the already-queued path as well. It is already idempotent
and already cheap: it returns immediately unless the tracked task is done, and
it routes through the same on_spawn handle tracking, so a live watcher is never
duplicated.

The queue entry itself is deliberately left untouched. Re-enqueueing would
reset attempts and next_retry, restarting the backoff ladder on every fatal
error and hammering a provider that is already refusing the connection.
2026-08-22 15:25:45 +05:30
Kshitij Kapoor 41e29a601e fix(gateway): clear resume_pending for all claimed ledger rows before any redelivery send
Follow-up to the salvaged #91986: the per-row clear still left rows the
loop had not reached exposed — a slow send ahead of them could hold the
loop past the inbound-gate timeout and let
_schedule_resume_pending_sessions replay those turns. Clearing every
claimed row up front closes the duplicate window; claiming already
spent the redelivery attempt, so the ledger retry path is unchanged.
2026-08-22 15:25:30 +05:30
HexLab98 ce944a5a55 fix(gateway): do not let boot-path sends hold the inbound gate
Restart notification and obligation redelivery ran before the
startup-restore gate opened, so one hung Telegram send queued inbound
on every platform. Bound those sends with the same timeout the resume
gate already uses, and clear resume_pending before send so a timed-out
redelivery cannot also replay the turn.
2026-08-22 15:25:30 +05:30
kshitijk4poor c45e2b19c3 fix(state): guard gateway FTS rebuild + comment early flag-set
Add the foreign-holder guard to gateway/session.py::_rebuild_fts_once(),
the third FTS rebuild path that was not covered by the original fix.
Also add a comment explaining why _fts_runtime_rebuild_attempted is set
before the foreign-holder check: the fail-open path that follows
persists FTS_STALE_KEY so the next startup retries via _recover_stale_fts.
2026-08-22 03:56:13 +05:30
kshitijk4poor 0b8a848754 perf(api): classify compaction rows once per message in run.completed transcript
_turn_transcript_messages pre-classified every message with
_is_compressed_summary_message (full content flatten + prefix scan), then
_message_response re-ran the same classifier inside its projection --
2x per non-summary row, 3x per summary row on every run.completed emit.
The outer guard was redundant: _message_response already yields
display_kind hidden for pure handoffs. One projection call per row now.
Surfaced by the post-merge simplify re-review of #91517/#91535.
2026-08-21 22:54:13 +05:30
kshitijk4poor 2cb8794f7a fix(browser): sweep orphan artifact files at store construction
Artifact receipts live only in memory, so files left behind by a dead
process were unreachable but persisted forever despite the advertised
300s TTL — a retention failure on the surface meant to be ephemeral
(blocker 4 of andrexibiza's #91535 review). A fresh ArtifactStore now
removes every artifact-id-shaped file and stale *.tmp with no index entry
(at construction the index is empty, so all such files are orphans).
Non-artifact-shaped names are untouched. Regression: store -> recreate
store over same root -> orphan+tmp gone, unrelated file kept.
2026-08-21 22:54:04 +05:30
kshitijk4poor c16c262d01 fix(browser): honor live Developer Mode for privileged capability selection
The global broker snapshotted browser.extension_control.developer_mode once
at construction, so flipping it OFF in config did not revoke raw CDP/eval
from already-attached controllers until process restart — a revocation
failure at the highest-privilege browser surface (blocker 3 of
andrexibiza's #91535 review). select() now consults the live config on
every privileged selection (explicit bool still pins for tests); off->on
also unlocks without restart. Regression test drives both directions
against an attached controller. Also drops the dead back-compat
_artifact_store property (zero readers).
2026-08-21 22:54:04 +05:30
kshitijk4poor 23a64a97ec fix(api): correct _handle_browser_control_frame return annotation
The frame handler returns reply dicts (heartbeat/detach acks) that the WS
reader loop sends back; the -> None annotation was the only new ty
diagnostic vs origin/main.
2026-08-21 22:33:45 +05:30
kshitijk4poor 847289864d fix(browser): make the artifact boundary compose end-to-end and scope stores per profile
Addresses both merge blockers from @andrexibiza's review of #85351:

1. HTTP-uploaded artifacts could never be consumed by broker dispatch:
   artifact_scope_key hashed (principal, session, family), the HTTP routes
   store with an EMPTY session (API-key auth has no server session) while
   broker validation carries a session-bearing ControllerScope — every
   real upload->dispatch journey died with ArtifactScopeMismatch
   (reproduced before fixing). Canonical ownership is now
   principal/transport-family (documented in the scope-key docstring);
   ids stay unguessable server-minted 32-hex and downloads one-shot.
   New composition regression: HTTP-shape upload -> registered controller
   scope -> broker artifact dispatch, mutation-checked (re-adding session
   to the key makes it fail).

2. The 'profile-scoped' artifact store was first-profile-wins process
   state: one adapter-level singleton pinned profile B to profile A's
   physical root on multiplex listeners (same frozen-handle class as
   #88734). Stores are now cached by resolved profile, and the broker
   selects the store from the controller scope's profile_id (default-slot
   fallback preserves single-profile/test behaviour). New A/B multiplex
   regression proves distinct physical roots regardless of touch order.

Also documents the advertised ticket_expires_at as best-effort wall clock
(broker enforces expiry monotonically) per review feedback.
2026-08-21 22:33:45 +05:30
kshitijk4poor 652e0a72d1 perf(browser): read the feature flags via load_config_readonly
browser_control_enabled()/browser_control_developer_mode() run on every
browser tool call and inside every check_fn evaluation (uncached for bound
sessions). Both are pure reads of nested dicts; load_config()'s defensive
deepcopy (~135us/call) is wasted there. Same pattern as the other read-only
config probes.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
kshitijk4poor 13f209d4fd refactor(browser): dedupe auth-flow names and sentinel identity
- Rename the broker's TicketInvalid to ControllerTicketInvalid: the same
  exception name already exists in hermes_cli/dashboard_auth/ws_tickets.py
  and BOTH are caught in the same WS auth flow this feature touches — two
  unrelated same-named exception types in one blast radius invited a wrong
  except clause.
- Import the 'server-internal' sentinel identity from its canonical
  definition (ws_tickets.INTERNAL_USER_ID/INTERNAL_PROVIDER) instead of
  re-declaring the strings; drift would have silently broken the
  internal-peer exclusion in _is_authenticated_identity.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
kshitijk4poor 45078eb941 fix(browser): offload broker lock acquisition off the event loop
attach/disconnect/detach acquire a per-controller threading.Lock that a
worker-thread dispatch can hold for up to 10s while blocking on the event
loop to transmit its command frame (run_coroutine_threadsafe +
result(timeout=10)). Acquiring that lock synchronously from loop context
(controller WS finally, frame handler, gateway WS teardown) could park the
ENTIRE gateway event loop behind the send bridge — a deterministic
multi-second global stall whenever controller teardown raced an in-flight
command. All loop-context broker calls now go through asyncio.to_thread,
matching the existing offload pattern for _close_sessions_for_transport.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
kshitijk4poor a5882058de fix(browser): bind the extension lane at controller registration, not transport auth
The router treated any server-stamped principal as a bound lane, so with the
flag ON every authenticated dashboard/API session lost the legacy browser
backend even when no extension controller ever registered (scope_for_session
returns None -> ControllerUnavailable, no fallback) — while check_fns still
advertised the tools via the legacy OR-gate.

New broker.lane_registered() distinguishes the two cases:
- lane never registered -> generic callers keep the legacy backend
- lane registered (controller offline/ambiguous) -> fail closed, unchanged —
  a control-this-tab session never silently jumps to another browser

Also makes the four non-allowlisted wrapped tools (cdp/console/vision/
get_images) behave correctly for never-registered lanes (legacy backend)
while staying fail-closed for registered lanes.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
abundantbeing 1977c3d2eb feat(browser): add scoped artifact endpoints, broker permission gates, and companion journal 2026-08-21 22:33:45 +05:30
abundantbeing 095a1d078c fix(browser): preserve controller work across reconnects
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.

Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
2026-08-21 22:33:45 +05:30
abundantbeing d524cc9a16 fix(browser): harden extension controller routing
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.

Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
2026-08-21 22:33:45 +05:30
abundantbeing c9fd5223f6 feat(browser): enable extension controller actions 2026-08-21 22:33:45 +05:30
abundantbeing 5df1d0e113 feat(browser): add authenticated control broker 2026-08-21 22:33:45 +05:30
kshitijk4poor b2c4f1f376 refactor(api): reuse _COMPACTION_INTERNAL_FIELDS from compaction_display
The 7-key internal-fields tuple was inlined twice (agent/compaction_display.py
and _project_client_message); a drift between the copies would silently leak
one internal field class through the API projection. Surfaced during review
of PR #85442.
2026-08-21 21:53:11 +05:30
abundantbeing a2a23a8f7e fix(clients): hide compaction carriers across surfaces 2026-08-21 21:53:11 +05:30
abundantbeing 97e32d49ac fix(api): hide compaction scaffolding from clients
Project client-visible session messages through the canonical compaction classifier. Hide standalone handoffs, unwrap merged carriers to their authentic prior-tail content, strip inherited internal fields, and keep model-facing recovery history unchanged.
2026-08-21 21:53:11 +05:30
Teknium 216c98aaed fix(gateway): scope draft-final finalize skip to fresh persistent sends
The salvaged skip condition keyed on _use_draft_streaming alone, which
also suppressed the explicit REQUIRES_EDIT_FINALIZE pass when a
draft-streaming run had degraded to edit-based delivery (draft failure
fallback sets _message_id). Key the skip on the got_done update being a
fresh persistent send through the native-draft transport (_message_id is
None), which is the only case where the update already carried its own
finalize.
2026-08-20 21:58:18 -07:00
Gille 790c850144 fix(telegram): preserve rich finals after DM drafts 2026-08-20 21:58:18 -07:00
Teknium 1d74833d8d feat(update): structured update receipts + post-update fleet version verification
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.

- hermes_cli/build_info.py: get_code_identity() — process-cached code
  identity (git sha for source installs, baked .hermes_build_sha for
  Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
  code_sha/code_version into gateway_state.json, so a running gateway's
  actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
  update run (steps, skips with reasons, gateway restart outcome, fleet
  snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
  pointer for the dashboard/desktop; plus collect_fleet_versions() /
  print_fleet_version_matrix() comparing every live profile gateway
  against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
  git, ZIP, and hard-failure paths; after the restart phase, prints the
  fleet version matrix and escalates provably-stale gateways into the
  existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
  gateways report 'unknown' and never fail the update (no false
  positives during rollout).

Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
2026-08-20 21:54:25 -07:00
Ben Barclay efb6b40f94 Merge pull request #91237 from NousResearch/fix/relay-env-exclusive-messaging
fix(gateway): GATEWAY_RELAY_URL env stamp disables direct messaging platforms
2026-08-21 14:33:29 +10:00
Ben Barclay 952d1cef18 fix(gateway): prompt-send failures keep their diagnostic detail; clarify caller path gets contract tests
Two review nits on the clarify ambiguity fix:

- _approval_send_outcome swallowed the failure detail the old inline
  callers logged (scheduling exception text / SendResult error). Both
  the approval and clarify lanes now share one warning with the detail,
  logged in the classifier itself.

- The clarify tests pinned the disposition helper but nothing proved
  the ambiguous branch actually reaches the bounded wait. The
  send-then-wait sequence is extracted to _clarify_send_then_wait (the
  callback closure now just binds context onto it) and the suite gains
  caller-path tests: ambiguous/sent -> wait_for_response with the
  generated clarify_id and configured timeout; definitive failure ->
  sentinel without waiting; no-response timeout sentinel preserved;
  plus caplog assertions that failed sends log their detail.

Relay + gateway sweep 289/289.
2026-08-20 19:56:01 -07:00
Ben Barclay a7cfe391f7 fix(gateway): clarify prompt sends get the same ambiguous-timeout handling as approvals
The clarify caller treated a send-scheduling timeout as a definitive
failure: clear_session() + '[clarify prompt could not be delivered]'.
Same physics as the approval card fixed earlier in this PR — the card
may well have posted with a late connector ack — so the teardown ran
out from under a rendered clarify card and the user's answer resolved
nothing.

New _clarify_send_disposition() routes the outcome through
_approval_send_outcome: only a DEFINITIVE failure (error result,
non-timeout exception, no future) clears the registration and aborts;
ambiguous logs a warning and falls through to wait_for_response, whose
existing bounded wait already handles the truly-lost-card case. This
makes the boundary rule stated in the ambiguity test docstring hold
for the clarify lane, not just approvals.

Tests: 5 disposition tests mirroring the approval suite, including
clear_session-not-called on timeout. Mutation-verified: folding
ambiguous into the failed branch sends
test_timeout_keeps_registration_armed_and_proceeds_to_wait red.
Relay + gateway sweep 283/283.
2026-08-20 19:56:01 -07:00
Victor Kyriazakos 424d07edac fix(relay): prompt-lifecycle acks are fire-and-forget — awaiting them ON the read loop self-deadlocked the transport
Round 2 of the approval-turn stuck-stream hunt. Round 1 (interim-marked
acks) fixed the draft-hijack-by-matching path — live logs confirm the
absorption fallback no longer fires — but the freeze persisted because
of a second, deeper defect on the same codepath:

_consume_prompt_response executes ON the transport read loop (inbound
frame -> _handle_frame -> _inbound handler). The handler awaited
self.send() for its '✅ Approved once' ack — but send() blocks on an
outbound_result future that ONLY the read loop can resolve, and the
read loop is blocked inside this very handler. Guaranteed self-deadlock
for the full outbound timeout (30s) on EVERY button tap. While wedged,
everything on the transport starved: draft appends (the frozen stream
right after approving), sibling approval-card sends (timed out into
'possibly-delivered' — the observed double-approval ambiguity), and the
turn's seal (timed out ambiguous -> plain-send fallback -> duplicate
final). Log signature was the tell: card-send timeout at tap time, no
absorption INFO, no seal-failed WARNING, no suppression line.

Fix: _send_lifecycle_ack() — acks ride a background task with strong
ref retention; the handler returns immediately and the read loop keeps
consuming, so the ack's own result frame resolves normally. Applied to
all six lifecycle sends (approval ack, slash-confirm ack + result text,
clarify acks, expiry notice). Acks are cosmetic by contract; failure
logs at debug and never breaks the reader.

Tests: new deadlock-shape test (gated transport send; handler must
return within 1s and the ack must still egress afterwards — RED on the
awaited version via TimeoutError at the exact deadlock), prior 3 tests
green with a yield for the background task. Targeted sweep 203/203.
2026-08-20 19:56:01 -07:00
Victor Kyriazakos 57fe01e367 fix(relay): prompt-lifecycle acks are interim sends — the approval ack was sealing the turn's own draft stream
Live finding (rc.4 staging, 100% reproducible on approval turns): after
resolving an exec-approval prompt_response, the adapter sends a short
ack ('Approved once'). _prompt_reply_metadata carried only placement
metadata (thread_id) — no per-turn identity, no interim marker — so
send()'s single-open-stream fallback (review B2) matched the approval
turn's OWN live draft and sealed it with the ack text. From there,
silently: every later append died on the post-seal tombstone (built for
millisecond stragglers, deliberately quiet), freezing the visible draft
mid-word; the turn-final found no open draft and fell through to a
plain send — the duplicate 'fallback' message. No suppression line, no
seal-failed warning: the log signature was pure absence.

Fix: _prompt_reply_metadata stamps _interim_send=True, which send()
already honors by bypassing draft matching. One source covers the whole
lifecycle class (approval ack, slash-confirm ack, prompt-expired
notice — all six call sites route through it).

Observability (the quiet parts, out loud):
- single-open-stream absorption now logs at INFO with the absorbed key;
- the FIRST post-seal tombstone swallow per draft key logs at WARNING
  (bounded FIFO dedup) — one swallow is the normal straggler race, a
  burst means a live stream was sealed mid-flight by someone else.

Tests (RED-first: both ack tests failed on the unfixed adapter at the
'draft still armed' assertion): approval ack leaves the open draft
armed and egresses as a plain send op; expiry notice same; regression
control pins the B2 contract — a real identity-less turn-final still
absorbs into its single open stream.
2026-08-20 19:56:01 -07:00
Victor Kyriazakos 5210dd48b8 fix(gateway+relay): approval prompts survive ambiguity without duplicates; streamed finals keep block formatting
Three live findings from rc.4 staging, all on the relay-fronted Slack
path, all with the failure observed in live logs before the fix:

1. Approval-send timeout is AMBIGUOUS, not failed (no re-ask).
   send_exec_approval through the connector can time out with the card
   already rendered — the connector may ack after the deadline (slow
   platform API call, transient backpressure, event-loop stall) — and
   the timeout-as-failure path re-sent and produced duplicate cards.
   The outcome is now tri-state: sent / failed / ambiguous. Ambiguous =
   no re-send, no text fallback; the prompt registration stays armed so
   a late tap still resolves. Only a definite send error falls back to
   text.

2. pending_approval tool results forbid re-issuing the command.
   With one card correctly armed, the agent could still mint a SECOND
   card by re-running a rephrased variant of the gated command after
   reading the pending_approval tool result (observed live: same
   command re-issued in a different form, two cards). The tool message
   now instructs: do not re-run/rephrase; wait or report pending.
   Applied to both the terminal and execute_code arms.

3. Draft interim AND seal frames carry format_hints.
   format_hints are stamped on send, edit, and send_for_platform, but
   both draft-frame builders (send_draft interim + _seal_open_draft
   seal) shipped bare metadata. A streamed final therefore arrived at
   the connector hintless and sealed as a plain code block while
   non-streamed sends rendered native markdown blocks (observed live:
   language-tagged block on send/edit, downgrade on streamed seal).
   Both sites now stamp _with_format_hints_for_chat
   (destination-resolved, same pattern as the existing lanes).
   Verified live after the fix against the platform's stored message
   payload: rich_text_preformatted with language field on a streamed
   seal.

Tests: tri-state outcome unit tests (5), draft/seal hint stamping + knobs-
off regression control (2, RED-first), existing format-hints suite intact
(14/14). Mutation-verified: reverting the adapter hunk sends
test_draft_interim_and_seal_frames_carry_hints red; restore -> green.

Boundary sweep (text egress lanes crossing the frame contract): send ✓
(pre-existing) edit ✓ (pre-existing) send_for_platform ✓ (pre-existing)
draft-interim ✓ (this PR) draft-seal ✓ (this PR); task_card lane carries
no text content — exempt.
2026-08-20 19:56:01 -07:00
Ben Barclay 2ef095ce46 fix(gateway): source-neutral log wording for relay-exclusive sweep
sol-reviewer MINOR: the _enabled_explicit marker is set by config.yaml,
gateway.json, and dashboard PUTs alike, so the WARNING no longer claims
config.yaml specifically. Also names the opt-out env var in the message
so an operator seeing the WARNING knows the escape hatch.
2026-08-21 12:42:13 +10:00
Ben Barclay 477a0222f9 fix(gateway): read relay-exclusive env vars through profile secret scope
sol-reviewer IMPORTANT: the relay trigger and opt-out read os.getenv()
directly, bypassing the profile secret scope that _apply_env_overrides
uses everywhere else. Under a multiplexed gateway a profile-scoped
GATEWAY_RELAY_URL was invisible, and a process-global one leaked into
every profile, disabling direct platforms in profiles that are not
relay-fronted. Both reads now go through the scope-aware getenv.

Also folds the NIT: opt-out truthiness now uses the shared
is_truthy_value helper instead of a local truthy tuple.
2026-08-21 12:42:06 +10:00
Ben Barclay 2d3f5c1554 feat(gateway): GATEWAY_RELAY_ALLOW_DIRECT_PLATFORMS opt-out for relay-exclusive mode
Deployments that intentionally mix connector-fronted and direct ingress
can set GATEWAY_RELAY_ALLOW_DIRECT_PLATFORMS=true to keep directly-
connected messaging adapters enabled beside the relay. Unset, the
GATEWAY_RELAY_URL env stamp keeps its exclusive behavior. Like the
trigger, the opt-out is a deploy-stamp env var, not config.yaml.
2026-08-21 12:26:41 +10:00
Ben Barclay 67292ec5b8 fix(gateway): GATEWAY_RELAY_URL env stamp disables direct messaging platforms
A GATEWAY_RELAY_URL set in the process environment marks a
connector-fronted deployment where the connector owns every platform
connection. A directly-connected messaging adapter in the same process
is a second, unmanaged ingress path: it causes duplicate deliveries and
split sessions, and its live socket disarms scale-to-zero.

At the end of _apply_env_overrides, after all enablement passes, the
env stamp now disables every other enabled messaging platform:

- Explicitly-enabled platforms (config.yaml enabled: true) are disabled
  with a WARNING that names the platform.
- Credential-auto-enabled platforms are disabled with an INFO line.
- Non-messaging surfaces (local, api_server, webhook) are untouched --
  the same exclusion set as the scale-to-zero arm gate.
- gateway.relay_url in config.yaml alone (no env stamp) keeps the old
  additive behavior: relay runs beside direct adapters.
2026-08-21 12:16:06 +10:00
Ben Barclay 48f15c1ece chore(gateway): drop the default scale-to-zero idle timeout to 2 minutes
With the gateway owning the suspend, the idle predicate covers every
work source (agent turns, cron jobs, API-server runs, background work,
fail-awake on unreadable sources) and the relay drains + flips before
the freeze, so a long timeout no longer buys safety — real work always
blocks the suspend and resume is sub-second. 5 idle minutes just bills
idle RAM. Per-instance override stays config.yaml
gateway.scale_to_zero.idle_timeout_minutes (D2).

New behavior-contract test: invalid config values degrade to the module
default (whatever it is), never zero/negative — asserts the RELATION,
not the literal, per the no-change-detector-tests rule.
2026-08-21 09:08:02 +10:00
Ben Barclay ee000768ce Merge pull request #90761 from NousResearch/fix/scale-to-zero-cron-aware-idle
fix(gateway): count cron and API-server work in the scale-to-zero idle predicate
2026-08-21 08:19:36 +10:00
Ben Barclay 0215930526 review: fail-awake work accounting + rename is_idle param to active_work_count
Address sol-reviewer findings:
- The shared shutdown-drain counters swallow exceptions to 0 — fine for
  a drain, unsafe for a suspend predicate (a transient read failure made
  live work look idle, reopening the mid-job freeze). The suspend path
  now reads both sources itself and treats an unreadable source as work
  (sentinel 1, fail-awake) with a debug log. A MISSING api_server
  adapter remains a normal not-work state.
- is_idle()'s parameter renamed running_agent_count -> active_work_count:
  it receives the broad aggregate, and the old name invited future
  callers to pass only agents again.
- New failure-path tests: unreadable cron source and unreadable API
  source each hold the machine awake (both fail against the fail-open
  shape); missing adapter stays idle-capable.
2026-08-21 07:37:11 +10:00
Chris Fontes 64505a2b83 fix(resolve_turn_limit): gateway bridge null handling, TUI resolver, docs
Addresses teknium1 sweeper review on PR #67696:

1. Gateway bridge: Skip str(None) bridging when YAML value is Python None
   (from  or bare ). Previously str(None) → None → unlimited
   instead of default 90. Now clears stale env var so resolver applies default.

2. TUI: Route _cfg_max_turns through resolve_turn_limit instead of bare
   int(). Old code crashed on none/unlimited and swallowed 0 via
   . HERMES_TUI_MAX_TURNS env var also routed through
   resolver.

3. Docs: Document unlimited spellings (none/unlimited/infinite/0/-1) in
   configuration.md.

4. Tests: Add TestGatewayBridgeNullHandling (4 tests) and TestTUIResolver
   (8 tests) covering null handling, string spellings, env var override,
   and legacy root-level config.

All 50 tests pass.
2026-08-20 04:50:39 -07:00
Chris Fontes 5046282867 feat(config): resolve_turn_limit — first-class 'none'/'unlimited' for agent.max_turns
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.

This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:

  - int/float → int(raw) (floats truncated)
  - numeric string ('120') → int(raw)
  - 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
    whitespace-tolerant) → sys.maxsize sentinel
  - YAML None/null → default (90)
  - bool/list/dict/garbage → default (with debug log)

All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.

The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.

Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
2026-08-20 04:50:39 -07:00
Teknium 95fa814269 feat(process): positive process identity — spawn tags, machine spawn ledger, Windows job-object self-attach
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:

- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
  spawn-ledger.json self-registration keyed on (pid, create_time) — PID
  reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
  and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
  the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
  themselves at startup and attach to the job; Desktop legacy
  HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
  works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
  _ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
  backends (purpose reapable + recorded spawner provably dead) in ANY update
  context. Ledger-unknown holders fall through to the existing rungs.

22 new tests, sabotage-verified.
2026-08-20 04:47:38 -07:00
Ben Barclay a1a1bee5c2 Merge branch 'main' into feat/relay-slack-parity
One conflict, gateway/relay/adapter.py send_for_platform: main added the
turn-final draft-seal interception (_sfp_metadata with the _interim_send
marker stripped, seal-or-fall-through); this branch added format-hint
stamping on the same frame. COMPOSED: the plain-send frame now stamps
_with_format_hints_for_platform over _sfp_metadata (the stripped copy),
so both the seal fall-through contract and the cron-lane block hints
hold. Note: the seal frame itself (op:draft final) does not stamp hints
— cron sends are never open drafts, so the flagship path is unaffected;
noted as a connector-PR follow-up for streamed interactive finals.
2026-08-20 20:56:03 +10:00
Ben Barclay 743dc935f5 fix(gateway): count cron and API-server work in the scale-to-zero idle predicate
_scale_to_zero_is_idle() consumed _running_agent_count(), but cron jobs
run through a standalone AIAgent on the scheduler's own thread pool and
API-server runs live on the adapter — both outside _running_agents (the
same blind spot the #60432 shutdown-drain fix addressed with
_active_work_count()). The idle predicate therefore read True DURING a
running cron job; a suspend at that moment freezes the job mid-flight.
Observed live on staging 2026-08-20: is_idle held True throughout the
10:45:04-22 cron run — only watcher-tick timing (next tick 9s after
completion) avoided a mid-job freeze.

Use _active_work_count() (agents + cron + API runs). New tests cover a
running cron job and an active API run each blocking idle, plus the
all-quiet True case; both blocking tests fail without the fix.
2026-08-20 20:53:16 +10:00
pierrenode a1ddb54840 fix(gateway): tag the loop-liveness and heartbeat-poll tasks as permanent supervised watchers (#84558)
#84327 excluded _spawn_supervised's permanent watchers (session-expiry,
kanban, reconnect, the scale-to-zero watcher itself, ...) from
_scale_to_zero_has_live_background_work() via a _hermes_supervised_watcher
tag, because counting them made an armed gateway consider itself busy
forever and never go dormant.

Two more permanent, infinite-loop tasks are added to _background_tasks
OUTSIDE _spawn_supervised and were untagged:

- _loop_heartbeat_task (loop_heartbeat_forever, #66892): a `while True`
  loop started unconditionally in start() on every gateway boot. Extracted
  the inline spawn block into _start_loop_heartbeat_task() so it's
  independently testable, matching the existing _start_heartbeat_poller()
  pattern.
- _heartbeat_poll_task (_poll_loop in _start_heartbeat_poller): also a
  `while True` loop, started the first time a session registers a
  heartbeat watch, and then permanent for the rest of the process.

Because _loop_heartbeat_task starts on every boot, it alone made
_scale_to_zero_has_live_background_work() return True forever on every
armed instance, regardless of the #84327 fix -- confirmed empirically
against the real method with the exact untagged-task shape this task has.

Two new regression tests spawn each task through its real production
entry point and assert the busy check returns False; both fail against
the unfixed code (missing method / real assertion failure).

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-08-20 20:24:40 +10:00
Ben Barclay 162b23c3e2 fix(relay): D6 in_channel capability gate resolves the destination platform's descriptor
RelayAdapter.supports_inchannel_continuable is a scalar adopted from the
PRIMARY identity's handshake descriptor, but one RelayAdapter fronts N
platforms and the connector advertises the bit per platform. Reading the
scalar for every logical platform both leaked a Slack-primary True onto
other fronted platforms (activating the flat surface their descriptor
never advertised) and suppressed a non-primary platform's advertised
True (forcing thread mode on capable Slack behind a Discord primary).

Add supports_inchannel_continuable_for_platform(platform): resolves the
platform's own negotiated descriptor via descriptor_for_platform (the
same Phase 1.5 seam max_message_length uses), scalar fallback only when
the per-platform descriptor is unavailable. The scheduler's D6 gate
prefers the query when the adapter provides it; native adapters keep
the class-attribute path byte-identically.

Tests: two-platform descriptor matrix (primary-True no-leak,
non-primary-True honored, unknown-platform scalar fallback).
2026-08-20 20:12:52 +10:00
Ben Barclay 79c39025c0 fix(relay): format hints resolve the DESTINATION platform, and stamp on send_for_platform
Two gaps in the block-formatting hint stamping:

1. Wrong descriptor: _format_hints gated on self.descriptor — the PRIMARY
   identity's scalar — while one RelayAdapter fronts N platforms. A
   Slack-primary adapter stamped Slack hints onto known Discord chats; a
   Discord-primary adapter suppressed hints for Slack chats whose own
   negotiated descriptor advertised the bit. Resolve per destination:
   send/edit use _descriptor_for_chat (the same seam max_message_length
   already uses) plus the chat's logical platform for the config
   sub-block; the knob lookup is now per-logical-platform
   (platforms.relay.extra.<platform>.*) instead of hardwired to slack.

2. Missing lane: send_for_platform — the scheduled/persisted-home lane
   (gateway/delivery.py), i.e. the CRON delivery path, the flagship
   consumer of the in_channel brief — never stamped hints at all. Stamp
   there too, resolving descriptor_for_platform(logical) off the
   transport; the scalar descriptor is used only when it belongs to that
   exact platform (fail closed).

Tests: Slack-primary/Discord-chat no-leak, Discord-primary/Slack-chat
still-stamps, send_for_platform stamps for capable platform and stays
clean for incapable — all against a two-platform negotiated-descriptor
transport. Existing single-platform suite unchanged and green.
2026-08-20 20:10:20 +10:00