Commit Graph

28271 Commits

Author SHA1 Message Date
fangliquanflq 80cec2785d fix(gateway): preserve routing state across recovery 2026-08-23 18:25:12 -07:00
fangliquanflq 4b659f0e33 fix(gateway): retry failed session database opens 2026-08-23 18:25:12 -07:00
Andrex Ibiza, MBA 31a01f373b fix(state): make automatic repair non-destructive
Reproduce the schema-btree failure where the in-place writable_schema/VACUUM ladder can reduce a 3,048-page canonical state.db to 113 pages and still return repaired=False.

Move all mutating strategies behind a complete SQLite online-backup snapshot, retain one exclusive SQLite guard from staging through transactional promotion, preserve committed WAL frames and the live inode, fail closed on environmental hazards, and add adversarial regression coverage for failed-repair preservation, post-stage writer races, interrupted copies, stale scratch, disk admission, attempt-ledger semantics, and durability routing.

Fixes #93064
Supersedes the delivery mechanics of #87409 while preserving its implementation provenance.

Co-authored-by: cervantesh <11169707+cervantesh@users.noreply.github.com>
2026-08-23 18:25:12 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Teknium a9e46229b2 fix: sniff fast-path keys on binary magic only, not NUL presence
A newer main-side sniff fast-path skipped any file with a NUL in its
head, short-circuiting before the magic-number check the salvaged fix
added — reintroducing the #77927 bypass. Key the fast-path on
executable magic only; NUL-bearing text falls through to the tail
logic (magic check, size-before-strip, NUL-strip, scan).
2026-08-23 18:24:36 -07:00
Meng Chee 92edb861be fix(cron): close NUL-padded script bypass in lifecycle guard
The #76762 binary check treats any NUL byte in the first chunk as "compiled
binary, nothing to scan":

    if b"\x00" in data:
        return None, False

"Contains a NUL" and "is a compiled binary" are different questions, and the
gap between them is a guard bypass. `bash` executes a *text* script straight
past an embedded NUL, so one pad byte disables the entire scan while the
script still runs:

    #!/bin/bash
    # pad<NUL>
    hermes gateway restart

    scan("bash padded.sh")  -> False   (not blocked)
    bash padded.sh          -> executes the lifecycle command

This shape was blocked before #76762, so the crash fix traded a loud failure
for a silent one.

Keying the check on a leading `#!` is not sufficient: a shebang-less file with
a NUL on any line but the first also executes normally. (A NUL on line 1 of a
shebang-less file is the one shape bash rejects, exit 126 — but that same file
is still executable via `. file`.)

Fix: identify binaries by MAGIC NUMBER — ELF, Mach-O (incl. byte-swapped and
universal/fat), PE/COFF, static archive, gzip, zip — with a shebang always
winning. A NUL-bearing *text* file is scanned with its NULs stripped;
stripping can only splice tokens together, never apart, so it fails closed.
File extensions are deliberately not consulted, so a suffixless shell script
is still scanned.

The size check now runs BEFORE the strip: stripping shrinks the buffer, so
checking afterwards would let an oversized file slip under the threshold and
skip the fail-closed branch. (Caught by
test_oversized_nul_bearing_text_still_fails_closed, which failed on the first
cut of this patch.)

Return values are unchanged, so this does not conflict with the in-flight
crash-class fixes to the same function.

Tests (tests/hermes_cli/test_gateway_restart_loop.py), 3 of which fail on main:

- test_nul_padded_script_is_still_scanned
- test_nul_padded_script_without_shebang_is_scanned
- test_oversized_nul_bearing_text_still_fails_closed
- test_elf_binary_is_not_scanned_as_script       (#76762 stays fixed)
- test_macho_binary_is_not_scanned_as_script     (incl. fat binary)
- test_clean_script_without_lifecycle_command_not_blocked
2026-08-23 18:24:36 -07:00
Meng Chee da30db8e8c fix(cron): scan dot-operator sourced scripts in lifecycle guard
`_iter_referenced_shell_scripts` recognises the `source` builtin so a script
pulled in with `source ./restart.sh` gets scanned for lifecycle commands. The
POSIX dot operator is the same builtin, but it was not caught:

    if executable_name in {".", "source"}:

`executable_name` is `Path(executable).name`, and `Path(".").name` is the
**empty string** -- pathlib normalises "." to the current directory, whose name
is "". So the set membership never matched for `.`, the sourced script was
never added to the reference walk, and its contents were never scanned.

Verified against current main:

    . /tmp/restart.sh        -> not blocked   (script never scanned)
    source /tmp/restart.sh   -> blocked
    bash /tmp/restart.sh     -> blocked

where /tmp/restart.sh contains a `hermes gateway restart` line. Sourcing runs
the script in the current shell, so the dot spelling is not merely equivalent
to `source` -- it is the more common form in practice.

Fix compares the raw token as well as the basename:

    if executable in {".", "source"} or executable_name == "source":

Keeping the `executable_name == "source"` arm preserves the existing behaviour
for a path-qualified spelling, while the raw-token test catches `.` without
relying on pathlib normalisation.

Tests (tests/hermes_cli/test_gateway_restart_loop.py):

- test_dot_operator_sourced_script_is_scanned -- the regression; fails on main
- test_source_builtin_sourced_script_is_scanned -- `source` stays blocked
- test_dot_operator_clean_script_not_blocked -- widening the check must not
  false-block an innocent `. ./activate.sh`

Found while auditing the guard after #76762. Scoped deliberately to this one
defect; the NUL-padded-script bypass I found in the same audit is a separate
PR.
2026-08-23 18:24:36 -07:00
Teknium b4d4167d42 fix(gateway): lazy/unpersisted resume also rebinds transport and cancels the pending reap
Live WS E2E after the #93361 merge (real web_server + tui_gateway, isolated
HERMES_HOME, 2s grace): drop socket -> re-resume stored id on a new socket
still produced a ws_orphan_reap reclaim. The lazy/unpersisted resume branch
(no state.db row yet -- every fresh Bot Chat) returned the sentinel-parked
live record without rebinding its transport or cancelling the armed reap
Timer, so the storm survived for exactly the Bot Mode sessions the cluster
targeted. The unit-covered paths (_live_session_payload, _reuse_live_response,
_claim_or_reuse_live) were all correct; this branch bypassed them.

Regression test drives the real session.resume RPC against a sentinel-parked
unpersisted record (sabotage-verified: fails without the fix). After the fix
the full live E2E passes 10/10 scenarios including a 4-cycle drop/resume storm
loop with zero reclaim broadcasts.
2026-08-23 18:24:26 -07:00
Teknium 65c58651b0 feat: review slot appears in every aux-model picker (desktop, dashboard, CLI)
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.

- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
  stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
  zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
  _AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
  of fallback-providers and the delegation /review section missed in
  #93339
2026-08-23 18:22:39 -07:00
Teknium 0c713049ef chore: add contributor email mapping for zgqq 2026-08-23 18:10:31 -07:00
zgqq 9aa0721b23 fix: inert heredoc bodies no longer trip the gateway lifecycle guard (#88336)
Runbook prose inside a quoted-delimiter heredoc feeding a data sink
(cat > file <<'EOF') is documentation, not a command this shell will
execute. Mask provably-inert heredoc bodies (tools/shell_heredoc's
conservative stripper, already used by terminal_tool) before scanning.
Fails open on any ambiguity: executable and unquoted-delimiter heredocs
stay scanned. Salvaged from PR #88336 by @zgqq (the heredoc half; its
Branch D boundary and dir-token halves already landed/were fixed).
2026-08-23 18:10:31 -07:00
Teknium c94ee2e06f chore: add contributor email mappings for arcimun and KeaneYan 2026-08-23 18:01:59 -07:00
Artur Hapantsou b34edd6b01 fix: execute_code and argv-list payloads no longer bypass the gateway lifecycle guard (#68289)
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
2026-08-23 18:01:59 -07:00
KeaneYan 1c791cbfe6 fix(gateway): resolve uninstall lifecycle guard conflict 2026-08-23 18:01:59 -07:00
Teknium ca8a598787 chore: add contributor email mapping for 03farren 2026-08-23 17:56:36 -07:00
BotUser 679e07a074 fix(gateway): close order-dependency + missing-verb gap in launchctl lifecycle guards
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.

A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:

    for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
      label=${item%%:*}; plist=${item#*:}
      launchctl bootout "gui/$uid/$label"
      launchctl bootstrap "gui/$uid" "$plist"
    done

The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.

cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).

`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.

We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.

Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.

Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 17:56:36 -07:00
Teknium fa63c3e1c2 chore: add contributor email mapping for acewong7 2026-08-23 17:54:20 -07:00
thenabbu acf8245607 fix(tools): pass single_query_deny_message to the ssh-config write approval gate
Commit 1596148ff made single_query_deny_message a required keyword-only
parameter of _run_approval_gate() and updated its two callers inside
tools/approval.py, but missed the third caller: the SSH-config write
guard in tools/file_tools.py (_check_approval_required_write,
pattern_key="ssh_config_write").

Any gated write to an SSH client config therefore raised
  TypeError: _run_approval_gate() missing 1 required keyword-only
  argument: single_query_deny_message
instead of routing through the human-approval flow.

- Pass the kwarg with a single-query-specific deny message that points
  operators at approvals.single_query_mode: approve.
- Add a regression test asserting the gate call passes every required
  kwarg (fails on unpatched main).

Fixes #93201
2026-08-23 17:54:20 -07:00
Kang Wang dd2b5172e4 fix: pass single_query_deny_message to approval gate for ssh config writes 2026-08-23 17:54:20 -07:00
Teknium 20e308fea7 fix: use lookbehind anchor so binary-decoded and remote-read content still scans
The separator-class anchor broke two fail-closed tests (binary bytes
decode to U+FFFD adjacent to the CLI name; remote head-c reads). A
negative lookbehind excluding path/word chars keeps the #77173 fix
while preserving every fail-closed content-scan path.
2026-08-23 17:50:48 -07:00
Teknium ec169b4885 chore: add contributor email mapping for eaglezzz0522-cloud 2026-08-23 17:50:48 -07:00
eaglezzz0522-cloud 180f981125 fix: lifecycle guard Branch A anchors the CLI name at command position (#77173 path false positive)
A file path with embedded spaces (/docs/... with lifecycle words in the
filename) matched Branch A via the path tail and hard-blocked innocent
commands. Anchor the CLI name at command position (start, separator, or
substitution opener). Salvaged from PR #77536 by @eaglezzz0522-cloud,
reapplied onto the current pattern with subshell coverage and tests.
2026-08-23 17:50:48 -07:00
Teknium 778c384120 chore: map JinUltimate1995 contributor email 2026-08-23 17:47:50 -07:00
liuhao1024 51239e8e2a fix(vision): forward the API key to the server-type probe and cache failed verdicts
The image-routing vision path calls detect_local_server_type without
the provider's API key. Against a remote API-keyed endpoint (sglang /
vLLM with --api-key) every leg of the 5-request probe waterfall came
back 401 — and because a failed verdict was never written to the
in-memory cache (only positive verdicts were), the waterfall re-ran on
EVERY image-bearing turn (#89863: 51 detail-less busy-acks observed in
one Slack channel while the probe sprayed the user's own server).

Two changes:

- image_routing._should_probe_ollama_vision now takes the API key and
  forwards it; a new _resolve_inference_api_key mirrors
  _resolve_inference_base_url's resolution order (runtime value,
  model.api_key, providers blocks) so the key always matches the URL
  being probed.

- detect_local_server_type caches a None verdict in memory with a short
  failure TTL (5 min, vs 1h for positives) so the next turn is served
  from the negative entry instead of re-running the waterfall — while
  a transient failure (server starting, key being fixed) recovers in
  minutes. Negative verdicts are deliberately not written to the
  cross-process disk cache.
2026-08-23 17:47:50 -07:00
JinUltimate 4ca993c746 fix(image_routing): stop fingerprint-probing remote OpenAI-compatible endpoints
Fixes #89863. With a custom: provider pointing at a remote, API-keyed
endpoint (sglang/vLLM/OpenAI-compat), every image turn triggered a 5-request
probe waterfall without Authorization, spraying 401s at the backend.

Two fixes:

1. _should_probe_ollama_vision now takes api_key and forwards it to
   detect_local_server_type so keyed local servers don't 401.

2. When provider != 'ollama', remote endpoints (per is_local_endpoint) are
   rejected early — server-fingerprint probing is only valid for local
   boxes. Non-Ollama remotes expose Ollama-compat endpoints that can
   misidentify and trigger unnecessary /api/show probes.

_lookup_supports_vision resolves the runtime api_key via
_runtime_main_value and forwards it to both helpers. New test class
TestShouldProbeOllamaVision covers the contract in both directions.
2026-08-23 17:47:50 -07:00
honor2030 fe483de4d3 fix(agent): keep max-iteration warnings out of quiet stdout
Route the max-iterations diagnostic through logging when quiet_mode is active so automation wrappers keep stdout machine-readable.

Add a regression test covering quiet max-iteration summary handling.
2026-08-23 17:45:58 -07:00
Teknium 36bbb41d6c fix(desktop): single-flight tolerates sync resume runners; delegate tests assert owner routing
CI sibling-test blast radius from the cluster branch:
- singleFlightSessionResume crashed on run() doubles that return
  non-promises (Cannot read 'finally'); wrap via Promise.resolve().then(run).
- Three use-session-tile-delegate tests pinned the pre-#92961 ambient
  dispatch for default-profile sessions; the routing-authority change
  intentionally routes every known owner through the profile router, so
  the tests now assert requestGatewayForProfile('default', ...) instead.
2026-08-23 17:43:39 -07:00
Teknium 0b2e8b8a3b docs: dashboard ws keepalive + orphan-reap grace config keys 2026-08-23 17:43:39 -07:00
Teknium b0af119963 fix(desktop): resolve the owning remote profile before a hint-less session read falls through to local
Fixes #85834 (Electron REST intercept fall-through). The
/api/sessions/{id}[/messages] intercept in electron/main.ts required an
explicit ?profile= (or request.profile) to route a read to its remote owner;
callers without a hint fell straight through to the LOCAL backend and 404'd
on its state.db even though the session lives on a configured remote — while
the list endpoints happily showed the row (remoteSessionList tags s.profile).

When no explicit profile resolves, consult the same remote session lists the
list endpoints use to find the owning profile (matching id or lineage root
id), memoized for 30s so a transcript+messages burst costs one sweep. Only
when the id is genuinely unknown remotely does the request fall through to
local, exactly as before. Pure lookup lives in profile-session-routing.ts
with unit tests (owner hit, lineage-root match, null on miss/dead
remotes/no remotes).

Maintainer commit (cluster salvage).
2026-08-23 17:43:39 -07:00
Teknium 09047ec69c fix(desktop): route approval.respond through the session's owner, not the ambient socket
Client half of #91684. The approval bar (approval.tsx) and the native
notification action path (native-notifications.ts) sent approval.respond on
the AMBIENT gateway socket. Ambient follows foreground focus; for an approval
raised by a cross-profile or tile-owned session it points at a backend that
never held the approval, so Run/Reject silently failed after a profile swap
or reconnect.

- New knownOwnerForSession/requestForOwnedSession in store/session-states.ts:
  resolve the owner sync (tile owner route -> known session profile via row or
  open-time hint; runtime ids translated to stored ids first) and dispatch via
  requestForSessionProfile. Ambient only when no owner is known — never a
  fall-back to "active".
- approval.tsx and native-notifications.ts respond through it, binding the
  ambient dispatcher so the no-owner path keeps the exact 2-arg call shape.
- Tests: owner resolution (tile route first, row-profile fallback,
  undefined for unknown/null) and ambient arity preservation in
  session-states.test.ts; existing approval + native-notification suites
  still pass unchanged on the ambient path.

Maintainer commit (cluster salvage).
2026-08-23 17:43:39 -07:00
Teknium 77dd069946 fix(desktop): single-flight session.resume per stored id + adopt-or-reuse on drift-abort
After sleep/wake or a reconnect, many surfaces discover the same dead runtime
at once (submit recovery, slash/rewind recovery, tile resumes, the target
resolver, session switch) and each fired its own session.resume — the gateway
minted a runtime per call and the losers fed the orphan reaper (#91276 storm).

- New use-prompt-actions/single-flight-resume.ts: module-level in-flight map
  keyed by storedSessionId; all resume call sites (utils.ts recovery, submit.ts
  direct rung, resolve-target-session.ts, use-session-actions switch resume,
  use-session-tile-delegate resumeTile) share one in-flight promise per stored
  id. Failed flights are not cached.
- Drift-abort paths no longer abandon a freshly-minted runtime: utils.ts
  SessionRecoveryAborted and submit.ts post-routed-resume / post-resume aborts
  register it in a stored->runtime recovery cache; the next action for that
  stored session adopts it (via onRecovered) or reuses it instead of minting
  another. Cache entries are take-once and skip a known-dead id.
- Unit tests: one RPC for two concurrent callers of the same stored id,
  drift-abort registers (not strands) the recovered runtime, independent
  stored ids resume independently, cached-runtime adoption.

Maintainer commit (cluster salvage, part of the session-not-found-after-
reconnect consolidation).
2026-08-23 17:43:39 -07:00
Teknium 18e941ac2f fix(desktop): make the explicit-queue-target recovery regression load-bearing
Follow-up to PR #91357 (salvaged, author enwaiax): the committed #90428
explicit-target regression fixture started foreground B with a valid active
runtime and a positive B->runtime cache entry, so routedSessionNeedsResume was
false and the formerly broken foreground-recovery branch was never exercised —
the test passed even on the broken head (b9df1f9c2).

Strengthen it per the review: B now starts with activeSessionIdRef null and an
empty ownership cache, resumeStoredSession(B) fully publishes B's runtime and
cache binding, and the assertions still require no high-level resume of B, an
authoritative session.resume(C), exactly one queued prompt.submit to C's
recovered runtime, and no mutation of foreground refs/cache.

Salvaged-from: PR #91357 (author enwaiax); fixture hardening by maintainer.
2026-08-23 17:43:39 -07:00
Shawn Wang 11fd82cf20 fix(desktop): isolate explicit queued submit recovery
Signed-off-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>
2026-08-23 17:43:39 -07:00
Shawn Wang 534719259c fix(desktop): recover routed submits after reconnect
Signed-off-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>
2026-08-23 17:43:39 -07:00
Teknium d12bc0c4d9 fix(desktop): active gateway is never a session-RPC routing authority
Step 2 of removing 'active gateway' as a routing input. A session's backend
is a property of the SESSION (its profile), never of whatever the window is
currently showing. The active-profile fallback was the root cause of Bot Mode
'session not found' / hangs: a hidden/unlisted session with an unknown owner
was silently dispatched to the active profile's backend, which never owned it.

- sessionRpcNeedsProfileRoute: drop the active-profile comparison entirely. A
  KNOWN owner (route or profile name) ALWAYS routes to its own profile's
  socket; only a null/empty owner (fresh draft, global chrome) routes ambient.
  A primary-profile owner collapses back to the primary socket inside
  gatewayForProfile, so the reauth-aware reconnect path is unchanged.
- session.ts: split knownSessionProfile (row -> hint, undefined when unknown)
  out of rememberedSessionProfile. rememberedSessionProfile keeps its active
  fallback but is now documented as PRESENTATION-only (navigation keying),
  never routing.
- wiring requestGateway: resolve the owner from the tile route -> known
  profile -> a cross-profile REST probe (resolveSessionProfile, stamps
  ownership) before dispatch; only a request with no session at all falls to
  ambient. Never the silent active fallback.

Tests updated to the new contract + knownSessionProfile coverage asserting it
returns undefined (not active) for an unknown session. tsc 0 errors.
2026-08-23 17:43:39 -07:00
wz-heng 319e77eea4 fix(desktop): clarify preserved session tile recovery 2026-08-23 17:43:39 -07:00
wz-heng 72f6127b98 fix(desktop): preserve live session tiles after reconnect 2026-08-23 17:43:39 -07:00
yu-xin-c febed060ab fix(desktop): force gateway reconnect after wake 2026-08-23 17:43:39 -07:00
Teknium 9b3f60c029 fix(gateway): resolve approval.respond by durable identity before failing 4001
Server half of #91684: the desktop can answer an approval prompt with a
stale live sid — its runtime record was re-minted after a reconnect while
the prompt stayed on screen. approval.respond now falls back, on 4001
only, to resolving the target session (1) by the unique approval
request_id across every live session's pending gateway approvals, then
(2) by treating session_id as a STORED session id mapped to its live
runtime record. Only when neither resolves does it return 4001.

Tests: request_id fallback, stored-id fallback, and 4001 when nothing
resolves.
2026-08-23 17:43:39 -07:00
Teknium fdd8d75ba0 fix(gateway): make ws keepalive and orphan-reap grace config-driven (#79635)
- New dashboard.ws_ping_interval / dashboard.ws_ping_timeout defaults
  (20.0/20.0) in DEFAULT_CONFIG; hermes_cli/web_server.py reads them for
  non-loopback binds. Loopback keeps ws_ping=None (event-loop stalls must
  never kill a healthy local connection).
- New dashboard.ws_orphan_reap_grace_s (20.0): tui_gateway/server.py's
  _WS_ORPHAN_REAP_GRACE_S now resolves from config via
  _resolve_ws_orphan_reap_grace(); the HERMES_TUI_WS_ORPHAN_REAP_GRACE_S
  env var is kept as an internal override for backward compat and wins
  when set.
- tests/test_ws_keepalive_config.py: real load_config against a temp
  HERMES_HOME yaml — defaults, propagation, deep-merge, env override,
  invalid-value fallback.
2026-08-23 17:43:39 -07:00
Teknium 4aa162b30d fix(gateway): cancel pending WS-orphan reaps on resume and supersede stale runtimes quietly
Storm killer for the reap->broadcast->auto-re-resume feedback loop:

- New _pending_ws_reaps registry (sid -> Timer): _schedule_ws_orphan_reap
  registers, _reap pops, and _cancel_ws_orphan_reap(sid) is called from
  every resume/reuse/rebind path — the session.resume fast-path reuse
  (methods_session.py), _claim_or_reuse_live winners, and the
  _live_session_payload live-transport rebind.
- When a resume mints a fresh runtime for stored session id S, any prior
  runtime for S still parked on the detached-WS sentinel is claimed under
  the resume lock, its reap Timer cancelled, and the record finalized
  quietly with end_reason superseded_by_resume — NOT in
  _RECLAIM_END_REASONS, so no session.reclaimed broadcast fires and the
  client's auto-re-resume can't storm.
- superseded_by_resume added to _RECOVERABLE_END_REASONS in
  hermes_state_common.py so canonical Bot Chat resurrection still applies.

Unit tests: resume cancels the reap timer, superseded runtimes finalize
without a reclaimed broadcast, and the normal orphan reap still fires
when nobody re-resumes.
2026-08-23 17:43:39 -07:00
Ayush Nangia f2dbd37ef9 fix(gateway): re-bind session transport to a surviving window on pop-out close
Live sessions hold one transport; a pop-out window's session.resume
rebinds it, and on pop-out close the disconnect path parked the session
on the drop sentinel — the original window never received stream events
again until a manual re-resume (#83716).

Sessions now track every transport that has shown them (viewers, stamped
in _live_session_payload). _close_sessions_for_transport re-binds to the
most recent surviving viewer instead of detaching when one exists; dead
viewers are filtered; the drop sentinel + grace reap remain the path for
the last viewer. Root cause and repro by CharlesR-sudo on #83716.
2026-08-23 17:43:39 -07:00
Kyzcreig 3dd0ed1d38 fix(tui-gateway): bound the interrupt-then-reap poll chain (review finding)
If an interrupted turn never settles (agent thread hung in a syscall,
supervisor lost), the 1s poll chain rescheduled forever — trading the old
leak-one-worker bug for leak-one-session-plus-timer-chain. After
_WS_ORPHAN_INTERRUPT_REAP_MAX_POLLS (60 = ~60s, 3x the default grace) the
reaper logs loudly and force-reaps, mirroring the pre-existing stuck-running
safety net's deadlock-breaking role. Focused suite 4 passed, 1 skipped.
2026-08-23 17:43:39 -07:00
Kyzcreig 14b50f5edd fix(tui-gateway): interrupt turns after websocket disconnect
After the existing reconnect grace, route a still-detached running session through the same interrupt mechanism as session.interrupt. Preserve delegation deferral, sidecar teardown, partial history, and single-owner reap semantics.

Verified on upstream main: RED 4 failed/2 passed without production changes; GREEN 603 related gateway/compute-host tests. Ruff and py_compile passed. Momus pass 2: APPROVE.
2026-08-23 17:43:39 -07:00
686f6c61 30a37668d9 fix(tui): reschedule WS orphan reap while a turn is still running
The grace timer treated mid-turn detached sessions as not-orphaned and
returned without arming another timer. Eviction paths also skip
running sessions, so the in-memory agent leaked until process restart.

Fixes #85578
2026-08-23 17:43:39 -07:00
hermes-seaeye[bot] 76c356c28a fmt(js): npm run fix on merge (#93381)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-24 00:43:15 +00:00
John Paul Soliva 3dc77acdc3 feat(desktop): OAuth sign-in for registry connections; keep profile picks on the browsed source (#92194)
* feat(desktop): add OAuth sign-in to the connections registry editor

A gated remote gateway (OAuth, or username/password) never accepts a
session token — it authenticates with a browser sign-in and the desktop
keeps whatever the flow mints. The registry editor only rendered a token
field for 'token' mode and nothing at all for 'oauth', so a gated
connection could be created but never authenticated: selecting OAuth left
an empty row, and Test failed with no way to fix it.

Render an Authentication row in the oauth branch that calls the existing
oauthLoginConnectionConfig IPC — the same one first-run-remote-form and
the gateway panel already use. The URL is probed (debounced) so the row
can name the provider and use password-specific copy when every
advertised provider supports passwords, matching gateway-settings.

No new i18n keys; all strings already exist under settings.gateway.
Test needed no change: testDesktopConnectionConfig already skips the
token for oauth and mints a ws-ticket from the session.

* fix(desktop): keep profile picks on the source being browsed

$profiles is the ACTIVE gateway's list, so a profile picked while a
registry source is live names one of THAT source's profiles. Both
selectProfile and newSessionInProfile sent it through the profile-only
path, which resolves the descriptor with a bare name — and
getConnection(profile) is answered against the primary. Picking
"researcher" while browsing a remote source therefore opened a LOCAL
backend of that name and snapped the gateway home, so the pick looked
like it never took: the user could reach the agent from Bot Mode but
never from the profile switcher.

Route both through the live source instead: a non-null
activeGatewayConnectionId means a registry source owns the current
gateway, so activate the (connection, profile) agent. A null id means
the primary is live, which is exactly the legacy path — single-source
users keep their existing behavior unchanged.

* fix(desktop): cancel the registry auth probe on unmount; reset the signed-in pill on mode flips

Review follow-ups on the OAuth sign-in row: the debounced probe sets a
cancelled flag in its effect cleanup (probeSeq covers staleness but not
unmount), and oauthConnected resets when the auth mode flips as well as on
URL changes — a saved row edited token -> oauth no longer reports a stale
'Signed in' from an earlier oauth stint.
2026-08-23 19:38:43 -05:00
Teknium 12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
xxxigm 0c1f1d2fe5 fix(desktop): stop Inbox-style session cards from clipping text (#93036)
* fix(desktop): stop Inbox-style session cards from clipping glyph ink

leading-none plus truncate (overflow:hidden) made the line box equal the
em-square, so Segoe UI on Windows shaved letter tops and bottoms. Give
truncated sidebar text 1.35 line-height and tighten card gaps so the
taller lines still fit.

* test(desktop): lock Inbox card lines to a line-height that fits glyph ink

Assert the workspace, title, and footer lines keep leading-[1.35] and
never fall back to leading-none, which is what clipped the screenshot.
2026-08-23 19:34:45 -05:00