#100540 added a REMOVED_BACKENDS startup warning keyed on tavily; with the
backend restored, that entry would warn on a working provider. The registry
stays (empty) for future removals; migration tests now pin the machinery via
a synthetic entry plus a guard asserting no live provider is ever listed as
removed.
GoalManager.set() on an event-loop thread only waits the bounded
_DB_BOOTSTRAP_INIT_WAIT_S window for the background SessionDB bootstrap
(deliberate: an unbounded init starved the gateway loop watchdog). On a
loaded CI runner the cold init overruns that window, the goal write is
silently dropped by design, and the test flakes downstream: /loop showed
no active-goal note and the goal continuation was never enqueued (both
FLAKY on main run 33455779041).
Fix the class: every async goal test fixture that clears goals._DB_CACHE
now pre-warms it via _get_session_db() from sync context (unbounded init
path), so the bounded-window degradation can never fire mid-test. Applied
to all four gateway goal/loop test files; sync-only goal tests are
unaffected by construction.
Live repro: slowing SessionDB.__init__ past the window reproduces the
dropped write deterministically without the pre-warm and never with it.
Path.rglob raises FileNotFoundError when a directory disappears between
listing and scandir — a sibling CI job creating/removing its sdist
extraction (hermes_agent-<ver>/) killed test_allowlist_has_no_stale_entries
on run 33531869442. Switch to os.walk (tolerates vanishing dirs) with
top-level pruning of exempt and packaging dirs; file set is byte-identical
(886 files verified old==new) and a 30-scan churn harness that reliably
exercised the window shows zero errors.
Same loaded-runner class: the 12-turn drain chain completed only 11
turns inside the 400x0.01s poll budget on main run 33455779041. The
loop still exits early on success, so the wider budget costs nothing
on healthy runs.
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).
Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).
The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).
Fixes#98028Fixes#100325
Draft frames set _last_sent_text for dedupe without setting
_already_sent (they are ephemeral); an ungated has_delivered_text match
let a draft-only preview count as durable delivery and regressed
test_relay_seal_failure's dead-transport guarantee on CI.
A record-less delivery flag (final_response_sent /
final_content_delivered set with no recorded turn-final payload) was
trusted blindly by delivered_final_matches (None -> legacy trust), so a
first-edit prefix or a truncated finalize suppressed the gateway's
corrective send — silent partial delivery.
- delivered_final_matches: record-less flags are now reconciled against
the FINAL content via has_delivered_text; only the explicitly-marked
ambiguous-timeout path (_delivery_ambiguous) keeps legacy trust.
- _try_fresh_final and the native-streaming optimistic finalize now
record their delivered payload (the last record-less flag setters);
the optimistic record rolls back on definitive dispatch failure.
- Discord adapter: dead-transport send failures (client gone, WS
closed/reset) are classified as send_path_degraded (retryable) so the
delivery-obligation ledger's reconnect sweep replays the stranded
final response instead of losing it until a process restart.
Fixes#95382; closes the #98552 false-positive class.
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).
- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
removed_backend_note(); selection_error() swaps in the specific
removal explanation (removed in v0.21.0, keyless alternatives) while
keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
search_backend / extract_backend against the registry and emits a
startup warning (deduped per stale value), surfaced by the existing
print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
per-capability keys, dedupe, healthy-config negative, live-backend
failure text preserved.
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.
Co-authored-by: Cursor <cursoragent@cursor.com>
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:
- _prune_replaced_custom_model_config_credentials skipped only the
preferred key, so a keyed provider's own legacy-named pool
(custom:b.ai) was false-pruned of its current model_config credential
when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
key, so a legacy-named pool stopped being seeded from model.api_key.
Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.
Follow-up to #100413.
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.
Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.
Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.
If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.
Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.
Co-authored-by: Cursor <cursoragent@cursor.com>
On the stateless api_server platform, a background delegate_task
completion was delivered by self-POSTing /v1/chat/completions with
role=user after the parent turn had already completed (event.complete,
finish_reason=stop). That starts an unauthorized new agent turn the
client never sent, persists the completion as an ordinary role=user row
(display_kind NULL), and can blow through a pending human-confirmation
gate — the exact skip-ahead reported in #85957.
Fix: async_delegation completions targeting a non-push api_server
session are now written into the session transcript as a durable
DELIVERY row (role=user + display_kind=async_delegation_complete +
display metadata — the same bookkeeping shape the TUI/desktop delivery
path persists). No agent turn runs; clients polling
GET /api/sessions/{id}/messages see the result immediately, and the
next real client turn carries it as context. Persist failures return
False so the durable claim is released and the completion retried.
Watch-pattern notifications keep the existing self-post wake behavior;
push-capable adapters are untouched.
Fixes#85957
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.
Two layers:
1. _wait_for_gateway_ready now treats the first hit as provisional: the
gateway must stay visible through a 2s confirmation window
(_confirm_gateway_stable) before it is reported ready; a death during
confirmation resumes polling until the deadline. Failure output is an
honest ✗ with the Job Object explanation and the schtasks /Run
recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
state/gateway.start-attestation.json with the vouched-for PIDs. The
next `gateway start`/`gateway status` invocation checks it — if the
attested PIDs are gone with no clean-exit record in the lifecycle
ledger, the CLI reports (once) that the previous ✓ was false and
prints the schtasks recovery hint. `gateway stop` and a clean
lifecycle-ledger exit clear the marker silently.
Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.
Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.
Fixes#91675
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.
Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.
6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.
Fixes#64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.
Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.
Fixes#99839
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.
That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.
The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).
Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).
Part of the #90837 sidecar-unlink audit (wave 6).
The test exercised _run_chrome_fallback_command against the REAL
/tmp agent-browser namespace, so any concurrent hermes/pytest process
running the orphan reaper could rmtree the fresh pidless socket dir
mid-command (FileNotFoundError on _stdout_open — the recurring CI
flake, 3 hits incl. two main runs on 2026-09-01). Route it through a
private tmp_path like every sibling browser test.
Post-merge review follow-ups on #100201:
- acquire(): drop the never-iterating 'while True' and the redundant
'existing is not generation' half of the race check — after the retire
path, generation is always None on the fresh-open leg, so 'existing is
not None' is the complete condition. Same behavior, flat flow.
- test_mirror: the conversion to the shared registry dropped the
cleanup assertion entirely; restore it by patching
hermes_state.release_or_close and asserting the handle is released
exactly once after _append_to_sqlite.
Installs converted by the reverted macOS TCC interpreter anchor
(#95425/#95541) are left with a real-file venv/bin/python copy, a
.tcc-anchor-source marker, and python3/python3.N aliases that die at
interpreter init ("No module named 'encodings'"). venv/bin/hermes execs
venv/bin/python3, so EVERY CLI entrypoint is dead — hermes doctor and
hermes update included — and the desktop hand-off loops on "Update failed
(exit 1)" forever. No Python-side heal can ever run on this class; the
hand-off shell is the last surface that still executes, so the heal lives
there.
posix.sh gains, before the update invocation:
* tcc_anchor_heal — probe-gated (only fires when venv/bin/python3 fails
a scrubbed-env `import encodings` boot probe), marker-validated
(absolute path, outside the venv), staged with per-attempt backups and
full rollback if the repaired interpreter still fails its probe.
Two repair shapes:
- alias-brick (#95541 class): the anchored copy boots — re-materialize
python3* as REAL FILES of the anchor (hardlink/copy; an alias symlink
onto the copy is the crash shape). Marker kept: this is exactly the
layout ensure_tcc_anchor marks "active", so no anchor ping-pong.
- full brick: restore python → symlink to the marker-recorded store
interpreter (if it boots) and aliases → symlinks; marker removed.
The unblocked `hermes update` then re-installs a boot-gated healthy
anchor — one-shot convergence, not a loop.
Fail-closed on missing/unbootable source (vanished uv store class),
missing marker, or relative/in-venv marker paths.
* tcc_pick_update_invoke — if aliases stay dead but venv/bin/python
boots (the launchd-gateway shape), drive the update via
`venv/bin/python -m hermes_cli.main` instead of the dead hermes shim.
* Honest terminal message: an unrecoverable dead interpreter is reported
as a venv repair problem instead of a generic "Update failed (exit 1)".
* --self-test-tcc-heal runs the real heal + invoke selection against an
--install-root and reports, for the test harness.
Tests (tests/test_desktop_update_tcc_heal.py) drive the REAL posix.sh
functions on Linux against synthetic venv trees: healthy no-op, alias
heal, symlink restore, fail-closed classes, rollback on failed
verification, invoke fallback, and an A/B of the reported loop (bricked
venv/bin/hermes fails with the exact field error before, boots after).
Sabotage-verified (re-introducing the alias-symlink bug fails 2 tests).
NOT mac-live-tested (no macOS runner); the heal is platform-independent
shell exercised through the self-test path without the uname gate.
Recovery design (validation/staging/rollback/probe pattern) after
@aeonsong's #96231; in-update heal intent from @liuhao1024's #95775
(its target function no longer exists on main and its heal point is
unreachable on dead-CLI installs); heal-point and ping-pong analysis by
@ahrazzle and @tokenfires on #95759.
Fixes#95759
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.
No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:
gateway waits on all in-flight work units (#77184 don't-amputate)
-> cron agent session waits on the \hermes update\ process to exit
-> \hermes update\ waits on the gateway to exit [back to A]
The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).
Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.
Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.
Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
(the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).
Fixes#100179
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).
- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
scheduled_at, dispatched_at, lateness_seconds, and kind
(on_time / late / catch_up, classified against the ticker tolerance and
the schedule's catch-up grace window). Manual triggers and one-shots are
not stamped (no scheduled instant to be late against / retired beyond
grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
catch-up, in both the built-in ticker and external provider paths.
CLI surface only — no new tools, no policy engine.
Addresses the visibility half of #99879.
Covers the reviewer-requested cases for #98555:
- successful streamed response emits the latest attempt's non-null
first_chunk_at (started_at <= first_chunk_at <= ended_at)
- non-streamed, failed-stream, and partial-stream-stub paths emit None
- a stale timestamp from a prior API call cannot leak into the next
call's post_api_request payload (per-attempt reset in the loop)
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.
Fixes#97343
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
The "Connecting to Telegram (attempt N/8)…" line logs at WARNING and
reaches the gateway's default stderr handler, but the matching
"Connected to Telegram (… mode)" line was INFO and went to the log file
only. A healthy startup therefore looked permanently hung at
"attempt 1/8" on the terminal — the logging-illusion half of #90835.
Promote the success line to WARNING so both sides of the connect
transition share the same console sink; a genuine hang is now the
absence of the success line. Adds an AST-level regression test pinning
the level pairing. Sibling adapters (homeassistant, wecom) log both
sides at INFO, so they don't have this asymmetry.
Fixes#90835
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.
Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.
Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.
Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.
Fixes#87694
Salvaged from #87745 with expiry mitigation added.
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.
Scope now travels with the unit end to end:
- the in-process systemd loop records a scope-qualified twin of
`restarted_services` (`restarted_scoped_units`) while the bare-name list
keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
stays unqualified and is read as scope-agnostic, and an unrecognized
scope drops the skip rather than honouring it: dropping a skip can only
cost one more restart-and-verify, honouring an unreadable one can leave
a stale generation running.
Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.
Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.
Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.
Refs #92145
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.
The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.
The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.
- restart active `hermes-serve*` systemd units from the fresh child,
enumerated from systemd rather than from the misclassifying inventory,
and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
process, and never kill one -- a manual or Desktop-owned serve has no
relaunch authority;
- require every runtime family, not just the gateway leg, before a
fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
Three kills at the shared chokepoints:
1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
while the active config.yaml is corrupt (AuthError code=corrupt_config).
A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
model.provider and tier-3/4 silently adopted the PAID openrouter provider
against the user's real (unparseable) intent. New probe:
hermes_cli.config.get_active_config_parse_failure(), recorded in the
existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
a fixed file clears the block immediately. Explicit provider requests
are untouched.
2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
(nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
honored untouched (paid-lane warning retained).
3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
process per provider) when a credential is newly ingested — ingestion
itself stays allowed.
Fixes#81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).