Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.
Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.
6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.
Fixes#64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
- curly: brace all single-line if statements in profile-migration.ts and
profile-migration.test.ts (17 errors in CI check:lint)
- perfectionist/sort-imports: node:fs builtin import before vitest external
- padding-line-between-statements: blank lines after block statements
- prettier: normalize formatting (fmt script style) in the three touched files
All 967 electron project tests still pass.
Polish from antigravity review of the rebase resolution (GPT-OSS):
the previous comment said "BEFORE the first primaryProfileKey() /
primaryBackendIsRemote() read" but those two calls live at different
points — primaryBackendIsRemote() is the very next line, primaryProfileKey()
is inside the connection IIFE. Be explicit about which is where so a
future reader who moves one of them knows what to preserve.
Addresses teknium1's review (#64195) finding #2: the multi-rung resolver
needs Electron tests covering precedence, stale-PID rejection, fallback
behavior, and the remote boot path. The pure decision helpers are now
covered by 29 unit tests in `profile-migration.test.ts` (vitest electron
project).
Coverage:
- precedence: legacy > single-running-gateway > state.db heuristic
- stale-PID rejection: recycled PIDs not owned by hermes are dropped
- malformed pid files: JSON parse errors, non-integer PIDs, zero/negative
- scoring edge cases: ancient files (recency floored at 0.1), tiny files
(size floored at MIN_SIZE), larger DB beats smaller at similar recency
- single-profile fallback: best === 'default' suppresses the write
- no-op cases: preference file already exists, missing profiles root
The remote boot path is verified by code review of the call-site move
(commit preceding this one) — `migrateActiveProfileIfMissing()` now runs
before `primaryProfileKey()` is first read in `startHermes()`.
The pure decision logic that the orchestrator relies on is covered end-
to-end below; this matches the repo's testable-helper pattern (see
`profile-delete-routing.test.ts`).
Addresses teknium1's review (#64195) finding #1: the previous PR placed
the migration inside the connection IIFE, AFTER
`resolveRemoteBackend(primaryProfileKey())`. When the preference file
was missing, `primaryProfileKey()` resolved to 'default' and the remote
branch returned immediately without ever reaching the migration. Remote-
mode users got no migration at all.
Move the call site to the top of `startHermes()`, before the connection
IIFE that reads `primaryProfileKey()`. Both remote and local branches now
flow through this path before any profile-dependent resolution, so the
migration runs on first boot regardless of mode.
The inlined implementation is replaced with a thin wrapper that builds a
`MigrationDeps` bag and delegates to `migrateActiveProfileIfMissing` from
`profile-migration.ts`. No production behavior change beyond the call-
site move.
Tests added in a separate commit.
Addresses teknium1's review (#64195) — the migration decision logic should
be unit-testable without Electron. Pull the ladder (legacy sticky file,
running-gateway scan, state.db heuristic) into pure helpers in a new
`profile-migration.ts` module, following the dep-injection pattern
already established by `profile-delete-routing.ts`.
The helpers take an injected `MigrationDeps` bag so tests can exercise
precedence, stale-PID rejection, fallback behavior, and the single-profile
case without touching `/proc`, `ps`, or the host filesystem. The default
profile is explicitly rejected from the legacy rung because the regex
matches `default` and accepting it would suppress the heuristic that is
the whole point of the migration.
The wrapper in main.ts is unchanged in behavior — the same `MigrationDeps`
fields get filled in from `fs`/`path` and `isHermesProcess`. The atomic
write + parent-dir-create that the wrapper performs matches
`writeActiveDesktopProfile`'s semantics so the migration produces a file
indistinguishable from a user-driven profile switch.
No production behavior change; pure code organization.
Tests added in a separate commit.
When active-profile.json does not exist (fresh install or first boot after
update), seed it from the best available signal so the Desktop launches
into the user's primary profile instead of always defaulting to "default".
Priority ladder:
1. Legacy ~/.hermes/active_profile (explicit CLI choice via hermes profile use)
2. Running gateway (gateway.pid with verified liveness + hermes identity check
via /proc/cmdline or ps -o args= to avoid PID recycling false positives)
3. state.db heuristics — hybrid recency×size score picks the primary workspace
(e.g. a 409MB coder DB beats a 28MB default DB even if touched at similar times)
The stored JSON includes _migrated:true for priority 3 (heuristic guess) so
the renderer can optionally surface a one-time notification. Priority 1 and 2
are higher-confidence signals and skip the flag.
The migration is a no-op once active-profile.json exists, and only writes
when a non-default profile is confidently identified — preserving the legacy
fallback-to-default behavior for single-profile users.
Fixes#64160 (active-profile half).
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.
Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.
Fixes#99839
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.
That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.
The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
The salvaged commit added the key to web/src/i18n/types.ts and en.ts
only; the 15 web locales typed as full Translations (fr/de/es/it/pt/
ru/tr/uk/ko/ja/zh/zh-hant/af/ga/hu) then failed 'tsc -b' (the real
build path) with TS2741. web/ar.ts uses defineLocale and falls back
to en, so it needs nothing. Follow-up for salvaged PR #98716.
The local expires_in countdown killed OAuth sessions with a bare
"Session expired" before the backend poller's enriched message (Portal
sign-in stalled in the opened tab, retry/API-key fallback) could reach
the UI, and the desktop onboarding poller had no local expiry at all —
a dead session polled forever. Both surfaces now lapse with guidance
naming the common cause, prefer the backend error_message when it has
one, and keep polling when the backend still reports pending (clock
skew).
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).
Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).
Part of the #90837 sidecar-unlink audit (wave 6).
The test exercised _run_chrome_fallback_command against the REAL
/tmp agent-browser namespace, so any concurrent hermes/pytest process
running the orphan reaper could rmtree the fresh pidless socket dir
mid-command (FileNotFoundError on _stdout_open — the recurring CI
flake, 3 hits incl. two main runs on 2026-09-01). Route it through a
private tmp_path like every sibling browser test.
_run_chrome_fallback_command creates agent-browser-<session> in the shared
tmpdir and then opens stdout/stderr files inside it, but never writes the
<session>.owner_pid marker. _reap_orphaned_browser_sessions rmtree's any
agent-browser-* dir that carries no live owner and is not tracked in the
calling process, so a second hermes process — or a parallel test worker —
deletes the directory between the makedirs and the first os.open, and the
command dies with FileNotFoundError on _stdout_open.
Write the owner marker immediately after creating the directory, which is
what every other socket-dir user already does.
Deterministic repro on main: run any test that exercises the fallback while
a second process calls _reap_orphaned_browser_sessions() in a loop —
5/5 fail before, 6/6 pass after.
Post-merge review follow-ups on #100201:
- acquire(): drop the never-iterating 'while True' and the redundant
'existing is not generation' half of the race check — after the retire
path, generation is always None on the fresh-open leg, so 'existing is
not None' is the complete condition. Same behavior, flat flow.
- test_mirror: the conversion to the shared registry dropped the
cleanup assertion entirely; restore it by patching
hermes_state.release_or_close and asserting the handle is released
exactly once after _append_to_sqlite.
Installs converted by the reverted macOS TCC interpreter anchor
(#95425/#95541) are left with a real-file venv/bin/python copy, a
.tcc-anchor-source marker, and python3/python3.N aliases that die at
interpreter init ("No module named 'encodings'"). venv/bin/hermes execs
venv/bin/python3, so EVERY CLI entrypoint is dead — hermes doctor and
hermes update included — and the desktop hand-off loops on "Update failed
(exit 1)" forever. No Python-side heal can ever run on this class; the
hand-off shell is the last surface that still executes, so the heal lives
there.
posix.sh gains, before the update invocation:
* tcc_anchor_heal — probe-gated (only fires when venv/bin/python3 fails
a scrubbed-env `import encodings` boot probe), marker-validated
(absolute path, outside the venv), staged with per-attempt backups and
full rollback if the repaired interpreter still fails its probe.
Two repair shapes:
- alias-brick (#95541 class): the anchored copy boots — re-materialize
python3* as REAL FILES of the anchor (hardlink/copy; an alias symlink
onto the copy is the crash shape). Marker kept: this is exactly the
layout ensure_tcc_anchor marks "active", so no anchor ping-pong.
- full brick: restore python → symlink to the marker-recorded store
interpreter (if it boots) and aliases → symlinks; marker removed.
The unblocked `hermes update` then re-installs a boot-gated healthy
anchor — one-shot convergence, not a loop.
Fail-closed on missing/unbootable source (vanished uv store class),
missing marker, or relative/in-venv marker paths.
* tcc_pick_update_invoke — if aliases stay dead but venv/bin/python
boots (the launchd-gateway shape), drive the update via
`venv/bin/python -m hermes_cli.main` instead of the dead hermes shim.
* Honest terminal message: an unrecoverable dead interpreter is reported
as a venv repair problem instead of a generic "Update failed (exit 1)".
* --self-test-tcc-heal runs the real heal + invoke selection against an
--install-root and reports, for the test harness.
Tests (tests/test_desktop_update_tcc_heal.py) drive the REAL posix.sh
functions on Linux against synthetic venv trees: healthy no-op, alias
heal, symlink restore, fail-closed classes, rollback on failed
verification, invoke fallback, and an A/B of the reported loop (bricked
venv/bin/hermes fails with the exact field error before, boots after).
Sabotage-verified (re-introducing the alias-symlink bug fails 2 tests).
NOT mac-live-tested (no macOS runner); the heal is platform-independent
shell exercised through the self-test path without the uname gate.
Recovery design (validation/staging/rollback/probe pattern) after
@aeonsong's #96231; in-update heal intent from @liuhao1024's #95775
(its target function no longer exists on main and its heal point is
unreachable on dead-CLI installs); heal-point and ping-pong analysis by
@ahrazzle and @tokenfires on #95759.
Fixes#95759
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.
No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:
gateway waits on all in-flight work units (#77184 don't-amputate)
-> cron agent session waits on the \hermes update\ process to exit
-> \hermes update\ waits on the gateway to exit [back to A]
The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).
Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.
Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.
Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
(the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).
Fixes#100179
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).
- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
scheduled_at, dispatched_at, lateness_seconds, and kind
(on_time / late / catch_up, classified against the ticker tolerance and
the schedule's catch-up grace window). Manual triggers and one-shots are
not stamped (no scheduled instant to be late against / retired beyond
grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
catch-up, in both the built-in ticker and external provider paths.
CLI surface only — no new tools, no policy engine.
Addresses the visibility half of #99879.
- Translate leftover English 'toolset(s)' strings in ru.ts
- Cover ru aliases/config value in languages.test.ts (parity with ar)
- List Russian among desktop UI languages in website/docs/user-guide/desktop.md
Full desktop UI translation (ru.ts, 3076 string/function leaves,
mirrors en.ts 1:1) plus locale registration in types, catalog and
language list with aliases (ru, ru-ru, ru_ru, ru-by).
Russian plurals (1 / 2-4 / 5+, 11-14 exception) via RU_PLURAL/RU_NOUN
helpers; count accepts number | string to match en.ts signatures.
Verified against current main: structural validation 3076/3076 with
placeholder parity, typecheck (renderer/electron/e2e) clean,
i18n vitest 28/28, prettier + eslint clean.
Covers the reviewer-requested cases for #98555:
- successful streamed response emits the latest attempt's non-null
first_chunk_at (started_at <= first_chunk_at <= ended_at)
- non-streamed, failed-stream, and partial-stream-stub paths emit None
- a stale timestamp from a prior API call cannot leak into the next
call's post_api_request payload (per-attempt reset in the loop)
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.
Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.
Backward compatible: existing hook subscribers ignore unknown kwargs.
The spawn-time output tail (#93608) puts child stdout into flowing mode at
spawn. Both backend spawn paths in main.ts then await claimBackendChild
(whose Windows Get-Process probe cold-starts in 2-8s) and boot-progress IPC
BEFORE waitForDashboardPortAnnouncement attaches its stdout listener. Node
streams never replay consumed chunks to late listeners, so a READY sentinel
printed during that window was lost forever — the wait hit its 90s timeout
and a healthy backend was killed (deterministic on Windows, racy on
macOS/Linux; still firing on v0.21.0 incl. concurrent multi-profile boots).
Fix (belt and suspenders, both spawn paths — primary and profile pool):
- create the port-announcement promise immediately after spawn, before any
await
- new bufferedOutput option on waitForDashboardPortAnnouncement: after
attaching its own listener, waitForDashboardPort scans the output tail's
already-buffered text for the sentinel, making listener-attach ordering
irrelevant regardless of call-site shape
The readyFile path was already ordering-safe (it polls a file, not the
stream). Approach follows stale PR #60986 by @ParaWheeler, rebased onto the
output-tail/readyFile plumbing added since.
Fixes#60323
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.
Fixes#97343
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
The "Connecting to Telegram (attempt N/8)…" line logs at WARNING and
reaches the gateway's default stderr handler, but the matching
"Connected to Telegram (… mode)" line was INFO and went to the log file
only. A healthy startup therefore looked permanently hung at
"attempt 1/8" on the terminal — the logging-illusion half of #90835.
Promote the success line to WARNING so both sides of the connect
transition share the same console sink; a genuine hang is now the
absence of the success line. Adds an AST-level regression test pinning
the level pairing. Sibling adapters (homeassistant, wecom) log both
sides at INFO, so they don't have this asymmetry.
Fixes#90835
contributors/emails/agent@Agents-Mac-mini.local collides on
case-insensitive filesystems (macOS/Windows checkouts) with its
lowercase sibling, breaking clean checkouts. The identity is a
local-machine artifact, not a real contributor address.
Fixes#88168Fixes#100047Fixes#99966Fixes#99821
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
Salvaged from PR #13996 (@giwaov, issue #13963), modernized onto current
main: config key renamed to skills.create_dir (per PR #81002's naming),
resolution centralized in agent/skill_utils.get_skill_create_dir() with
~/${VAR} expansion and HERMES_HOME-relative paths, and the directory is
folded into get_all_skills_dirs() so created skills are discovered,
trusted, findable, and patchable like local ones. Out-of-root creations
report their absolute path instead of crashing relative_to().
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.
Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.
Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.
Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.
Fixes#87694
Salvaged from #87745 with expiry mitigation added.
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.
Scope now travels with the unit end to end:
- the in-process systemd loop records a scope-qualified twin of
`restarted_services` (`restarted_scoped_units`) while the bare-name list
keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
stays unqualified and is read as scope-agnostic, and an unrecognized
scope drops the skip rather than honouring it: dropping a skip can only
cost one more restart-and-verify, honouring an unreadable one can leave
a stale generation running.
Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.
Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.
Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.
Refs #92145
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.
The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.
The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.
- restart active `hermes-serve*` systemd units from the fresh child,
enumerated from systemd rather than from the misclassifying inventory,
and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
process, and never kill one -- a manual or Desktop-owned serve has no
relaunch authority;
- require every runtime family, not just the gateway leg, before a
fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
Three kills at the shared chokepoints:
1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
while the active config.yaml is corrupt (AuthError code=corrupt_config).
A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
model.provider and tier-3/4 silently adopted the PAID openrouter provider
against the user's real (unparseable) intent. New probe:
hermes_cli.config.get_active_config_parse_failure(), recorded in the
existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
a fixed file clears the block immediately. Explicit provider requests
are untouched.
2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
(nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
honored untouched (paid-lane warning retained).
3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
process per provider) when a credential is newly ingested — ingestion
itself stays allowed.
Fixes#81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
on every surface