De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):
- web dashboard CronPage: last_status was never rendered at all — a
delivery_failed job showed a green 'scheduled' badge and only a small red
'delivery: ...' line. New pure cronLastResult() helper maps the closed
literal set to tones (ok=success, delivery_failed/blocked_config=warning,
error/unknown=destructive) and the card now shows an amber
'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
literal; routineLastResult() spells out each one ('Ran, but delivery
failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
on this branch; no consumer compared == 'ok' for success apart from the
cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
detail field carries the reason.
Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
A successful agent run whose delivery failed used to persist
last_status=ok and bury the failure in last_delivery_error. CLI list
painted that as green and the run looked identical to a quiet success.
Record last_status=delivery_failed instead, keep last_delivery_error,
do not increment failure_streak, and teach cron list/doctor not to
treat it as ok.
Fixes#83993
Widen the two salvaged fixes (#100490, #100493) to the whole class:
- match_runtime_outcomes: serve/dashboard rows never borrow gateway
bookkeeping at ANY site — not just the bare hermes-gateway unit name
(#100490) but also relaunched_profiles / externally_supervised_profiles
and the profile-substring unit match (hermes-gateway-work credited the
'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
units (exact names, scope prefix tolerated) or, when the caller passes
the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
feed the Phase-2 reconciliation, so a surviving unmanaged serve is
'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
remedy instead of 'hermes gateway restart', which cannot reach it.
Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
match_runtime_outcomes() treats any default-profile runtime as covered
once the bare "hermes-gateway" unit restarts, regardless of the
runtime's own kind. An sshd-spawned `serve --isolated` backend (no
systemd unit, supervisor "manual-serve") shares the default profile
and gets silently marked "restarted" even though its own PID was never
touched — so the #91277 Phase 2 unaccounted-runtime tripwire never
fires for it and `hermes update` reports success while it keeps
running pre-update code (#100479).
Restrict the "hermes-gateway" special case to kind == "gateway" so a
serve/dashboard runtime under the same profile falls through to
"unaccounted" instead of borrowing the gateway's outcome.
All seven TestLoopTickWitness cases that need real UNIX-domain sockets
(socket.AF_UNIX socket nodes or asyncio.start_unix_server producers)
fail on native Windows, where neither primitive exists. Mark exactly
those cases with a shared skipif so a Windows run reports SKIPPED
instead of erroring, while the platform-independent witness-absent
contracts (mocked probes, file-only heartbeats) keep running there.
Split the legacy two-witness-contract test in two: its stale-file arm
is file-only and keeps running on Windows; its dead-listener-node arm
needs a real socket node and is skipped with the rest.
Follow-up to @zoser69's #78111 cherry-pick:
- lift the redirect into _redirect_platform_display_key() and apply it
BEFORE _validate_config_key / type coercion, so the unknown-key hint and
the string-vs-bool coercion both see the canonical path
- widen to the sibling surfaces: config get resolves the canonical key
(previously echoed the dead top-level value — the misleading half of the
report) and config unset removes the canonical leaf
- regression tests: get mirrors gateway resolve_display_setting, unset
removes the redirected leaf, note printed, helper touches ONLY
OVERRIDEABLE_KEYS (connection keys / 4-segment / already-canonical
paths untouched)
- docs: configuration.md per-platform section names the canonical CLI
path and the accepted shorthand
Problem A of #71047: 'hermes config set platforms.telegram.streaming false'
wrote to a key the gateway never reads. The connection config
(gateway/config.py) reads only token/extra/overrides from the top-level
platforms.<name> block, while per-platform display settings (streaming,
show_reasoning, tool_progress, ...) are resolved from
display.platforms.<name>.<setting> (gateway/display_config.py).
Redirect a platforms.<name>.<setting> key to
display.platforms.<name>.<setting> only when <setting> is a known per-platform
display setting (gateway.display_config.OVERRIDEABLE_KEYS), leaving real
connection keys (token, extra, channel_overrides, ...) untouched. The
gateway.display_config import is lazy/try-guarded to avoid a circular import
and to keep the CLI working where gateway is not importable.
Adds tests/hermes_cli/test_config_set_platforms_redirect.py covering the
redirect, connection-key non-redirect, and the no-stray-top-level-platforms
case.
- test_shutdown_watchdog: the non-POSIX arm now asserts the witness ARMS
over TCP (port published, no AF_UNIX call, no warning, no socket node)
instead of pinning the old witness-absent fail-safe.
- test_update_wedged_gateway: TestLoopTickTcpWitness exercises the
consumer probe against a real loopback listener — stale file + answering
witness stays ALIVE (#90502 shape), stale + silent is WEDGED, fresh +
silent is UNKNOWN, garbage port never counts as armed. Runs on the
Linux lane so the TCP path is not Windows-CI-only.
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.
Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.
Addresses #100313
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.
Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.
Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.
On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.
Fixes#100436
Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
* fix(linux): install Lanczos-resized panel icons, not a PNG in scalable
Cinnamon's panel is ~24px. v2026.8.31 dropped the 1024px asset into
hicolor/scalable (SVG-only), so the Mint panel nearest-neighbor scaled
it into a mangled blob. Decode the PNG and write 24/32/48/256 rasters;
undecodable bytes still copy into one indexed dir. Drop the leftover
scalable file.
* test(linux): cover resized hicolor panel icons and stale scalable cleanup
Pin that a decodeable PNG lands as 24×24/256×256 rasters (not scalable),
a leftover scalable copy from v2026.8.31 is deleted, and truncated
PNGs still fall back to an indexed copy.
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).
Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.
Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes#97994.
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
provider label never carries the vendor token to the alias host (#28660);
route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
Regression for #58576: _profile_scope holds _SKILLS_PROFILE_LOCK across
the payload build, which can block up to 15s on a models.dev cache miss
and starve concurrent /api/config on the same lock. The test records
which scope the handler enters for a selected profile and asserts only
the config-only (contextvar) scope is used.
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.
Three surgical changes:
1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
The inlined watcher now routes the respawned gateway's stray
stdout/stderr to the same sidecar log gateway_windows._spawn_detached
uses (DEVNULL only as fallback), so a gateway killed moments after
respawn leaves a trace. Direct implementation of the 4th repro's
hardening suggestion (1).
2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
canonical _spawn_detached, so the respawned gateway's exit-diag /
lifecycle records show whether it escaped the parent Job Object — a
job-teardown kill is no longer indistinguishable from any other silent
death.
3. Post-update resume verifies liveness before vouching
(hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
runs the same provisional-hit + 2s-confirmation liveness poll every
other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
with all_profiles= for the fleet) before printing ✓, writes the #91675
start attestation for the verified PIDs, and fails the resume with a
"restart could not be verified" warning + recovery hint when no stable
gateway appears. Suggestion (2) of the 4th repro; closes the last
silent-success hole in the family (#84185 fixed the cold-start leg,
#91675 the direct-start leg; this is the relaunch leg).
Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.
Fixes the Bug-1 relaunch-trust leg of #48820.
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).
Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:
- _prune_replaced_custom_model_config_credentials skipped only the
preferred key, so a keyed provider's own legacy-named pool
(custom:b.ai) was false-pruned of its current model_config credential
when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
key, so a legacy-named pool stopped being seeded from model.api_key.
Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.
Follow-up to #100413.
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.
Co-authored-by: Cursor <cursoragent@cursor.com>
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.
Two layers:
1. _wait_for_gateway_ready now treats the first hit as provisional: the
gateway must stay visible through a 2s confirmation window
(_confirm_gateway_stable) before it is reported ready; a death during
confirmation resumes polling until the deadline. Failure output is an
honest ✗ with the Job Object explanation and the schtasks /Run
recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
state/gateway.start-attestation.json with the vouched-for PIDs. The
next `gateway start`/`gateway status` invocation checks it — if the
attested PIDs are gone with no clean-exit record in the lifecycle
ledger, the CLI reports (once) that the previous ✓ was false and
prints the schtasks recovery hint. `gateway stop` and a clean
lifecycle-ledger exit clear the marker silently.
Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.
Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.
Fixes#91675
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.
Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.
6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.
Fixes#64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.
Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.
Fixes#99839
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).
Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).
Part of the #90837 sidecar-unlink audit (wave 6).
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.
No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:
gateway waits on all in-flight work units (#77184 don't-amputate)
-> cron agent session waits on the \hermes update\ process to exit
-> \hermes update\ waits on the gateway to exit [back to A]
The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).
Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.
Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.
Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
(the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).
Fixes#100179
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).
- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
scheduled_at, dispatched_at, lateness_seconds, and kind
(on_time / late / catch_up, classified against the ticker tolerance and
the schedule's catch-up grace window). Manual triggers and one-shots are
not stamped (no scheduled instant to be late against / retired beyond
grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
catch-up, in both the built-in ticker and external provider paths.
CLI surface only — no new tools, no policy engine.
Addresses the visibility half of #99879.
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.
Fixes#97343
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.
Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.
Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.
Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.
Fixes#87694
Salvaged from #87745 with expiry mitigation added.
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.