On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
The live-checkout git mutation guard blocked history-rewriting git ops
(checkout, reset --hard, rebase, cherry-pick, ...) in the running source
checkout and its worktrees on every platform. The hazard it protects
against is only real on Windows, where NTFS locks loaded module files and
an in-place rewrite can corrupt the running process. On POSIX, open file
handles pin the old inodes, so a checkout swap under a running process is
safe, and the guard mostly taxed normal dev/salvage workflows with clone
workarounds.
- tools/self_repo_guard.py: add guard_active() -> os.name == "nt"
- tools/terminal_tool.py: consult guard_active() before running the
detector; detector logic and block message unchanged for Windows
- tests: wiring tests force the guard on; new tests cover the POSIX
pass-through and the platform predicate
/simplify-code residual. The note hard-coded "'commits' and 'dirty' are
UNKNOWN", but the two probes fail independently: a bad base_commit fails
rev-list while `git status` still succeeds, so `dirty` is a REAL measurement
being reported as unknown. Safety was never affected (the worktree is preserved
either way), but telling the parent a measured value is untrustworthy is its own
kind of misreport — and it would push a human toward re-inspecting something
already proven.
`mark_worktree_payload_unproven()` now takes an `unmeasured` argument, and
finalize tracks which probe actually failed. The raising path still disclaims
both, because which probe raised is unknowable there.
Validation: 22/22 tests/tools/test_subagent_worktree.py; ruff + ty clean. New
guard mutation-checked (hard-coding "commits/dirty" back fails it).
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.
- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
`unproven_worktree_payload()`. Both producers of this schema now call them,
so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
hand-building the dict (-16 lines). The re-import is guarded: the outer
`except` can be entered because the `from tools import subagent_worktree`
itself failed, in which case the name is unbound -- an inline fallback keeps
the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
compares its key set against live `finalize_subagent_worktree()` output. Same
contract, no source reading, refactor-proof, and it actually executes the
code.
Also folded from the same review:
- Fail-closed on an unmeasurable commit count. With no `base_commit` the
rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
tree still reached `git worktree remove --force` + `git branch -D` -- the
exact bug class #88113 is about, on a public function that takes a
caller-supplied dict. Now returns un-inspected instead, with a test driving a
real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
names WHICH probe failed, its exit code, and a bounded git stderr tail, so
the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
removed the duplicated index-corruption block in favor of the existing
`_break_git_index()` helper.
Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.
- The distinguishability test asserted the failure payload was byte-identical
to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
That freezes the AMBIGUITY as a required property: emitting `commits: None`
for "unknown" would improve exactly what #88113 is about and fail the test.
Now asserts what the parent actually depends on -- both keep the worktree,
and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
happy path, forbidding an always-present-but-False flag (a legitimately
better JSON contract: stable key set for serializers). Now
`assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
was not even a cross-producer contract: delegate_tool's note said "state
unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
assert the note names the worktree AND branch -- the actionable part for a
human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
before any git call would keep it green while proving nothing). Now checks
`call_count` and mirrors the branch-survival + note-names-path legs its
sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
parent reads, but delegate_tool's fallback -- the second producer -- was
verified only by reading. It now AST-parses the real fallback dict literal
and compares against live `finalize_subagent_worktree()` output, so the two
producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
raising, handled in delegate_tool), and the module docstring listed
`inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
`_break_git_index()` beside the file's other module-level helpers.
Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
The preserved worktree is invisible to the only consumer that can act on it.
Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":
inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
inspected OK, child produced nothing -> {commits: 0, dirty: False, pruned: False}
The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.
Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
naming the worktree/branch, warns, and returns the payload. Both unproven
exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
non-numeric rev-list stdout) produced the same unproven payload but logged at
DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
schema missing commits/dirty/pruned. It now emits the same flagged shape, and
logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
by restoring the unconditional prune and reintroducing this P1.
Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.
Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
work intact on disk; proven-clean still prunes (pruned=true).
finalize_subagent_worktree() treated a non-zero exit from its rev-list
or status probes as proof of the payload defaults (commits=0, clean),
then pruned on them: git worktree remove --force plus branch -D
permanently deleted a child's uncommitted work whenever git could not
inspect the tree (e.g. a corrupted index) (#88113).
A destructive cleanup now requires affirmative proof of zero commits
plus a clean tree. Any non-zero inspection result keeps the worktree
and branch for manual review, with a warning naming both.
The rotation path flushes its un-persisted transcript to the parent (#47202)
and only then calls publish_compression_child. The abort handler rolls back
the in-memory transcript and keeps agent.session_id on the parent - its own
comment says "keep the parent live and discard the stale compacted snapshot" -
but the rows the flush just wrote are not part of what it discards. Every
failed rotation therefore leaves the parent transcript longer than it found
it, whatever the failure was.
That is survivable for a one-off failure and pathological for a sticky one.
A parent row carrying ended_at fails the publish on every attempt and nothing
in this path clears it, so each auto-compaction appends another copy of the
current turn to the transcript it was supposed to shrink. Worse, the growth
then satisfies conversation_compression's own len(durable_parent) >
len(messages) check, so the next attempt adopts the inflated snapshot as if it
were genuine concurrent activity and the in-memory transcript doubles too.
Check that one precondition before writing. It is a plain read of the row the
publish is about to read anyway, and it raises the publish's own message, so
split_status=aborted, failure_class=session_split_failed and the rollback path
are all unchanged; a live parent reaches the flush exactly as before.
Deliberately not extended to the compression lease, which is re-acquirable - a
transient miss there would abort a rotation that would otherwise have
committed. old_session_id moves above the flush so a failure raised from here
takes the same in-memory rollback as any other pre-publish failure.
Scope: this fixes the amplification for every abort cause. It does not fix
what marks a live session as ended in the first place (#88197 Bug 1), which
needs a maintainer decision on end-reason taxonomy and is tracked on the
issue; an affected session still aborts every attempt, it just stops making
itself larger while it does.
Refs #88197
/simplify-code finding: only one-shots carry a run_claim, yet the three
dispatch-failure paths called clear_run_claim unconditionally — each call
acquires _jobs_lock (blocking cross-process flock) and does a full
load_jobs read just to return False for any non-'once' job. The trigger
is exactly a failure storm (interpreter shutdown, EMFILE with N due
jobs): N serialized flock+file reads at the moment the process can least
afford I/O, all guaranteed no-ops for the majority job kind.
Gate at the call site on schedule.kind == 'once'; new mutation-checked
test proves recurring dispatch failures skip the claim I/O entirely.
9/9 tests green; ruff clean.
Follow-ups on the #87591 salvage:
- cron/scheduler.py: wrap the three clear_run_claim call sites in a
best-effort helper — clear_run_claim does load_jobs/save_jobs file I/O,
and on the interpreter-shutdown path (or with a corrupt store) it could
itself raise, defeating the skip-cleanly purpose of these early exits.
A claim that can't be cleared simply expires at the TTL, as before.
- tests/cron/test_oneshot_dispatch_failure_run_claim.py (new): 8 tests —
clear_run_claim unit contract (one-shot cleared / already-clear noop /
recurring never touched / unknown id), all three dispatch-failure paths
through a real tick() clear the claim, and a raising clear_run_claim
does not crash the tick. Mutation-verified: reverting the fix makes the
suite fail.
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
Follow-ups on the #87259 salvage:
- cron/scheduler.py: the ledger-terminal reconciliation now requires the
terminal execution row's claimed_at to be >= the in-memory claim's
registration time (_running_since). Without this, the latest terminal
row for a recurring job is usually the PREVIOUS run's outcome — a fresh
claim in the try_register_running_job -> create_execution window (or a
finished run whose worker finally block hasn't released yet) would be
force-released and the job double-dispatched. Unparseable/missing
claimed_at fails closed to the age-based bound.
- cron/scheduler.py: take the _running_job_ids snapshot for the ledger
query under _running_lock — list() over a set concurrently mutated by
try_register/release_running_job can raise RuntimeError.
- tests: existing reconciliation tests updated to the claimed_at contract;
two new race-guard tests (previous-run terminal row never releases a
fresh claim; missing claimed_at fails closed). Mutation-verified:
removing the ownership guard fails both.
The age-only stale-claim sweep (t_3778a491, already on main) force-releases
an in-memory _running_job_ids claim only once it is older than
max(2*interval, 30m). A leaked claim that is YOUNG (inside its allowance)
while the durable executions ledger already proves the last run ended stays
wedged: the job is returned as due every tick, _submit_with_guard short-
circuits on 'already running', and next_run_at keeps fast-forwarding with no
execution — the exact 2026-08-14 recurring-router incident (t_20e23f84),
which survived a gateway restart because the in-memory age bound alone could
not see a run the ledger had already finished.
sweep_stale_inflight now reconciles each in-flight claim against the durable
executions ledger (cron/executions.db): if the job's MOST RECENT execution
row is terminal (completed/failed/unknown), the run provably ended, so the
claim is stale by construction regardless of its in-memory age and is force-
released. This is a persisted-state recovery path: the ledger is written by
the worker that ran the job and read by ANY ticker process (including one
that started AFTER the leak), so a leaked claim is recoverable without
force-run/resume and without depending on which process holds it in memory.
A ledger-terminal release is authoritative — it does not write a synthetic
mark_job_run failure (the ledger already records the outcome).
Added TestLedgerTerminalReconciliation (4 tests): young+terminal -> released
(RED on main, GREEN here), no-ledger-row -> not released, running-row -> not
released, old+terminal -> released once without synthetic failure.
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.
- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
Follow-ups on the #87261 salvage:
- cron/jobs.py: the persisted-error re-arm now respects schedule legality.
Re-arming to `now` fired CRON jobs at times their expression excludes —
a weekday-only 9am job whose Friday run errored would fire on SATURDAY
(croniter measures a 24h cadence on Saturday, so 27h > cadence+grace and
the guard tripped). Cron jobs re-arm to compute_next_run(schedule, now)
— the next LEGAL occurrence — and only when that actually moves
next_run_at earlier; interval jobs (the 2026-08-14 incident class) keep
the immediate now re-arm, which is always legal for intervals.
- cron/jobs.py: cache _schedule_cadence_seconds' croniter measurement per
expr (mirrors scheduler.py's _cron_interval_cache) — it runs inside
_jobs_lock on every tick for every stale-errored job.
- tests/cron/test_persisted_error_rearm_legality.py (new): weekday job
errored Friday re-arms to Monday (not Saturday), correctly-parked cron
value untouched, interval job still due immediately.
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.
cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.
Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.
A blackholed IPv6 path to api.telegram.org never errors, so
_await_with_thread_deadline never fires and connect hangs at
"attempt 1/8". Known A-record IPs connect over IPv4 immediately.
DoH timeout now fail-opens to the seed IPv4 list instead of the
hostname. Hostname stays last for IPv6-only hosts.
Closes#87015
Follow-up to the salvaged #70734 fix:
- test_sanitize_dedup_drops_tool_calls_key_when_all_removed encoded the old
global-uniqueness assumption (its second assistant call reused the id AFTER
the first call was answered, which is now a legitimate new call). The
replayed call now precedes the result, making it a true duplicate of a
still-outstanding call, preserving the intended empty-tool_calls key-drop
coverage from #64335.
- New test: Hermes' own deterministic local counter ids repeating across
turns (the #76632 scenario) survive sanitization.
- New test: the 50-step constant-id field repro from #70724 (Kimi K3 /
llama.cpp) — stock main kept 1/50 tool results, now 50/50.
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).
Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).
Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.
repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.
Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
CI git consolidates during incremental pack creation differently per
build (4 packs from 6 attempts on ubuntu-latest, 6 locally, 3 on the
previous run) — even pack-objects counts drift with auto-maintenance.
The fixture now only guarantees strictly-more-packs-than-threshold and
the test asserts consolidation strictly decreases the count.
Incremental 'git repack' consolidates small packs on newer git builds
(CI produced 3 packs from 6 commits), making the sprawl fixture count
nondeterministic. pack-objects with an explicit sha per commit creates
exactly one pack each on every git version.
Two gaps from the Aug 2026 'hermes -w timed out after 30s' incident:
1. Atomic failure cleanup: a timed-out/failed `git worktree add` left a
partially-materialized directory plus a LOCKED admin entry under
.git/worktrees/ (lock pid = the live hermes process that timed out),
which the startup pruner's dead-pid unlock never reaps — retries of
the same name fail forever. _cleanup_failed_worktree_add sweeps dir,
admin entry, and orphaned branch on every failure path (timeout,
nonzero exit, remote-base retry).
2. Pack maintenance: nothing consolidated the object store; on a
multi-agent box packs sprawl (39 packs / 638MB at the incident) and
every object lookup scans all pack indexes until worktree creation
blows its timeout. _maintain_pack_health repacks (niced, background,
fail-soft) when *.pack count reaches 15, wired into the existing
startup maintenance thread on both the CLI (-w) and TUI paths.
gc --auto doesn't cover this: its threshold is 50 packs.
Both sabotage-verified; full repack on the incident box: 39 packs ->
2, 638MB -> 287MB, worktree add 30s-timeout -> 0.5s.
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):
- Protocol-era negotiation (_negotiate_session): per-server `protocol`
config key — auto (default, handshake-first with server/discover
fallback on -32022/-32601), stateless (discover-first), legacy
(handshake only). Auto is handshake-first deliberately: zero extra
round-trips and zero behavior change for the entire existing server
fleet, while 2026-07-28-only servers now connect via the fallback.
All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
route through the one choke point, so the CLI/desktop probe path
inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
during discovery and bound to the lazy-startup schema cache — TTL'd
entries expire and force a live re-probe; hint-less (pre-2026)
servers keep the never-expires behavior. Pagination continuation now
speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
(config-overridable), with a fallback for 1.x-era metadata models.
(RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
native to SDK 2.0's OAuthClientProvider — verified, no client-side
gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
Sampling feature as upstream-deprecated (12-month window) — kept
fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):
- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
(health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
are allowed in the generic loop while native anthropic stays skipped
E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
Matrix m.audio/m.file/m.video events populate content.body with the
uploaded filename when the sender adds no caption. The adapter already
blanks that for m.image (PR #16821, issue #13482) but not for audio,
file, or video msgtypes, so the filename survives into event.text and
is appended after the transcript, where the model reads it as the
user message rather than as transport noise.
Extend the existing adapter-level blanking: add
_looks_like_matrix_media_filename() with the same conservative
heuristic (single token, no whitespace, no path separators, known
media suffix or mimetypes audio/video match) and apply it at the
media message handler for m.audio, m.file, and m.video.
Salvage of #87968 by @AiwendilInTheWoods — reworked from shared
gateway path to adapter-level fix for consistency with the existing
m.image blanking.
tests/test_lazy_secrets_dispatch.py::TestUpdatePathE2E ran the real
`hermes update --check` bare, so the child's `git fetch origin main`
hit github.com on every CI run. During the 2026-08-17 GitHub incident
the fetch stalled past the 30s subprocess timeout and both update tests
went red on main for hours with zero code change (slices seen on runs
32003200300 and 32011520139).
These tests assert the lazy-crypto / no-self-lock dispatch invariants,
not update connectivity. Rewrite all git remote URLs in the child to an
unreachable file:// path via GIT_CONFIG_* env overrides: the update path
still exercises its full parser/dispatch/fetch code, but the fetch now
fails in milliseconds, deterministically, offline. Exit code 1 was
already accepted by the assertions.
Before/after under a blackhole proxy simulating the outage:
OLD: TimeoutExpired after 20s (reproduces the CI failure)
NEW: completes in 0.2s, exit=1
External review (Fable) caught a real false-positive widening in the
original commit: the new argv[1] script-name check reused the loose
`script_name == "hermes" or script_name.startswith("hermes")` pattern
(copy-pasted from the exe_name check above it), but argv[1] can be ANY
user-invoked python script path when argv[0] is a bare interpreter --
unlike a directly-resolved executable name, where a false match on the
substring is rare. A user's own script named e.g. "hermes-notes.py" or
"hermes-unrelated-tool" run via `python3 <script>` would be misidentified
as the console-script shim and become killable by profile delete.
Match against the actual known console-script entry points instead
(pyproject.toml [project.scripts]: hermes, hermes-agent, hermes-acp),
stripping the script's extension before comparing.
Added 2 regression tests: one confirms the false-positive case is now
rejected (fails against the pre-fix loose-match code, confirmed via a
scripted revert), the other confirms the other two real entry points
(hermes-agent, hermes-acp) still match via the shebang-exec path.
Tests: tests/hermes_cli/test_profiles.py -- 158 passed (156 previous + 2
new).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two independent bugs let a deleted profile reappear / leave orphaned
resources on next launch:
1. hermes_cli/profiles.py's backend-process scanner required argv[0] to
resolve to an executable literally named "hermes". Electron's
pool-backend spawn resolves the hermes console-script shim's path and
execs it via the interpreter directly (python3 /path/to/hermes ...), so
argv[0] reports as "python3" and the scanner never matched the running
backend -- delete removed the profile's files but left its live backend
process running (still bound to a port via uvicorn), which
accumulates across repeated delete/recreate cycles.
2. The desktop sidebar's ProfileRail only refreshed its cached profile
list once, on mount, so a delete/create/rename from another surface
(another window, or the CLI) left a stale ghost entry until something
unrelated triggered a refetch. Note: a delete via this window's own
Manage-Profiles view already refreshes the shared $profiles atom
ProfileRail subscribes to (confirmed by reading refreshProfiles() and
handleConfirmDelete()) -- this fix only covers the cross-window/cross-
process staleness gap, not a duplicate of the already-merged
#57329's Manage-Profiles rail-refresh work.
Fix 1: recognize a python-interpreter argv[0] exec'ing a hermes-named
console-script shim via argv[1]. Fix 2: refresh the profile list on window
focus/visibilitychange, matching the existing pattern used elsewhere in
the sidebar (sidebar/index.tsx, use-background-sync.ts, star-map.tsx,
use-gateway-boot.ts all use the same focus+visibilitychange pattern).
## Related work already on main
PR #57329 (merged) fixed the *headline* symptom from issue #52279
(deleted profile respawns) via a different, non-overlapping mechanism:
routing profile-delete through the primary backend instead of spawning a
fresh pool backend, plus a separate recreation guard in
ensure_hermes_home() (#49435, merged) that makes a backend spawned into a
deleted profile's directory raise FileNotFoundError instead of silently
recreating it.
This PR is NOT a duplicate of that fix. Verified: even with both of those
merged, a backend process that survives because of gap #1 above still
holds a bound port via uvicorn -- it just can no longer resurrect the
profile directory. That's real resource-hygiene, not a symptom already
covered. Gap #2 touches a different file/component (ProfileRail /
profile-switcher.tsx) than #57329's rail-refresh half (which touched the
Manage-Profiles view's own $profiles.ts / index.tsx) and covers a
distinct staleness path (cross-window/cross-process, not same-window
delete-then-refresh).
Tests: tests/hermes_cli/test_profiles.py -- 156 passed (existing +
regression coverage for the argv[0] python-interpreter detection case).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Same behavior, same coverage, less boilerplate (test file 1691 -> 1512
lines; PR diff unchanged in semantics).
Production (mechanical only):
- Extract _strip_hermes_owned_pythonpath_and_runtime_markers(): the three
builders (_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env)
ran the identical strip-then-pop-markers sequence in the same order
(ordering is load-bearing for VIRTUAL_ENV validation); the helper makes
that explicit once instead of three times.
Tests:
- Non-owned preservation: 11 single-shape tests -> one parametrized matrix
(user/Nix/other-version/python2.7/pythonX.Y-contained/raw spelling/empty
component/empty PYTHONPATH) + one runtime-shaped matrix (other-version SP,
venv-SP descendant, repo direct child, repo deep child).
- Owned stripping: venv SP, repo root (independent parents[2] computation),
duplicates, all-owned key removal, mixed ordering -> one matrix.
- Builder integration: _make_run_env/_sanitize_subprocess_env/
hermes_subprocess_env venv-SP stripping -> one parametrized test;
same for the four PYTHONHOME builders (incl. build_subprocess_env).
- Junction: same-named non-owned negative control now covers both the
configured-root location and an unrelated location; shared
_physical_repo_root helper; profile resolution matrix (root->named,
profile-shaped->named no nesting, profile-shaped->default, custom root).
- Every independent proof preserved: home-level junction, repo-level
junction, profile interaction, negative identity control, uv-base lexical
VIRTUAL_ENV, validated/unrelated VIRTUAL_ENV, no-scrub escape hatch,
#84500 same-env/external-env composition, PYTHONHOME removal, real
Windows-only semantics, POSIX fail-closed backslash paths.
Second real-world topology reported and confirmed on native Windows 11:
the repository itself is a cross-drive junction (D:\hermes\hermes-agent ->
C:\...\hermes-agent) under a real HERMES_HOME directory. The editable
import spelling resolves to the physical location, so _hermes_repo_root is
physical while the launcher writes the lexical spelling into PYTHONPATH.
The home-relative mapping cannot express a cross-drive link (commonpath
raises on different drives), so the lexical repo root survives stripping;
and with the repo alias missing, a lexical VIRTUAL_ENV
(D:\hermes\hermes-agent\venv) also fails _validated_runtime_venv, so the
venv site-packages survives too (uv-base gateway: both entries survive).
Fix: after the existing home/profile-root mapping, try the single
deterministic candidate <lexical root>/<repo dirname> for every trusted
home candidate (configured home, plus the profile root when the configured
home is a profile path) and accept it only when strict resolve proves it is
the exact physical repo root (fail-closed: missing paths, real directories
that are not the known repo, and unrelated spellings are never aliased).
This also re-enables the VIRTUAL_ENV validation for lexical venv spellings,
so uv-base gateway site-packages cleanup follows the repo alias.
Tests: repo-level junction positive + negative control (same-named real
directory preserved), profile-home + repo-level junction combination,
lexical VIRTUAL_ENV validation after recovery (root + site-packages
stripped, user entries kept), and a no-provenance lookalike preserved.
The execute_code composition test now compares composed paths with
os.path.normcase so a Windows case-only spelling difference (resolve() vs
abspath() casing) can never fail the composition contract.
The junction fix made resolve_profile_env preserve the configured
HERMES_HOME spelling as the launch root. Cover the four pre-existing
resolution invariants so the spelling-preservation never regresses them:
- root env + named profile -> <root>/profiles/<name>
- profile-shaped env + named profile -> <root>/profiles/<name> (no nesting)
- profile-shaped env + default -> <root>
- custom root env never falls back to the platform default
Plus existence/validation semantics (missing named profile still raises
FileNotFoundError) and the unset-env fallback contract.
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling. tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).
Two narrow changes, no heuristics, no new env vars:
- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
the configured spelling IS the launch root (junction-transparent,
physically identical dirs); keep it instead of re-deriving the native
default. This is the same producer contract _preserve_hermes_home_path
already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
configured home is a profile home (<root>/profiles/<name>), also derive
the root spelling lexically (parent of the profiles component, same
rule get_default_hermes_root uses) and run the exact-ownership mapping
against it, so the launcher's lexical root is recovered after re-home
without ever matching arbitrary descendants of HERMES_HOME.
Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
The PYTHONPATH/PATH sanitization suite was written POSIX-centric and
failed on real Windows 11 (reproduced natively: 4 failures before this
change). Fix the tests to express the true per-platform contract:
- test_other_major_version_site_packages_preserved /
test_make_run_env_injects_hermes_bin_dir: build inputs with
os.pathsep instead of hardcoded ':'.
- test_make_run_env_appends_homebrew_on_minimal_path: split on
os.pathsep, neutralise Git Bash dir prepending, and assert the
documented Windows passthrough (_append_missing_sane_path_entries is
a no-op off POSIX) instead of the Homebrew append.
- test_make_run_env_real_launchd_path_gains_homebrew: mark
macos_only per repo OS-marker policy (the regression is the macOS
launchd PATH; the merge is a passthrough on Windows).
- test_configured_home_alias_matches_launcher_output: create the
configured-home link via a helper that falls back to an unprivileged
directory junction (cmd /c mklink /J) when symlink creation raises
WinError 1314, and skips with a clear reason if no mechanism exists.
Also correct a stale comment in execute_code: the child is not always
the same Python as Hermes (project mode can select an external venv),
so the strip is about compatibility, not redundancy.
Integration test for the #84500 + #82581 intersection: seeds a
contaminated inherited PYTHONPATH (Hermes repo root + Hermes venv
site-packages + user entries) through os.environ and drives execute_code
to Popen. Asserts the staging tmpdir stays first, inherited Hermes
site-packages never survive, the repo root is re-added exactly once for
a same-env child (proving the inherited copy was stripped) and stays
absent for an external-env child, and user entries survive in order.
Adversarial review of the previous two commits (and #78917 itself)
found three ownership-boundary issues; this commit addresses them:
1. Repo direct-child over-strip (Finding A)
No launcher injects <repo>/tools or another direct child as an
independent PYTHONPATH entry - audited all four producers (Electron
electron-main.mjs, gateway/run.py::_ensure_windows_gateway_venv_imports,
cron/scheduler.py::_windows_cron_python_invocation,
tui_gateway/host_supervisor.py). The depth<=1 rule deleted user paths
that merely live under the repo directory; only the EXACT repo root is
now stripped.
2. Windows junction/symlink alias (Finding B)
The gateway launcher renders Hermes-owned paths under the configured
HERMES_HOME spelling (gateway_windows.py::_preserve_hermes_home_path),
which may be a junction to another drive, so it differs lexically from
the resolved repo root. _hermes_repo_root_aliases now carries both the
resolved and unresolved spellings; both are recognized as Hermes-owned.
3. Stale abstraction rename (Phase 4)
_strip_mismatched_site_packages -> _strip_hermes_owned_pythonpath:
the cross-version heuristic is gone, so the old name misdescribes the
behavior (ownership-based, not version-based).
Tests: direct-child now preserved; junction alias stripped (lexical pair
monkeypatched); Windows-only real-semantics test added (POSIX test remains
a safety test); mixed-ordering, duplicate-Hermes, and no-scrub PYTHONHOME
contract tests added. Full file: 52 passed / 16 failed (identical failure
set to base, all isolation-venv environment issues).
The gateway runs inside its own venv; if its PYTHONHOME leaks into
subprocesses (terminal commands, cron no_agent scripts, TTS providers),
any child interpreter redirects its stdlib search to the Hermes venv and
crashes with version-mismatch errors before importing anything.
PYTHONHOME is now part of _ACTIVE_VENV_MARKER_VARS so all env builders
(_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env, and
build_subprocess_env used by cron) drop it, consistent with Hermes'
existing PYTHONHOME handling in managed_uv.py and sqlite_runtime.py.
execute_code already scrubbed it via _SAFE_ENV_PREFIXES.
Tests cover all four builders plus the marker constant.
Remove the cross-version heuristic from _strip_mismatched_site_packages:
the subprocess env builder cannot know which Python version a child will
run, so judging user PYTHONPATH entries against the backend interpreter's
version deletes legitimate paths meant for a different child Python
(e.g. /custom/lib/python3.13/site-packages while Hermes runs 3.11).
Also fix over-strip: entries merely containing a pythonX.Y path component
(e.g. /opt/tools/python3.13/bin) were stripped even though they are not
site-packages. Hermes-owned entries (repo root, own venv site-packages)
are now identified by path ownership, not by version.
Regression tests cover both cases; user paths with any pythonX.Y
component are preserved.
`ElicitationHandler` read `params.requested_schema`, but on the pinned
`mcp==1.28.1` the model field is spelled `requestedSchema`. The getattr
always missed and returned its `{}` default, so
`_format_elicitation_schema_summary` took its no-properties branch and the
approval prompt collapsed to the generic
Approval requested by MCP server '<name>'.
for every request. The field names, types, and descriptions the summary
exists to surface never reached the user, so an elicitation asking for a
card number rendered identically to one asking for a nickname — consent
without the substance of what was being consented to.
Read both spellings rather than just correcting to the 1.x name: mcp 2.0
renames this field to `requested_schema` (it renamed every model field to
snake_case and kept camelCase only as a serialization alias, which
pydantic does not expose to attribute access), so a dual read is correct
on either SDK generation and does not go wrong again on the next bump.
Verified against real 1.28.1 and 2.0.0 installs.
Every existing test in tests/tools/test_mcp_elicitation.py builds a
duck-typed `SimpleNamespace` stand-in, which carries whatever field name
the test wrote and therefore cannot detect a mismatch with the real model.
Add one test that constructs the actual `ElicitRequestFormParams` and
asserts the requested field name reaches the consent description; it fails
on the unfixed tree. The cheap stand-ins are left alone elsewhere.
Found while porting the tree to the mcp 2.x SDK in #76736, but independent
of it: this reproduces on the current pin with no other changes, #76736
does not touch this line, and the two branches merge cleanly in either
order.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses three gaps found in review of the memory-guard PR:
P1a — standalone daemon was the one uncapped entry point. run_daemon()
now resolves kanban.max_in_progress every tick (explicit config wins,
else the memory-derived default) exactly like the gateway dispatcher and
`hermes kanban dispatch`. New shared parser configured_max_in_progress()
so all three entry points agree on what "explicitly configured" means.
P1b — max_in_progress was enforced per board while the gateway ticks
every active board, multiplying the host budget by the number of boards
(2 boards x cap 2 = 4 workers on a host sized for 2). The cap is now
host-level: _dispatch_once_locked() adds count_running_tasks_other_boards()
to the running count before deriving the tick's spawn budget. Enforced in
the shared locked path, so gateway, CLI, and daemon all inherit it.
max_spawn deliberately keeps its historical per-board semantics.
Fails open per board so one corrupt board can't brick dispatch on the rest.
P2 — the ready loop consumed the entire shared spawn budget before the
review loop ran, so a sustained ready backlog starved autonomous reviews
indefinitely. When spawnable review work exists (assigned + real profile,
mirroring the review loop's own gate) and the tick has budget, one slot
is held back from the ready lane. Reservation is per-tick and
self-releasing; the review lane still spends from the shared budget —
it gains fairness, not extra capacity.
11 new tests in tests/hermes_cli/test_kanban_host_cap.py. Existing
kanban suites: 278 passed (15 failures pre-existing, identical on clean
main baseline). ruff clean.
Two production incidents (OOF-77 "larrikin-lollies", OOF-30
"synclare-task-manager") followed the same shape: no
kanban.max_in_progress configured, a busy board, and a 1 GiB hosted VM.
The dispatcher fanned out 26-31 concurrent workers, the host went into
swap-thrash/OOM, and the whole machine — dashboard included — became
unreachable. NAS restart loops then masked the problem: each restart
"recovered" briefly before the kanban dispatcher immediately respawned
unbounded workers.
Building on the cherry-picked max_in_progress-across-both-lanes fix
(PR #28695, credit @Dusk1e), this adds two complementary safeguards to
hermes_cli/kanban_db.py:
1. Memory-DERIVED default concurrency cap. When kanban.max_in_progress
is unset, resolve_max_in_progress() derives a default of
clamp(MemTotal / 512 MiB, 2, 8) — e.g. 2 workers on a 1 GiB VM,
8 on 4 GiB+. Explicit config always wins in either direction. On
hosts where total memory can't be read (macOS/Windows dev machines),
the default stays None (no cap — unchanged behaviour). Wired into
both dispatch entry points (gateway/kanban_watchers.py and
hermes kanban dispatch) so behaviour matches regardless of path.
2. Live memory-PRESSURE guard inside dispatch_once. A static cap can't
see the host's actual memory state (other tenants, bloated
long-lived workers). The dispatcher now samples system memory each
tick via gateway.lifecycle_ledger.sample_memory() and classifies it
with gateway.memory_status.classify_pressure() (same thresholds as
the dashboard memory banner and OOM-suspicion heuristics from
NS-608/NS-656): critical -> spawn nothing this tick; elevated ->
at most one new worker; unknown -> no restriction (fail-open).
Reclaim/promotion bookkeeping still runs under pressure, and
deferred tasks stay queued — nothing is dropped. Restriction is
surfaced on DispatchResult.memory_pressure and logged.
Tests: tests/hermes_cli/test_kanban_memory_guard.py (14 tests) covers
the derived cap (floor/ceiling/fail-open/explicit-config-wins), the
pressure classifier, and dispatch behaviour under critical/elevated/
unknown pressure including defer-not-drop and bookkeeping-still-runs.
An autouse fixture in tests/conftest.py pins the memory sample to
"no data" suite-wide so existing dispatch tests don't depend on the
CI runner's live memory state (opt-out marker: real_memory_guard).
clean_registry cleared the connection registry only on teardown, so it
protected the tests that ran after it but not the test holding it. A leak
from earlier in the session — a failed test that never reached its
close(), or any test that does not take this fixture — would leave a stale
entry behind, and read_header_bytes_preopen would then refuse for that
stale reason instead of the one under test. The refusal assertions would
still pass, but for the wrong reason, which is the failure mode a
regression test can least afford.
Clearing on entry as well makes the fixture independent of what ran
before it, and the teardown clear now runs under try/finally so a failing
test cannot skip it.
read_header_bytes_preopen answers None for a live connection, a missing
file and an unreadable file alike, so the error string doctor prints is
now chosen rather than inherited from the OSError. These cases pin that
choice: the missing file keeps its errno text, and the chmod-000 file is
still reported as a permission problem rather than collapsing into the
generic message — the behaviour the raw open() gave before.
test_reason_does_not_open_the_file is the load-bearing one. It patches
builtins.open to raise and asserts _unreadable_reason still answers,
which fixes the constraint that makes the helper safe to call on a
database path at all: stat() and access() read metadata and take no file
descriptor, so no close() of ours can cancel the file's advisory locks. A
future edit that reached for open() here to get a better message would
reintroduce the original bug on the error path, and this test fails
loudly if it does.
The root check is written as hasattr(os, "geteuid") and os.geteuid() == 0
rather than the bare call the surrounding tests use. skipif conditions are
evaluated at collection time and os.geteuid is POSIX-only, so the bare
form raises AttributeError and takes the whole module down on Windows.
The pre-existing occurrences are left alone — #81926 and #84073 are
already open against exactly those lines, and this only avoids adding a
third instance of the same defect.
Locks the invariant the probe now honours: while this process holds a
registered connection to a database, _read_journal_mode reports it as
unreadable instead of taking a descriptor whose close() would cancel that
connection's POSIX advisory locks.
Against the previous implementation the four regression cases fail with
`assert 'wal' is None` — it read the header straight out of a live
database — and pass once the read is routed through
read_header_bytes_preopen. Coverage is both the registry API
(track_connection) and connect_tracked, the path SessionDB actually takes,
plus the _report_database_journal_modes output so the degraded row is
asserted end to end.
Two cases deliberately hold in both directions and are guards rather than
probes:
- an untracked sqlite3.connect holding BEGIN EXCLUSIVE must NOT block the
read. Only connections this process registered can be cancelled by a
close() we make; another process's locks are irrelevant. Without this,
a later "just refuse whenever the file looks busy" change would silently
turn every doctor row into "could not be read".
- the refusal creates no new -wal/-shm sidecars, which is the property the
function's docstring promises and the reason it byte-probes rather than
opening a connection in the first place.
/simplify-code review found _notify_single_query_session_finalize was
missing the _handed_off_session_ids guard that _should_emit_cleanup_session_finalize
and _emit_interrupted_session_end already had. One-shot CLI queries that
somehow handed off would still finalize the session via this path.
Added guard + test.