/simplify-code findings on the salvage stack:
- the classify+reclaim+counter block was pasted verbatim into both ticker
loops (_start and _start_multiplex) along with duplicated function-local
imports — extracted _note_tick_failure() next to _backoff_wait_seconds
so both loops share one implementation.
- hermes_cli/cron.py's EMFILE hint reimplemented the text half of
_is_fd_exhaustion with a case-SENSITIVE variation (drift risk) — split
_is_fd_exhaustion_text() out and use it from both.
11 EMFILE tests + 54 provider/ticker tests green; ruff clean.
Follow-ups on the #87796 salvage:
- cron/scheduler.py: drop the _reclaim_fds_best_effort call at tick()'s
lock-failure raise site — the ticker loop's except handler already runs
reclamation once per failed tick, so the raise-site call doubled the
gc.collect() pause on every EMFILE failure.
- cron/scheduler_provider.py: extract the exponential-backoff math
duplicated verbatim in start() and _start_multiplex() into a module-level
_backoff_wait_seconds() helper.
- hermes_cli/cron.py: `hermes cron tick` now reports a propagated OSError
cleanly (exit 1) instead of dumping a traceback — tick() raising on real
lock-acquisition failures is new behavior from this fix.
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.
- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
Two gaps from the Aug 2026 'hermes -w timed out after 30s' incident:
1. Atomic failure cleanup: a timed-out/failed `git worktree add` left a
partially-materialized directory plus a LOCKED admin entry under
.git/worktrees/ (lock pid = the live hermes process that timed out),
which the startup pruner's dead-pid unlock never reaps — retries of
the same name fail forever. _cleanup_failed_worktree_add sweeps dir,
admin entry, and orphaned branch on every failure path (timeout,
nonzero exit, remote-base retry).
2. Pack maintenance: nothing consolidated the object store; on a
multi-agent box packs sprawl (39 packs / 638MB at the incident) and
every object lookup scans all pack indexes until worktree creation
blows its timeout. _maintain_pack_health repacks (niced, background,
fail-soft) when *.pack count reaches 15, wired into the existing
startup maintenance thread on both the CLI (-w) and TUI paths.
gc --auto doesn't cover this: its threshold is 50 packs.
Both sabotage-verified; full repack on the incident box: 39 packs ->
2, 638MB -> 287MB, worktree add 30s-timeout -> 0.5s.
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):
- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
(health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
are allowed in the generic loop while native anthropic stays skipped
E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
External review (Fable) caught a real false-positive widening in the
original commit: the new argv[1] script-name check reused the loose
`script_name == "hermes" or script_name.startswith("hermes")` pattern
(copy-pasted from the exe_name check above it), but argv[1] can be ANY
user-invoked python script path when argv[0] is a bare interpreter --
unlike a directly-resolved executable name, where a false match on the
substring is rare. A user's own script named e.g. "hermes-notes.py" or
"hermes-unrelated-tool" run via `python3 <script>` would be misidentified
as the console-script shim and become killable by profile delete.
Match against the actual known console-script entry points instead
(pyproject.toml [project.scripts]: hermes, hermes-agent, hermes-acp),
stripping the script's extension before comparing.
Added 2 regression tests: one confirms the false-positive case is now
rejected (fails against the pre-fix loose-match code, confirmed via a
scripted revert), the other confirms the other two real entry points
(hermes-agent, hermes-acp) still match via the shebang-exec path.
Tests: tests/hermes_cli/test_profiles.py -- 158 passed (156 previous + 2
new).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two independent bugs let a deleted profile reappear / leave orphaned
resources on next launch:
1. hermes_cli/profiles.py's backend-process scanner required argv[0] to
resolve to an executable literally named "hermes". Electron's
pool-backend spawn resolves the hermes console-script shim's path and
execs it via the interpreter directly (python3 /path/to/hermes ...), so
argv[0] reports as "python3" and the scanner never matched the running
backend -- delete removed the profile's files but left its live backend
process running (still bound to a port via uvicorn), which
accumulates across repeated delete/recreate cycles.
2. The desktop sidebar's ProfileRail only refreshed its cached profile
list once, on mount, so a delete/create/rename from another surface
(another window, or the CLI) left a stale ghost entry until something
unrelated triggered a refetch. Note: a delete via this window's own
Manage-Profiles view already refreshes the shared $profiles atom
ProfileRail subscribes to (confirmed by reading refreshProfiles() and
handleConfirmDelete()) -- this fix only covers the cross-window/cross-
process staleness gap, not a duplicate of the already-merged
#57329's Manage-Profiles rail-refresh work.
Fix 1: recognize a python-interpreter argv[0] exec'ing a hermes-named
console-script shim via argv[1]. Fix 2: refresh the profile list on window
focus/visibilitychange, matching the existing pattern used elsewhere in
the sidebar (sidebar/index.tsx, use-background-sync.ts, star-map.tsx,
use-gateway-boot.ts all use the same focus+visibilitychange pattern).
## Related work already on main
PR #57329 (merged) fixed the *headline* symptom from issue #52279
(deleted profile respawns) via a different, non-overlapping mechanism:
routing profile-delete through the primary backend instead of spawning a
fresh pool backend, plus a separate recreation guard in
ensure_hermes_home() (#49435, merged) that makes a backend spawned into a
deleted profile's directory raise FileNotFoundError instead of silently
recreating it.
This PR is NOT a duplicate of that fix. Verified: even with both of those
merged, a backend process that survives because of gap #1 above still
holds a bound port via uvicorn -- it just can no longer resurrect the
profile directory. That's real resource-hygiene, not a symptom already
covered. Gap #2 touches a different file/component (ProfileRail /
profile-switcher.tsx) than #57329's rail-refresh half (which touched the
Manage-Profiles view's own $profiles.ts / index.tsx) and covers a
distinct staleness path (cross-window/cross-process, not same-window
delete-then-refresh).
Tests: tests/hermes_cli/test_profiles.py -- 156 passed (existing +
regression coverage for the argv[0] python-interpreter detection case).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling. tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).
Two narrow changes, no heuristics, no new env vars:
- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
the configured spelling IS the launch root (junction-transparent,
physically identical dirs); keep it instead of re-deriving the native
default. This is the same producer contract _preserve_hermes_home_path
already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
configured home is a profile home (<root>/profiles/<name>), also derive
the root spelling lexically (parent of the profiles component, same
rule get_default_hermes_root uses) and run the exact-ownership mapping
against it, so the launcher's lexical root is recovered after re-home
without ever matching arbitrary descendants of HERMES_HOME.
Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
Addresses three gaps found in review of the memory-guard PR:
P1a — standalone daemon was the one uncapped entry point. run_daemon()
now resolves kanban.max_in_progress every tick (explicit config wins,
else the memory-derived default) exactly like the gateway dispatcher and
`hermes kanban dispatch`. New shared parser configured_max_in_progress()
so all three entry points agree on what "explicitly configured" means.
P1b — max_in_progress was enforced per board while the gateway ticks
every active board, multiplying the host budget by the number of boards
(2 boards x cap 2 = 4 workers on a host sized for 2). The cap is now
host-level: _dispatch_once_locked() adds count_running_tasks_other_boards()
to the running count before deriving the tick's spawn budget. Enforced in
the shared locked path, so gateway, CLI, and daemon all inherit it.
max_spawn deliberately keeps its historical per-board semantics.
Fails open per board so one corrupt board can't brick dispatch on the rest.
P2 — the ready loop consumed the entire shared spawn budget before the
review loop ran, so a sustained ready backlog starved autonomous reviews
indefinitely. When spawnable review work exists (assigned + real profile,
mirroring the review loop's own gate) and the tick has budget, one slot
is held back from the ready lane. Reservation is per-tick and
self-releasing; the review lane still spends from the shared budget —
it gains fairness, not extra capacity.
11 new tests in tests/hermes_cli/test_kanban_host_cap.py. Existing
kanban suites: 278 passed (15 failures pre-existing, identical on clean
main baseline). ruff clean.
Two production incidents (OOF-77 "larrikin-lollies", OOF-30
"synclare-task-manager") followed the same shape: no
kanban.max_in_progress configured, a busy board, and a 1 GiB hosted VM.
The dispatcher fanned out 26-31 concurrent workers, the host went into
swap-thrash/OOM, and the whole machine — dashboard included — became
unreachable. NAS restart loops then masked the problem: each restart
"recovered" briefly before the kanban dispatcher immediately respawned
unbounded workers.
Building on the cherry-picked max_in_progress-across-both-lanes fix
(PR #28695, credit @Dusk1e), this adds two complementary safeguards to
hermes_cli/kanban_db.py:
1. Memory-DERIVED default concurrency cap. When kanban.max_in_progress
is unset, resolve_max_in_progress() derives a default of
clamp(MemTotal / 512 MiB, 2, 8) — e.g. 2 workers on a 1 GiB VM,
8 on 4 GiB+. Explicit config always wins in either direction. On
hosts where total memory can't be read (macOS/Windows dev machines),
the default stays None (no cap — unchanged behaviour). Wired into
both dispatch entry points (gateway/kanban_watchers.py and
hermes kanban dispatch) so behaviour matches regardless of path.
2. Live memory-PRESSURE guard inside dispatch_once. A static cap can't
see the host's actual memory state (other tenants, bloated
long-lived workers). The dispatcher now samples system memory each
tick via gateway.lifecycle_ledger.sample_memory() and classifies it
with gateway.memory_status.classify_pressure() (same thresholds as
the dashboard memory banner and OOM-suspicion heuristics from
NS-608/NS-656): critical -> spawn nothing this tick; elevated ->
at most one new worker; unknown -> no restriction (fail-open).
Reclaim/promotion bookkeeping still runs under pressure, and
deferred tasks stay queued — nothing is dropped. Restriction is
surfaced on DispatchResult.memory_pressure and logged.
Tests: tests/hermes_cli/test_kanban_memory_guard.py (14 tests) covers
the derived cap (floor/ceiling/fail-open/explicit-config-wins), the
pressure classifier, and dispatch behaviour under critical/elevated/
unknown pressure including defer-not-drop and bookkeeping-still-runs.
An autouse fixture in tests/conftest.py pins the memory sample to
"no data" suite-wide so existing dispatch tests don't depend on the
CI runner's live memory state (opt-out marker: real_memory_guard).
read_header_bytes_preopen returns None for every failure, so routing the
probe through it flattened "[Errno 2] No such file or directory: …" and
"[Errno 13] Permission denied: …" into one opaque "file could not be
read". doctor exists to name the problem, so that detail is worth keeping:
_report_database_journal_modes prints the string verbatim, and on a
vulnerable SQLite it is the only clue the user gets about why WAL exposure
could not be ruled out.
_unreadable_reason recovers it from metadata only. stat() reports the
missing file, the dangling symlink and the unsearchable parent directory;
os.access(..., R_OK) reports the unreadable file that stat() can still see.
Neither call takes a file descriptor, so neither can cancel the POSIX
advisory locks the previous commit was about — the invariant holds.
_read_journal_mode opened each Hermes database with a bare open(db_path,
"rb") to read header byte 18. The read itself is harmless; the close() is
not. Per sqlite.org/howtocorrupt.html, close() on *any* descriptor for a
file cancels every POSIX advisory lock this process holds on it — so the
close at the end of that with-block drops the locks a live connection is
holding, including the EXCLUSIVE lock a VACUUM holds while it rewrites the
whole file. Another process is then free to write into a file its writer
still believes it owns, which is the documented route to "database disk
image is malformed".
This is reachable. run_doctor is not only a standalone CLI process: the
dashboard console registers "doctor" (console_engine.py:570) and calls
run_doctor directly, in-process (console_engine.py:1297), on the web
server's console thread pool — in a process that holds live SessionDB
connections (web_server.py:11673, :11689). Typing "doctor" there
raw-opened and closed state.db, projects.db, response_store.db,
cron/executions.db and every board's kanban.db while those connections
were live. The HTTP route at /api/ops/doctor deliberately spawns a
subprocess instead; the console path did not.
hermes_cli.sqlite_safe_read exists to prevent exactly this, and its
read_header_bytes_preopen is documented as "the ONLY sanctioned
byte-level read of a database file". It performs the registry check and
the open/read/close together under the connection-lifecycle lock, so it
refuses once any connection to the path is live. The audit that converted
the other byte-probes (hermes_state.py:2750, backup.py:436,
kanban_db.py:1861) landed in 95fb477856 on 2026-07-25;
_read_journal_mode was added in 6583297086 on 2026-08-06 and reintroduced
the pattern, so this is a regression against an invariant the tree already
states, not a refactor preference.
The helper is a plain byte read, so the docstring's stated property is
preserved: no SQLite engine open, and no -wal/-shm sidecars are created.
Only the acquisition of `header` changes; the empty / not-a-database /
unrecognized-format-version branches are untouched.
Consolidation follow-up on top of #59182's cherry-picked base:
- Add _looks_structured_value(): triggers a yaml.safe_load structured
parse only when the value starts with '[' / '{' or spans multiple
lines with YAML list-item ('- x') or mapping-entry ('key: v') shaped
lines. Deliberately avoids the over-broad leading '-' trigger from
#88066 so '-5' and '--flag' stay strings.
- Stays folded INSIDE the string-typed-key guard: keys whose
DEFAULT_CONFIG type is str (e.g. approvals.mode) are never coerced.
- Tests: multi-line YAML list/dict, string-typed key given '[x]' and
'-5' stays string, dash-prefixed scalars stay strings, plain
multi-line prose stays a string, load_config round-trip.
Sabotage-verified: 7 of the suite's tests fail on main without the fix.
Fold the list/mapping parser INSIDE the existing string-typed-value coercion guard (the `not isinstance(_default_value_for_key(key), str)` block from e4ea0a0ed) instead of running it unconditionally, so a genuinely string-typed setting whose value merely starts with '[' or '{' is left untouched while non-string keys get JSON/YAML flow literals parsed to real lists/dicts.
Update website/docs/user-guide/configuring-models.md: the `config set only writes scalar values` note is no longer accurate; document the list/mapping support with a quoted example.
Fixes#40545#50168
paperclip#10978 made destructive replacement an explicit caller choice
in their skill-sync and package-import paths: a rerun must never remove
operator edits by default. Our hub-skill updater had the same hazard --
'hermes skills update' calls do_install(force=True), which rmtree-replaces
the skill directory even when the user edited it after install.
do_update now compares the on-disk content hash against the hash the
lockfile recorded at install time; drifted skills are skipped with a
notice and only overwritten with the new --force flag (CLI + /skills
slash path). Bundled skills already had this protection via the
user-modified manifest in hermes update; this brings hub-installed
skills to parity.
Sabotage-verified: disabling the drift check makes the new skip test fail.
Copilot CLI 1.0.79-3 added /worktree new (start a session in a new
worktree). Hermes already has hermes -w launch-time isolation; this adds
the mid-session counterpart: /worktree new [name] creates a tree under
.worktrees/ (remote-tip base, worktree_sync honored), retargets
TERMINAL_CWD + process cwd, and registers the same keep-if-unpushed exit
cleanup. /worktree shows the active tree; /worktree list lists them.
Named trees skip the hermes- prefix so the startup pruner ages them on
the slower named-tree schedule.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
Wire the continuity flag through every cron-creation surface, not just the
model tool:
- dashboard (web/): checkbox in the cron job editor; form state round-trips
the stored reserved 'self' entry into the toggle and strips it from the
context_from textarea; web_server dashboard validator skips 'self'
(create precedes the job's existence)
- Bot Mode Routines tab (hermes-bots plugin): Continuity checkbox in the
New Cronjob dialog, forwarded through cron.manage
- tui_gateway cron.manage RPC: optional continuity param on action=add
vitest cron-job suite 10/10 (4 new), tsc app project clean, py_compile clean.
Perplexity Computer's July update let its agent manage sessions
conversationally from any surface — pin, archive, rename, fork — treating
session organization as operational infrastructure rather than a GUI
nicety. Hermes already has the durable pinned flag in state.db (Desktop
sidebar writes it; auto-archive honors it), but no CLI access existed:
GUI-only management was a single point of failure and blocked scripting
(issue #52955).
- hermes sessions pin <id...> / unpin <id...>: set/clear the durable keep
flag via SessionDB.set_session_pinned (whole compression lineage,
prefix resolution, multi-id, exit 1 on any miss)
- hermes sessions pinned [--json]: list all pinned conversations via the
include_pinned back-fill (old pins can't fall off a paging window);
--json enables backup/restore scripting
- docs: user-guide/sessions.md section
- tests: 6 tests covering prefix resolution, multi-id partial failure,
pinned-only filtering, JSON shape, empty hint
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:
- tools/checkpoint_manager.py: per-project agent-write ledger
(sha256 of every landed write_file/patch), safe_restore_plan()
classifier, and restore(safe=True) that reverts only agent-authored
changes, deletes agent-created files, and preserves user hand-edits.
Empty ledger (pre-existing stores) falls back to the classic full
restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
tweaks, agent-created file removal, empty-ledger fallback.
Claude Cowork (Aug 6, 2026) added skill & plugin security scanning:
third-party skills and plugins are automatically checked for malicious
content on upload/edit, returning pass/warn/fail. Hermes already scans
hub-installed skills (tools/skills_guard.py), but `hermes plugins
install` cloned and activated arbitrary Git repos completely unscanned —
and plugins run Python in-process, making them the more dangerous
surface.
- tools/plugin_guard.py: plugin-adapted scanner reusing the skills_guard
pattern engine. Exempts the documented provider-plugin patterns (own
requires_env API-key reads, HTTP calls with keys) on code files while
keeping true threat signals (foreign credential-store access, reverse
shells, destructive/persistence/obfuscation patterns, prompt injection
in docs). Plugin-sized structural limits; VCS/venv dirs excluded.
- hermes_cli/plugins_cmd.py: scan the temp clone before it is moved into
~/.hermes/plugins/. safe=install, caution=confirm (interactive prompt
or --force), dangerous=blocked (--force does NOT override). Re-scan on
`hermes plugins update`; a dangerous updated tree is deactivated until
the user reviews the findings. Dashboard install path returns
structured scan_blocked/scan_findings.
- Config gate: plugins.scan_on_install (default true) in config.yaml.
- Validated against all 60 bundled plugins: 57 safe, 3 caution (real
sudo / curl|sh content in their docs), 0 false-positive blocks.
- 15 new tests incl. E2E through _install_plugin_core with real git
clones.
'✓ Update complete!' now shows what the update actually delivered:
'✓ Update complete! (v0.19.4 → v0.20.0)' when the pyproject version
changed, '(v0.20.0)' when commits landed within one release, and the
plain message when the version cannot be read. Reads the on-disk
pyproject.toml (not importlib.metadata, which still describes the old
install after a pull). Applied to both the git and Windows-ZIP paths.
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.
- tools/skill_ledger.py: append/list/get, blob store, actor derivation
(curator|agent|user), single-entry rollback that takes a pre-rollback
safety entry first and FAILS CLOSED when that capture fails (consistent
with the whole-run tarball rollback hardening from #63366). Path
containment check so a hand-edited ledger can't write outside
HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
delete intent recorded via absorbed_into/archived evidence),
archive_skill()/restore_skill(), and curator auto-transitions (tagged
actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
hermes curator rollback <entry-id> (whole-tree snapshot rollback
unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
(default 0 = never) + explicit hermes curator purge, recorded in the
ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
archive TTL purge.
Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).
Closes#45778, #50875. Tests adapted from #50261 by @yu-xin-c.
The 'Configure auxiliary models' menu under 'hermes model' now includes a
Delegation entry so the delegate_task subagent model is discoverable and
configurable interactively, instead of requiring hand-edited
delegation.provider / delegation.model keys in config.yaml.
Delegation is not an auxiliary_client task — subagents are full child
agents resolved via tools/delegate_tool.py — so the picker entry writes to
the top-level delegation.* section rather than auxiliary.*. 'auto' (inherit
the parent agent) is persisted as empty strings, never the literal 'auto',
which delegate_tool would try to resolve as a provider name. 'Reset all to
auto' clears only the four delegation routing fields and preserves
non-routing settings like max_concurrent_children.
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:
agent:
empty_response_guard:
enabled: true # false = legacy fixed 3-retry behaviour
cost_threshold_usd: 0.25 # per-attempt cost that halves the budget
- hermes_cli/config_defaults.py: new documented subsection under agent
(additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
section to (enabled, threshold) with fail-open tolerance for
malformed values; guard_enabled()/_cost_threshold_usd() now read the
init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
covering malformed sections, YAML string booleans, bad thresholds,
and a DEFAULT_CONFIG sync check; new integration test proving
enabled:false restores the legacy 1+3-call behaviour.
Requested by isak-ialogics on PR #75115.
Kanban worktree workspaces were never removed by anything: _cleanup_workspace
preserved them by design, the CLI startup pruner explicitly defers t_* trees
to 'hermes kanban gc', and gc only sweeps scratch — so every worktree task
leaked its checkout forever (measured ~130GB on one estate).
- _cleanup_workspace now dispatches worktree workspaces to a new
_cleanup_worktree_workspace, which removes the worktree and its
auto-generated wt/<task-id> branch only when the tree is clean AND every
commit is reachable from a remote-tracking ref (reusing cli.py's
_worktree_is_dirty / _worktree_has_unpushed_commits predicates). Any
doubt preserves the worktree. dir workspaces stay untouched.
- The #33774 active-children deferral now covers worktree parents, and
_try_cleanup_parent_workspaces reaps deferred worktree parents when the
last child reaches a terminal state.
- archive_task reaps workspaces too; tasks archived without completing
previously leaked forever.
- 'hermes kanban gc' gains a backstop sweep for archived worktree tasks
that predate these hooks.
Tests: tests/hermes_cli/test_kanban_worktree_teardown.py (10 cases: removal,
dirty/unpushed/custom-branch/main-checkout/non-git preservation, complete/
archive integration, deferred-parent handoff).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to #80451. /api/status cleared gateway_platforms whenever the
gateway process was down — correct for a clean stop (stale 'connected'
states are noise) but wrong for startup_failed, where the fatal entries
ARE the diagnosis: per-profile credential collisions and auth failures
(multiplex '<profile>:<platform>' keys) that the single exit_reason
string cannot express. #80451's writer-identity and freshness filters
already drop entries from other/older processes, so preserving
fatal-state entries here cannot leak another gateway's live state.
Live-validated shape: a real multiplex gateway (2 secondary profiles,
rejected tokens) persists telegram / alpha:telegram / beta:telegram
fatals in gateway_state.json; /api/status previously reported {} for
platforms while state was startup_failed.
Cmd/Ctrl+Shift+B worktree flows on a remote gateway route through the
backend's /api/git mirror (hermes_cli/web_git.py), but that mirror had
drifted behind the Electron-local git ops the same UI drives locally, so
the flows broke exactly and only on remote connections:
- Convert-a-branch: the picker offers remote-tracking refs, and the
Electron op turns "origin/feature" into a local tracking branch. The
mirror ran `git worktree add <dir> origin/feature` verbatim, which
either fails or detaches HEAD. It now resolves the ref's remote via
git (never assuming "origin"), fetches best-effort, and creates the
worktree with `--track -b <short-name>`.
- branch_list omitted remote-tracking refs entirely and never set the
`isRemote` flag the renderer's HermesGitBranch contract requires —
the convert picker on a remote gateway couldn't reach a teammate's
branch and mislabeled every row's action.
- Branching off an `origin/…` base silently wired the new branch to the
remote upstream; the mirror now passes `--no-track` like the Electron
op does.
Renderer side, replace the silent degradation with a capability gate:
when a remote backend predates the /api/git worktree routes, worktree
creation failed with an opaque "Expected JSON … got HTML" toast. The
route-missing shapes now surface a clear "update the Hermes backend"
message (isGitEndpointMissingError, mirroring the sidebar batch-endpoint
detector); real git errors still pass through untouched.
Sibling audit (documented, no code change needed): repo status / review /
file-diff / git-root / default-cwd already route through desktopGit()'s
REST bridge or /api/fs on remote; repo scan is deliberately a no-op there.
Stale comments claiming "empty/false on a remote backend" in projects.ts
and coding-status.ts updated to describe the backend-routed reality.
Fixes#81724
The freshness window (updated_at >= live process create_time - 2s) had a
P1 boundary hole: a stale failure written by the PREVIOUS process
immediately before a fast restart landed inside the slack and was
aggregated; if that platform was then removed, the new process never
replaces the entry and NAS stays degraded indefinitely.
Replace clock heuristics with persisted writer identity:
- write_runtime_status now stamps every platform entry with the writing
process's (writer_pid, writer_start_time) — the same PID-reuse
fingerprint the liveness checks use, so a recycled PID never
masquerades as the original writer.
- The aggregation ownership filter requires exact equality between an
entry's stamp and the profile's validated live gateway process
(get_runtime_status_running_pid + _get_process_start_time). No slack,
no timestamps. Legacy entries without a stamp fail closed.
- Writer stamps are process recon (same class as the auth-gated
gateway_pid) and are stripped from all /api/status projections, both
active-profile and merged cross-profile entries.
Near-boundary regression test: prior-process entry stamped 100ms before
restart is excluded; recycled-pid-different-fingerprint excluded;
legacy no-stamp excluded; current-process entry kept.
Gateway startup deliberately preserves plain platform entries in
gateway_state.json across restarts, and the active-profile endpoint
compensates by filtering against current configuration. The cross-profile
aggregation copied raw maps, so a fatal entry for a platform the operator
had since disabled/removed could keep NAS reporting the instance degraded
indefinitely.
The aggregation has no cheap per-profile config context (platform sets
depend on tokens in each profile's .env behind its secret scope), so use
freshness instead: an entry is aggregatable only when its updated_at is
at/after the live gateway process's create time (validated PID via
get_runtime_status_running_pid + psutil create_time; the record's own
start_time field is a PID-reuse fingerprint in clock ticks, not a
timestamp). Config changes require a restart to take effect, so
restart-anchored freshness is exactly the config filter's semantics.
Fail closed: unparseable timestamps or no live process exclude the entry
— a false 'degraded forever' is the worse failure mode.
- /api/status now folds LIVE independent per-profile gateways' platform
failures (gateway_mode == 'multiple', the OOF-3 deployment mode) into
gateway_platforms under the validated <profile>:<platform> grammar, so
NAS fleet health sees them without a schema change. ?profile= requests
stay unmerged (single-profile view).
- Namespaced-key validation no longer fails open: colon-containing keys
are grammar-checked even when configured-platform loading throws.
- Platform key segment now accepts hyphens, matching plugin platform IDs
(plugins/platforms/<dir> names, e.g. foo-bar).
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).
Split the unreachable path by durable operator intent:
- desired_state == "stopped" (written only by the s6 lifecycle commands;
the same intent signal container-boot reconciliation trusts) -> drop
the fire with 200 + a structured log line, mirroring NAS's own
instance_stopped drop. Jobs are not lost: the Chronos provider
reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
file without desired_state) -> keep the retryable 503, now stamped
with Retry-After: 60 so a scheduler that honors it spaces retries
past the wake/restart window instead of exhausting them inside it.
The gateway's own pass-through 503s (draining) get the same hint.
The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.
- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
(process, home), keyed off the agent's OWN home (not ambient
HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
agent.bot_mode_protocol (default True), stable tier, byte-stable
across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
the bundled plugin gates ALL SOUL protocol writes on it (backfill,
composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere
Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
A same-day version-floor bump (0.20 runtime contract) left every install
with an older cua-driver hard-failing on all computer_use calls: the
start() gate fails closed, while the `hermes update` refresh defers to the
driver's own check-update verb — whose ~20h cache routinely answers "no
update available" right after we raise the floor. Hermes knew it required
0.20+ but never acted on that knowledge.
Two changes:
- tools_config.install_cua_driver(): a contract-failed installed driver is
repaired on the upgrade=True path too (previously only upgrade=False).
The contract failure itself is the confirmation, so the
require_confirmed_update gate and the check-update short-circuit are
bypassed for repairs — an indeterminate or stale-cached check can no
longer pin users on an unusable driver.
- cua_backend.CuaDriverBackend.start(): when the contract gate fails on an
installed binary, attempt one automatic repair per process via the
standard install path, then re-probe. HERMES_CUA_DRIVER_CMD overrides
are never repaired (explicit override is authoritative even when broken)
and a missing binary still just reports the install hint. A failing
installer can't loop: the second start() surfaces the original error.
Tests: contract-repair coverage in test_computer_use.py (auto-repair
success, failed repair surfaces the original error, once-per-process
guard, override never repaired, missing binary never repaired) and
test_install_cua_driver.py (incompatible driver repairs despite an
indeterminate check-update, check-update not consulted). All new tests
verified to fail against the unfixed source (sabotage run).
Live-testing the Cua Driver 0.20 convergence on Windows 11 (session 2,
cua-driver 0.20.0) surfaced three defects in the existing-profile browser
path and in install status.
1. The config grant was silently nullified by an approval bypass.
`--yolo` / `-z` map onto a private unrestricted daemon, which answers every
browser_prepare. Because the host delegated the entire existing-profile
decision to the driver, that bypass also nullified
`computer_use.grant_existing_profile: false`: a plain `hermes -z` attached
to the user's real Chrome profile and read live page content over CDP, with
the driver reporting it as "the approved existing Chromium profile". It was
never approved.
An approval bypass is consent to skip prompts, not consent to read an
existing profile's pages, cookies, and storage. CuaTypedBrowserRoute.prepare
now enforces the key itself, regardless of permission mode. bounded stays
exempt - its reviewed capability manifest is the authorization boundary.
The authorization inputs are resolved in the backend from config and the
backend's immutable mode, never from model-supplied kwargs.
2. The grant, once set, still could not be used.
With `grant_existing_profile: true` the runtime is launched
`--grant existing-profile` correctly, but cua_browser_prepare then hit a
runtime approval prompt anyway - re-asking the user to authorize what the
config already authorized, and making the documented opt-in unusable on any
non-interactive run, where the prompt has nobody to answer it and the call
dies on approval timeout. The durable, file-backed grant now stands in for
that prompt. Scope is narrow: only the existing-profile prepare, only when
the grant is present; isolated launches still prompt and any resolution
failure falls closed to prompting.
3. `computer-use status` hid a custom override and spliced its output.
With HERMES_CUA_DRIVER_CMD pointed at cmd.exe, status printed the child's
multi-line banner and prompt inside the one-line version field, never
mentioned the override, and advised `hermes computer-use install` - which
install itself (correctly) refuses to run against an overridden path. It now
names the override and mirrors install's update-or-unset guidance, and
version output is reduced to one bounded line.
Verified on the reported host: `-z` existing-profile attach now refuses and
names the key; `grant: true` no longer prompts (33s vs a 300s approval
timeout); status names the override and prints one line. No change to the
reconciliation path - driver SHA256 unchanged end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Mirror the strict unit-name shape from the hermes-serve gate (review on
PR #83595) on the gateway side too: the discovery gate and the SIGUSR1
eligibility helper now accept only `hermes-gateway.service` or the
`hermes-gateway-<profile>` family, so a near-prefix unit like
`hermes-gatewayd.service` can neither enter the restart path nor be sent
a SIGUSR1 it does not handle.
Review on #83595 flagged two service-lifecycle gaps in the hermes-serve
restart support:
- The unit-name gate accepted anything starting with "hermes-serve",
which also matched the unrelated hermes-server.service. Require the
exact base unit or the hyphenated profile family instead.
- The fleet-restart loop and _finish_dashboard_update_cleanup() could
both restart the same hermes-serve unit — the loop restarts it
directly, then cleanup's PID scan finds the fresh process and
restarts its owning unit again. Thread the fleet loop's restarted
unit names through to _kill_stale_dashboard_processes() so it skips
units already handled.
hermes update discovered and restarted hermes-gateway* systemd units but
never looked for hermes-serve* — the Desktop app's backend — so it kept
running stale pre-update code until the user restarted it by hand (#83438).
Extend the systemd unit discovery/restart loop to also match hermes-serve*
units. They don't wire SIGUSR1 to a graceful drain (only gateway/run.py
does), so restart eligibility for the graceful path is now gated on unit
name via a small, directly-tested helper; hermes-serve units fall straight
to the existing blunt systemctl restart path, matching the workaround the
issue already documents.