Port the stale PR #61665 behavior to current main. The original two-file contribution is by yungchentang; this candidate preserves its scoped Codex OAuth fallback intent.
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
Live incident 2026-08-17: the source checkout was parked on a stale feature
branch (claude-code-inspired/local-terminal-memory-limit, days behind main),
left there by earlier tooling. 'hermes update' autostashed, refreshed lazy
backends, synced skills, and printed '✓ Code updated!' / '✓ Update complete!'
while the checkout stayed on the stale branch with none of main's new code.
Two sessions burned time on 'the fix is missing' confusion.
- Parked-branch guard: auto-switch back to the update target ONLY when the
parked branch is clean and fully merged (git cherry origin/<target> shows
nothing unmerged); the checkout then STAYS on the target instead of being
re-parked. Otherwise: loud CODE UPDATE SKIPPED block naming the branch,
behind-count, and resolution commands; exit 1; branch untouched.
- The up-to-date (commit_count == 0) path no longer switches back to a
fully-merged parked branch either.
- Post-pull gate additionally refuses to print '✓ Code updated!' when HEAD
ends up attached to a non-target branch.
- Summary lines now carry the actual branch + HEAD short-sha:
'✓ Update complete! [main @ 30fcf9580]' — drift visible at a glance.
- New config toggle updates.auto_switch_parked_branch (default true).
- Real-git-fixture regression tests (init/clone/branch, no subprocess
mocks): clean+merged auto-switch, dirty skip, unmerged skip, cherry-picked
equivalence, config opt-out, unverifiable ref, on-main fast path,
up-to-date no-repark, summary branch/sha assertions.
Backend: /api/status now carries a stable random install_id persisted once
under the root HERMES_HOME, shared by every profile of the install.
Desktop: roster enumeration captures it per connection (TTL-cached probe),
buildAgentRoster collapses same-install rows with a deterministic canonical
pick (active > local > ssh > remote > cloud > earliest), the @name-device
handle rule runs after the collapse, and Settings → Gateways shows a
display-only 'Same backend as' hint. Backends without install_id bypass the
collapse (fully backward compatible).
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):
- Promote the media-send timeout to the standard resolution pattern:
HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
cron.media_send_timeout_seconds in config.yaml, then 300s default
(mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
(environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
in delivery_errors (post-#88631 the reason reaches the run status, not
just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
Review finding on the off-loop bootstrap: returning None on every cold-
cache loop-thread call silently dropped the first goal/heartbeat
persistence op even when the DB was perfectly healthy. The loop-thread
path now waits up to 250ms on the bootstrap event - a healthy init
(tens of ms) completes inside the window and the caller gets the real
DB; a contended init (the crash-loop scenario) exceeds it and degrades
to None with a bounded, watchdog-safe stall.
SessionDB.__init__ runs schema init, and a migration against a contended
state.db blocks for seconds. The goal/heartbeat path reached it
synchronously on the gateway's event-loop thread (GoalManager() ->
load_goal -> _get_session_db -> SessionDB()), so a contended DB starved
the loop-liveness watchdog, which hard-exited with code 75 and the
supervisor restarted straight back into the same state - an unbounded
crash loop reported from an enterprise fleet.
_get_session_db now detects a running loop on the calling thread: on a
cache miss it kicks a one-shot background bootstrap thread and returns
None immediately (every caller already degrades gracefully on None);
the cached instance serves all later calls. Worker threads construct
inline as before, with a lock-guarded cache so a bootstrap race keeps
one instance and closes the loser. The heartbeat module shares this
boundary via the same _get_session_db.
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):
- Incomplete-validator findings are now PRESERVED as partial evidence;
only the validator's pass/fail verdict is excluded from the advisory
verdict. A report with findings from an incomplete check no longer
reads as clean.
- Clean-report wording is now "no findings from completed checks"
whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
discard-pinning test, plus the completed-checks wording case.
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.
- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
static (no measurable cost); security invokes NVIDIA SkillSpector in
its keyless static-rules mode (~+1.2s per install). schema/quality
stay excluded: hygiene signal ("author not specified" is
high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
or failing, the security check reports status="incomplete" and the
adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
This also absorbs a live upstream inconsistency: SkillEvaluator's
anti-tamper cross-check on SkillSpector's risk score currently trips
on moderate-finding skills (fail verdict with zero findings, e.g.
github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
NVIDIA separately; either way an evidence-free fail must not render
as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
not-run formatting.
E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).
- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
over the quarantined bundle, parses the JSON report, classifies
secrets-class findings (private keys, tokens, credentialed connection
strings) apart from advisory PII findings. Every failure mode
(binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
guard's policy decision and before the install confirmation. Findings
are shown with file:line; secrets-class findings render red with a
loud warning. Warn-and-continue by design — the built-in skills guard
remains the only enforcement layer, because the upstream PII scanner
has known false-positive classes (git@github.com, docs example
emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
config toggle.
Scanner install (optional):
uv tool install --python 3.13 \
"skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"
E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
Two bugs found by a REAL two-gateway live test (two isolated HERMES_HOMEs,
bravo running the api_server platform, alpha's agent autonomously running
`hermes peer dm` from its Bot Chat protocol; reply relayed correctly and
persisted in bravo's canonical Bot Chat):
1. peer dm parsed the session-create response flat, but api_server wraps
the row: {"object": "hermes.session", "session": {...}} — every first DM
to a fresh peer failed with "Peer did not return a session id" (and the
orphaned Bot Chat then 400'd retries with duplicate-title). Parse the
wrapped shape; the test fake now mirrors the real response shape so this
class can't pass green again.
2. api_server: a session created with no model persists the advertised
virtual model ("hermes-agent") on the row; session chat then replayed it
as a REAL model id and the provider 400'd ("hermes-agent is not a valid
model ID"). _request_agent_overrides already filters the virtual model
for per-request bodies — apply the same filter to the stored session
model at both chat sites (sync + stream), so it means "gateway default"
exactly like the request-body path.
Live E2E transcript (bravo's Bot Chat, via /api/sessions/{id}/messages):
user: Message from 🤖 alpha (@alpha): What is your callsign?
assistant: CALLSIGN-BRAVO-7
peer cmd unit suite 10/10 with the corrected fake.
The desktop's main serve process opens memory_store.db for every known
profile and nothing closed those connections before delete_profile's
rmtree — on Windows the open SQLite handles make the removal fail with
WinError 32 for both the CLI and the DELETE /api/profiles/<name> route
(#88347). POSIX unlinking of open files hid the same leak.
MemoryStore.close() is refcount-driven, so a live holder keeps the
handle forever; add MemoryStore.release_all_under(directory) to
force-close every shared connection under a directory, and call it in
delete_profile after stopping the profile backends. Inside serve the
handles live in that very process and get released; from the CLI it is
a no-op.
Fixes#88347
Bots could message teammates on their own machine (hermes -p <bot> chat) and
the desktop could relay user mentions over Connections, but a bot had NO
transport to a bot on another gateway. This adds one, with zero new server
surface: the peer's existing api_server platform is the wire.
- hermes_cli/subcommands/peer.py: `hermes peer add/list/remove/dm`.
`dm <peer>[/<agent>]` resolves the remote agent's canonical "Bot Chat"
(list by title, create when missing), runs one synchronous agent turn via
POST /api/sessions/{id}/chat, and prints the reply on stdout — the exact
cross-machine twin of the local bot-messaging command, so the Bot Mode
protocol composes over it unchanged. Named profiles route via the peer's
/p/<profile>/ multiplex mirror. Peer URLs live in config.yaml
(`bot_peers`); the peer's API_SERVER_KEY is a credential and lives in
~/.hermes/.env as HERMES_PEER_<NAME>_KEY.
- hermes_cli/main.py: parser wiring + fast-path/session-flag command sets.
- tools/bot_mode_probe.py: when peers are registered, the injected Bot Chat
messaging protocol gains a cross-machine paragraph (peer roster +
`hermes peer dm` pattern) so agents discover remote teammates on their
own; peers join the capability fingerprint so registering/removing one
refreshes eternal Bot Chat prompts on the next message (loud, one-time,
user-initiated — no per-turn cache drift).
- Docs: Bot Mode guide (bot-initiated DMs across machines) + cli-commands
reference (`hermes peer` section + summary row).
Tests: tests/hermes_cli/test_peer_cmd.py (target parsing, /p/ scoping,
registry round-trip in isolated config, real-loopback-HTTP dm flow incl.
Bot Chat create-vs-reuse and bearer auth), bot_mode_probe peer-paragraph +
epoch tests. E2E: real `python -m hermes_cli.main peer ...` against a live
fake peer over HTTP with isolated HERMES_HOME (config/.env persistence,
bare + /p/<profile> routing, stdin, --json). 23 passed; ruff clean.
Implement Claude Opus review findings for Meta API support:
- Document in agent/agent_init.py that provider="meta" without an api.meta.ai URL falls through to chat_completions by design (URL-driven wire selection).
- Comment on suppression guard in hermes_cli/runtime_provider.py noting api.meta.ai is handled by _detect_api_mode_for_url.
- Replace inline __import__ with top-of-module import in tests/hermes_cli/test_model_switch_openai_api_mode.py.
- Rename test_meta_retention_not_sent_when_overridden -> test_meta_retention_override_wins in tests/agent/transports/test_meta_codex_cache.py.
- Add test in tests/agent/test_meta_agent_init.py for provider="meta" fallback without api.meta.ai URL.
- Add test in tests/agent/test_auxiliary_client.py for prompt_cache_retention: "24h" under _CodexCompletionsAdapter.
Source: Claude Opus review findings for feat/meta-api-support.
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
(after explicit api_mode wins, before provider-name specials) via lazy
import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
usage cache reporting, model-switch override, and config roundtrip; extend
test_model_switch_openai_api_mode with meta cases.
Surface meta/muse-spark-1.2 in the Hermes model selector (CLI, desktop,
gateway) via the curated OpenRouter list and regenerated model-catalog.json.
The model is live on OpenRouter with tool calling; the picker intentionally
does not show the full OpenRouter catalogue.
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.
Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.
- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
[project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
skill_view resolves cross-tier collisions in favor of the project tier
(same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)
Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
/simplify-code findings on the salvage stack:
- the classify+reclaim+counter block was pasted verbatim into both ticker
loops (_start and _start_multiplex) along with duplicated function-local
imports — extracted _note_tick_failure() next to _backoff_wait_seconds
so both loops share one implementation.
- hermes_cli/cron.py's EMFILE hint reimplemented the text half of
_is_fd_exhaustion with a case-SENSITIVE variation (drift risk) — split
_is_fd_exhaustion_text() out and use it from both.
11 EMFILE tests + 54 provider/ticker tests green; ruff clean.
Follow-ups on the #87796 salvage:
- cron/scheduler.py: drop the _reclaim_fds_best_effort call at tick()'s
lock-failure raise site — the ticker loop's except handler already runs
reclamation once per failed tick, so the raise-site call doubled the
gc.collect() pause on every EMFILE failure.
- cron/scheduler_provider.py: extract the exponential-backoff math
duplicated verbatim in start() and _start_multiplex() into a module-level
_backoff_wait_seconds() helper.
- hermes_cli/cron.py: `hermes cron tick` now reports a propagated OSError
cleanly (exit 1) instead of dumping a traceback — tick() raising on real
lock-acquisition failures is new behavior from this fix.
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.
- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
Two gaps from the Aug 2026 'hermes -w timed out after 30s' incident:
1. Atomic failure cleanup: a timed-out/failed `git worktree add` left a
partially-materialized directory plus a LOCKED admin entry under
.git/worktrees/ (lock pid = the live hermes process that timed out),
which the startup pruner's dead-pid unlock never reaps — retries of
the same name fail forever. _cleanup_failed_worktree_add sweeps dir,
admin entry, and orphaned branch on every failure path (timeout,
nonzero exit, remote-base retry).
2. Pack maintenance: nothing consolidated the object store; on a
multi-agent box packs sprawl (39 packs / 638MB at the incident) and
every object lookup scans all pack indexes until worktree creation
blows its timeout. _maintain_pack_health repacks (niced, background,
fail-soft) when *.pack count reaches 15, wired into the existing
startup maintenance thread on both the CLI (-w) and TUI paths.
gc --auto doesn't cover this: its threshold is 50 packs.
Both sabotage-verified; full repack on the incident box: 39 packs ->
2, 638MB -> 287MB, worktree add 30s-timeout -> 0.5s.
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):
- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
(health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
are allowed in the generic loop while native anthropic stays skipped
E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
External review (Fable) caught a real false-positive widening in the
original commit: the new argv[1] script-name check reused the loose
`script_name == "hermes" or script_name.startswith("hermes")` pattern
(copy-pasted from the exe_name check above it), but argv[1] can be ANY
user-invoked python script path when argv[0] is a bare interpreter --
unlike a directly-resolved executable name, where a false match on the
substring is rare. A user's own script named e.g. "hermes-notes.py" or
"hermes-unrelated-tool" run via `python3 <script>` would be misidentified
as the console-script shim and become killable by profile delete.
Match against the actual known console-script entry points instead
(pyproject.toml [project.scripts]: hermes, hermes-agent, hermes-acp),
stripping the script's extension before comparing.
Added 2 regression tests: one confirms the false-positive case is now
rejected (fails against the pre-fix loose-match code, confirmed via a
scripted revert), the other confirms the other two real entry points
(hermes-agent, hermes-acp) still match via the shebang-exec path.
Tests: tests/hermes_cli/test_profiles.py -- 158 passed (156 previous + 2
new).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two independent bugs let a deleted profile reappear / leave orphaned
resources on next launch:
1. hermes_cli/profiles.py's backend-process scanner required argv[0] to
resolve to an executable literally named "hermes". Electron's
pool-backend spawn resolves the hermes console-script shim's path and
execs it via the interpreter directly (python3 /path/to/hermes ...), so
argv[0] reports as "python3" and the scanner never matched the running
backend -- delete removed the profile's files but left its live backend
process running (still bound to a port via uvicorn), which
accumulates across repeated delete/recreate cycles.
2. The desktop sidebar's ProfileRail only refreshed its cached profile
list once, on mount, so a delete/create/rename from another surface
(another window, or the CLI) left a stale ghost entry until something
unrelated triggered a refetch. Note: a delete via this window's own
Manage-Profiles view already refreshes the shared $profiles atom
ProfileRail subscribes to (confirmed by reading refreshProfiles() and
handleConfirmDelete()) -- this fix only covers the cross-window/cross-
process staleness gap, not a duplicate of the already-merged
#57329's Manage-Profiles rail-refresh work.
Fix 1: recognize a python-interpreter argv[0] exec'ing a hermes-named
console-script shim via argv[1]. Fix 2: refresh the profile list on window
focus/visibilitychange, matching the existing pattern used elsewhere in
the sidebar (sidebar/index.tsx, use-background-sync.ts, star-map.tsx,
use-gateway-boot.ts all use the same focus+visibilitychange pattern).
## Related work already on main
PR #57329 (merged) fixed the *headline* symptom from issue #52279
(deleted profile respawns) via a different, non-overlapping mechanism:
routing profile-delete through the primary backend instead of spawning a
fresh pool backend, plus a separate recreation guard in
ensure_hermes_home() (#49435, merged) that makes a backend spawned into a
deleted profile's directory raise FileNotFoundError instead of silently
recreating it.
This PR is NOT a duplicate of that fix. Verified: even with both of those
merged, a backend process that survives because of gap #1 above still
holds a bound port via uvicorn -- it just can no longer resurrect the
profile directory. That's real resource-hygiene, not a symptom already
covered. Gap #2 touches a different file/component (ProfileRail /
profile-switcher.tsx) than #57329's rail-refresh half (which touched the
Manage-Profiles view's own $profiles.ts / index.tsx) and covers a
distinct staleness path (cross-window/cross-process, not same-window
delete-then-refresh).
Tests: tests/hermes_cli/test_profiles.py -- 156 passed (existing +
regression coverage for the argv[0] python-interpreter detection case).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling. tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).
Two narrow changes, no heuristics, no new env vars:
- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
the configured spelling IS the launch root (junction-transparent,
physically identical dirs); keep it instead of re-deriving the native
default. This is the same producer contract _preserve_hermes_home_path
already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
configured home is a profile home (<root>/profiles/<name>), also derive
the root spelling lexically (parent of the profiles component, same
rule get_default_hermes_root uses) and run the exact-ownership mapping
against it, so the launcher's lexical root is recovered after re-home
without ever matching arbitrary descendants of HERMES_HOME.
Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
Addresses three gaps found in review of the memory-guard PR:
P1a — standalone daemon was the one uncapped entry point. run_daemon()
now resolves kanban.max_in_progress every tick (explicit config wins,
else the memory-derived default) exactly like the gateway dispatcher and
`hermes kanban dispatch`. New shared parser configured_max_in_progress()
so all three entry points agree on what "explicitly configured" means.
P1b — max_in_progress was enforced per board while the gateway ticks
every active board, multiplying the host budget by the number of boards
(2 boards x cap 2 = 4 workers on a host sized for 2). The cap is now
host-level: _dispatch_once_locked() adds count_running_tasks_other_boards()
to the running count before deriving the tick's spawn budget. Enforced in
the shared locked path, so gateway, CLI, and daemon all inherit it.
max_spawn deliberately keeps its historical per-board semantics.
Fails open per board so one corrupt board can't brick dispatch on the rest.
P2 — the ready loop consumed the entire shared spawn budget before the
review loop ran, so a sustained ready backlog starved autonomous reviews
indefinitely. When spawnable review work exists (assigned + real profile,
mirroring the review loop's own gate) and the tick has budget, one slot
is held back from the ready lane. Reservation is per-tick and
self-releasing; the review lane still spends from the shared budget —
it gains fairness, not extra capacity.
11 new tests in tests/hermes_cli/test_kanban_host_cap.py. Existing
kanban suites: 278 passed (15 failures pre-existing, identical on clean
main baseline). ruff clean.
Two production incidents (OOF-77 "larrikin-lollies", OOF-30
"synclare-task-manager") followed the same shape: no
kanban.max_in_progress configured, a busy board, and a 1 GiB hosted VM.
The dispatcher fanned out 26-31 concurrent workers, the host went into
swap-thrash/OOM, and the whole machine — dashboard included — became
unreachable. NAS restart loops then masked the problem: each restart
"recovered" briefly before the kanban dispatcher immediately respawned
unbounded workers.
Building on the cherry-picked max_in_progress-across-both-lanes fix
(PR #28695, credit @Dusk1e), this adds two complementary safeguards to
hermes_cli/kanban_db.py:
1. Memory-DERIVED default concurrency cap. When kanban.max_in_progress
is unset, resolve_max_in_progress() derives a default of
clamp(MemTotal / 512 MiB, 2, 8) — e.g. 2 workers on a 1 GiB VM,
8 on 4 GiB+. Explicit config always wins in either direction. On
hosts where total memory can't be read (macOS/Windows dev machines),
the default stays None (no cap — unchanged behaviour). Wired into
both dispatch entry points (gateway/kanban_watchers.py and
hermes kanban dispatch) so behaviour matches regardless of path.
2. Live memory-PRESSURE guard inside dispatch_once. A static cap can't
see the host's actual memory state (other tenants, bloated
long-lived workers). The dispatcher now samples system memory each
tick via gateway.lifecycle_ledger.sample_memory() and classifies it
with gateway.memory_status.classify_pressure() (same thresholds as
the dashboard memory banner and OOM-suspicion heuristics from
NS-608/NS-656): critical -> spawn nothing this tick; elevated ->
at most one new worker; unknown -> no restriction (fail-open).
Reclaim/promotion bookkeeping still runs under pressure, and
deferred tasks stay queued — nothing is dropped. Restriction is
surfaced on DispatchResult.memory_pressure and logged.
Tests: tests/hermes_cli/test_kanban_memory_guard.py (14 tests) covers
the derived cap (floor/ceiling/fail-open/explicit-config-wins), the
pressure classifier, and dispatch behaviour under critical/elevated/
unknown pressure including defer-not-drop and bookkeeping-still-runs.
An autouse fixture in tests/conftest.py pins the memory sample to
"no data" suite-wide so existing dispatch tests don't depend on the
CI runner's live memory state (opt-out marker: real_memory_guard).
read_header_bytes_preopen returns None for every failure, so routing the
probe through it flattened "[Errno 2] No such file or directory: …" and
"[Errno 13] Permission denied: …" into one opaque "file could not be
read". doctor exists to name the problem, so that detail is worth keeping:
_report_database_journal_modes prints the string verbatim, and on a
vulnerable SQLite it is the only clue the user gets about why WAL exposure
could not be ruled out.
_unreadable_reason recovers it from metadata only. stat() reports the
missing file, the dangling symlink and the unsearchable parent directory;
os.access(..., R_OK) reports the unreadable file that stat() can still see.
Neither call takes a file descriptor, so neither can cancel the POSIX
advisory locks the previous commit was about — the invariant holds.
_read_journal_mode opened each Hermes database with a bare open(db_path,
"rb") to read header byte 18. The read itself is harmless; the close() is
not. Per sqlite.org/howtocorrupt.html, close() on *any* descriptor for a
file cancels every POSIX advisory lock this process holds on it — so the
close at the end of that with-block drops the locks a live connection is
holding, including the EXCLUSIVE lock a VACUUM holds while it rewrites the
whole file. Another process is then free to write into a file its writer
still believes it owns, which is the documented route to "database disk
image is malformed".
This is reachable. run_doctor is not only a standalone CLI process: the
dashboard console registers "doctor" (console_engine.py:570) and calls
run_doctor directly, in-process (console_engine.py:1297), on the web
server's console thread pool — in a process that holds live SessionDB
connections (web_server.py:11673, :11689). Typing "doctor" there
raw-opened and closed state.db, projects.db, response_store.db,
cron/executions.db and every board's kanban.db while those connections
were live. The HTTP route at /api/ops/doctor deliberately spawns a
subprocess instead; the console path did not.
hermes_cli.sqlite_safe_read exists to prevent exactly this, and its
read_header_bytes_preopen is documented as "the ONLY sanctioned
byte-level read of a database file". It performs the registry check and
the open/read/close together under the connection-lifecycle lock, so it
refuses once any connection to the path is live. The audit that converted
the other byte-probes (hermes_state.py:2750, backup.py:436,
kanban_db.py:1861) landed in 95fb477856 on 2026-07-25;
_read_journal_mode was added in 6583297086 on 2026-08-06 and reintroduced
the pattern, so this is a regression against an invariant the tree already
states, not a refactor preference.
The helper is a plain byte read, so the docstring's stated property is
preserved: no SQLite engine open, and no -wal/-shm sidecars are created.
Only the acquisition of `header` changes; the empty / not-a-database /
unrecognized-format-version branches are untouched.
Two data-loss bugs reported by users:
1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
cleanup called finalize_session on the session the gateway just
reopened. This set end_reason on a row the gateway was actively
writing to, causing the handoff leg to vanish from session history
and breaking session_search recall. Fix: add _handed_off_session_ids
module-level set (mirrors _single_query_finalize_attempted_session_ids
pattern). _handle_handoff_command registers the session_id on
completion; _should_emit_cleanup_session_finalize and
_emit_interrupted_session_end check it before firing.
2. state.db corruption silent failure (#88235): When SessionDB init
failed at gateway startup, the error stayed in logs — messages
flowed but nothing was persisted, with no user-visible indication.
Fix: store _session_db_init_error on GatewayRunner, broadcast a
recovery-guidance message to all home channels via
_send_session_db_warning_notifications() after the gateway connects.
Also improved the 'corrupt' persistence cause wording in
_format_turn_completion_explanation to include the full recovery
path (hermes doctor --fix, sqlite3 .recover, backups).
Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.
Consolidation follow-up on top of #59182's cherry-picked base:
- Add _looks_structured_value(): triggers a yaml.safe_load structured
parse only when the value starts with '[' / '{' or spans multiple
lines with YAML list-item ('- x') or mapping-entry ('key: v') shaped
lines. Deliberately avoids the over-broad leading '-' trigger from
#88066 so '-5' and '--flag' stay strings.
- Stays folded INSIDE the string-typed-key guard: keys whose
DEFAULT_CONFIG type is str (e.g. approvals.mode) are never coerced.
- Tests: multi-line YAML list/dict, string-typed key given '[x]' and
'-5' stays string, dash-prefixed scalars stay strings, plain
multi-line prose stays a string, load_config round-trip.
Sabotage-verified: 7 of the suite's tests fail on main without the fix.
Fold the list/mapping parser INSIDE the existing string-typed-value coercion guard (the `not isinstance(_default_value_for_key(key), str)` block from e4ea0a0ed) instead of running it unconditionally, so a genuinely string-typed setting whose value merely starts with '[' or '{' is left untouched while non-string keys get JSON/YAML flow literals parsed to real lists/dicts.
Update website/docs/user-guide/configuring-models.md: the `config set only writes scalar values` note is no longer accurate; document the list/mapping support with a quoted example.
Fixes#40545#50168
paperclip#10978 made destructive replacement an explicit caller choice
in their skill-sync and package-import paths: a rerun must never remove
operator edits by default. Our hub-skill updater had the same hazard --
'hermes skills update' calls do_install(force=True), which rmtree-replaces
the skill directory even when the user edited it after install.
do_update now compares the on-disk content hash against the hash the
lockfile recorded at install time; drifted skills are skipped with a
notice and only overwritten with the new --force flag (CLI + /skills
slash path). Bundled skills already had this protection via the
user-modified manifest in hermes update; this brings hub-installed
skills to parity.
Sabotage-verified: disabling the drift check makes the new skip test fail.
Copilot CLI 1.0.79-3 added /worktree new (start a session in a new
worktree). Hermes already has hermes -w launch-time isolation; this adds
the mid-session counterpart: /worktree new [name] creates a tree under
.worktrees/ (remote-tip base, worktree_sync honored), retargets
TERMINAL_CWD + process cwd, and registers the same keep-if-unpushed exit
cleanup. /worktree shows the active tree; /worktree list lists them.
Named trees skip the hermes- prefix so the startup pruner ages them on
the slower named-tree schedule.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
Wire the continuity flag through every cron-creation surface, not just the
model tool:
- dashboard (web/): checkbox in the cron job editor; form state round-trips
the stored reserved 'self' entry into the toggle and strips it from the
context_from textarea; web_server dashboard validator skips 'self'
(create precedes the job's existence)
- Bot Mode Routines tab (hermes-bots plugin): Continuity checkbox in the
New Cronjob dialog, forwarded through cron.manage
- tui_gateway cron.manage RPC: optional continuity param on action=add
vitest cron-job suite 10/10 (4 new), tsc app project clean, py_compile clean.
Perplexity Computer's July update let its agent manage sessions
conversationally from any surface — pin, archive, rename, fork — treating
session organization as operational infrastructure rather than a GUI
nicety. Hermes already has the durable pinned flag in state.db (Desktop
sidebar writes it; auto-archive honors it), but no CLI access existed:
GUI-only management was a single point of failure and blocked scripting
(issue #52955).
- hermes sessions pin <id...> / unpin <id...>: set/clear the durable keep
flag via SessionDB.set_session_pinned (whole compression lineage,
prefix resolution, multi-id, exit 1 on any miss)
- hermes sessions pinned [--json]: list all pinned conversations via the
include_pinned back-fill (old pins can't fall off a paging window);
--json enables backup/restore scripting
- docs: user-guide/sessions.md section
- tests: 6 tests covering prefix resolution, multi-id partial failure,
pinned-only filtering, JSON shape, empty hint
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:
- tools/checkpoint_manager.py: per-project agent-write ledger
(sha256 of every landed write_file/patch), safe_restore_plan()
classifier, and restore(safe=True) that reverts only agent-authored
changes, deletes agent-created files, and preserves user hand-edits.
Empty ledger (pre-existing stores) falls back to the classic full
restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
tweaks, agent-created file removal, empty-ledger fallback.