session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.
Now a named session composes with whatever browser source is configured:
- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
too; previously a named daemon ignored it and fell back to scanning
local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
keys its own provider browser via the shared _get_session_info cache
(bu-named-<name>), so each name gets its own cloud browser, the same
name reuses one across calls and tasks, and unnamed calls keep the
per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
(provider resolution would double-session and double-bill).
Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.
E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
The installer clones with --depth 1, so every default install is shallow.
In a shallow repo, an older worktree HEAD (a past snapshot of main) is
disconnected from current origin/main by the shallow boundary, so
'git log HEAD --not --remotes' misreports thousands of already-public
commits as unpushed. The fail-safe unpushed guard then preserves every
aged 'hermes -w' worktree forever, and the git-cherry squash-merge
escape hatch never rescues them (22k 'ahead' >> max_ahead=20).
Real incident: 21 of 25 hermes-* worktrees stuck on one install.
Fix at the root, one owner:
- _deepen_shallow_repo(): one-time blobless unshallow
(fetch --unshallow --filter=blob:none; plain --unshallow fallback)
run from the background startup pruner thread before classification,
so history verdicts become correct and the backlog self-clears on the
next 'hermes -w' startup. Fail-soft offline: keep preserving.
- _cleanup_worktree(): when the unpushed verdict comes from a shallow
clone, say 'Shallow clone — cannot verify push state' instead of the
misleading 'has unpushed commits' message.
- Document the shallow caveat on _worktree_has_unpushed_commits (the
primitive stays conservative on purpose).
Tests: real shallow clone over file:// reproducing the disconnect shape,
covering detection, deepen+verdict flip, pruner E2E reap, offline
fail-soft preserve, full-clone noop, and genuine-unpushed-work survival.
Sabotage-verified: the E2E test fails with the deepen call disabled.
Follow-up to the salvaged #86823: the guard queried a hardcoded
"HermesGateway" task, but `hermes gateway install` registers
Hermes_Gateway (Hermes_Gateway_<profile> for named profiles) via
gateway_windows.get_task_name(). Query that name so the supervisor
guard is active on standard installs; fall back to the default literal
if the module import fails. Test now asserts the profile-aware name is
what reaches the task-state query.
Also corrects the cherry-picked commit's placeholder author email to
the contributor's GitHub noreply address.
The orphan-reap sweep (_reap_unsupervised_gateway_orphans) must not kill a
gateway that Windows Task Scheduler is actively managing. The existing
services.exe parent-chain backstop fails open: when the Task-launched conhost
bootstrap has already exited, Windows does not reparent the gateway, the
chain breaks, and the supervised gateway is treated as an orphan. The reaper
then writes the planned-stop marker, the gateway exits cleanly with code 0,
and the scheduler never restarts it (RestartCount only fires on non-zero
exit) — silently killing A2A/messaging on every desktop-app launch.
Querying the task's own state is the authoritative signal and closes the gap
without depending on process ancestry: if HermesGateway is Running, skip the
reap entirely. Uses PowerShell Get-ScheduledTask (English State enum,
locale-stable) rather than schtasks (localized output + codepage mangling).
Fixes#86721.
`hermes cron run <job_id>` (a one-shot CLI invocation) dispatches
manual runs via the same background-delegation path as an agent's
`cronjob(action='run')` tool call (tools/cronjob_tools.py's
_try_dispatch_background_run -> dispatch_async_delegation(role=
"cron_run", runner=_runner, ...)). The runner thread lives in the
calling process's shared daemon executor. When the one-shot process
exits right after printing "Triggered job: ...", the in-flight runner
dies mid-execution, leaving its cron/executions.db row permanently
stuck at status='claimed' -- every subsequent `hermes cron run` on the
same job then reports "Ran now: failed" because of the still-claimed
row.
cron/executions.py already has the exact self-heal this needs:
recover_interrupted_executions() correctly identifies and reclassifies
'claimed'/'running' rows whose owner process has provably exited
(_owner_is_live checks PID existence AND matches process start-time,
so a reused PID isn't mistaken for the original live owner) to
'unknown', unblocking the job for a fresh claim. But it was only ever
called once, at the long-lived scheduler ticker's own startup
(cron/scheduler.py:379's self.recover_interrupted()) -- a one-shot CLI
invocation has no equivalent "startup" moment of its own, so this
self-heal never ran for it.
Added a call to recover_interrupted_executions() at the top of
_try_dispatch_background_run, right after the async-delivery-supported
gate and before any claim attempt for the current job -- mirroring
exactly what the long-lived scheduler already does at its own
startup, just triggered per one-shot invocation instead of once at
daemon startup. Wrapped in try/except: pass (best-effort; a failure
here must not block the actual dispatch this function exists for).
Traced (but did not attempt to fix) the deeper "why does the runner
die with the process at all" question -- that's the harder problem
options 1/2 in the issue describe (route to the persistent scheduler,
or block the one-shot process until completion). This fix addresses
the more urgent, more clearly-scoped symptom: a stranded stale claim
permanently blocking ALL future manual runs of the affected job, which
is option 3 from the issue and the one with an existing, already-
correct implementation just needing to be wired into this call site.
Added 3 regression tests to a new file, following the established
real-subprocess dead-owner pattern already used in
tests/cron/test_execution_ledger.py (a genuinely-dead PID, not a
mock, matching the real-world failure mode exactly): a sanity test
confirming the stale claim sits unrecovered without the fix; a direct
test of recover_interrupted_executions() reaping such a claim; and a
unit test on _try_dispatch_background_run itself confirming recovery
is called before any claim attempt. Verified as a genuine regression
by reverting the fix and confirming the unit test fails with recovery
never having been called.
35/35 pass across the new test file plus tests/cron/test_execution_ledger.py
and tests/tools/test_cronjob_run_background.py (no regression).
Follow-ups on the #85006 salvage:
- A key_cmd token with no advertised expiry was cached for the life of the
process. The "refresh on 401" contract it relied on has no implementation
(SDK retries cover 429/5xx only), so an expired no-TTL token would 401
every request until restart. Cache on a bounded 15-minute window instead;
helpers that want a longer cache can advertise their real expiry.
- Test for the no-TTL path updated to pin the bounded-window contract;
the remint test's $RANDOM (bash-only, empty under dash) replaced with
date +%s%N so it exercises remint under any /bin/sh.
- website/docs/integrations/providers.md: document key_cmd in the named
custom providers section (contract, precedence, secrets.command contrast).
Custom providers could only authenticate from a static credential (inline
api_key or a key_env env var). Enterprise gateways -- SSO/OIDC brokers, cloud
IAM, internal auth proxies -- issue short-lived bearers instead, so a value
copied into .env is stale within the hour: long sessions start returning 401s
and the user has to restart or run an external cron that rewrites .env.
The existing `secrets.command` source does not cover this: it runs once per
process at startup (subsequent calls are no-ops by design), so it cannot
re-mint a credential mid-session.
Add providers.<name>.key_cmd: a command that prints a token, wrapped at
resolution in a zero-argument callable. Both wire clients already accept a
callable api_key and invoke it per request (the Entra ID path established
this), so chat_completions, codex_responses and anthropic_messages all work
unchanged and always send a fresh credential. The callable also routes the
Anthropic client through its per-request Authorization hook, which is what
OAuth-gated gateway routes require -- so no per-vendor auth wiring is needed
anywhere in core.
- cached until shortly before the advertised expiry (60s leeway), so the
helper runs about once per token lifetime rather than once per request
- expiry is read from the OAuth 2.0 relative `expires_in` when present, and
otherwise from an absolute ISO 8601 deadline (`expiry`, `expiresOn`), which
is what CLI token helpers commonly print. Reading only `expires_in` treated
those helpers as advertising no TTL at all, cached their token for the life
of the process, and returned 401 on every request once the real deadline
passed. ISO parsing reuses hermes_cli.auth._parse_iso_timestamp rather than
adding another datetime parser.
- no synthetic expiry: when no TTL is advertised, or the advertised one is
unparseable or already past, the token is used and refreshed on 401 instead
of re-minted on an invented schedule
- stdout contract matches OAuth 2.0 token endpoints and existing agent
helpers (bare token or JSON access_token/expires_in); multi-line output is
rejected rather than guessed at, so a misconfigured helper surfaces as a
clear error instead of a corrupt-credential 401
- precedence: explicit --api-key still wins; otherwise key_cmd beats a
static api_key/key_env on the same entry
- failures never include the helper's output (may hold a partial token) or
the command string (may embed a client secret)
Resolution happens on two paths. agent/auxiliary_client.py resolves named
custom providers itself rather than calling _resolve_named_custom_runtime, so
key_cmd is honoured in both: wiring only the runtime resolver leaves the main
agent turn working while every auxiliary call (title generation, compression,
vision, embedding) falls back to the no-key-required placeholder and 401s.
Precedence is identical on both paths, so one config entry cannot yield two
different credentials depending on which resolver the caller reached.
Closes#84162
Signed-off-by: LordMelkor <kray@block.xyz>
Compaction re-injects the todo list verbatim (TODO_INJECTION_HEADER +
TodoStore.format_for_injection) while skill instructions are pruned down
to [SKILL_PRUNED: ...] markers — the imperative crosses the boundary
without the policy that governed it, and the agent keeps executing
preserved tasks with the guidance deleted (#84718's T6 pattern).
Close the retention asymmetry at the injection site: when the compressed
transcript carries [SKILL_PRUNED: ...] markers AND a todo snapshot is
being re-injected, append a bounded reload notice to the snapshot naming
each pruned skill with its exact skill_view() reload call, plus a
one-line instruction to re-check that preserved tasks are still
justified. Skill guidance recovery now travels in the SAME boundary
artifact as the imperative — same message, same stale-snapshot strip
lifecycle, so repeated compactions refresh rather than accumulate.
Properties:
- deterministic: derived only from the compressed transcript (same input,
same bytes) — no per-turn nondeterminism in the rebuilt prompt
- zero recurring cost when nothing was pruned (clean sessions unchanged)
- bounded: shares _MAX_PRUNED_SKILL_MARKERS with the summary re-injection
cap; the notice text never contains the canonical marker prefix, so it
can never feed the marker extractor at the next boundary
- rides after TODO_INJECTION_HEADER, so _strip_stale_todo_snapshot
removes snapshot + notice together and the synthetic-row classifier
(_is_synthetic_compression_user_turn) is unaffected
Tests: tests/agent/test_skill_todo_retention_parity.py — unit contract of
the notice builder (naming, dedup/order, cap, determinism, no marker
self-feed) and behavioral compaction runs through the real
_compress_context path (notice travels with the snapshot, absent when
nothing pruned, synthetic-row classification unbroken, strip lifecycle
across repeated boundaries). Sabotage-verified: disabling the append
flips the 3 behavioral tests red.
Part of #84718
A gateway whose asyncio event loop is stalled (e.g. an in-loop
compression pass, #72707) cannot process SIGTERM/SIGUSR1 shutdown.
The updater's drain wait then burned the full 180s budget, warned
"Gateway PID X still running after 180.0s — restart may fail", and
`hermes update` could deadlock behind the wedged process — the user
cannot update their way out of the stall.
Fix: before any drain wait, read the loop-liveness heartbeat file the
gateway rewrites every 30s (#66892). Classification:
- alive (fresh heartbeat): busy-but-alive loop — take the normal
graceful drain, honoring the in-flight cron drain floor (#86684).
- wedged (heartbeat for this PID stale >90s = 3 missed beats): the
loop is provably dead; drain is pointless. Bounded escalation:
SIGTERM + 5s grace, then SIGKILL + 5s wait, then proceed (~10s
worst case, far under the 180s drain budget).
- unknown (missing/corrupt file, PID mismatch): never escalate on
ambiguity — full drain path.
Wired into launchd_restart, systemd_restart, and both updater
gateway-shutdown sites (systemd unit drain + manual profile
gateways). The probe is a local stat + JSON read (well inside the
10s query tier of the subprocess timeout tiering).
The cron drain floor from #86684 is bypassed ONLY when the loop is
provably dead — a merely busy gateway still refreshes its heartbeat
and keeps the full drain budget.
Root cause of the loop stall itself (compression blocking the loop)
is #72707 territory and deliberately out of scope here.
Fixes#81642
Part of #86647.
During the #79278 persistent-mute investigation the decisive evidence was
an absence: a Desktop request left no INFO record in agent.log OR
gateway.log ("832 platform=webhook, 194 platform=telegram, 0
platform=desktop"), so the muted 13:15-13:19 window — 12 non-idempotent
Qdrant snapshots, zero results returned — was structurally
indistinguishable from a request that never arrived. The issue calls out
fixing this observability gap as the first actionable step.
This adds the two INFO records to _run_prompt_submit, the single choke
point every Desktop/TUI turn passes through (user submits, queued
prompts, auto-continue, goal follow-ups, watch upgrades):
- "tui prompt accepted": emitted before the turn thread starts, carrying
the UI session id, the gateway session_key, and the agent's live
session_id — the id triple a rotation-mute trace needs, since
compression rotates agent.session_id independently of the other two.
No prompt content is logged (length only).
- "tui turn finished": emitted in the turn's finally on every path
(success, returned error, exception, interrupt), re-reading
agent.session_id so a mid-turn compression rotation shows up as an
accepted/finished pair with different agent ids. A missing finished
record now positively identifies a turn thread that died before its
finally.
Placement follows @Adolanium's note on the issue: in _run_prompt_submit,
logging sid + session_key + agent.session_id, NOT another platform= line
in gateway.log (Desktop does not use the messaging gateway).
tui_gateway is under COMPONENT_PREFIXES["gui"], so the records land in
agent.log (root catch-all) and gui.log when running under the dashboard.
Tests (tests/tui_gateway/test_prompt_accept_logging.py): accepted+finished
pair on success with the full id triple and no prompt content leaked;
mid-turn rotation visible as differing agent ids across the pair;
finished record fires on the exception path and the returned-error path.
Sabotage-verified: removing the accepted record fails the suite.
End-to-end #86632 reproduction: a real AIAgent child (mocked LLM) with the
post-turn skill-review trigger armed, dispatched through
delegate_task(background=True) on a session runtime where async delivery is
unsupported and no origin session id is bound (cron, post-#66617) — forcing
the synchronous fallback. Asserts (1) delegate_task returns the child's
result, and (2) the automatic background-review fork never spawns inside the
delegated child (the wedge site: the fork replayed the conversation on the
child's finalize path, and _child_future.result(timeout=None) never returned;
heartbeat went stale after 15 idle cycles and the cron watchdog killed the
job).
Verified RED on pre-fix main (fork spawns and wedges), GREEN with the
_delegate_depth guard in AIAgent._spawn_background_review.
Fixes#86632
The post-turn background review fork (`agent/background_review.py`) inherits
the parent agent's live runtime by default. That is a cost win when the parent
IS the main chat model (warm prompt cache, cheap), but the fork also fires
inside delegation subagents, where it inherits the *subagent's* model. When a
subagent runs a premium delegation model, the review silently replays the whole
conversation and emits skill/memory-update output at premium rates, with
nothing in the log or config flagging it (#85859).
Subagents are already barred from writing shared MEMORY.md
(`DELEGATE_BLOCKED_TOOLS`) and are spawned with `skip_memory=True`, so an
automatic review here has little to persist. Guard `_spawn_background_review`
(the single choke point both the turn-finalizer and codex-runtime callers pass
through) to return early when `_delegate_depth > 0`. An explicit `/refine`
(`focus` set) is a deliberate user request and still runs; the top-level path
is unchanged.
Fixes#85859
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.
Two-part fix:
1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
channel stateless (scoped ContextVar set/reset around the call) before
invoking the cron API, so `async_delivery_supported()` gates off
`_try_dispatch_background_run` and the run executes synchronously to
completion in the CLI process — the same behavior `hermes -z` already
gets via declare_stateless_channel().
2. cron/scheduler.py: tick() now periodically invokes
recover_interrupted_executions() (previously only run at scheduler
startup), so execution rows whose exact owner process is provably dead
(pid + process start time check in _owner_is_live) are reaped to
'unknown' by the long-lived gateway ticker without a restart.
Throttled to once per 300s so idle 60s ticks don't pay a ledger
connection every cycle.
Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.
Fixes#86721
The #86687 self-lock preflight fired on every Windows `hermes update`:
bitwarden.py's module-level cryptography import (fixed in #86782 /
#86826-class change) meant cryptography._rust was ALWAYS mapped by the
time the preflight ran, so the update exited 2 before even fetching and
looped forever — including the Desktop in-app update (#86780).
Two structural fixes so the guard can never re-brick the flow it protects:
1. Version-gated detection: _detect_self_loaded_native_modules() now
consults _dependency_sync_would_rewrite(dist) — installed version vs
the on-disk pyproject pins (base deps + all extras, env markers
honored). A loaded module whose distribution the sync will not touch
is no lock risk and is not reported. Unknown → fail closed.
2. Relocated deferral: the check no longer runs pre-fetch. It runs via
_abort_dependency_sync_if_self_locked() immediately before each venv
rewrite (git-path dep sync, ZIP-path dep sync, current-checkout venv
repair) — AFTER the code swap. A deferral now leaves the user on NEW
code with only the dependency install pending (completed by the next
launch's marker recovery), instead of stranding them on the old
checkout in an exit-2 loop.
PyYAML's _yaml extension (loaded by every CLI process) joins the
registry — with version gating it is now safe to list.
Tests: version-gate unit coverage (no-change skip, stale pin, missing
dist, extras, markers, fail-closed None), deferral wiring (marker +
gateway resume + exit 2), placement guards (no detector call pre-fetch;
guard present at git/ZIP sync), and subprocess-verified import hygiene
(import hermes_cli.main and the update --check dispatch never load
cryptography._rust).
Follow-up to #86687 (Halldrix's #83590 salvage — the preflight's intent
stands as defence-in-depth; this makes it fire only when true).
Fixes#86735Fixes#86780Fixes#86781
A corruption class the repair strategies cannot heal (b-tree page
damage) failed repair_state_db_schema on every process start, forever:
_claim_repair_attempt's in-memory set only bounds one process, so each
restart re-ran the full surgery AND took a fresh ~900MB forensic backup
of the same damaged bytes — 105 attempts / 89GB of dead
state.db.malformed-backup-* files over 11 days in the reporting install.
Three bounded behaviors, all sidecar-file based (no schema changes):
1. Persistent attempt ledger (<db>.repair-attempts.json): after 3 failed
repair passes against the same file fingerprint (size + mtime_ns),
repair_state_db_schema refuses with a terminal, actionable error
(restore a backup / `sqlite3 state.db ".recover"` / delete the ledger
to force a retry) instead of re-running surgery. Success clears the
ledger; a replaced or restored file re-keys it and gets fresh
attempts. Missing/corrupt ledger fails open (never blocks a first
repair).
2. Backup dedupe: _backup_db_file reuses the newest existing forensic
backup when it is byte-identical to the damaged file (size+mtime
match, preserved by copy2) instead of copying another ~900MB.
3. Retention cap: only the 3 newest malformed-backup copies (plus
sidecars) are kept; older ones are pruned after each new backup.
Also fixes a same-second timestamp collision that silently
overwrote an earlier forensic copy.
Tests cover ledger accumulation, terminal refusal (surgery not called,
no new backup), budget reset on file change, success-clears-ledger,
corrupt-ledger tolerance, dedupe, distinct-state backups, retention
prune incl. sidecars, and the end-to-end one-backup invariant.
Fixes#86747
The hermetic conftest now exports HERMES_TEST_ISOLATION (value = the tmp
isolation root) before any test module imports, and re-pins it per test in
_hermetic_environment. hermes_state._running_under_pytest() honors the
marker as a test-context signal alongside PYTEST_CURRENT_TEST /
PYTEST_VERSION.
Why a third layer: PYTEST_* belongs to pytest, and tests that spawn
children routinely rebuild the child env and strip it ("the subprocess
must look like a real CLI" — tests/cli/test_exit_watchdog_signal_arm.py,
tests/hermes_cli/test_config_loader_e2e.py do exactly this on purpose).
Such a child loses the HERMES_HOME redirect and the guard's arming signal
in one step, which is how 700+ zero-message fixture rows (dm:123, chat-1,
wx-chat, ...) landed in a developer's production state.db. The marker is
OURS: stripping it is never required to make a child "look real" (no
production code branches on it except the guard), it inherits by default,
and children that genuinely need a real DB use the sanctioned
HERMES_STATE_DB_GUARD_BYPASS=1 hatch instead.
The ancestry-walk layer (previous commits) stays: it covers children whose
env was rebuilt from a completely empty dict. The marker layer covers the
common **os.environ-derived rebuilds cheaply (one dict lookup, no psutil),
and — unlike ancestry — also covers detached/daemonized children that
escape the process tree.
tests/hermes_state/test_isolation_marker_env.py pins: the conftest export,
the marker-alone arming, the rebuilt-env child refusing the production
path, and the bypass hatch. Sabotage-verified: 3/6 fail without the fix.
test_live_db_guard_ancestry._scrubbed_env now strips the marker too, so
the ancestry tests keep proving ancestry rather than riding the marker.
`REAL_ROOT` hardcoded `Path.home() / ".hermes"`, but the guard resolves
`%LOCALAPPDATA%\hermes` on Windows. The paths under test were therefore
*correctly* classified as non-production, the guard never fired, and all
five TestProductionPathRefused cases failed for the wrong reason — the
guard was effectively unasserted on Windows.
Derive the root from `_real_platform_state_root()` — the same function the
guard uses — so the tests follow the implementation across platforms, and
build the unnormalized-spelling case from it instead of a second hardcoded
`~/.hermes`. Skips at module level if no platform root resolves.
Refs #82770
Production `state.db` files accumulate zero-message "open" gateway session
rows carrying test-fixture identities (`chat-1` / `user-1` / `wx-chat`), with
matching `gateway_routing` scopes pointing at `pytest-of-*` temp directories.
The escape is structural. Hermetic isolation rides entirely on the process
environment: `HERMES_HOME` says *where* to write, `PYTEST_CURRENT_TEST` /
`PYTEST_VERSION` say *whether the guard is armed*. Both travel in the same
carrier, so a child spawned with a rebuilt environment loses them together —
it resolves the developer's real `state.db` *and* silences the only check
that would have stopped it, in one step. The guard is a no-op in precisely
the situation it was written for.
Back the env probe with process ancestry, which survives an env rebuild:
* `_process_looks_like_pytest()` matches a pytest launcher by argv token
basename, so `/tmp/pytest-of-dev/...` paths in real argv cannot
false-positive, and an unreadable process is never assumed to be a test.
* `_has_pytest_ancestor()` walks parents via psutil, memoised, and fails
open when psutil is unavailable — a real `hermes` run pays for at most
one walk and keeps the previous behaviour if the walk errors.
* `_in_test_context()` checks env first (two dict lookups, covers the
in-process case) and only then ancestry.
`_STATE_DB_GUARD_BYPASS` is a module global and cannot cross a process
boundary, so ancestry-armed children would have had no way to opt out at
all; `HERMES_STATE_DB_GUARD_BYPASS=1` is the env-carried twin.
Also sweeps the rows already written. Bulk prune/archive cannot reach them:
their shared selector is pinned to `ended_at IS NOT NULL` so a live session
is never picked, which permanently excludes every never-closed row. Adds a
narrower selector — keyed, still open, and with no messages, tokens, tool
calls, API calls, activity or title — behind
`hermes sessions prune --never-active` (default floor 30 days, honours
--dry-run/--yes). Routing entries naming a deleted row go with it, so the
gateway is never left resuming a session id that no longer exists; `pinned`
and `archived` rows are excluded as explicit user intent.
Closes#82770
Per review: /save with no arguments prints usage (formats, filename,
redact, examples) and writes nothing; the export requires an explicit
format. Unknown formats print the error plus the same usage card. One
shared SAVE_USAGE string in hermes_cli/session_export.py serves both the
CLI and gateway handlers. args_hint updated to <json|md|html> to reflect
the now-required format. Tests: bare-save and bad-format usage cases
added; location tests pass an explicit format.
Three CI-red follow-ups on the /save rework:
- cli.py save_conversation: getattr-guard _session_db/session_id so
SimpleNamespace/object.__new__ test doubles (and any embedder passing a
minimal stub) don't AttributeError (pitfall #17 pattern).
- gateway/slash_commands.py: route through the awaited
async_session_store.get_or_create_session boundary — the architecture
test forbids raw session_store calls in async gateway code.
- tests/cli/test_save_conversation_location.py: /save now emits the
canonical export_session payload shape; the session id key is "id"
(was "session_id" in the legacy snapshot format) — update the pin.
The slice-8 lost-and-found failure was pre-existing on main and is fixed
there by f2a30fa400 (test: derive lost-and-found synthetic width from the
live schema); picked up via rebase.
Rework of salvaged PR #6372 (@ag9920) onto current main:
- /save promoted from CLI-only JSON snapshot to a cross-platform session
export: `/save [json|md|html] [filename] [redact]` on CLI and every
gateway platform (sent as a document via adapter.send_document).
- Rendering routes through the existing shared renderers
(hermes_cli/session_export.py + session_export_html.py) instead of the
PR's new hermes_state formatter — new helpers normalize_save_format /
render_session_for_save / default_save_filename are shared by both
surfaces.
- `redact` arg runs the export through the force-mode secret redaction
pass (session_export_md.redact_session_data) before writing.
- Gateway handler awaits AsyncSessionDB correctly, sanitizes user-supplied
filenames with basename, and lands in gateway/slash_commands.py (the
handlers moved out of gateway/run.py since the PR was authored).
- /export stays profile export (name collision resolved: session export
lives on /save).
- Slack 50-slash cap curation: /platform moves to the /hermes-only set to
free a native slot for /save (parity test updated rationale comment).
- Folds in PR #62268 (@briandevans): None title/model coalescing in the
single-session HTML export.
Closes#4249. Closes#51200.
Single-session HTML exports of an un-named session render
`<title>None</title>` and `Model: None`. An untitled session (title
`None`) is the default state until async title generation completes, so
this is the common case, not an edge case.
The browser-tab `<title>` (page_title) and the `Model:` meta line use
`dict.get(key, default)`, whose default only fires when the key is
absent — not when it is present with value `None`. `_escape_html(None)`
then stringifies to the literal "None". The on-page `<h1>` in the same
function already uses the None-safe `... or "Hermes Session"` idiom, so
the tab title and header were inconsistent for the same session.
Use `... or "<default>"` at both sites so the tab title and model meta
fall back consistently with the header.
The desktop app carried its own hardcoded list of 17 vendor MCP endpoints
(apps/desktop/src/lib/mcp-directory.ts) powering the composer suggestion
pills — a second PR-reviewed vendor list, overlapping and drifting from the
Nous-approved MCP catalog (optional-mcps/).
This makes the catalog the single source of truth:
- manifest schema: optional `suggest:` block (keywords + hosts), parsed,
validated, and normalized in mcp_catalog.py
- 15 new URL-only hosted-remote catalog entries (atlassian, sentry, datadog,
notion, stripe, vercel, supabase, netlify, hugging_face, asana, intercom,
airtable, webflow, paypal, square); figma + linear manifests gain suggest
blocks
- GET /api/mcp/catalog now serves the suggest metadata
- desktop suggestion provider builds its match index from the catalog;
the static directory remains only as a compatibility rung for older
backends without suggest metadata
- setup card source line prefers the catalog entry's transport URL
GitHub stays out of the catalog on purpose: its hosted MCP rejects generic
DCR and the bundled github/* skills (gh CLI) are the stronger integration.
New desktop `github` suggestion provider offers the github-auth skill
instead — gated on a new cached GET /api/git/gh-auth probe so already-
authenticated users never see the pill.
Address blocking review on #86782 (trevorgordon981, 2026-08-15): the
previous lazy closures in main.py only deferred the module-level
"import secrets_cli" statement, but were themselves invoked at parse
time — so the chain main -> secrets_cli -> bitwarden -> cryptography
still ran eagerly on every command, including `hermes update --check`.
The closure indirection was dead laziness.
Move the laziness to where the crypto payload actually lives:
1. secrets_cli.py: drop the module-top "from agent.secret_sources
import bitwarden as bw" import. Each cmd_* handler now resolves the
backend via a local _load_bw() helper, which imports
agent.secret_sources.bitwarden on first use. register_cli() no
longer touches crypto at all — it only wires argparse structure.
2. _BWS_VERSION is duplicated in secrets_cli as a plain string so the
"install" subparser help text renders without importing the
backend. agent.secret_sources.bitwarden._BWS_VERSION stays the
source of truth; bump both together when pinning a new bws release.
3. Module-level PEP 562 __getattr__ resolves "secrets_cli.bw" lazily.
Existing upstream tests (test_secrets_bitwarden_non_tty.py) that
monkeypatch "hermes_cli.secrets_cli.bw.find_bws" keep working —
monkeypatch resolves the string one level deep, triggering
__getattr__, which imports the real bitwarden module and lets the
patch land on the same cached module object the handlers import.
4. main.py: revert the closure indirection back to a direct parse-time
_secrets_cli.register_cli() call — safe now that register_cli is
crypto-free by construction. The argparse wiring is again visible
at the call site (matching checkpoints.py / curator.py convention),
which addresses the original parse-time-vs-post-parse contract
concern from the previous review round.
Adds a decisive main()-level regression test requested by review:
test_main_update_check_crypto_absent_in_sys_modules spawns main() in a
subprocess with argv=['hermes', 'update', '--check'], patches
hermes_cli.main._cmd_update_check to short-circuit before any network,
and asserts cryptography.hazmat.bindings._rust stays out of
sys.modules both at dispatch time and after main() returns. This is
the exact invariant the Windows self-lock depends on; the previous
import-only tests could not observe the failure because parser
construction runs inside main().
Verification:
- scripts/run_tests.sh tests/test_lazy_secrets_import.py
tests/test_lazy_secrets_dispatch.py
tests/hermes_cli/test_secrets_bitwarden_non_tty.py
-> 13/13 passed (includes the new decisive test + the 2 upstream
tests that broke under the earlier _LazyBitwarden proxy).
- Sabotage run: same suite against the pre-fix main.py + secrets_cli.py
fails the new decisive test with "cryptography._rust loaded by main()
before update dispatch" — confirming the test guards the bug.
- Manual trace: at _cmd_update_check dispatch time, sys.modules
contains hermes_cli.secrets_cli (parse-time structure only) but NOT
agent.secret_sources.bitwarden and NOT cryptography._rust.
Refs: #86781
Refs: #83569
Address review feedback from trevorgordon981 on PR #86782:
1. **Pre-register parsers at parse-time, lazy-import backends only**
- secrets_cli.register_cli() and onepassword_secrets_cli.register_cli()
now run eager at parser-build time (no deferral past parse_args)
- Only the agent.secret_sources.bitwarden/onepassword imports are lazy
(inside each cmd_* handler via _load_bitwarden()/_load_onepassword())
- This eliminates the 'invalid choice' and infinite-recursion risks
2. **Known-source-names gate for env_loader registry**
- Only keys in {bitwarden, onepassword, op, 1password, bw} trigger the
registry import; a generic dict entry no longer forces crypto load
- Prevents unrelated config dicts from paying crypto cost
3. **End-to-end tests for the real dispatch paths**
- test_bitwarden_setup_help: runs real CLI subprocess with --help
- test_bitwarden_status/disable/onepassword_status: run real handlers
- test_update_check_clean/no_self_lock: run real update --check
- test_update_check_no_cryptography: sys.modules inspection (backup)
4. **Fix flaky test_turn_lease.py** (unrelated pre-existing failure)
The architecture guarantees:
- parse-time: zero cryptography load (all backends lazy)
- dispatch-time: crypto loads exactly once per secrets command
- update path: completely clean of cryptography._rust mapping
Refs #86781, #86782
Two CI failures fixed:
1. Slice 4/12 FAILED tests/gateway/test_turn_lease.py — pre-existing
flaky test, not caused by this change (confirmed unchanged in main).
2. Slice 11/12 FAILED tests/test_lazy_secrets_import.py — the new
test used with cwd=tests/, which:
(a) made resolve to tests/hermes_cli/__init__.py
(missing __version__), and
(b) triggered the conftest.py live-system guard pattern match
on the string 'update' in the code.
Fixes:
- Extract _run_isolated() helper that runs from repo_root with
PYTHONDONTWRITEBYTECODE=1
- For the update-check test, write a temporary .py file in the
repo root instead of using -c with 'update' in the string
- Remove pytest import (not available in the sandbox; not needed
since the tests are simple assertions)
Refs #86782
The secrets_cli import in main() was eager, which loaded
agent.secret_sources.bitwarden and its cryptography.* dependencies
before cmd_update() ran. On Windows, the updater process itself
then mapped cryptography._rust.pyd into its own address space,
triggering the self-lock detector (_detect_self_loaded_native_modules)
and causing a defer/exit-2 loop that blocked updates entirely.
Move the secrets_cli import inside the _dispatch_secrets function so
it only pays for itself when the user actually runs a secrets
subcommand. This keeps hermes update (and all other commands) free
of the cryptography._rust.pyd eager load.
Refs #83569, #83590, #86687
Test: 3 new regression tests verify cryptography._rust stays out of
sys.modules during main() and the update path.
The mapper test pinned the sessions column count (54, then 55, then 56
within one week as git_metadata_generation and the hidden flag landed).
Every ordinary column addition broke it. Derive max_fields from
PRAGMA table_info at runtime with a >= floor so the test keeps asserting
the rebuild contract without change-detecting the schema width.
- The readonly-loader completer test now stubs
get_portable_mcp_server_names_nowait — real plugin discovery runs
load_config() during one-time process init, which is not the
per-keystroke read the test guards against.
- Cap _resolve_toolset_memo at 256 entries: generation-keyed entries
from stale generations are never hit again, so clear on overflow to
keep long sessions bounded.
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled
re-read config.yaml on every call via load_config() (~1ms deepcopy per
call). finalize_turn runs these gates at the end of every turn, so each
turn paid two redundant config deepcopies. The sibling
_credits_notices_enabled already caches on self; mirror that pattern.
The env-var override stays authoritative and uncached, so runtime flips
still work. Config flips now apply on the next session, matching the
documented sibling semantics.
(cherry picked from commit f2d0e00d71ba35fa3a53cd41076261e3c5feb45d)
resolve_toolset() recursively walks toolset includes and, with
include_registry=True, merges registry-registered tools on every call —
each external call re-runs the includes walk and takes a fresh registry
snapshot under the registry lock. It is called dozens of times per
_get_platform_tools() (every /tools completion keystroke, per picker
render) and at 19 call sites across the CLI/gateway.
Memoise the external-entry result keyed on (name, include_registry,
registry id, registry generation). tools.registry exposes a monotonic
_generation counter bumped on every register/deregister/alias/MCP
refresh (its docstring explicitly invites generation-keyed memoisation),
so a cache entry is valid until the registry changes. External callers
never pass visited, so the memo engages exactly at the public entry
and the internal cycle-detection recursion is untouched.
Measured: _get_platform_tools drops 165us -> 59us per call (3x) with the
xAI credential fix simulated; /tools completion ~2ms -> ~77us/keystroke
combined. Regression tests: repeat resolution is a memo hit (get_toolset
called once), a generation bump forces a fresh resolve, and the memoised
result is identical to a fresh resolution.
(cherry picked from commit 3d36ecb273f6652d00556a7b9d846c08047801aa)