Commit Graph

14240 Commits

Author SHA1 Message Date
Teknium bb4f680f22 fix(browser): named browser_exec sessions compose with every backend
session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.

Now a named session composes with whatever browser source is configured:

- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
  pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
  too; previously a named daemon ignored it and fell back to scanning
  local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
  keys its own provider browser via the shared _get_session_info cache
  (bu-named-<name>), so each name gets its own cloud browser, the same
  name reuses one across calls and tasks, and unnamed calls keep the
  per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
  (provider resolution would double-session and double-bill).

Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.

E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
2026-08-15 04:05:57 -07:00
Teknium bfd9cef389 fix(worktree): deepen shallow clones so worktree cleanup can verify push state
The installer clones with --depth 1, so every default install is shallow.
In a shallow repo, an older worktree HEAD (a past snapshot of main) is
disconnected from current origin/main by the shallow boundary, so
'git log HEAD --not --remotes' misreports thousands of already-public
commits as unpushed. The fail-safe unpushed guard then preserves every
aged 'hermes -w' worktree forever, and the git-cherry squash-merge
escape hatch never rescues them (22k 'ahead' >> max_ahead=20).
Real incident: 21 of 25 hermes-* worktrees stuck on one install.

Fix at the root, one owner:
- _deepen_shallow_repo(): one-time blobless unshallow
  (fetch --unshallow --filter=blob:none; plain --unshallow fallback)
  run from the background startup pruner thread before classification,
  so history verdicts become correct and the backlog self-clears on the
  next 'hermes -w' startup. Fail-soft offline: keep preserving.
- _cleanup_worktree(): when the unpushed verdict comes from a shallow
  clone, say 'Shallow clone — cannot verify push state' instead of the
  misleading 'has unpushed commits' message.
- Document the shallow caveat on _worktree_has_unpushed_commits (the
  primitive stays conservative on purpose).

Tests: real shallow clone over file:// reproducing the disconnect shape,
covering detection, deepen+verdict flip, pruner E2E reap, offline
fail-soft preserve, full-clone noop, and genuine-unpushed-work survival.
Sabotage-verified: the E2E test fails with the deepen call disabled.
2026-08-15 04:02:07 -07:00
fangliquanflq d528f4da00 fix(updater): defer native parser imports after recovery 2026-08-15 03:54:07 -07:00
fangliquanflq 97051703ae fix(updater): keep recovered retries native-safe 2026-08-15 03:54:07 -07:00
fangliquanflq 8f5e5e49a2 fix(updater): recover deferred installs on update retry 2026-08-15 03:54:07 -07:00
Teknium efdd715f8a fix(gateway): resolve the profile-aware scheduled-task name in the reaper guard
Follow-up to the salvaged #86823: the guard queried a hardcoded
"HermesGateway" task, but `hermes gateway install` registers
Hermes_Gateway (Hermes_Gateway_<profile> for named profiles) via
gateway_windows.get_task_name(). Query that name so the supervisor
guard is active on standard installs; fall back to the default literal
if the module import fails. Test now asserts the profile-aware name is
what reaches the task-state query.

Also corrects the cherry-picked commit's placeholder author email to
the contributor's GitHub noreply address.
2026-08-15 03:53:51 -07:00
hutao562 2795b2ab9f fix(gateway): treat a Running HermesGateway scheduled task as a supervisor
The orphan-reap sweep (_reap_unsupervised_gateway_orphans) must not kill a
gateway that Windows Task Scheduler is actively managing. The existing
services.exe parent-chain backstop fails open: when the Task-launched conhost
bootstrap has already exited, Windows does not reparent the gateway, the
chain breaks, and the supervised gateway is treated as an orphan. The reaper
then writes the planned-stop marker, the gateway exits cleanly with code 0,
and the scheduler never restarts it (RestartCount only fires on non-zero
exit) — silently killing A2A/messaging on every desktop-app launch.

Querying the task's own state is the authoritative signal and closes the gap
without depending on process ancestry: if HermesGateway is Running, skip the
reap entirely. Uses PowerShell Get-ScheduledTask (English State enum,
locale-stable) rather than schtasks (localized output + codepage mangling).
2026-08-15 03:53:51 -07:00
ygd58 22e638db7c fix(cron): reap stale execution claims before a one-shot hermes cron run dispatch
Fixes #86721.

`hermes cron run <job_id>` (a one-shot CLI invocation) dispatches
manual runs via the same background-delegation path as an agent's
`cronjob(action='run')` tool call (tools/cronjob_tools.py's
_try_dispatch_background_run -> dispatch_async_delegation(role=
"cron_run", runner=_runner, ...)). The runner thread lives in the
calling process's shared daemon executor. When the one-shot process
exits right after printing "Triggered job: ...", the in-flight runner
dies mid-execution, leaving its cron/executions.db row permanently
stuck at status='claimed' -- every subsequent `hermes cron run` on the
same job then reports "Ran now: failed" because of the still-claimed
row.

cron/executions.py already has the exact self-heal this needs:
recover_interrupted_executions() correctly identifies and reclassifies
'claimed'/'running' rows whose owner process has provably exited
(_owner_is_live checks PID existence AND matches process start-time,
so a reused PID isn't mistaken for the original live owner) to
'unknown', unblocking the job for a fresh claim. But it was only ever
called once, at the long-lived scheduler ticker's own startup
(cron/scheduler.py:379's self.recover_interrupted()) -- a one-shot CLI
invocation has no equivalent "startup" moment of its own, so this
self-heal never ran for it.

Added a call to recover_interrupted_executions() at the top of
_try_dispatch_background_run, right after the async-delivery-supported
gate and before any claim attempt for the current job -- mirroring
exactly what the long-lived scheduler already does at its own
startup, just triggered per one-shot invocation instead of once at
daemon startup. Wrapped in try/except: pass (best-effort; a failure
here must not block the actual dispatch this function exists for).

Traced (but did not attempt to fix) the deeper "why does the runner
die with the process at all" question -- that's the harder problem
options 1/2 in the issue describe (route to the persistent scheduler,
or block the one-shot process until completion). This fix addresses
the more urgent, more clearly-scoped symptom: a stranded stale claim
permanently blocking ALL future manual runs of the affected job, which
is option 3 from the issue and the one with an existing, already-
correct implementation just needing to be wired into this call site.

Added 3 regression tests to a new file, following the established
real-subprocess dead-owner pattern already used in
tests/cron/test_execution_ledger.py (a genuinely-dead PID, not a
mock, matching the real-world failure mode exactly): a sanity test
confirming the stale claim sits unrecovered without the fix; a direct
test of recover_interrupted_executions() reaping such a claim; and a
unit test on _try_dispatch_background_run itself confirming recovery
is called before any claim attempt. Verified as a genuine regression
by reverting the fix and confirming the unit test fails with recovery
never having been called.

35/35 pass across the new test file plus tests/cron/test_execution_ledger.py
and tests/tools/test_cronjob_run_background.py (no regression).
2026-08-15 03:53:42 -07:00
Teknium 4ef56cef4c fix(auth): bound the no-TTL key_cmd token cache; docs for key_cmd
Follow-ups on the #85006 salvage:

- A key_cmd token with no advertised expiry was cached for the life of the
  process. The "refresh on 401" contract it relied on has no implementation
  (SDK retries cover 429/5xx only), so an expired no-TTL token would 401
  every request until restart. Cache on a bounded 15-minute window instead;
  helpers that want a longer cache can advertise their real expiry.
- Test for the no-TTL path updated to pin the bounded-window contract;
  the remint test's $RANDOM (bash-only, empty under dash) replaced with
  date +%s%N so it exercises remint under any /bin/sh.
- website/docs/integrations/providers.md: document key_cmd in the named
  custom providers section (contract, precedence, secrets.command contrast).
2026-08-15 03:16:21 -07:00
LordMelkor 6efab28726 feat(auth): add key_cmd credential source for custom providers
Custom providers could only authenticate from a static credential (inline
api_key or a key_env env var). Enterprise gateways -- SSO/OIDC brokers, cloud
IAM, internal auth proxies -- issue short-lived bearers instead, so a value
copied into .env is stale within the hour: long sessions start returning 401s
and the user has to restart or run an external cron that rewrites .env.

The existing `secrets.command` source does not cover this: it runs once per
process at startup (subsequent calls are no-ops by design), so it cannot
re-mint a credential mid-session.

Add providers.<name>.key_cmd: a command that prints a token, wrapped at
resolution in a zero-argument callable. Both wire clients already accept a
callable api_key and invoke it per request (the Entra ID path established
this), so chat_completions, codex_responses and anthropic_messages all work
unchanged and always send a fresh credential. The callable also routes the
Anthropic client through its per-request Authorization hook, which is what
OAuth-gated gateway routes require -- so no per-vendor auth wiring is needed
anywhere in core.

- cached until shortly before the advertised expiry (60s leeway), so the
  helper runs about once per token lifetime rather than once per request
- expiry is read from the OAuth 2.0 relative `expires_in` when present, and
  otherwise from an absolute ISO 8601 deadline (`expiry`, `expiresOn`), which
  is what CLI token helpers commonly print. Reading only `expires_in` treated
  those helpers as advertising no TTL at all, cached their token for the life
  of the process, and returned 401 on every request once the real deadline
  passed. ISO parsing reuses hermes_cli.auth._parse_iso_timestamp rather than
  adding another datetime parser.
- no synthetic expiry: when no TTL is advertised, or the advertised one is
  unparseable or already past, the token is used and refreshed on 401 instead
  of re-minted on an invented schedule
- stdout contract matches OAuth 2.0 token endpoints and existing agent
  helpers (bare token or JSON access_token/expires_in); multi-line output is
  rejected rather than guessed at, so a misconfigured helper surfaces as a
  clear error instead of a corrupt-credential 401
- precedence: explicit --api-key still wins; otherwise key_cmd beats a
  static api_key/key_env on the same entry
- failures never include the helper's output (may hold a partial token) or
  the command string (may embed a client secret)

Resolution happens on two paths. agent/auxiliary_client.py resolves named
custom providers itself rather than calling _resolve_named_custom_runtime, so
key_cmd is honoured in both: wiring only the runtime resolver leaves the main
agent turn working while every auxiliary call (title generation, compression,
vision, embedding) falls back to the no-key-required placeholder and 401s.
Precedence is identical on both paths, so one config entry cannot yield two
different credentials depending on which resolver the caller reached.

Closes #84162

Signed-off-by: LordMelkor <kray@block.xyz>
2026-08-15 03:16:21 -07:00
poisdahl d65b7d0760 Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final 2026-08-15 12:05:28 +02:00
poisdahl 3f075d41dd Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tui_gateway/methods_prompt.py
#	tui_gateway/server.py
2026-08-15 12:05:14 +02:00
Teknium 7a16840add fix(compression): couple pruned-skill reload instruction to the preserved todo snapshot
Compaction re-injects the todo list verbatim (TODO_INJECTION_HEADER +
TodoStore.format_for_injection) while skill instructions are pruned down
to [SKILL_PRUNED: ...] markers — the imperative crosses the boundary
without the policy that governed it, and the agent keeps executing
preserved tasks with the guidance deleted (#84718's T6 pattern).

Close the retention asymmetry at the injection site: when the compressed
transcript carries [SKILL_PRUNED: ...] markers AND a todo snapshot is
being re-injected, append a bounded reload notice to the snapshot naming
each pruned skill with its exact skill_view() reload call, plus a
one-line instruction to re-check that preserved tasks are still
justified. Skill guidance recovery now travels in the SAME boundary
artifact as the imperative — same message, same stale-snapshot strip
lifecycle, so repeated compactions refresh rather than accumulate.

Properties:
- deterministic: derived only from the compressed transcript (same input,
  same bytes) — no per-turn nondeterminism in the rebuilt prompt
- zero recurring cost when nothing was pruned (clean sessions unchanged)
- bounded: shares _MAX_PRUNED_SKILL_MARKERS with the summary re-injection
  cap; the notice text never contains the canonical marker prefix, so it
  can never feed the marker extractor at the next boundary
- rides after TODO_INJECTION_HEADER, so _strip_stale_todo_snapshot
  removes snapshot + notice together and the synthetic-row classifier
  (_is_synthetic_compression_user_turn) is unaffected

Tests: tests/agent/test_skill_todo_retention_parity.py — unit contract of
the notice builder (naming, dedup/order, cap, determinism, no marker
self-feed) and behavioral compaction runs through the real
_compress_context path (notice travels with the snapshot, absent when
nothing pruned, synthetic-row classification unbroken, strip lifecycle
across repeated boundaries). Sabotage-verified: disabling the append
flips the 3 behavioral tests red.

Part of #84718
2026-08-15 02:45:50 -07:00
Teknium 8fdfe04371 fix(update): probe gateway loop liveness before drain; bounded escalation for wedged gateways
A gateway whose asyncio event loop is stalled (e.g. an in-loop
compression pass, #72707) cannot process SIGTERM/SIGUSR1 shutdown.
The updater's drain wait then burned the full 180s budget, warned
"Gateway PID X still running after 180.0s — restart may fail", and
`hermes update` could deadlock behind the wedged process — the user
cannot update their way out of the stall.

Fix: before any drain wait, read the loop-liveness heartbeat file the
gateway rewrites every 30s (#66892). Classification:

- alive (fresh heartbeat): busy-but-alive loop — take the normal
  graceful drain, honoring the in-flight cron drain floor (#86684).
- wedged (heartbeat for this PID stale >90s = 3 missed beats): the
  loop is provably dead; drain is pointless. Bounded escalation:
  SIGTERM + 5s grace, then SIGKILL + 5s wait, then proceed (~10s
  worst case, far under the 180s drain budget).
- unknown (missing/corrupt file, PID mismatch): never escalate on
  ambiguity — full drain path.

Wired into launchd_restart, systemd_restart, and both updater
gateway-shutdown sites (systemd unit drain + manual profile
gateways). The probe is a local stat + JSON read (well inside the
10s query tier of the subprocess timeout tiering).

The cron drain floor from #86684 is bypassed ONLY when the loop is
provably dead — a merely busy gateway still refreshes its heartbeat
and keeps the full drain budget.

Root cause of the loop stall itself (compression blocking the loop)
is #72707 territory and deliberately out of scope here.

Fixes #81642
2026-08-15 02:45:27 -07:00
Teknium be083358b2 feat(tui_gateway): INFO 'prompt accepted / turn finished' records on the Desktop/TUI turn path
Part of #86647.

During the #79278 persistent-mute investigation the decisive evidence was
an absence: a Desktop request left no INFO record in agent.log OR
gateway.log ("832 platform=webhook, 194 platform=telegram, 0
platform=desktop"), so the muted 13:15-13:19 window — 12 non-idempotent
Qdrant snapshots, zero results returned — was structurally
indistinguishable from a request that never arrived. The issue calls out
fixing this observability gap as the first actionable step.

This adds the two INFO records to _run_prompt_submit, the single choke
point every Desktop/TUI turn passes through (user submits, queued
prompts, auto-continue, goal follow-ups, watch upgrades):

- "tui prompt accepted": emitted before the turn thread starts, carrying
  the UI session id, the gateway session_key, and the agent's live
  session_id — the id triple a rotation-mute trace needs, since
  compression rotates agent.session_id independently of the other two.
  No prompt content is logged (length only).
- "tui turn finished": emitted in the turn's finally on every path
  (success, returned error, exception, interrupt), re-reading
  agent.session_id so a mid-turn compression rotation shows up as an
  accepted/finished pair with different agent ids. A missing finished
  record now positively identifies a turn thread that died before its
  finally.

Placement follows @Adolanium's note on the issue: in _run_prompt_submit,
logging sid + session_key + agent.session_id, NOT another platform= line
in gateway.log (Desktop does not use the messaging gateway).

tui_gateway is under COMPONENT_PREFIXES["gui"], so the records land in
agent.log (root catch-all) and gui.log when running under the dashboard.

Tests (tests/tui_gateway/test_prompt_accept_logging.py): accepted+finished
pair on success with the full id triple and no prompt content leaked;
mid-turn rotation visible as differing agent ids across the pair;
finished record fires on the exception path and the returned-error path.
Sabotage-verified: removing the accepted record fails the suite.
2026-08-15 02:45:03 -07:00
Teknium e729055a56 test(delegate): regression for #86632 — cron sync fallback must return after child completes
End-to-end #86632 reproduction: a real AIAgent child (mocked LLM) with the
post-turn skill-review trigger armed, dispatched through
delegate_task(background=True) on a session runtime where async delivery is
unsupported and no origin session id is bound (cron, post-#66617) — forcing
the synchronous fallback. Asserts (1) delegate_task returns the child's
result, and (2) the automatic background-review fork never spawns inside the
delegated child (the wedge site: the fork replayed the conversation on the
child's finalize path, and _child_future.result(timeout=None) never returned;
heartbeat went stale after 15 idle cycles and the cron watchdog killed the
job).

Verified RED on pre-fix main (fork spawns and wedges), GREEN with the
_delegate_depth guard in AIAgent._spawn_background_review.

Fixes #86632
2026-08-15 02:44:38 -07:00
PRATHAMESH75 ece3bc7d52 fix(agent): skip the automatic background review inside delegation subagents
The post-turn background review fork (`agent/background_review.py`) inherits
the parent agent's live runtime by default. That is a cost win when the parent
IS the main chat model (warm prompt cache, cheap), but the fork also fires
inside delegation subagents, where it inherits the *subagent's* model. When a
subagent runs a premium delegation model, the review silently replays the whole
conversation and emits skill/memory-update output at premium rates, with
nothing in the log or config flagging it (#85859).

Subagents are already barred from writing shared MEMORY.md
(`DELEGATE_BLOCKED_TOOLS`) and are spawned with `skip_memory=True`, so an
automatic review here has little to persist. Guard `_spawn_background_review`
(the single choke point both the turn-finalizer and codex-runtime callers pass
through) to return early when `_delegate_depth > 0`. An explicit `/refine`
(`focus` set) is a deliberate user request and still runs; the top-level path
is unchanged.

Fixes #85859
2026-08-15 02:44:38 -07:00
Teknium 0fc2a10d82 fix(cron): stop one-shot CLI cron run from orphaning the job; reap dead-owner claims on tick
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.

Two-part fix:

1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
   channel stateless (scoped ContextVar set/reset around the call) before
   invoking the cron API, so `async_delivery_supported()` gates off
   `_try_dispatch_background_run` and the run executes synchronously to
   completion in the CLI process — the same behavior `hermes -z` already
   gets via declare_stateless_channel().

2. cron/scheduler.py: tick() now periodically invokes
   recover_interrupted_executions() (previously only run at scheduler
   startup), so execution rows whose exact owner process is provably dead
   (pid + process start time check in _owner_is_live) are reaped to
   'unknown' by the long-lived gateway ticker without a restart.
   Throttled to once per 300s so idle 60s ticks don't pay a ledger
   connection every cycle.

Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.

Fixes #86721
2026-08-15 02:44:14 -07:00
Teknium 8711f2c005 fix(update): make the self-lock deferral honest — fire only when the swap is at risk (#86735, #86780, #86781)
The #86687 self-lock preflight fired on every Windows `hermes update`:
bitwarden.py's module-level cryptography import (fixed in #86782 /
#86826-class change) meant cryptography._rust was ALWAYS mapped by the
time the preflight ran, so the update exited 2 before even fetching and
looped forever — including the Desktop in-app update (#86780).

Two structural fixes so the guard can never re-brick the flow it protects:

1. Version-gated detection: _detect_self_loaded_native_modules() now
   consults _dependency_sync_would_rewrite(dist) — installed version vs
   the on-disk pyproject pins (base deps + all extras, env markers
   honored). A loaded module whose distribution the sync will not touch
   is no lock risk and is not reported. Unknown → fail closed.

2. Relocated deferral: the check no longer runs pre-fetch. It runs via
   _abort_dependency_sync_if_self_locked() immediately before each venv
   rewrite (git-path dep sync, ZIP-path dep sync, current-checkout venv
   repair) — AFTER the code swap. A deferral now leaves the user on NEW
   code with only the dependency install pending (completed by the next
   launch's marker recovery), instead of stranding them on the old
   checkout in an exit-2 loop.

PyYAML's _yaml extension (loaded by every CLI process) joins the
registry — with version gating it is now safe to list.

Tests: version-gate unit coverage (no-change skip, stale pin, missing
dist, extras, markers, fail-closed None), deferral wiring (marker +
gateway resume + exit 2), placement guards (no detector call pre-fetch;
guard present at git/ZIP sync), and subprocess-verified import hygiene
(import hermes_cli.main and the update --check dispatch never load
cryptography._rust).

Follow-up to #86687 (Halldrix's #83590 salvage — the preflight's intent
stands as defence-in-depth; this makes it fire only when true).

Fixes #86735
Fixes #86780
Fixes #86781
2026-08-15 02:26:38 -07:00
Teknium 0d00ebef79 fix(state): bound the state.db repair loop and stop dead-backup accumulation (#86747)
A corruption class the repair strategies cannot heal (b-tree page
damage) failed repair_state_db_schema on every process start, forever:
_claim_repair_attempt's in-memory set only bounds one process, so each
restart re-ran the full surgery AND took a fresh ~900MB forensic backup
of the same damaged bytes — 105 attempts / 89GB of dead
state.db.malformed-backup-* files over 11 days in the reporting install.

Three bounded behaviors, all sidecar-file based (no schema changes):

1. Persistent attempt ledger (<db>.repair-attempts.json): after 3 failed
   repair passes against the same file fingerprint (size + mtime_ns),
   repair_state_db_schema refuses with a terminal, actionable error
   (restore a backup / `sqlite3 state.db ".recover"` / delete the ledger
   to force a retry) instead of re-running surgery. Success clears the
   ledger; a replaced or restored file re-keys it and gets fresh
   attempts. Missing/corrupt ledger fails open (never blocks a first
   repair).

2. Backup dedupe: _backup_db_file reuses the newest existing forensic
   backup when it is byte-identical to the damaged file (size+mtime
   match, preserved by copy2) instead of copying another ~900MB.

3. Retention cap: only the 3 newest malformed-backup copies (plus
   sidecars) are kept; older ones are pruned after each new backup.
   Also fixes a same-second timestamp collision that silently
   overwrote an earlier forensic copy.

Tests cover ledger accumulation, terminal refusal (surgery not called,
no new backup), budget reset on file change, success-clears-ledger,
corrupt-ledger tolerance, dedupe, distinct-state backups, retention
prune incl. sidecars, and the end-to-end one-backup invariant.

Fixes #86747
2026-08-15 02:22:58 -07:00
Teknium 62d3ef683c fix(tests): subprocess-surviving isolation marker closes the #82770 fixture escape
The hermetic conftest now exports HERMES_TEST_ISOLATION (value = the tmp
isolation root) before any test module imports, and re-pins it per test in
_hermetic_environment. hermes_state._running_under_pytest() honors the
marker as a test-context signal alongside PYTEST_CURRENT_TEST /
PYTEST_VERSION.

Why a third layer: PYTEST_* belongs to pytest, and tests that spawn
children routinely rebuild the child env and strip it ("the subprocess
must look like a real CLI" — tests/cli/test_exit_watchdog_signal_arm.py,
tests/hermes_cli/test_config_loader_e2e.py do exactly this on purpose).
Such a child loses the HERMES_HOME redirect and the guard's arming signal
in one step, which is how 700+ zero-message fixture rows (dm:123, chat-1,
wx-chat, ...) landed in a developer's production state.db. The marker is
OURS: stripping it is never required to make a child "look real" (no
production code branches on it except the guard), it inherits by default,
and children that genuinely need a real DB use the sanctioned
HERMES_STATE_DB_GUARD_BYPASS=1 hatch instead.

The ancestry-walk layer (previous commits) stays: it covers children whose
env was rebuilt from a completely empty dict. The marker layer covers the
common **os.environ-derived rebuilds cheaply (one dict lookup, no psutil),
and — unlike ancestry — also covers detached/daemonized children that
escape the process tree.

tests/hermes_state/test_isolation_marker_env.py pins: the conftest export,
the marker-alone arming, the rebuilt-env child refusing the production
path, and the bypass hatch. Sabotage-verified: 3/6 fail without the fix.
test_live_db_guard_ancestry._scrubbed_env now strips the marker too, so
the ancestry tests keep proving ancestry rather than riding the marker.
2026-08-15 02:20:13 -07:00
joaomarcos 3f026fe210 test(state): assert the live-DB guard against the real platform root
`REAL_ROOT` hardcoded `Path.home() / ".hermes"`, but the guard resolves
`%LOCALAPPDATA%\hermes` on Windows. The paths under test were therefore
*correctly* classified as non-production, the guard never fired, and all
five TestProductionPathRefused cases failed for the wrong reason — the
guard was effectively unasserted on Windows.

Derive the root from `_real_platform_state_root()` — the same function the
guard uses — so the tests follow the implementation across platforms, and
build the unnormalized-spelling case from it instead of a second hardcoded
`~/.hermes`. Skips at module level if no platform root resolves.

Refs #82770
2026-08-15 02:20:13 -07:00
joaomarcos 174ce8770d fix(state): arm the live-DB guard by process ancestry, not env alone
Production `state.db` files accumulate zero-message "open" gateway session
rows carrying test-fixture identities (`chat-1` / `user-1` / `wx-chat`), with
matching `gateway_routing` scopes pointing at `pytest-of-*` temp directories.

The escape is structural. Hermetic isolation rides entirely on the process
environment: `HERMES_HOME` says *where* to write, `PYTEST_CURRENT_TEST` /
`PYTEST_VERSION` say *whether the guard is armed*. Both travel in the same
carrier, so a child spawned with a rebuilt environment loses them together —
it resolves the developer's real `state.db` *and* silences the only check
that would have stopped it, in one step. The guard is a no-op in precisely
the situation it was written for.

Back the env probe with process ancestry, which survives an env rebuild:

* `_process_looks_like_pytest()` matches a pytest launcher by argv token
  basename, so `/tmp/pytest-of-dev/...` paths in real argv cannot
  false-positive, and an unreadable process is never assumed to be a test.
* `_has_pytest_ancestor()` walks parents via psutil, memoised, and fails
  open when psutil is unavailable — a real `hermes` run pays for at most
  one walk and keeps the previous behaviour if the walk errors.
* `_in_test_context()` checks env first (two dict lookups, covers the
  in-process case) and only then ancestry.

`_STATE_DB_GUARD_BYPASS` is a module global and cannot cross a process
boundary, so ancestry-armed children would have had no way to opt out at
all; `HERMES_STATE_DB_GUARD_BYPASS=1` is the env-carried twin.

Also sweeps the rows already written. Bulk prune/archive cannot reach them:
their shared selector is pinned to `ended_at IS NOT NULL` so a live session
is never picked, which permanently excludes every never-closed row. Adds a
narrower selector — keyed, still open, and with no messages, tokens, tool
calls, API calls, activity or title — behind
`hermes sessions prune --never-active` (default floor 30 days, honours
--dry-run/--yes). Routing entries naming a deleted row go with it, so the
gateway is never left resuming a session id that no longer exists; `pinned`
and `archived` rows are excluded as explicit user intent.

Closes #82770
2026-08-15 02:20:13 -07:00
Teknium ca5c04818b feat: bare /save shows a usage card instead of exporting
Per review: /save with no arguments prints usage (formats, filename,
redact, examples) and writes nothing; the export requires an explicit
format. Unknown formats print the error plus the same usage card. One
shared SAVE_USAGE string in hermes_cli/session_export.py serves both the
CLI and gateway handlers. args_hint updated to <json|md|html> to reflect
the now-required format. Tests: bare-save and bad-format usage cases
added; location tests pass an explicit format.
2026-08-15 02:04:49 -07:00
Teknium e5a2c80c0e fix: harden /save for test doubles, async session-store boundary, snapshot shape pin
Three CI-red follow-ups on the /save rework:

- cli.py save_conversation: getattr-guard _session_db/session_id so
  SimpleNamespace/object.__new__ test doubles (and any embedder passing a
  minimal stub) don't AttributeError (pitfall #17 pattern).
- gateway/slash_commands.py: route through the awaited
  async_session_store.get_or_create_session boundary — the architecture
  test forbids raw session_store calls in async gateway code.
- tests/cli/test_save_conversation_location.py: /save now emits the
  canonical export_session payload shape; the session id key is "id"
  (was "session_id" in the legacy snapshot format) — update the pin.

The slice-8 lost-and-found failure was pre-existing on main and is fixed
there by f2a30fa400 (test: derive lost-and-found synthetic width from the
live schema); picked up via rebase.
2026-08-15 02:04:49 -07:00
Teknium 26aa12337a feat: /save exports the current session as json, md, or html on all platforms
Rework of salvaged PR #6372 (@ag9920) onto current main:

- /save promoted from CLI-only JSON snapshot to a cross-platform session
  export: `/save [json|md|html] [filename] [redact]` on CLI and every
  gateway platform (sent as a document via adapter.send_document).
- Rendering routes through the existing shared renderers
  (hermes_cli/session_export.py + session_export_html.py) instead of the
  PR's new hermes_state formatter — new helpers normalize_save_format /
  render_session_for_save / default_save_filename are shared by both
  surfaces.
- `redact` arg runs the export through the force-mode secret redaction
  pass (session_export_md.redact_session_data) before writing.
- Gateway handler awaits AsyncSessionDB correctly, sanitizes user-supplied
  filenames with basename, and lands in gateway/slash_commands.py (the
  handlers moved out of gateway/run.py since the PR was authored).
- /export stays profile export (name collision resolved: session export
  lives on /save).
- Slack 50-slash cap curation: /platform moves to the /hermes-only set to
  free a native slot for /save (parity test updated rationale comment).
- Folds in PR #62268 (@briandevans): None title/model coalescing in the
  single-session HTML export.

Closes #4249. Closes #51200.
2026-08-15 02:04:49 -07:00
briandevans edf6d2a074 fix(cli): coalesce None title/model in single-session HTML export
Single-session HTML exports of an un-named session render
`<title>None</title>` and `Model: None`. An untitled session (title
`None`) is the default state until async title generation completes, so
this is the common case, not an edge case.

The browser-tab `<title>` (page_title) and the `Model:` meta line use
`dict.get(key, default)`, whose default only fires when the key is
absent — not when it is present with value `None`. `_escape_html(None)`
then stringifies to the literal "None". The on-page `<h1>` in the same
function already uses the None-safe `... or "Hermes Session"` idiom, so
the tab title and header were inconsistent for the same session.

Use `... or "<default>"` at both sites so the tab title and model meta
fall back consistently with the header.
2026-08-15 02:04:49 -07:00
ag9920 23ad35b7fa feat: Add /export command for session export (Markdown/JSON) to CLI and Gateway 2026-08-15 02:04:49 -07:00
Teknium fc8ebff6d8 feat(mcp): unify the desktop MCP suggestion directory into the catalog
The desktop app carried its own hardcoded list of 17 vendor MCP endpoints
(apps/desktop/src/lib/mcp-directory.ts) powering the composer suggestion
pills — a second PR-reviewed vendor list, overlapping and drifting from the
Nous-approved MCP catalog (optional-mcps/).

This makes the catalog the single source of truth:

- manifest schema: optional `suggest:` block (keywords + hosts), parsed,
  validated, and normalized in mcp_catalog.py
- 15 new URL-only hosted-remote catalog entries (atlassian, sentry, datadog,
  notion, stripe, vercel, supabase, netlify, hugging_face, asana, intercom,
  airtable, webflow, paypal, square); figma + linear manifests gain suggest
  blocks
- GET /api/mcp/catalog now serves the suggest metadata
- desktop suggestion provider builds its match index from the catalog;
  the static directory remains only as a compatibility rung for older
  backends without suggest metadata
- setup card source line prefers the catalog entry's transport URL

GitHub stays out of the catalog on purpose: its hosted MCP rejects generic
DCR and the bundled github/* skills (gh CLI) are the stronger integration.
New desktop `github` suggestion provider offers the github-auth skill
instead — gated on a new cached GET /api/git/gh-auth probe so already-
authenticated users never see the pill.
2026-08-15 02:03:24 -07:00
Halldrix 3f9150e5c4 fix(secrets_cli): defer bitwarden backend import to first attribute access
Address blocking review on #86782 (trevorgordon981, 2026-08-15): the
previous lazy closures in main.py only deferred the module-level
"import secrets_cli" statement, but were themselves invoked at parse
time — so the chain main -> secrets_cli -> bitwarden -> cryptography
still ran eagerly on every command, including `hermes update --check`.
The closure indirection was dead laziness.

Move the laziness to where the crypto payload actually lives:

1. secrets_cli.py: drop the module-top "from agent.secret_sources
   import bitwarden as bw" import. Each cmd_* handler now resolves the
   backend via a local _load_bw() helper, which imports
   agent.secret_sources.bitwarden on first use. register_cli() no
   longer touches crypto at all — it only wires argparse structure.

2. _BWS_VERSION is duplicated in secrets_cli as a plain string so the
   "install" subparser help text renders without importing the
   backend. agent.secret_sources.bitwarden._BWS_VERSION stays the
   source of truth; bump both together when pinning a new bws release.

3. Module-level PEP 562 __getattr__ resolves "secrets_cli.bw" lazily.
   Existing upstream tests (test_secrets_bitwarden_non_tty.py) that
   monkeypatch "hermes_cli.secrets_cli.bw.find_bws" keep working —
   monkeypatch resolves the string one level deep, triggering
   __getattr__, which imports the real bitwarden module and lets the
   patch land on the same cached module object the handlers import.

4. main.py: revert the closure indirection back to a direct parse-time
   _secrets_cli.register_cli() call — safe now that register_cli is
   crypto-free by construction. The argparse wiring is again visible
   at the call site (matching checkpoints.py / curator.py convention),
   which addresses the original parse-time-vs-post-parse contract
   concern from the previous review round.

Adds a decisive main()-level regression test requested by review:
test_main_update_check_crypto_absent_in_sys_modules spawns main() in a
subprocess with argv=['hermes', 'update', '--check'], patches
hermes_cli.main._cmd_update_check to short-circuit before any network,
and asserts cryptography.hazmat.bindings._rust stays out of
sys.modules both at dispatch time and after main() returns. This is
the exact invariant the Windows self-lock depends on; the previous
import-only tests could not observe the failure because parser
construction runs inside main().

Verification:
- scripts/run_tests.sh tests/test_lazy_secrets_import.py
  tests/test_lazy_secrets_dispatch.py
  tests/hermes_cli/test_secrets_bitwarden_non_tty.py
  -> 13/13 passed (includes the new decisive test + the 2 upstream
     tests that broke under the earlier _LazyBitwarden proxy).
- Sabotage run: same suite against the pre-fix main.py + secrets_cli.py
  fails the new decisive test with "cryptography._rust loaded by main()
  before update dispatch" — confirming the test guards the bug.
- Manual trace: at _cmd_update_check dispatch time, sys.modules
  contains hermes_cli.secrets_cli (parse-time structure only) but NOT
  agent.secret_sources.bitwarden and NOT cryptography._rust.

Refs: #86781
Refs: #83569
2026-08-15 01:55:22 -07:00
Halldrix 5a76c8a978 fix(update): lazy-import secrets backends + defer registry import — break Windows self-lock loop
Address review feedback from trevorgordon981 on PR #86782:

1. **Pre-register parsers at parse-time, lazy-import backends only**
   - secrets_cli.register_cli() and onepassword_secrets_cli.register_cli()
     now run eager at parser-build time (no deferral past parse_args)
   - Only the agent.secret_sources.bitwarden/onepassword imports are lazy
     (inside each cmd_* handler via _load_bitwarden()/_load_onepassword())
   - This eliminates the 'invalid choice' and infinite-recursion risks

2. **Known-source-names gate for env_loader registry**
   - Only keys in {bitwarden, onepassword, op, 1password, bw} trigger the
     registry import; a generic dict entry no longer forces crypto load
   - Prevents unrelated config dicts from paying crypto cost

3. **End-to-end tests for the real dispatch paths**
   - test_bitwarden_setup_help: runs real CLI subprocess with --help
   - test_bitwarden_status/disable/onepassword_status: run real handlers
   - test_update_check_clean/no_self_lock: run real update --check
   - test_update_check_no_cryptography: sys.modules inspection (backup)

4. **Fix flaky test_turn_lease.py** (unrelated pre-existing failure)

The architecture guarantees:
- parse-time: zero cryptography load (all backends lazy)
- dispatch-time: crypto loads exactly once per secrets command
- update path: completely clean of cryptography._rust mapping

Refs #86781, #86782
2026-08-15 01:55:22 -07:00
Halldrix a85a981107 test(lazy-secrets): fix CI compatibility — run from repo root, avoid live-system guard
Two CI failures fixed:

1. Slice 4/12 FAILED tests/gateway/test_turn_lease.py — pre-existing
   flaky test, not caused by this change (confirmed unchanged in main).

2. Slice 11/12 FAILED tests/test_lazy_secrets_import.py — the new
   test used  with cwd=tests/, which:
   (a) made  resolve to tests/hermes_cli/__init__.py
       (missing __version__), and
   (b) triggered the conftest.py live-system guard pattern match
       on the string 'update' in the code.

Fixes:
- Extract _run_isolated() helper that runs from repo_root with
  PYTHONDONTWRITEBYTECODE=1
- For the update-check test, write a temporary .py file in the
  repo root instead of using -c with 'update' in the string
- Remove pytest import (not available in the sandbox; not needed
  since the tests are simple assertions)

Refs #86782
2026-08-15 01:55:22 -07:00
Halldrix 3dc3186873 fix(main): lazy-import secrets_cli to prevent cryptography._rust self-lock on Windows
The secrets_cli import in main() was eager, which loaded
agent.secret_sources.bitwarden and its cryptography.* dependencies
before cmd_update() ran. On Windows, the updater process itself
then mapped cryptography._rust.pyd into its own address space,
triggering the self-lock detector (_detect_self_loaded_native_modules)
and causing a defer/exit-2 loop that blocked updates entirely.

Move the secrets_cli import inside the _dispatch_secrets function so
it only pays for itself when the user actually runs a secrets
subcommand. This keeps hermes update (and all other commands) free
of the cryptography._rust.pyd eager load.

Refs #83569, #83590, #86687

Test: 3 new regression tests verify cryptography._rust stays out of
sys.modules during main() and the update path.
2026-08-15 01:55:22 -07:00
Benjamin Ang ad42ecfc06 fix(desktop): stream remote media without renderer credentials 2026-08-15 01:53:34 -07:00
Nicolas Formenton c75d835555 feat(desktop): mark a session as unread/read with a persisted watermark 2026-08-15 01:21:40 -07:00
Moisés Valero fef9c537d7 fix(cli): convert Alt key shortcuts to sequence tuple for prompt_toolkit (#74169) 2026-08-15 01:05:39 -07:00
John Lussier b3447c2129 test: align submit bindings with multiline default 2026-08-15 01:05:39 -07:00
John Lussier 2ae7884ffa fix: make CLI multiline shortcuts work by default 2026-08-15 01:05:39 -07:00
doncazper 6a375a24a5 fix(gateway): timestamp shared stderr output 2026-08-15 01:04:19 -07:00
doncazper 1db9273584 fix(gateway): timestamp launchd error log lines 2026-08-15 01:04:19 -07:00
coe0718 7f7aefe5cb fix: restore complete message timestamp coverage 2026-08-15 01:04:19 -07:00
Tuck 03636ab33f test: add message_metadata unit tests 2026-08-15 01:04:19 -07:00
Teknium f2a30fa400 test: derive lost-and-found synthetic width from the live schema
The mapper test pinned the sessions column count (54, then 55, then 56
within one week as git_metadata_generation and the hidden flag landed).
Every ordinary column addition broke it. Derive max_fields from
PRAGMA table_info at runtime with a >= floor so the test keeps asserting
the rebuild contract without change-detecting the schema width.
2026-08-15 00:55:20 -07:00
Paul BlackSwan d5a865882a fix(desktop): save remote gateway files natively 2026-08-15 00:37:00 -07:00
Teknium 46fe44fa8d test(cli): stub portable-MCP lookup in completer read test; bound resolve_toolset memo
- The readonly-loader completer test now stubs
  get_portable_mcp_server_names_nowait — real plugin discovery runs
  load_config() during one-time process init, which is not the
  per-keystroke read the test guards against.
- Cap _resolve_toolset_memo at 256 entries: generation-keyed entries
  from stale generations are never hit again, so clear on overflow to
  keep long sessions bounded.
2026-08-15 00:36:03 -07:00
jackoconner55 2f54ad4023 perf(cli): skip launcher-side plugin discovery for TUI handoff
(cherry picked from commit e2e0edd6b8ec10d02e867ff11dd1f9b4961a1ba7)
2026-08-15 00:36:03 -07:00
spfcraze b9f7525a1b test(agent): pin measured-work regression for display-flag config cache
(cherry picked from commit 57a7044c20965150e74b0ecadb2550f0ed81ee8e)
2026-08-15 00:36:03 -07:00
spfcraze e25cafc83e perf(agent): cache per-turn display-flag config reads
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled
re-read config.yaml on every call via load_config() (~1ms deepcopy per
call). finalize_turn runs these gates at the end of every turn, so each
turn paid two redundant config deepcopies. The sibling
_credits_notices_enabled already caches on self; mirror that pattern.

The env-var override stays authoritative and uncached, so runtime flips
still work. Config flips now apply on the next session, matching the
documented sibling semantics.

(cherry picked from commit f2d0e00d71ba35fa3a53cd41076261e3c5feb45d)
2026-08-15 00:36:03 -07:00
spfcraze 7b45d1d049 perf(toolsets): memoise resolve_toolset keyed on registry generation
resolve_toolset() recursively walks toolset includes and, with
include_registry=True, merges registry-registered tools on every call —
each external call re-runs the includes walk and takes a fresh registry
snapshot under the registry lock. It is called dozens of times per
_get_platform_tools() (every /tools completion keystroke, per picker
render) and at 19 call sites across the CLI/gateway.

Memoise the external-entry result keyed on (name, include_registry,
registry id, registry generation). tools.registry exposes a monotonic
_generation counter bumped on every register/deregister/alias/MCP
refresh (its docstring explicitly invites generation-keyed memoisation),
so a cache entry is valid until the registry changes. External callers
never pass visited, so the memo engages exactly at the public entry
and the internal cycle-detection recursion is untouched.

Measured: _get_platform_tools drops 165us -> 59us per call (3x) with the
xAI credential fix simulated; /tools completion ~2ms -> ~77us/keystroke
combined. Regression tests: repeat resolution is a memo hit (get_toolset
called once), a generation bump forces a fresh resolve, and the memoised
result is identical to a fresh resolution.

(cherry picked from commit 3d36ecb273f6652d00556a7b9d846c08047801aa)
2026-08-15 00:36:03 -07:00
spfcraze 47400fe2af test: make memo pins pre-fix-safe (raising=False resets)
(cherry picked from commit 4822daed5d9238348c50bfbdf8c3c4795adc4986)
2026-08-15 00:36:03 -07:00