Commit Graph

14240 Commits

Author SHA1 Message Date
Teknium 7cd91114b4 fix(web): drop tavily from the removed-backend registry after the restore
#100540 added a REMOVED_BACKENDS startup warning keyed on tavily; with the
backend restored, that entry would warn on a working provider. The registry
stays (empty) for future removals; migration tests now pin the machinery via
a synthetic entry plus a guard asserting no live provider is ever listed as
removed.
2026-09-01 10:56:49 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium f8f4d056f5 test(gateway): pre-warm the goals SessionDB cache in async goal tests
GoalManager.set() on an event-loop thread only waits the bounded
_DB_BOOTSTRAP_INIT_WAIT_S window for the background SessionDB bootstrap
(deliberate: an unbounded init starved the gateway loop watchdog). On a
loaded CI runner the cold init overruns that window, the goal write is
silently dropped by design, and the test flakes downstream: /loop showed
no active-goal note and the goal continuation was never enqueued (both
FLAKY on main run 33455779041).

Fix the class: every async goal test fixture that clears goals._DB_CACHE
now pre-warms it via _get_session_db() from sync context (unbounded init
path), so the bounded-window degradation can never fire mid-test. Applied
to all four gateway goal/loop test files; sync-only goal tests are
unaffected by construction.

Live repro: slowing SessionDB.__init__ past the window reproduces the
dropped write deterministically without the pre-warm and never with it.
2026-09-01 10:52:58 -07:00
Teknium 75bf672a78 test: managed-runtime source scan survives a vanishing sdist dir (TOCTOU)
Path.rglob raises FileNotFoundError when a directory disappears between
listing and scandir — a sibling CI job creating/removing its sdist
extraction (hermes_agent-<ver>/) killed test_allowlist_has_no_stale_entries
on run 33531869442. Switch to os.walk (tolerates vanishing dirs) with
top-level pruning of exempt and packaging dirs; file set is byte-identical
(886 files verified old==new) and a 30-scan churn harness that reliably
exercised the window shows zero errors.
2026-09-01 10:52:42 -07:00
Teknium 83cde7f31d test(gateway): widen pending-drain chain wait budget (4s -> 20s)
Same loaded-runner class: the 12-turn drain chain completed only 11
turns inside the 400x0.01s poll budget on main run 33455779041. The
loop still exits early on success, so the wider budget costs nothing
on healthy runs.
2026-09-01 10:52:42 -07:00
Teknium 28834a2098 test: raise tight wall-clock bounds that flaked on loaded CI runners
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).

Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
2026-09-01 10:52:42 -07:00
Teknium 67de93862c fix(tui-gateway): gate the ws-orphan interrupt of running turns on activity staleness
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).

The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).

Fixes #98028
Fixes #100325
2026-09-01 10:52:23 -07:00
Teknium 894fc35337 fix(state): break provably-orphaned repair/FTS-rebuild locks left by dead holders (#100108) 2026-09-01 10:52:06 -07:00
Teknium 09b88bab88 fix(state): stop the on-write identity probe cancelling our own POSIX locks (#100368) 2026-09-01 10:51:52 -07:00
teknium1 46e7ad8e12 fix(gateway): gate record-less visible-text match on _already_sent
Draft frames set _last_sent_text for dedupe without setting
_already_sent (they are ephemeral); an ungated has_delivered_text match
let a draft-only preview count as durable delivery and regressed
test_relay_seal_failure's dead-transport guarantee on CI.
2026-09-01 10:51:37 -07:00
teknium1 fd998120c1 fix(gateway): judge delivery success against final content, not flag trust (#95382, #98552)
A record-less delivery flag (final_response_sent /
final_content_delivered set with no recorded turn-final payload) was
trusted blindly by delivered_final_matches (None -> legacy trust), so a
first-edit prefix or a truncated finalize suppressed the gateway's
corrective send — silent partial delivery.

- delivered_final_matches: record-less flags are now reconciled against
  the FINAL content via has_delivered_text; only the explicitly-marked
  ambiguous-timeout path (_delivery_ambiguous) keeps legacy trust.
- _try_fresh_final and the native-streaming optimistic finalize now
  record their delivered payload (the last record-less flag setters);
  the optimistic record rolls back on definitive dispatch failure.
- Discord adapter: dead-transport send failures (client gone, WS
  closed/reset) are classified as send_path_degraded (retryable) so the
  delivery-obligation ledger's reconnect sweep replays the stranded
  final response instead of losing it until a process restart.

Fixes #95382; closes the #98552 false-positive class.
2026-09-01 10:51:37 -07:00
Teknium 043c258ac2 fix(web): stale removed-backend config warns at startup and errors by name
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).

- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
  removed_backend_note(); selection_error() swaps in the specific
  removal explanation (removed in v0.21.0, keyless alternatives) while
  keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
  search_backend / extract_backend against the registry and emits a
  startup warning (deduped per stale value), surfaced by the existing
  print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
  per-capability keys, dedupe, healthy-config negative, live-backend
  failure text preserved.
2026-09-01 10:43:20 -07:00
rainbowgits 622883bad7 fix(agent): accept marker-only finish_reason after stream supersession
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:27:06 -07:00
kshitijk4poor b81383ec21 fix(auth): compare pool-identity callers against all candidate keys
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:

- _prune_replaced_custom_model_config_credentials skipped only the
  preferred key, so a keyed provider's own legacy-named pool
  (custom:b.ai) was false-pruned of its current model_config credential
  when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
  key, so a legacy-named pool stopped being seeded from model.api_key.

Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.

Follow-up to #100413.
2026-09-01 22:42:27 +05:30
xxxigm 0bee5ff408 fix(auth): look up keyed custom providers by durable pool slug
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
2026-09-01 22:42:27 +05:30
xxxigm 43470980bf test(auth): cover keyed providers.<key> credential-pool lookup
New-style providers store keys under the durable config slug, but
runtime still looks up custom:<display-name> and sends a placeholder.
2026-09-01 22:42:27 +05:30
Teknium 6ddafd34f0 test(proxy): cover multi-line SSE data joins, truthy lastOne, EOF-without-blank-line dispatch 2026-09-01 10:12:21 -07:00
Jon Nielsen d304422b3d fix(streaming): extract finish_reason/usage before content-shape continues
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.

Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.

Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
2026-09-01 10:12:21 -07:00
loulanyue 66d42e0dba fix(stream): do not misclassify stream with final usage chunk as mid-stream drop (#91373)
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.

If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.

Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
2026-09-01 10:12:21 -07:00
rainbowgits ce7f805869 fix(proxy): append SSE [DONE] when Nous streams omit the sentinel
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:12:21 -07:00
Hermes Agent 219cb714a1 fix(gateway): persist api_server async-delegation completions as deliveries, not user-turn wakes
On the stateless api_server platform, a background delegate_task
completion was delivered by self-POSTing /v1/chat/completions with
role=user after the parent turn had already completed (event.complete,
finish_reason=stop). That starts an unauthorized new agent turn the
client never sent, persists the completion as an ordinary role=user row
(display_kind NULL), and can blow through a pending human-confirmation
gate — the exact skip-ahead reported in #85957.

Fix: async_delegation completions targeting a non-push api_server
session are now written into the session transcript as a durable
DELIVERY row (role=user + display_kind=async_delegation_complete +
display metadata — the same bookkeeping shape the TUI/desktop delivery
path persists). No agent turn runs; clients polling
GET /api/sessions/{id}/messages see the result immediately, and the
next real client turn carries it as context. Persist failures return
False so the durable claim is released and the completion retried.

Watch-pattern notifications keep the existing self-post wake behavior;
push-capable adapters are untouched.

Fixes #85957
2026-09-01 10:11:07 -07:00
Teknium e9541d213f fix(gateway): never print ✓ for a Windows gateway that dies after the liveness poll
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.

Two layers:

1. _wait_for_gateway_ready now treats the first hit as provisional: the
   gateway must stay visible through a 2s confirmation window
   (_confirm_gateway_stable) before it is reported ready; a death during
   confirmation resumes polling until the deadline. Failure output is an
   honest ✗ with the Job Object explanation and the schtasks /Run
   recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
   state/gateway.start-attestation.json with the vouched-for PIDs. The
   next `gateway start`/`gateway status` invocation checks it — if the
   attested PIDs are gone with no clean-exit record in the lifecycle
   ledger, the CLI reports (once) that the previous ✓ was false and
   prints the schtasks recovery hint. `gateway stop` and a clean
   lifecycle-ledger exit clear the marker silently.

Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.

Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.

Fixes #91675
2026-09-01 10:06:06 -07:00
Teknium 045865377c fix(update): restore user model settings config.yaml rewrites drop during update
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.

Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.

6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.

Fixes #64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
2026-09-01 09:54:22 -07:00
Konstantin Khlopkov aea2daee29 fix(backup): import into the active HERMES_HOME instead of the root home
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.

Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.

Fixes #99839
2026-09-01 09:28:46 -07:00
joaomarcos ecdbcef7af fix(compression): roll the live transcript back when an in-place compaction commit fails (#99477)
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.

That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.

The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 09:28:31 -07:00
fangliquanflq 26476a978a test(installer): cover managed Python timeout 2026-09-01 09:28:01 -07:00
fangliquanflq c454fbc5bd fix(installer): harden managed Python resolution 2026-09-01 09:28:01 -07:00
fangliquanflq a2895e9968 fix(installer): preserve managed Python fallback version 2026-09-01 09:28:01 -07:00
fangliquanflq 60c7c0c667 fix(installer): require Hermes-managed Python on Windows 2026-09-01 09:28:01 -07:00
Teknium ada3c28d45 fix(restore): fail closed on in-process holders before unlinking state.db sidecars
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).

Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).

Part of the #90837 sidecar-unlink audit (wave 6).
2026-09-01 09:27:47 -07:00
Teknium 4789fa1066 test(browser): isolate chrome-fallback sandbox test from the shared tmpdir
The test exercised _run_chrome_fallback_command against the REAL
/tmp agent-browser namespace, so any concurrent hermes/pytest process
running the orphan reaper could rmtree the fresh pidless socket dir
mid-command (FileNotFoundError on _stdout_open — the recurring CI
flake, 3 hits incl. two main runs on 2026-09-01). Route it through a
private tmp_path like every sibling browser test.
2026-09-01 09:27:00 -07:00
Hermes Agent 4cca38be86 fix(browser): preserve starting sessions during orphan reap 2026-09-01 09:27:00 -07:00
kshitijk4poor 58472d803a refactor(state): flatten registry acquire flow; restore mirror cleanup assertion
Post-merge review follow-ups on #100201:

- acquire(): drop the never-iterating 'while True' and the redundant
  'existing is not generation' half of the race check — after the retire
  path, generation is always None on the fresh-open leg, so 'existing is
  not None' is the complete condition. Same behavior, flat flow.
- test_mirror: the conversion to the shared registry dropped the
  cleanup assertion entirely; restore it by patching
  hermes_state.release_or_close and asserting the handle is released
  exactly once after _append_to_sqlite.
2026-09-01 21:12:35 +05:30
Teknium 28044757aa fix(desktop-update): self-heal a TCC-anchor-bricked venv in the posix hand-off (#95759)
Installs converted by the reverted macOS TCC interpreter anchor
(#95425/#95541) are left with a real-file venv/bin/python copy, a
.tcc-anchor-source marker, and python3/python3.N aliases that die at
interpreter init ("No module named 'encodings'"). venv/bin/hermes execs
venv/bin/python3, so EVERY CLI entrypoint is dead — hermes doctor and
hermes update included — and the desktop hand-off loops on "Update failed
(exit 1)" forever. No Python-side heal can ever run on this class; the
hand-off shell is the last surface that still executes, so the heal lives
there.

posix.sh gains, before the update invocation:

* tcc_anchor_heal — probe-gated (only fires when venv/bin/python3 fails
  a scrubbed-env `import encodings` boot probe), marker-validated
  (absolute path, outside the venv), staged with per-attempt backups and
  full rollback if the repaired interpreter still fails its probe.
  Two repair shapes:
  - alias-brick (#95541 class): the anchored copy boots — re-materialize
    python3* as REAL FILES of the anchor (hardlink/copy; an alias symlink
    onto the copy is the crash shape). Marker kept: this is exactly the
    layout ensure_tcc_anchor marks "active", so no anchor ping-pong.
  - full brick: restore python → symlink to the marker-recorded store
    interpreter (if it boots) and aliases → symlinks; marker removed.
    The unblocked `hermes update` then re-installs a boot-gated healthy
    anchor — one-shot convergence, not a loop.
  Fail-closed on missing/unbootable source (vanished uv store class),
  missing marker, or relative/in-venv marker paths.
* tcc_pick_update_invoke — if aliases stay dead but venv/bin/python
  boots (the launchd-gateway shape), drive the update via
  `venv/bin/python -m hermes_cli.main` instead of the dead hermes shim.
* Honest terminal message: an unrecoverable dead interpreter is reported
  as a venv repair problem instead of a generic "Update failed (exit 1)".
* --self-test-tcc-heal runs the real heal + invoke selection against an
  --install-root and reports, for the test harness.

Tests (tests/test_desktop_update_tcc_heal.py) drive the REAL posix.sh
functions on Linux against synthetic venv trees: healthy no-op, alias
heal, symlink restore, fail-closed classes, rollback on failed
verification, invoke fallback, and an A/B of the reported loop (bricked
venv/bin/hermes fails with the exact field error before, boots after).
Sabotage-verified (re-introducing the alias-symlink bug fails 2 tests).

NOT mac-live-tested (no macOS runner); the heal is platform-independent
shell exercised through the self-test path without the uname gate.

Recovery design (validation/staging/rollback/probe pattern) after
@aeonsong's #96231; in-update heal intent from @liuhao1024's #95775
(its target function no longer exists on main and its heal point is
unreachable on dead-CLI installs); heal-point and ping-pong analysis by
@ahrazzle and @tokenfires on #95759.

Fixes #95759
2026-09-01 08:36:03 -07:00
Teknium 9f377264aa refactor(update): consolidate the gateway drain triage into one shared helper
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.

No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
2026-09-01 08:34:51 -07:00
salch-cred 49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Teknium 71c4bcf7af fix(cron): surface missed-fire catch-up lateness in hermes cron list / status
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).

- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
  scheduled_at, dispatched_at, lateness_seconds, and kind
  (on_time / late / catch_up, classified against the ticker tolerance and
  the schedule's catch-up grace window). Manual triggers and one-shots are
  not stamped (no scheduled instant to be late against / retired beyond
  grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
  show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
  catch-up, in both the built-in ticker and external provider paths.

CLI surface only — no new tools, no policy engine.

Addresses the visibility half of #99879.
2026-09-01 08:31:37 -07:00
Teknium ddd8065232 test: regression coverage for first_chunk_at in post_api_request hook
Covers the reviewer-requested cases for #98555:
- successful streamed response emits the latest attempt's non-null
  first_chunk_at (started_at <= first_chunk_at <= ended_at)
- non-streamed, failed-stream, and partial-stream-stub paths emit None
- a stale timestamp from a prior API call cannot leak into the next
  call's post_api_request payload (per-attempt reset in the loop)
2026-09-01 08:30:45 -07:00
CoolStar aaca343110 fix(agent): preserve Bedrock redacted reasoning replay 2026-09-01 08:30:26 -07:00
Justin Wilson fb18fedf29 fix(updater): rebuild desktop on Windows hand-off repair path
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.

Fixes #97343
2026-09-01 08:27:58 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium e2521fdef6 fix(telegram): promote connect success line to WARNING so healthy startup is terminal-visible
The "Connecting to Telegram (attempt N/8)…" line logs at WARNING and
reaches the gateway's default stderr handler, but the matching
"Connected to Telegram (… mode)" line was INFO and went to the log file
only. A healthy startup therefore looked permanently hung at
"attempt 1/8" on the terminal — the logging-illusion half of #90835.

Promote the success line to WARNING so both sides of the connect
transition share the same console sink; a genuine hang is now the
absence of the success line. Adds an AST-level regression test pinning
the level pairing. Sibling adapters (homeassistant, wecom) log both
sides at INFO, so they don't have this asymmetry.

Fixes #90835
2026-09-01 08:24:25 -07:00
Teknium f709bd88b6 feat(skills): render the configured create dir in every instruction that names the path
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
2026-09-01 07:32:45 -07:00
JoaoMarcos44 86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
Teknium 35f4ababe1 test: the managed-dashboard restart now continues the serve scan
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
2026-09-01 07:00:54 -07:00
Teknium 5a677479e9 test: pin the no-authority contract with the serve probe disabled
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.
2026-09-01 07:00:54 -07:00
joaomarcos 27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
joaomarcos 2d70d67b15 fix(update): stop the managed dashboard restart from skipping serve backends
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.

The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
2026-09-01 07:00:54 -07:00
joaomarcos f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00