Composition of #92581 on #100985: the hard cap is a constructor constant,
so raising the pool max in Settings would have left new spawns queued
behind the launch-time value. Add setLimit(); slot hand-off now goes
through a single #drain that respects the current cap, which also fixes
the original release path handing a slot to the next waiter even when the
cap had just been lowered (test: lowering never revokes granted slots; new
requests queue until under cap). main.ts constructs from
poolLimits.maxBackends and pushes changes from setPoolLimits(); pinned by
a wiring test.
Hover-intent prewarm sweeps across the Bots rail spawned past the pool cap,
LRU-evicting the backend the user was about to click — an evict/respawn
cascade that made profile switching progressively slower (#91545).
- prewarmProfileBackend skips speculative spawns once every pool slot holds
an open socket; the real click still spawns on demand.
- Pool max/idle become a device preference (Settings -> Advanced), persisted
atomically in userData (pool-limits.json) and applied live over IPC; the
HERMES_DESKTOP_POOL_* env vars remain the initial fallback. Defaults are
unchanged (3 backends / 10 min idle).
Squash of the 3-commit PR #92581 branch (a00dc088c5..783899d12f) applied
via diff onto the spawn-coordinator salvage; import + constant-block
conflicts resolved so the coordinator is constructed from, and follows,
the live preference (setLimit added in the next commit).
Follow-up to the spawn coordinator: the queued ticket waited up to
POOL_IDLE_MS (10 min) for a free local slot, but the renderer gives up on
a backend boot after 45 s. A user clicking a 4th profile with 3 fresh
backends open would see the generic "backend didn't come up" error while
the ticket kept the pool key hostage, so every later click joined the same
stale wait. Cap the wait at 30 s, log the slot pressure when it happens, and
pin the relationship to BACKEND_BOOT_WAIT_TIMEOUT_MS with a wiring test
(fails when the timeout is reverted). Also eslint --fix on the salvaged
files (import order was a lint error).
Desktop could spawn a local hermes serve per profile with no hard cap on
starting+running children: LRU eviction spares keepalive-fresh entries, so
a roster refresh across many profiles became a process wave (40+ backends,
load 30-50 reported).
LocalBackendSpawnCoordinator: at most POOL_MAX_BACKENDS local backends may
be starting or running. Remote descriptors never take a slot. Queue tickets
are per request; a slot is released only after process exit is proven
(exitCode/signalCode). A rejected wait keeps the slot occupied. Pool entries
re-assert ownership at each await so an evicted entry cannot spawn a zombie.
Squash of PR #100985 (6017abbbc4 + merge), applied via diff onto current
main. Original commits were authored as 'Motor (Hermes AI) <ceo@xtremagency.com>';
attributed here to the PR author's GitHub identity.
The ownership file accumulates one record per profile per launch, and each
record can cost up to two identity probes (parent + backend) plus a stop.
On Windows those shell out to PowerShell, whose 5.1 cold starts are slow.
Without a bound, a large roster could stall boot for minutes while the
renderer's 45s backend-boot budget expires and the user stares at the
connecting screen.
- Add REAP_PROBE_TIMEOUT_MS (5s) for the orphan-reap path; the claim path
keeps the full 30s headroom for a freshly spawned backend's marker.
- Add reapDeadlineMs (5s default) as an overall budget for one reap sweep;
when exhausted, unprocessed records are preserved for the next launch.
- stopOwnedBackend now throws when the identity probe fails (not confirmed
gone) so the record is preserved instead of leaking the backend.
Verified: two cold starts complete in ~29s (was 5+ min); renderer connects
immediately after backend ready.
The salvaged #88217 test replaced Path.stat on the class with a lambda
taking one positional arg; pathlib.exists() passes follow_symlinks= and
pytest's own tmp_path teardown crashed with INTERNALERROR TypeError.
Delegate to the real stat for every path other than state.db.
Composition bug between the salvaged #101266 (in-place v29 startup
migration: swap the trigram view/triggers, FTS5 'rebuild') and #88217
(FTS_STORAGE_VERSION 2 drops the tool_calls column from the trigram
vtable, opt-in via optimize-storage). On an install still carrying the v1
vtable, the startup migration replaced the view with one that has no
tool_calls, then 'rebuild' failed with 'no such column: T.tool_calls' and
SessionDB.__init__ raised — reproduced by opening a real main-built DB.
Gate the in-place migration on the vtable not projecting tool_calls; such
installs are already offered optimize-storage, which recreates the vtable
from FTS_TRIGRAM_SQL (cron-filtered view included). E2E: main-built DB ->
opens on this branch, optimize_fts_storage() yields v2 columns and purges
the cron row. Test fixture now builds a real external-content vtable for
both layouts; new test mutation-checked against the missing guard.
Clear zero-row rebuild markers so empty databases can finish teardown and
stamp the new layout. Extend the migration regression through close/reopen
recovery for both empty and populated databases.
Keep structured tool_calls searchable through the standard FTS index while
removing their repetitive JSON from the trigram projection. Reuse the
existing optimize-storage rebuild path for deployed v1 layouts.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Replace the requires_wal-gated unit assertion (which never runs on WAL-reset-
vulnerable runtimes such as macOS 3.46) with an end-to-end test through
repair_state_db_schema that traces every _connect_repair_durable call against
the guard's enter/exit and fails on main's shape (a connect after guard-exit).
Runs on every platform. Docstring now names the actual hazard.
apply_wal_with_fallback keeps a store in DELETE where the linked SQLite has
the WAL-reset bug; the connection-reuse contract the test guards is unaffected
but its final journal_mode assertion is not satisfiable there.
Review follow-ups on the live-inode import:
- A `.db` member is now page-restored into the live file, but a paired
`-wal`/`-shm`/`-journal` member from an old or hand-built archive still went
through the rename publish — installing a foreign WAL beside the restored
database (and over a live sidecar's inode). Skip them; current backups never
ship them (_EXCLUDED_SUFFIXES), now shared as _SQLITE_SIDECAR_SUFFIXES.
- restore_quick_snapshot ignored _safe_restore_db's False and counted a
refused restore as success; honour it like run_import does.
- _safe_restore_db docstring described the pre-#90950 unconditional fallback.
`hermes import` published every zip member, including `state.db`, with
`_extract_member_atomically` — a rename that swaps the file's inode. Any
gateway, dashboard, or WebUI process holding the database open keeps its
descriptor on the now-unlinked inode: it goes on serving pre-import pages
and writing sessions no other process can see, while the sidecar WAL left
beside the new file describes the database that was just unlinked. Nothing
raises, so the import prints "Import complete" and the sessions are simply
absent from the database everyone opens next.
The live-safe path already exists: `/snapshot restore` has routed `.db`
files through `_safe_restore_db()` since #65942, writing snapshot pages
into the existing file so every open connection converges. `hermes import`
— the disaster-recovery path, reached by users who already lost something
once — never got that treatment.
Route `.db` members through it. A target that does not exist yet has no
holders and no inode worth preserving, so it keeps the ordinary atomic
publish. A refused or failed live-safe restore now raises, so the import
reports a skipped file instead of counting a silent success, and the
existing database is left untouched.
Importing an older backup over newer work stays allowed but no longer
silent: the summary reports the session/message counts the import replaced,
the same before/after evidence `restore_cron_jobs_if_emptied` uses for
`cron/jobs.json`.
Closes#100960
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ShTVU941HYygvypY9JMYE
(cherry picked from commit 8ff260a312341cb85bdf7cdd570de5ff888c6013)
Completes the constant + samples + noise-regex trio teknium named as the bar
for this heartbeat on #98371; the Telegram noise-filter and TUI retag
parametrised suites now iterate the heartbeat wording too.
The seed put the retry deadline 900s in the future, so on unguarded code the
method short-circuited on the backoff check and the protection assertions
passed anyway; only the field-reset assertions failed. Seed the deadline in the
past so a rebuild is due, and condense the guard comment.
retry_deferred_fts_recovery() is called unconditionally on every
housekeeping tick for the life of a long-running gateway process
(#100108). It checks _fts_stale, read_only, and _conn is None, but never
_db_corrupt (bcc2e65818, #101095/#101224): a handle that observed
structural corruption is supposed to stop being touched entirely (see
_try_wal_checkpoint's identical guard, and close()'s skip of the
checkpoint), but this method has no such check.
If a handle both has a deferred stale-FTS breadcrumb AND later trips
quarantine (both are plausible on the same corrupted file — the field
incidents motivating the quarantine feature describe corruption touching
FTS shadow tables and canonical btrees together), every subsequent
housekeeping tick runs a real FTS rebuild (DROP TABLE / CREATE VIRTUAL
TABLE / bulk INSERT) against the file the code has explicitly decided to
stop touching — exactly what quarantine exists to prevent.
Fixed by returning False immediately when _db_corrupt is set, mirroring
_try_wal_checkpoint's "quarantined: never touch a damaged image" guard.
The method's own contract ("never raises") is preserved — no
StateDbCorruptError is raised here, this is a quiet skip like the other
corrupt-aware call sites.
Also resets the backoff bookkeeping (_fts_stale_retry_after,
_fts_stale_retry_interval) in the same early-return, mirroring the
success path's own reset a few lines down (review feedback from
Baophan00 on the PR). Verified empirically before making this change:
_db_corrupt is set to False nowhere in the codebase outside __init__, and
the shared registry's file-replace path always constructs a genuinely new
SessionDB instance rather than clearing the flag on a live one — so no
code path today revives a quarantined handle in place, and leaving the
backoff fields untouched is inert in practice. The reset is still cheap,
harmless, and closes a real footgun for whoever adds an un-quarantine
path later: without it, a handle quarantined mid-backoff would carry a
doubled multi-minute interval into any future retry instead of starting
from the default.
Added a regression test that marks a handle stale, forces the open-time
recovery to defer via a real held rebuild lock (so _fts_stale survives
construction), sets _db_corrupt plus a pre-existing multi-minute backoff,
and asserts the retry is a no-op with both backoff fields reset to 0.0.
Mutation-verified: reverting hermes_state_schema.py makes the retry
actually run the rebuild and return True, and separately makes the
backoff-reset assertions fail with the stale pre-quarantine values still
in place.
(cherry picked from commit 3445da1d98d84bb60cb3799ef59e1fa4c619100f)
Partial salvage of #100887 by @jwilson411. The restore-path half (raise
TranscriptReadError from load_transcript, fail the turn closed) already landed
via #100910; this keeps the complementary half: every slash-command handler and
platform helper that reads the transcript now catches TranscriptReadError and
tells the user the history exists but is unreadable, instead of letting the
exception reach the dispatch wrapper, which logs it and sends no reply.
(cherry picked from commit 2a132c903f51c8d21f6eea64ddeef688c9619a18, run.py/session.py
hunks dropped as already on main; notice-path tests replaced accordingly)
The gateway broadcast and turn-failure explanation printed a literal
~/.hermes/state.db; now that the line is a command the operator is meant to
run as-is, interpolate _default_db_path() so profile / HERMES_HOME installs are
pointed at the store that actually failed. Also: split a comment that a merge
fused onto the logger line in session_lost_and_found.py, and fix an inverted
test docstring.
Review blocker on e62940d: every state-db guidance site printed
hermes sessions recover --source <db>
but cmd_sessions rejects that shape with exit 2 ("--output is required
unless --inspect-only is used") before any snapshot is taken — the user
follows the instruction during a corruption incident and gets nothing.
All five state-db sites now print the established two-stage operator
contract (the same shape `sessions repair` failure output and
docs/state-db-recovery.md already use):
hermes sessions recover --source <db> --inspect-only
hermes sessions recover --source <db> --output recovered-state.db
with the stop-the-gateway precondition stated for the gateway/turn
banners, and --inspect-only leading in the hermes_state refusal strings
(inspection before writing anything).
New TestEmittedCommandsSatisfyCliContract dispatches the exact emitted
flag shapes through the real cmd_sessions and asserts they pass the
contract gate (rc != 2) on a scratch DB, plus a premise test pinning
that the v1 no-flag shape is still rejected with rc 2 — so a guidance
string can never again pass a source-substring test while the command
it prints deterministically fails.
Noted for merge order: #101423 and #101168 also touch
hermes_cli/session_recovery.py. They are complementary recovery-integrity
work, not duplicates of this guidance/gate fix; whichever lands second
should rebase and rerun the lost_and_found + session-recovery suites.
(cherry picked from commit 34dc59a284509e76a0342c36d03a2a437aa8a3b9)
Refs #100368. The forensics thread established that a sqlite3 CLI with
the WAL-reset opener bug (fixed 3.51.3+ / backports 3.50.7 / 3.44.6;
Debian/Ubuntu system shells 3.45.1/3.46.1 are in the vulnerable band)
unlinks the live -wal/-shm pair when pointed at a live state.db whose
writer's DMS lock has been cancelled, splitting the store into two
concurrent generations whose acknowledged writes vanish while both
report integrity_check ok. Hermes' own corruption banners instructed
exactly that command.
- gateway corruption broadcast, run_agent corrupt-cause explanation,
hermes_state repair-budget and forensic-backup refusals, and the
kanban manual-recovery hint now route operators to
`hermes sessions recover --source <db>` (which snapshots the damaged
bundle before any shell touches it) and warn against a raw sqlite3
shell on the live file
- find_sqlite3_cli() now refuses a WAL-reset-vulnerable shell for the
page-level salvage lane even on the snapshot, reusing the canonical
gate from hermes_cli.sqlite_runtime so the embedded runtime and the
salvage shell can never disagree
- find_sqlite3_cli_refusal() records why a shell was refused so the
lost_and_found lane can tell the operator exactly what to install
instead of a generic "not found"
- regression tests cover the version gate (vulnerable/fixed matrix, the
mirror check), every refusal reason, and each guidance site
Test plan:
- scripts/run_tests.sh tests/hermes_cli/test_sqlite3_cli_salvage_gate.py
tests/test_state_db_repair_loop_cap.py
tests/run_agent/test_corruption_recovery_guidance.py
tests/hermes_cli/test_session_recovery_lost_and_found.py
tests/hermes_cli/test_session_recovery.py tests/test_sqlite_wal_reset_gate.py
tests/hermes_cli/test_sqlite_runtime.py - 91 passed, 1 skipped locally
(cherry picked from commit e62940d1021e80e9b7d6423ced1cbdfe7dd0c37d)
Replaces the hand-transcribed regex fragment with re.escape of COMPACTION_STATUS
and COMPACTION_HEARTBEAT_STATUS, matching the COMPACTION_DONE_STATUS precedent
two lines down, so wording drift cannot desync the chat-platform gate.
Review follow-ups: the heartbeat opened a visible compacting phase even when
the context engine suppressed the routine start status (no terminal edge would
ever close it), and its start() emitted a second start line milliseconds after
the routine one (two chat messages on adapters without send_or_update_status).
Gate the client-visible heartbeat on the start status having been emitted, drop
the start-time emit, and route ticks through agent._emit_status so CLI print
and gateway filtering match every other compaction status.
The contributor heartbeat called status_callback("compacting", <ad-hoc text>).
Every other compaction status uses the "lifecycle" key: the TUI gateway
re-tags lifecycle statuses to kind="compacting" via
is_compaction_progress_status, Telegram edits one bubble per status key, and
gateway/run.py's chat-platform filter only recognises the registered
templates — so the ad-hoc text leaked to Telegram/Discord on every tick with
compression.progress_notices off (verified: _prepare_gateway_status_message
passed it through). Register COMPACTION_HEARTBEAT_STATUS next to
COMPACTION_STATUS, emit it under "lifecycle", and cover the filter.
Context compression can stream for minutes with no deltas, tool events,
or status lines reaching remote transports. Idle-progress watchdogs on
those clients treat the silence as a dead turn and interrupt it — the
Android relay app's 180s turn watchdog fires session.interrupt, killing
a healthy compression mid-flight and rolling back its work. On sessions
near the context ceiling this loops forever: every new prompt retriggers
preflight compression, which dies at exactly +180s again (observed
telemetry: attempts aborted at 180469ms/180233ms/180219ms with
failure_class=explicit_interrupt).
Fix: the existing _CompressionActivityHeartbeat (which today only
refreshes the SessionDB activity tracker) now also emits a 'compacting'
status through agent.status_callback — once at compression start and on
every heartbeat tick. The gateway already routes status_callback to
status.update events, and clients already reset their watchdogs on any
received event, so each heartbeat re-arms them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6d2e0e860d6fdb44ce549d1fa11a06662b1a4b0b)
Review follow-ups on the in-flight replay:
- Run _sanitize_tool_pairs BEFORE the re-append. Its trailing-in-flight
exemption (#79278) walks back from the list end; with the replay user row
there, a genuinely pending assistant(tool_calls) looked orphaned and had its
calls stripped, so the executor's late tool result was dropped.
- When the restatement is merged onto a user-pinned summary carrier, flag the
carrier (_inflight_replay_merged). The carrier's metadata marks it synthetic,
so conversation_compression._ensure_compressed_has_user_turn inserted a second
copy of the same request; it now treats the flag as intent-present, and the
next cycle recognises the carrier as the task instead of losing it.
- Header idempotency: restate the text after the last header so a task that
survives several compactions carries one header and one copy.
- Exclude metadata-flagged scaffolding (_todo_snapshot_synthetic, recovery
nudges) from the in-flight scan via the shared _is_real_user_message.
Tests cover all four; each fails with its fix reverted.
The in-flight replay checked only compressed[-1] before choosing between
appending a user row and merging onto the summary carrier. A user-pinned
summary followed by an exempt tool_calls/tool tail therefore gained a second
visible user turn and broke the Mistral-style alternation pre-flight
(tests/agent/test_summary_role_template_alternation.py::test_zero_user_guard_still_forces_user).
Look through the exempt tail to the last template-visible role instead.
Compaction of a single-prompt (cron) session left no user message after
the handoff, so SUMMARY_PREFIX ordered the model to do nothing and the
scheduler recorded success. Re-append the in-flight task after the
summary.
Fixes#100818
(cherry picked from commit c0ea50ac02fd7553b1ef7deb0162483cb9a1b6ac)
Session kernels: a kernel mid-spawn (proc=None) read as dead, so every
concurrent cell for one owner replaced the registry entry and the
winner's process leaked outside the registry (110 live kernels, 1.1 GB,
330 threads under one 4-capped process). Reap/evict also tore down
kernels with cells attached, rmtree-ing the staging dir under the
spawner. Kernels now track attached cells: only settled kernels are
reaped/evicted, a kernel dropped while busy is torn down by its last
cell, and in-cell registry pops never remove a replacement.
LSP: a directory holding __init__.py is a package, not a project root.
hermes_cli/setup.py matched the python marker list and gave every
worktree a second pyright rooted at hermes_cli/ (70 of 105 reaped
clients in one session, ~40 servers / 8.7 GB live).
OpenCode pins requests sharing an x-opencode-session value to one upstream
backend, which is what keeps its prompt cache warm across a conversation.
Hermes never sent it, so cache ratios on OpenCode traffic were poor.
- agent/opencode_affinity.py: single owner of the header — target detection
(built-in zen/go/free, custom opencode-* providers, any opencode.ai URL)
and the key (affinity scope → conversation root → session id, cron
timestamp stripped), same resolution as OpenRouter/xAI affinity hints.
- build_api_kwargs: merged once after the per-mode builder, so
chat_completions, codex_responses and anthropic_messages all carry it.
- auxiliary _build_call_kwargs: same key from the runtime-main session so
compression/title/vision calls stay on the conversation's backend; the
aux Codex and Anthropic adapters now forward extra_headers.
Closes#81584, #81832 (deepseek-v4-flash 400 without the header).
- Register meta-ai as a live-first picker provider so the /v1/models catalog
leads the picker; new models appear without a PR
- Override fetch_models to exclude non-chat models (muse-image-*, muse-voice-*)
from the picker; new chat model families pass through automatically
- Slim fallback_models to a single safety-net entry (muse-spark-1.2), shown
only when the live fetch fails
- Make data-policy contributor warning model-generic (not hardcoded to 1.2)
so it covers any future -contributor model
- Update test assertion to match generic warning text
LOCAL ONLY — pre-launch, not for push.
tui_gateway.server kicks off prefetch_update_check() at import, which runs
`git rev-parse`/`rev-list` on a daemon thread. The test patched
subprocess.run with a fake that returned None and recorded every call, so
whenever that thread landed inside the test window the assertion saw a
`git` argv and the thread crashed on `.returncode` (red on main since
6e775907d7 made the fetch slower). Filter to non-git spawns and return a
real-shaped proc for everything.
The stampede fix added a `stale_access_token` hint to
resolve_nous_runtime_credentials() so a process whose bearer just 401'd
adopts a token a sibling already rotated instead of re-POSTing the shared
grant — but only the credential-pool caller passed it. The main agent's
401 path (run_agent._try_refresh_nous_client_credentials), the auxiliary
client rebuild, and the proxy adapter all called force_refresh=True with
no hint, so `_already_rotated_by_peer` could never fire: N subagents
hitting hourly expiry still issued N serialized refreshes, each one
invalidating the token a sibling had just adopted.
Live 12-process A/B against a fake Portal: 12 refresh POSTs / 9 distinct
final tokens before, 1 POST / 1 token after.