After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).
Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):
1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
writable open, typically the user's first NEW session. The dashboard
backend now schedules one writable open of its own state.db from the
lifespan (daemon thread, never blocks the ready-probe socket, never
raises), so the store is brought current before the first session-
list poll on every `hermes serve` / `hermes dashboard` / Desktop
headless entrypoint.
2. _reconcile_columns caught sqlite3.OperationalError around every
ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
orphaned sibling backends made the ALTER fail silently — startup
"succeeded" with a half-reconciled schema, and the open-time lock
patience (#74478) never saw the error because it was swallowed
inside first. Now: "duplicate column" races stay at DEBUG,
locked/busy re-raises so _connect_and_init_with_lock_patience
retries the whole idempotent init with jittered backoff, and any
other failure (e.g. un-ADDable NOT NULL) logs at WARNING.
Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.
Fixes#79531Fixes#80037
Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
Persist backend ownership for reliable cleanup, park inactive panes, and
evict unreferenced transcripts so Desktop stays responsive over long sessions.
💘 Generated with Crush
Assisted-by: Crush:gpt-5.6
* feat(dashboard-auth): extend RFC 8252 native sign-in to password providers
The desktop app runs password sign-in for gated gateways in an embedded
Electron BrowserWindow, where OS password managers (macOS Passwords /
iCloud Keychain autofill) cannot reach the form — Chromium-in-Electron
has no bridge to them, so users retype credentials by hand even though
the /login form already carries the right autocomplete attributes.
The existing RFC 8252 native flow (system browser + loopback + PKCE)
solves exactly this for OAuth providers, but was explicitly disabled for
password providers on the grounds that they have "no IDP round trip to
broker". The brokering is still worth having: it moves the credential
form into the system browser, where password-manager autofill just works.
Gateway-only change; the desktop needs no changes (runNativeLogin is
already page-agnostic), and older desktop builds pick the capability up
automatically once the gateway advertises it:
* /auth/native/authorize now accepts a supports_password provider:
register the pending broker authorization as usual, then 302 the
system browser to the interactive /login form with the opaque
broker_state in the gateway's PKCE cookie (the same server-controlled
channel the OAuth branch uses) instead of an IDP redirect.
* /auth/password-login: when the server-set PKCE cookie carries a
broker handle, a successful credential check completes the pending
authorization exactly like the /auth/callback native branch — mint
the one-time loopback code, return the loopback redirect (validated
loopback-only at authorize time) as `next`, clear the PKCE cookie,
and set NO session cookies. A lapsed broker is a clean 400 telling
the user to restart sign-in; a failed credential attempt leaves the
pending entry intact so the user can retype.
* /api/status now advertises "native_pkce" whenever any interactive
session provider is registered (previously only for non-password
providers), so the desktop selects the system-browser strategy for
password-only gateways.
Security posture is unchanged from the existing flow: loopback-literal
redirect_uri enforcement, PKCE S256 binding, single-use short-TTL codes,
constant-time comparison, and the same rate limiter on password attempts.
Tests: full authorize → /login → password-login → loopback → token →
bearer round trip, wrong-password keeps the pending entry, lapsed broker
→ 400, no-broker browser login keeps minting cookies, and the /api/status
advertisement for password-only gateways.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(dashboard-auth): bind native password completion to the authorize-time provider
Review follow-ups for #75808:
* /auth/password-login now enforces that body.provider matches the
provider recorded in the server-set PKCE cookie by
/auth/native/authorize before completing a pending native
authorization. /login renders a form for every session provider, so
without this a native flow started for provider A could be completed
with provider B's credentials, binding B's session into A's pending
entry. The mismatch is rejected BEFORE credential verification (no
session minted, no oracle) and preserves both the pending entry and
the cookie, so the user can still submit the correct provider's form.
Covered by a two-password-provider E2E regression test.
* Update the two docs spots that still said password-only providers do
not advertise native_pkce (website desktop-native-signin guide and the
auth_flows type comment in web/src/lib/api.ts).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: map contributor email for #75808 (buffpesos)
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
* fix(dashboard): suppress Ctrl+C shutdown traceback
* fix(dashboard): extend clean Ctrl+C exit to the Windows serve branch
The Windows loop-factory branch (and its pre-0.36 asyncio.run fallback)
runs under the same uvicorn capture_signals() re-raise as the POSIX
path, so console Ctrl+C leaked the identical KeyboardInterrupt
traceback there. Guard both serve calls with the same clean-exit
contract, keeping the import-resolution try/except comment accurate
(genuine serve-time errors still propagate).
Also ports the reworded POSIX-test docstring (the serve path is no
longer 'byte-for-byte unchanged'), wraps the POSIX KI test in
pytest.fail so a regression reports red instead of aborting the pytest
session, and adds the windows_only sibling test.
Extends #52970 to the whole bug class.
* chore: map contributor email for @wangs1203
* test(dashboard): actually exercise the pre-0.36 Windows fallback KI contract
Copilot review caught that patching uvicorn._compat.asyncio_run with
raising=False makes the import succeed, so _runner is non-None and the
extra asyncio.run patch never covered the fallback. Split it out: a
dedicated windows_only test halts the _compat import (None in
sys.modules) so the fallback branch is genuinely selected, then asserts
the same clean-KI contract on bare asyncio.run.
---------
Co-authored-by: Emiya·Leon <wangs.coder@gmail.com>
Adds hermes_cli/model_selection_guards.py: a single evaluation point that
runs every selection guard (cost + the new data-policy guard) and returns
the warnings that fired. All seven model-selection surfaces (CLI picker,
cli.py TUI modal, gateway typed /model, dashboard web_server, TUI gateway,
Telegram and Discord pickers) now call the registry instead of importing
model_cost_guard directly — so the data-training-tier warning from
PR #81416 fires everywhere at once, and future guards need zero surface
wiring.
Guard modules keep their public APIs; existing mock patch points
(hermes_cli.model_cost_guard.expensive_model_warning) remain valid.
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).
- gateway/disk_status.py (new): collect_disk_status() samples
shutil.disk_usage(HERMES_HOME) and classifies pressure
(critical: <256 MB free or >=95% used; elevated: <512 MB free).
Never raises — degrades to pressure="unknown" with null telemetry,
same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
resource banner with worst-first triggers (disk critical > memory
critical > OOM restart > disk elevated > memory elevated) and
cascading dismissals — hiding the top trigger surfaces the next one
instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
with English fallback per existing pattern).
Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.
- web_server CONFIG_SCHEMA: fold the one-field models_dev category
(models_dev.url) into the agent tab via _CATEGORY_MERGE, matching the
established pattern for single-field categories (slice 7,
test_no_single_field_categories).
- image_routing._lookup_supports_vision: pass allow_network=True to
get_model_capabilities. The vision-capability lookup runs when an
image actually needs routing (not per conversation turn), and the
#31179 text-only-main guard depends on catalog data — with the new
allow_network=False default a cold cache returned 'unknown', which
falls back to attempting the call and reintroduced the #31179
failure shape (slice 8, test_text_only_main_skipped_when_no_
aggregator). This preserves that path's historical
network-on-cold-cache behavior; the fetch stays 4h-TTL cached and
backoff-limited.
Renames the openai-codex provider's display label across the CLI
(hermes model picker, provider labels), the dashboard OAuth accounts
catalog, and the Desktop onboarding + settings provider pickers.
Slug, aliases, and auth flows are unchanged.
* fix(gateway): pass live adapters to cron fire webhook's fire_due
The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).
Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.
Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).
* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)
The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.
Restore the invariant that the GATEWAY owns cron execution:
- Dashboard route: after verifying the NAS JWT and resolving the job's
profile, FORWARD the fire to the gateway api_server's own
/api/cron/fire on loopback, NAS bearer preserved (the gateway
re-verifies the JWT — defense in depth, no new trust link), and pass
the gateway's response through. Gateway unreachable → 503 so NAS
retries per the Chronos contract (non-2xx = retryable; the store CAS
de-dupes the eventual double fire). Deliberately NO local-execution
fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
per target profile (config.yaml extra.port → API_SERVER_PORT from
process env or the profile's .env → 8642), with /p/<profile>/ prefix
routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
first boot when absent (never overwrites an operator value), so the
loopback api_server passes its startup guard on hosted images. The
fire route itself is NAS-JWT-authed; the key gates the rest of the
api_server surface. The listener binds 127.0.0.1 by default and the
Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
topology and the 503-retry semantics.
Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.
* fix(cron): read the profile api_server port via the canonical config loader
CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).
Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.
* fix(gateway): only messaging platforms count for the scale-to-zero arm gate
The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).
The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
The custom-endpoint REST handlers ran bare load_config/save_config, so
every add/activate/delete landed in the process-level default profile
regardless of which profile the desktop settings UI was targeting. A
provider added under a non-default profile silently went to default:
visible only in default-bound sessions, absent everywhere else, and
un-addable to another profile without hand-editing its config.yaml.
Scope all four handlers (list/upsert/activate/delete) to the requested
profile via _config_profile_scope, matching /api/config, and spread the
active profile into the four hermes.ts wrappers alongside their existing
validateCustomEndpoint sibling.
An unreadable profile dir made (entry / '.env').exists() raise
PermissionError out of the sidebar fallback, 500ing /api/profiles.
Found by hostile fixture during live E2E of the scandir conversion.
Two-part fix for the dashboard fd exhaustion reported in #81547:
1. Raise RLIMIT_NOFILE soft limit on startup (before uvicorn binds).
macOS defaults to 256 for LaunchAgent processes — too tight for the
dashboard which opens 3 fds (db+wal+shm) per SessionDB per request
across all profiles. After days of polling the soft limit exhausts
and every os.listdir/open raises OSError [Errno 24]. The helper raises
to the hard limit (or minimum 4096), matching the reporter's ulimit
workaround. No-op on Windows (no resource module).
2. Replace bare Path.iterdir() with context-managed os.scandir() in four
dashboard hot paths: _fallback_profile_dicts, file manager list,
checkpoint listing, and plugin discovery. iterdir() returns a
generator that holds an open directory fd until fully consumed; if
an exception interrupts iteration the fd leaks. os.scandir() is an
explicit context manager that guarantees close on exit, following
the same idiom already used in /api/fs/list.
Tests: 6 passed, 3 skipped (resource-module tests skip on Windows).
_spawn_gateway_restart() now calls _reap_unsupervised_gateway_orphans()
before spawning a new `hermes gateway restart` child. On desktop-app
restart the old serve exits but its gateway child gets reparented to
launchd (PPID=1) and keeps its platform connection alive. The new
serve then spawns a fresh gateway, resulting in two live gateways
racing the same connection.
The reap was already implemented for the CLI restart path (#75936) but
the dashboard's _spawn_gateway_restart path was not covered.
Fixes#77276
On Desktop serve startup, reap orphan gateway processes (PPID=1) left
behind by a previous serve session that exited abnormally. This prevents
the old and new gateways from racing for the same QQ WebSocket
credential, which splits messages across parallel session trees (#77276).
Production incident: the orphan reap killed a legitimate SSH remote backend
started by another client machine. Its process sat at ppid 1 with the same
cmdline shape as a genuine orphan, and the exclusion list only covered THIS
app instance's children — ownership by OTHER clients was invisible.
The reap now treats every backend.lock.json under ~/.hermes/desktop-ssh/*/
as an ownership claim: lock payloads are schema-validated (mirroring
remote-lifecycle.ts) and their PIDs are excluded both before the scan and
re-checked after it (defense in depth against a lock written mid-scan).
Regression tests cover the exact incident shape: a lock-owned PID and a
genuine orphan with identical process shapes — only the orphan is reaped.
Also: fold the new single-field `runtime` config category into `agent`
(_CATEGORY_MERGE) and fix an env leak in the serve-startup test
(HERMES_SERVE_HEADLESS restored via monkeypatch) so the combined suites
run green in any order.
When Desktop exits uncleanly, leftover `hermes serve --host 127.0.0.1 --port 0`
processes can be reparented to pid 1 and keep full MCP trees alive. The next
boot then stacks another backend on top of the corpses until EMFILE kills
sidebar/session APIs and tabs disappear.
- Detect Desktop-local serve shape (loopback + ephemeral port 0)
- Only reap processes whose ppid is 0/1 (true orphans)
- Spare fixed-port remote serves (e.g. --port 9119) and HERMES_DESKTOP_CHILD_PID
- Run at Desktop backend start (HERMES_DESKTOP=1) before parent-death watchdog
Complements parent-death watchdog (prevents future orphans) and configurable
nofile soft limit (capacity floor). Together these stop the multi-backend
pile-up cascade observed on macOS Desktop SSH/local installs.
An unclean desktop exit (crash / SIGKILL / update handoff) stranded every
`hermes serve` profile backend as an orphan (ppid=1) still serving, each
holding its MCP child subtree — 31 orphans / ~1.3 GiB RSS on one install.
Root causes + fixes:
- serve had no parent-death watchdog: add _start_parent_death_watchdog() in
web_server.py (mirrors slash_worker.py), gated on HERMES_PARENT_PID; os._exit
cascades to MCP watchdogs. No-op for standalone `hermes serve`.
- desktop passes HERMES_PARENT_PID in both serve spawn env blocks (main.ts).
- POSIX teardown now group-kills (process.kill(-pid, ...)) so MCP grandchildren
die too (backend-child.ts + waitForBackendExit SIGKILL fallback).
Windows path unchanged (forceKillProcessTree). Tests updated + passing.
Two follow-ups to the off-loop move, from external review (both verified,
the second larger than reported):
- Config read-modify-write handlers moved to worker threads could now
interleave — _CONFIG_LOCK covers each load/save individually, never the
span between them; the event loop used to serialize these accidentally.
New _CONFIG_MUTATION_LOCK (worker-threads only, so it can never block
the loop) held across the whole load→mutate→save span in all seven RMW
handlers. update_config_raw skipped: it's a full-document replace with
no server-side read, so a lock cannot close its client-side window.
- The review flagged two skills routes still taking _SKILLS_PROFILE_LOCK
on the event loop; a systematic audit of hermes_cli/web_routers/ found
24 on-loop routes (skills 5, mcp 9, tools 10, cron 1). All moved to the
same inner-_run + asyncio.to_thread pattern, mutating ones under the
mutation lock, uniform lock order (_SKILLS_PROFILE_LOCK →
_CONFIG_MUTATION_LOCK). Await-safe _config_profile_scope routes, plain
def routes, and already-threaded routes unchanged.
Regression tests: concurrent theme+font updates both survive (fails with
the lock nulled: "theme write lost to a concurrent font write"); event
loop stays responsive while the profile lock is held during GET
/api/skills. 214 tests passing across the touched suites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The diagnostics loop watchdog caught GET /api/config freezing the gateway
event loop for >1s, stack-sampled blocking on _SKILLS_PROFILE_LOCK inside
_profile_scope. Any async handler that entered _profile_scope (process-wide
threading lock) or called load_config()/save_config() on-loop could stall
every chat and WebSocket at once while a slow lock-holder ran.
Move 28 such handlers to the existing inner-_run + asyncio.to_thread
pattern (contextvar-safe: the whole scope enter/body/exit stays inside one
worker thread). Handlers using the await-safe _config_profile_scope, plain
def endpoints (FastAPI threadpool), and tui_gateway's contextvar-only
decorator are unaffected and unchanged.
Regression test holds _SKILLS_PROFILE_LOCK in a thread while calling
GET /api/config and asserts an event-loop heartbeat keeps ticking; it fails
against the pre-fix code.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to the salvaged registration contract:
- share one _raise_if_cron_registration_error() helper for the two
byte-identical dashboard 424 except-blocks (web_server + cron router,
via the existing late() seam)
- add endpoint-level 424 coverage for /api/cron/blueprints/instantiate
(previously only the sync worker was tested)
- give chat/CLI surfaces a human-facing user_message() (job name, no
exception class name) and add a recovery hint (pause/resume or update
re-registers via provider reconcile) to the model/REST message
- consolidate five inline provider test doubles into one ABC-subclassing
make_cron_provider conftest factory; the web_server test double now
subclasses CronScheduler so an ABC rename fails loudly
- narrow the wrapper facade to keyword-only (**kwargs) and route the
tool's partial-failure return through tool_error()
After `hermes update`, the desktop sidebar showed "No sessions yet" until
the user's first message. #72424 added sessions.last_activity_at, which
list_sessions_rich now selects — but column adds only land through
_reconcile_columns() in the writable _init_schema, and read-only opens
skip that by design. Every sidebar read path opens state.db read-only, so
each poll raised "no such column: s.last_activity_at" until the first
prompt's lazy session-row persist forced a writable open and reconciled.
A heal for exactly this class already existed (_open_session_db_for_profile
probes the read-only handle and does a one-time writable reopen on
staleness), but its probe was a hand-written four-column list that never
learned last_activity_at — it went stale three days after shipping. And the
batched sidebar route (/api/profiles/sessions/sidebar) bypassed the helper
entirely, swallowing per-profile failures into an errors array the desktop
never surfaces, so the incident produced an empty sidebar with clean logs.
The fix removes the maintenance burden instead of paying it once more:
- hermes_state_schema.schema_read_probe_statements() derives one
`SELECT <every declared column> FROM <table> LIMIT 0` per table from
SCHEMA_SQL via the existing _parse_schema_columns() — the same source of
truth the writable reconciler diffs against, so any future ADD COLUMN is
probed with no list to update. Column references are table-qualified:
an unqualified double-quoted identifier that fails to resolve silently
degrades to a string literal (SQLite's double-quoted-string misfeature)
and would make the probe pass on exactly the store it exists to catch.
- web_server splits the heal into a path-level _open_session_db_at_path
(semantics unchanged) so the cross-profile session routes can share it;
both profiles.py loops and _count_status_active_sessions (the remaining
raw read-only sibling) now open through it. The heal stays a helper
rather than a SessionDB classmethod on purpose: escalation-to-writable
must remain an explicit caller decision — update_cmd.py opens read-only
mid-update and must never write.
- Exhaustion guard: if the writable heal SUCCEEDS and the re-probe still
fails (a schema problem ADD COLUMN cannot express), the store is marked
exhausted — warn once, skip the probe, serve reads probe-less — instead
of re-running the full writable init on every poll against a possibly
live DB. A FAILED writable open (transient lock) is deliberately not
recorded, so the next poll retries the heal.
- The per-profile swallow sites in profiles.py now also log a deduplicated
warning, so a persistent read failure is loud in errors.log even though
the response errors array stays invisible to the sidebar.
Tests: probe/SCHEMA_SQL coverage invariants (tests/test_schema_read_probe.py),
last_activity_at added to the /api/sessions heal parametrize, a sidebar-route
heal test reproducing the shipped symptom (errors == [] and the session
returned against a store missing the column), and an exhaustion test pinning
exactly one writable open. The sidebar and last_activity_at tests fail on
main.
The dashboard now mints an action_id per backend update, hands it to the
spawned `hermes update` via HERMES_ACTION_ID, and reuses an in-flight
update action instead of spawning a duplicate. The updater prints a
bounded `=== hermes-update completed <id> ===` receipt on every success
path — normal, zip, dependency-repair, and the no-op "Already up to
date!" path that previously ended with no terminal marker at all
(#58764) — so the Desktop can prove completion across the dashboard
restart boundary instead of guessing from stale log text.
Co-authored-by: Vitor Cepeda Lopes <vitor@vitorcepedalopes.com>
Co-authored-by: doncazper <caztronics@yahoo.com>
Three fixes for the Desktop/TUI cold-start stall where the event loop
is blocked for ~14s between HERMES_BACKEND_READY and the first
prompt (#60800):
1. copilot_auth: skip subprocess fallback when any
Copilot env var is explicitly set (even if invalid). The user
expressed token intent via env var; silently substituting a CLI
token is surprising and the subprocess adds up to 5s on Windows.
2. tui_gateway/ws: run resolve_skin() via asyncio.to_thread so config
loading + skin engine init do not block the WS read loop during
the cold-start RPC burst.
3. web_server: extend _warm_gateway_module to pre-import the heavy
module chains (auth, copilot_auth, runtime_provider, skin_engine,
inventory, model_switch) that the first WS connection + RPC burst
would otherwise import on the loop thread. These trigger .pyc
compilation and Defender scans on Windows (15-30s per the existing
comment) and were not covered by the original gateway-only warm.
Tests: 5 new tests in test_cold_start_gil_stall.py + 2 new tests in
test_copilot_auth.py. All 36 copilot_auth tests + 16 ws/web_server
tests pass.
Every hashed bundle chunk under /assets/ was served with no caching
directives, so each dashboard load re-fetched (or at best revalidated)
every JS/CSS chunk. Those filenames carry a Vite content hash — the
bytes behind a given URL can never change; a rebuild mints new
filenames referenced by a freshly served index.html.
Mark them Cache-Control: public, max-age=31536000, immutable:
- the /assets StaticFiles mount, via a subclass that stamps the header
on 200s only (404s stay uncached — a rebuild can create the file),
- serve_css, preserving its X-Forwarded-Prefix url() rewrites for
/fonts/, /fonts-terminal/, /ds-assets/, /assets/.
index.html keeps no-store, no-cache, must-revalidate — it is the
mutable entry point that binds users to the current hashes.
The original PR also added hand-rolled per-request gzip compression of
asset responses; that part is deliberately dropped. This server is a
localhost-default dashboard backend: compressing every response on the
CPU to save loopback bandwidth is a pessimization, and callers that
front it with a real proxy already get compression there.
Salvaged from PR #28543 (idea by @sea-monsters; gzip groups dropped as
described above).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Re-derivation of aydnOktay's twin clamp PRs onto current main (the
session-list endpoints moved into web_routers/; the analytics endpoints
gained asyncio.to_thread wrappers since the originals):
- limit le=100 on /api/sessions, /api/sessions/search and the
/api/profiles/sessions fan-out (one unbounded request could drag every
session row + correlated-subquery preview work out of SQLite, times
every profile's state.db on the fan-out).
- days ge=1 le=365 on /api/analytics/usage + /api/analytics/models
(huge or non-positive values force full-history InsightsEngine work or
inverted windows; the UI only offers 7/30/90 presets).
FastAPI Query bounds reject at the validation layer (422). 8 new tests;
both clamp classes mutation-checked (clamp removed -> its tests fail).
On dashboard-only sessions nothing else executes check_fn warmers (they
live only in the tool-schema build), so the hub's read-only cache lookup
would report auth_required=False forever. On a cache miss, schedule a
deduplicated daemon-thread probe off the request path; the short hub TTL
surfaces the verdict on the next fetch.
/api/status is the desktop's boot liveness probe (polled ~1/s) but since
#60537 every call ran a full topology scan — per-profile yaml.safe_load
(pure-Python loader), psutil process probes, realpath walks — in the
default executor. On multi-profile installs concurrent polls pile up and
hold the GIL 14-16s, starving the event loop: the WS sidecar cannot
flush gateway.ready, the desktop times out into the next stall, and boot
escalates to the 'Hermes couldn't start' overlay (#60800).
Memoize the scan behind a 10s TTL with a collapse lock so concurrent
polls share one scan. Topology only changes on gateway start/stop, so a
<=10s stale badge is an acceptable trade for not starving the loop. The
cache also keys on the collector's identity: tests monkeypatch
_collect_profile_gateway_topology per case, and the identity check keeps
them hermetic (a swapped collector is a miss) without a reset hook.
py-spy captures during a failing boot land in _profile_platform_ports ->
yaml.safe_load on executor threads (7 profiles, Windows). After: one
cold-start scan, zero recurring stalls, desktop boots.
Extract _resolve_session_token() in hermes_cli/web_server.py so tests can
exercise token resolution directly instead of importlib.reload(ws), which
re-executed the whole module mid-suite (fresh FastAPI app + token) and
split module identity between test and app state.
Salvaged from PR #39038 by @rodboev (maintainer-endorsed direction);
rebased onto the rewritten test_web_server.py — dropped the PR's hunks for
test_falls_back_to_random_token's old body (test deleted in the prune,
re-added here in the PR's new form).
Co-authored-by: Rod Boev <rod.boev@gmail.com>