_denormalize_config_from_web ran the new switch_model validation inside the
pre-existing 'except Exception: pass' disk-read fallback. Any rejection
(offline models.dev, no OpenRouter key, unlisted model) left config['model'] a
flat string, and PUT /api/config's deep-merge wrote that string over the on-disk
model: dict, destroying provider/base_url/api_mode/context_length/model_slots.
Only the load_config() read keeps its fallback; validation propagates as the
HTTPException(400) the caller's http_failure passes through. Invariant test PUTs
a rejected model against a real config.yaml and asserts 400 + byte-identical file.
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
The dashboard's `_eager_reconcile_own_session_db` did an unconditional
writable `acquire()` at every startup. When the gateway shares that
state.db the dashboard became a second long-lived writable SessionDB
owner: a close-time WAL checkpoint plus a possible FTS rebuild in
`_init_fts`, the two-writer vector behind the corruption reports in
#107688 and #100896 ("5 live SessionDB handles" precursor, gateway +
dashboard both holding the WAL).
Route the startup reconcile through `_open_session_db_at_path(...,
read_only=True)`, which already bootstraps a missing store and heals a
stale/malformed schema through exactly ONE writable open before
reopening read-only. A healthy store now gets zero writable opens from
the dashboard while the #79531/#80037 "bring schema current before the
first poll" contract is kept (existing heal test unchanged).
Live repro (healthy store, count writable SessionDB.__init__ calls from
the startup worker): before=1 after=0.
Reported-by: #107688, #100896 (@kokhlo diagnosis)
Refs #107688#100896
/model <name> on provider A, where the name is only known to provider B
(static catalog or OpenRouter), switched the session to B even when B had
no key: an immediate 401 for most vendors, and for OpenRouter — whose
runtime resolves with an EMPTY key instead of raising — a silent switch
onto a metered aggregator. The dashboard's flat Model field had two more
copies of the same guess ("vendor/model on a native provider" → openrouter).
detect_provider_for_model() now walks its ladder as candidates and skips any
target without credentials (env/.env key, auth-store login, or a usable
credential pool entry). Exceptions: the user NAMED the provider (/model nous)
or there is no current provider yet ("auto") — then the guess is handed back
so the credential step fails loudly instead of silently ignoring input. A
vendor/ prefix naming a provider declared in `providers:` is a selection, not
a guess, and always routes. The dashboard fallbacks apply the same gate.
Tests that pinned "switch to OpenRouter/vendor with no key" now grant the
credential they assumed; two new invariants cover the gate.
Preserve the two contributor fixes, slim them to two behavioral invariants, and enter explicitly requested homes even inside a nested scope. Real native remote Desktop changes Disabled to gateway_stopped for default and named profiles; direct API controls preserve explicit disable and empty-profile isolation. Unit A/B and regression suites remain queued under the shared campaign lock.
The scoped branch of _platform_enablement consulted only config.yaml's
platforms: section, but the `hermes gateway setup` wizard writes .env
credentials and never a platforms: entry. The desktop always sends
?profile=default (normalizeProfileKey maps the primary profile to
`default`), so the Settings - Messaging page showed a working bot as
"Disabled" while /api/status reported it connected (#104614).
Mirror _enable_from_env (gateway/config_env.py): env credentials alone
enable a platform, an explicit enabled: false still wins. Only the
profile's own .env (env_on_disk) is consulted, so the root install's
os.environ credentials still never leak into a profile's state.
Fixes#104614
The salvaged #95487 adds an expression index over sessions.last_activity_at.
Two web_server tests emulate a legacy store with ALTER TABLE ... DROP COLUMN
last_activity_at, which SQLite refuses while an index references the
column ('error in index ... after drop column'). Drop the index first —
the same adjustment the PR already made in test_schema_read_probe.py. A
real pre-column store has neither the column nor the index, so the healed
path under test is unchanged.
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.
Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.
Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.
On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.
Fixes#100436
Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:
Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
rejected the session token.
The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.
Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.
Fixes#96490
_denormalize_config_from_web only wrote model_context_length into the
on-disk model dict inside the branch gated on `model` also being present
in the payload. That was harmless when the frontend always sent the full
config, but the prior commit switched Settings autosave to send only the
diff (diffConfig), so editing the Context Window control alone omits
`model` from the payload and the context-length edit is silently thrown
away. The mirror case regressed too: editing `model` alone now omits
model_context_length from the diff, and the old code treated that missing
key the same as an explicit 0, wiping an existing context_length override
that the user never touched.
Track whether model_context_length was actually present in the payload
and only mutate context_length when it was, independent of whether
`model` also changed.
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:
- web_server._open_session_db_at_path: the one-writable-open heal only
caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
to decode SQLite's own error message over corrupt file bytes) bypassed
it, so the heal documented for malformed schema never fired (#98924
Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
vtables whose probe raised UnicodeDecodeError, the same too-narrow
catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
could not open, so prompt.submit streamed the turn while persisting
nothing (#98924 Failure 2). It now returns False and prompt.submit
fails the RPC with code 5072 so desktop maps it to a toast, mirroring
the disk-full/5070 convention. session.create stays silent per its
pinned degraded-mode contract.
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
The Electron shell boots by fetching / and extracting
window.__HERMES_SESSION_TOKEN__ to authenticate /api/ws
(dashboard-token.ts adoptServedDashboardToken). Headless serve 404'd
every path, so when the renderer's spawn token drifted from the
backend's live token — e.g. hermes update replaced the backend and the
env pin no longer matched — the renderer had no way to adopt the served
token, the WebSocket handshake failed, and the primary window
white-screened (#95575).
Serve a minimal token-only HTML page at the exact root path in
mount_spa()'s headless branch, matching the renderer's extraction regex.
Gate it on app.state.auth_required read at request time: a gated
(non-loopback / remote public_url) serve keeps returning the 404 JSON so
the session token never leaks past the loopback boundary. Every other
path stays 404 JSON — the SPA remains unserved.
Regression tests: TestHeadlessServeTokenPage (3 cases) — verified to
fail against the pre-fix headless branch.
The apt/docker CLI tests pinned exit 1; refusals are now exit 2
(refused-by-contract, distinct from errors). The web_server guards
patched the module-local detect_install_method alias, which the shared
admission gate no longer consults — patch hermes_cli.config directly.
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".
- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.
Closes#82614
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
Update the sibling tests that pinned the old use_gateway-writing
contract: image/video selector and reconfigure rows now assert the
single provider string ('nous' managed / 'fal' BYOK) plus legacy-key
popping, the stt/video picker writes drop the use_gateway expectation,
the web_server managed-browser select asserts the persisted 'nous'
cloud_provider, and explicit-local STT pins no-cloud-fallback against
a stored raw-config selection.
Backend: /api/status now carries a stable random install_id persisted once
under the root HERMES_HOME, shared by every profile of the install.
Desktop: roster enumeration captures it per connection (TTL-cached probe),
buildAgentRoster collapses same-install rows with a deterministic canonical
pick (active > local > ssh > remote > cloud > earliest), the @name-device
handle rule runs after the collapse, and Settings → Gateways shows a
display-only 'Same backend as' hint. Backends without install_id bypass the
collapse (fully backward compatible).
When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.
Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.
Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.
Fixes#87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
After `hermes update`, an existing state.db on an old schema made every
GET /api/sessions poll fail with sqlite3.OperationalError "no such
column: s.last_read_at" (or s.last_activity_at) until something
unrelated forced a writable open — the desktop sidebar showed "No
sessions yet" while every row sat intact on disk (#79531, #80037).
Two remaining root causes (the stale hand-written read probe was
already replaced by the SCHEMA_SQL-derived probe on main, prototyped in
draft PR #80030 by @Tilly-YL):
1. Migrations ran lazily: _init_schema/_reconcile_columns only ran on a
writable open, typically the user's first NEW session. The dashboard
backend now schedules one writable open of its own state.db from the
lifespan (daemon thread, never blocks the ready-probe socket, never
raises), so the store is brought current before the first session-
list poll on every `hermes serve` / `hermes dashboard` / Desktop
headless entrypoint.
2. _reconcile_columns caught sqlite3.OperationalError around every
ALTER TABLE ADD COLUMN and logged at DEBUG. Lock contention from
orphaned sibling backends made the ALTER fail silently — startup
"succeeded" with a half-reconciled schema, and the open-time lock
patience (#74478) never saw the error because it was swallowed
inside first. Now: "duplicate column" races stay at DEBUG,
locked/busy re-raises so _connect_and_init_with_lock_patience
retries the whole idempotent init with jittered backoff, and any
other failure (e.g. un-ADDable NOT NULL) logs at WARNING.
Regression tests: a store missing sessions.last_read_at is healed by
the eager startup reconcile and serves list_sessions_rich; a locked
ALTER propagates and is retried to success by the open lock patience;
duplicate-column races stay quiet; other ALTER failures warn.
Fixes#79531Fixes#80037
Reported-by: @yenhunghuang (#79531) and @FLOW3R0111 (#80037)
Root-cause analysis: @wangyi0177-eng (stale read probe) and
@www654cc-pixel (_reconcile_columns DEBUG-swallow under lock
contention); draft PR #80030 by @Tilly-YL prototyped the probe fix.
Dedupe key now includes tool_call_id/tool_calls/tool_name: compaction
copies carry those fields verbatim, so identical tool messages across
generations still collapse, while distinct tool calls sharing
role/content/timestamp are never merged. Add endpoint-level coverage
for the desktop's real read path (limit + order=latest +
include_compacted=true).
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).
- gateway/disk_status.py (new): collect_disk_status() samples
shutil.disk_usage(HERMES_HOME) and classifies pressure
(critical: <256 MB free or >=95% used; elevated: <512 MB free).
Never raises — degrades to pressure="unknown" with null telemetry,
same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
resource banner with worst-first triggers (disk critical > memory
critical > OOM restart > disk elevated > memory elevated) and
cascading dismissals — hiding the top trigger surfaces the next one
instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
with English fallback per existing pattern).
Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.
Choosing Computer Use should be a config flip, not a hunt for
'hermes computer-use install'. Three provisioning rungs:
- install.sh / install.ps1 pre-install cua-driver (best-effort,
non-fatal, time-boxed at 660s above the upstream installer's 600s
lock window; --skip-computer-use / -SkipComputerUse to opt out;
Termux and unwritable-/Applications skipped cleanly)
- PUT /api/tools/toolsets/{name} (dashboard + desktop toggle) spawns
the background 'hermes tools post-setup cua_driver' action when the
toolset is enabled while the binary is missing — previously the
toggle 'saved' but the tool never appeared in the schema because
check_computer_use_requirements() couldn't find the binary
- hermes tools interactive flow already installed via
_toolset_needs_configuration_prompt/_POST_SETUP_INSTALLED (unchanged)
Docs: computer-use.md enabling section rewritten around the new flow;
installation.md documents --skip-computer-use.
get_env_value/load_config read through the shared os.environ mirror that
save_env_value writes, so a reader-based assertion cannot prove which
profile's store actually received the write. Read the two profiles'
config.yaml and .env directly instead, and cover the credential path.
The custom-endpoint REST handlers ran bare load_config/save_config, so
every add/activate/delete landed in the process-level default profile
regardless of which profile the desktop settings UI was targeting. A
provider added under a non-default profile silently went to default:
visible only in default-bound sessions, absent everywhere else, and
un-addable to another profile without hand-editing its config.yaml.
Scope all four handlers (list/upsert/activate/delete) to the requested
profile via _config_profile_scope, matching /api/config, and spread the
active profile into the four hermes.ts wrappers alongside their existing
validateCustomEndpoint sibling.
OFFSET paging made the streaming export O(n^2) on huge transcripts;
after_id keyset paging keeps each page seek O(1). Adds after_id to
SessionDB.get_messages (ascending-only, guarded against latest/offset
combos).
After `hermes update`, the desktop sidebar showed "No sessions yet" until
the user's first message. #72424 added sessions.last_activity_at, which
list_sessions_rich now selects — but column adds only land through
_reconcile_columns() in the writable _init_schema, and read-only opens
skip that by design. Every sidebar read path opens state.db read-only, so
each poll raised "no such column: s.last_activity_at" until the first
prompt's lazy session-row persist forced a writable open and reconciled.
A heal for exactly this class already existed (_open_session_db_for_profile
probes the read-only handle and does a one-time writable reopen on
staleness), but its probe was a hand-written four-column list that never
learned last_activity_at — it went stale three days after shipping. And the
batched sidebar route (/api/profiles/sessions/sidebar) bypassed the helper
entirely, swallowing per-profile failures into an errors array the desktop
never surfaces, so the incident produced an empty sidebar with clean logs.
The fix removes the maintenance burden instead of paying it once more:
- hermes_state_schema.schema_read_probe_statements() derives one
`SELECT <every declared column> FROM <table> LIMIT 0` per table from
SCHEMA_SQL via the existing _parse_schema_columns() — the same source of
truth the writable reconciler diffs against, so any future ADD COLUMN is
probed with no list to update. Column references are table-qualified:
an unqualified double-quoted identifier that fails to resolve silently
degrades to a string literal (SQLite's double-quoted-string misfeature)
and would make the probe pass on exactly the store it exists to catch.
- web_server splits the heal into a path-level _open_session_db_at_path
(semantics unchanged) so the cross-profile session routes can share it;
both profiles.py loops and _count_status_active_sessions (the remaining
raw read-only sibling) now open through it. The heal stays a helper
rather than a SessionDB classmethod on purpose: escalation-to-writable
must remain an explicit caller decision — update_cmd.py opens read-only
mid-update and must never write.
- Exhaustion guard: if the writable heal SUCCEEDS and the re-probe still
fails (a schema problem ADD COLUMN cannot express), the store is marked
exhausted — warn once, skip the probe, serve reads probe-less — instead
of re-running the full writable init on every poll against a possibly
live DB. A FAILED writable open (transient lock) is deliberately not
recorded, so the next poll retries the heal.
- The per-profile swallow sites in profiles.py now also log a deduplicated
warning, so a persistent read failure is loud in errors.log even though
the response errors array stays invisible to the sidebar.
Tests: probe/SCHEMA_SQL coverage invariants (tests/test_schema_read_probe.py),
last_activity_at added to the /api/sessions heal parametrize, a sidebar-route
heal test reproducing the shipped symptom (errors == [] and the session
returned against a store missing the column), and an exhaustion test pinning
exactly one writable open. The sidebar and last_activity_at tests fail on
main.
The dashboard now mints an action_id per backend update, hands it to the
spawned `hermes update` via HERMES_ACTION_ID, and reuses an in-flight
update action instead of spawning a duplicate. The updater prints a
bounded `=== hermes-update completed <id> ===` receipt on every success
path — normal, zip, dependency-repair, and the no-op "Already up to
date!" path that previously ended with no terminal marker at all
(#58764) — so the Desktop can prove completion across the dashboard
restart boundary instead of guessing from stale log text.
Co-authored-by: Vitor Cepeda Lopes <vitor@vitorcepedalopes.com>
Co-authored-by: doncazper <caztronics@yahoo.com>
Every hashed bundle chunk under /assets/ was served with no caching
directives, so each dashboard load re-fetched (or at best revalidated)
every JS/CSS chunk. Those filenames carry a Vite content hash — the
bytes behind a given URL can never change; a rebuild mints new
filenames referenced by a freshly served index.html.
Mark them Cache-Control: public, max-age=31536000, immutable:
- the /assets StaticFiles mount, via a subclass that stamps the header
on 200s only (404s stay uncached — a rebuild can create the file),
- serve_css, preserving its X-Forwarded-Prefix url() rewrites for
/fonts/, /fonts-terminal/, /ds-assets/, /assets/.
index.html keeps no-store, no-cache, must-revalidate — it is the
mutable entry point that binds users to the current hashes.
The original PR also added hand-rolled per-request gzip compression of
asset responses; that part is deliberately dropped. This server is a
localhost-default dashboard backend: compressing every response on the
CPU to save loopback bandwidth is a pessimization, and callers that
front it with a real proxy already get compression there.
Salvaged from PR #28543 (idea by @sea-monsters; gzip groups dropped as
described above).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 2-line alias had zero production consumers (web_server calls
get_usage_breakdown directly). Tests rewired onto the real API; the
contracts they pin are unchanged. Stale test docstring fixed.