purge_stale_done_notify_subs only matched status='done', so a task the
circuit breaker parked in 'blocked' kept its notify-sub rows forever on
boards that never archive. Widen the predicate to done OR blocked while
keeping the existing age clause; backlog/ready cards are idle, not
abandoned, and stay exempt (test_gc_spares_reopened_task_even_when_old).
Watcher comment/log and docs updated to say done/blocked.
Closes#100955
Co-authored-by: itsflownium <itsflownium@users.noreply.github.com>
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
`hermes profile delete` read the target profile's gateway.pid raw and
SIGTERMed it. When that pid file was poisoned by a sibling profile's gateway
(the #89315 shape), deleting profile A killed profile B's running gateway.
- gateway/status.py: `_pid_record_belongs_to_profile()` helper — a pid
record whose recorded home differs from the expected profile home is not
ours; legacy records without a home prove nothing and are left alone.
- hermes_cli/profiles.py: `_stop_gateway_process` refuses (and says so)
when the record belongs to another profile; still stops its own gateway.
The stop/restart paths in hermes_cli/gateway.py did not need a guard:
`get_running_pid()` already filters cross-profile records and unlinks the
poisoned pid file before any kill can happen — verified live; the test for
that path now pins the real contract (returns False, other process alive,
poisoned pid file gone).
Live repro (unpatched main): `_stop_gateway_process(tim_home)` -> "Gateway
stopped (PID ...)" and the OTHER profile's process exits -15. After: "Refusing
to stop PID ..." and the process stays alive. 8 tests; sabotage (guard
removed) fails 1.
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.
- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.
Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
NousDashboardAuthProvider._verify_jwt (and the identical hunk in the
self-hosted OIDC provider) folded EVERY PyJWKClient failure into
ProviderError, which the gate translates to HTTP 503
{"detail":"Auth provider 'nous' unreachable"}. That branch fires for
jwt.DecodeError('Not enough segments') — i.e. the bearer is not a JWT at all
(an opaque peer key, a legacy token, garbage) — and for PyJWKSetError (JWKS
fetched fine, foreign kid). Neither involves reaching Portal, which is why
the hosted sjc agents in #94558 returned a fast, well-formed 503 that
survived token re-mint and instance restart while Portal was healthy.
Add one shared classifier, hermes_cli.dashboard_auth.classify_jwks_lookup_error:
only PyJWKClientConnectionError (transport) and an unexpected bare
PyJWKClientError stay ProviderError; DecodeError / PyJWKSetError /
InvalidTokenError become InvalidCodeError so verify_session() returns None
and the middleware proceeds to the next provider / refresh / 401 exactly as
the protocol documents. Both providers now use it.
Live repro (real NousDashboardAuthProvider against a local reachable JWKS
server; and the real gated web_server app): before — opaque bearer ->
ProviderError "JWKS lookup failed: DecodeError('Not enough segments')" ->
503 unreachable; after — verify_session() -> None, gated GET /api/auth/me
with the opaque bearer -> 401; a real JWT against an unreachable JWKS still
-> ProviderError (503).
This does not add /api/v1/message to the public-path allowlist (#94579):
that route has no verifier in this repo, so bypassing the gate would leave a
state-changing ingress fail-open. The correct fix is classification, which
also covers every other opaque-bearer surface.
Refs #94558
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.
Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
Widen #100802 to the two explicit review entry points. The CLI and gateway
/refine handlers built their own snapshot with a shallow list(), which
aliases the nested tool_calls/content containers of the live history. The
review fork sanitizes its transcript in place (sanitize_tool_call_arguments
rewrites function["arguments"]), so a /refine could rewrite the parent's
persisted transcript exactly like the automatic review could (#100795).
Both sites now use _clone_background_review_messages, the same structural
clone the automatic review uses. Regression tests drive the real handlers
and assert the snapshot shares no containers with the live transcript.
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.
`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.
Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.
Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:
1. `hermes profile create --clone-all` and the dashboard/TUI
`mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
`providers.<id>` device-code blocks, and the PKCE singleton file; API
keys are still copied. The clone reads the root grant through the
existing credential-pool root fallback.
2. A named profile with no local rows BORROWS the root grant via
`read_credential_pool()`'s fallback, but every persist
(`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
rows into the profile's own auth.json — materializing a fork on the first
rotation. `persist_pool_entries()` now routes borrowed single-use rows
back to the root store (update-only, under the root lock; never falls
back to a local copy). A borrowed `hermes_pkce` rotation commits its
singleton to the root `.anthropic_oauth.json`, the borrower never prunes
root-seeded rows it cannot see the backing file for, and
`hermes -p <profile> auth add` persists only the profile's own rows.
Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.
Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).
Closes#100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):
- web dashboard CronPage: last_status was never rendered at all — a
delivery_failed job showed a green 'scheduled' badge and only a small red
'delivery: ...' line. New pure cronLastResult() helper maps the closed
literal set to tones (ok=success, delivery_failed/blocked_config=warning,
error/unknown=destructive) and the card now shows an amber
'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
literal; routineLastResult() spells out each one ('Ran, but delivery
failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
on this branch; no consumer compared == 'ok' for success apart from the
cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
detail field carries the reason.
Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
A successful agent run whose delivery failed used to persist
last_status=ok and bury the failure in last_delivery_error. CLI list
painted that as green and the run looked identical to a quiet success.
Record last_status=delivery_failed instead, keep last_delivery_error,
do not increment failure_streak, and teach cron list/doctor not to
treat it as ok.
Fixes#83993
Widen the two salvaged fixes (#100490, #100493) to the whole class:
- match_runtime_outcomes: serve/dashboard rows never borrow gateway
bookkeeping at ANY site — not just the bare hermes-gateway unit name
(#100490) but also relaunched_profiles / externally_supervised_profiles
and the profile-substring unit match (hermes-gateway-work credited the
'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
units (exact names, scope prefix tolerated) or, when the caller passes
the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
feed the Phase-2 reconciliation, so a surviving unmanaged serve is
'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
remedy instead of 'hermes gateway restart', which cannot reach it.
Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
match_runtime_outcomes() treats any default-profile runtime as covered
once the bare "hermes-gateway" unit restarts, regardless of the
runtime's own kind. An sshd-spawned `serve --isolated` backend (no
systemd unit, supervisor "manual-serve") shares the default profile
and gets silently marked "restarted" even though its own PID was never
touched — so the #91277 Phase 2 unaccounted-runtime tripwire never
fires for it and `hermes update` reports success while it keeps
running pre-update code (#100479).
Restrict the "hermes-gateway" special case to kind == "gateway" so a
serve/dashboard runtime under the same profile falls through to
"unaccounted" instead of borrowing the gateway's outcome.
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.
- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
streak raises a halt (identical_call_streak_halt) at
hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
stay exempt; a changed result resets the streak; warning-only sessions are
unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
for pollers, or when results change.
Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
identical failing read_file main: 602 API calls, budget exhausted
branch: 8 calls, repeated_exact_failure_block
identical successful terminal main: 602 API calls, budget exhausted
branch: 5 calls, identical_call_streak_halt
Problem A of #71047: 'hermes config set platforms.telegram.streaming false'
wrote to a key the gateway never reads. The connection config
(gateway/config.py) reads only token/extra/overrides from the top-level
platforms.<name> block, while per-platform display settings (streaming,
show_reasoning, tool_progress, ...) are resolved from
display.platforms.<name>.<setting> (gateway/display_config.py).
Redirect a platforms.<name>.<setting> key to
display.platforms.<name>.<setting> only when <setting> is a known per-platform
display setting (gateway.display_config.OVERRIDEABLE_KEYS), leaving real
connection keys (token, extra, channel_overrides, ...) untouched. The
gateway.display_config import is lazy/try-guarded to avoid a circular import
and to keep the CLI working where gateway is not importable.
Adds tests/hermes_cli/test_config_set_platforms_redirect.py covering the
redirect, connection-key non-redirect, and the no-stray-top-level-platforms
case.
`asyncio.start_unix_server` does not exist on Windows (no AF_UNIX event-loop
support in asyncio), so arming the loop-tick witness in
`loop_heartbeat_forever` raised AttributeError on every native-Windows
gateway start. The broad except swallowed it and recorded
`loop_tick_socket=False`, so every stale-heartbeat probe classified the
gateway as UNKNOWN — never WEDGED, never ALIVE-with-stalled-write. The
two-witness interlock from a1c83ef9 (issue #90502 follow-up) has been
effectively disabled on Windows since it landed: a wedged native-Windows
gateway could never be detected, and an alive one could never be
distinguished from a stalled heartbeat write.
On non-POSIX platforms the witness now arms over a TCP loopback server on
127.0.0.1 (OS-assigned dynamic port) instead:
- same protocol — connect, read one byte "1"
- same semantics — pure in-memory, zero disk I/O, answered only while the
loop is dispatching, armed by the loop task itself (an awaited
`asyncio.start_server` is structurally loop-owned exactly like the Unix
variant, so a wedged loop cannot keep answering pings)
- the assigned port is published in the heartbeat payload as
`loop_tick_tcp_port`, and `probe_gateway_loop_liveness` prefers the TCP
witness when the producer published a port, falling back to the AF_UNIX
socket for POSIX/legacy producers
POSIX behavior is unchanged: the AF_UNIX arm (including the stale-node
sweep) stays gated behind `os.name == "posix"` so the missing attribute can
never raise on Windows again. Legacy heartbeats without `loop_tick_tcp_port`
keep the existing socket-node contract untouched.
Tested end-to-end on native Windows: witness arms, port is published,
`_probe_loop_tick_tcp` answers from an external thread while the loop
dispatches, and the existing loop-liveness suite passes unchanged (the
AF_UNIX structural test still passes — the Unix arm text is preserved
inside the POSIX branch).
Adds two tests pinning the new behavior: an E2E test that arms the TCP
witness and probes it (skipped on POSIX, where the Unix arm is the real
witness), and a structural test that the TCP arm stays awaited on the loop
task and the AF_UNIX arm stays POSIX-gated.
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.
Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.
Addresses #100313
The lazy gateway outbox was missing from the recovery inventory, so a
verified salvage could drop owed replies even when the rows were still
readable. Initialize the destination schema and copy the table.
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.
Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):
- CompressionCommitFence gains mark_commit_watermark_fenced() /
commit_watermark_fenced; compress_context marks the fence right after
capturing get_active_message_watermark() under the durable compression
lock (#75316/#87484) — the property that makes a LATE commit safe: rows
appended after compression start survive both commit paths verbatim as
cloned concurrent tail (archive_and_compact watermark= and
publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
the detached worker (already kept alive via
_defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
user's turn proceeds on the uncompressed transcript at the same 10s
budget, and the summary is adopted at the worker's own watermark-fenced
commit boundary. Unfenced workers are cancelled exactly as before —
never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
block preflight adoption via the same-session cooldown); re-attempt
spacing is covered by the durable compression lock
(_session_has_compression_in_flight). If the worker ends WITHOUT
committing, a done-callback restores the flat non-escalating 60s
retry-after; a successful adoption resets the hygiene failure streak.
The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
to describe deferred adoption and the thinking-model case;
config_defaults.py comment updated. Knob stays config.yaml-only.
Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
start watermark; the fence still gates admission and unfenced/late
results are discarded.
New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
turn still released at the budget, no cooldown while running,
streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.
Fixes#97963
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.
Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.
Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.
On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.
Fixes#100436
Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
* fix(linux): install Lanczos-resized panel icons, not a PNG in scalable
Cinnamon's panel is ~24px. v2026.8.31 dropped the 1024px asset into
hicolor/scalable (SVG-only), so the Mint panel nearest-neighbor scaled
it into a mangled blob. Decode the PNG and write 24/32/48/256 rasters;
undecodable bytes still copy into one indexed dir. Drop the leftover
scalable file.
* test(linux): cover resized hicolor panel icons and stale scalable cleanup
Pin that a decodeable PNG lands as 24×24/256×256 rasters (not scalable),
a leftover scalable copy from v2026.8.31 is deleted, and truncated
PNGs still fall back to an indexed copy.
PR re-review caught that cli-config.yaml.example and the setup wizard
still described the OLD day-stamp gate ("period starts on or after the
day you opted in"). The actual gate (CONSENT_GATE_SQL) requires the
entire package period to be contained in one recorded consent window -
stricter, and privacy-significant at opt-in/revocation boundaries: a
package straddling a revocation starts after opt-in yet is correctly
held back. Both surfaces now state the containment rule literally;
docs A.1 already did.
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).
Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.
Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes#97994.
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:
- hermes_cli/auth.py Codex OAuth clients (token refresh at
auth.openai.com/oauth/token, device-code login, token exchange, usage
probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
every connect eats the full timeout per AAAA before IPv4 is tried, so
auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
had no explicit racing wired.
Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
installs the existing _HappyEyeballsSyncBackend on a ready-built sync
httpx.Client's direct transports (default transport + mounts), skipping
proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
five Codex OAuth/probe endpoints. Best-effort: falls back to default
serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
no custom backend needed. Documented in build_keepalive_http_client and
pinned by tests (contract test on the anyio signature + a live
regression test where a blackholed 100::1 IPv6 addr hangs and local
IPv4 wins in ~250ms instead of the serial connect timeout).
network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.
Refs #13834; follows #94388 (9cce8725).
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
provider label never carries the vendor token to the alias host (#28660);
route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
_profile_scope holds _SKILLS_PROFILE_LOCK (threading.RLock) across the
entire context-manager yield. When get_model_options' worker thread
blocks on fetch_models_dev → requests.get() (up to 15s on a models.dev
cache miss), the lock stays held for the full duration. Concurrent
requests to /api/config (get_config also enters _profile_scope) then
block the main event-loop thread on the RLock, freezing the server.
Switch to _config_profile_scope which uses only the contextvar-based
HERMES_HOME override (thread-safe, no lock) — sufficient for the config
reads + credential checks that build_model_options_payload needs, and
already used by other await-safe endpoints.
Refs #58576
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.
Three surgical changes:
1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
The inlined watcher now routes the respawned gateway's stray
stdout/stderr to the same sidecar log gateway_windows._spawn_detached
uses (DEVNULL only as fallback), so a gateway killed moments after
respawn leaves a trace. Direct implementation of the 4th repro's
hardening suggestion (1).
2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
canonical _spawn_detached, so the respawned gateway's exit-diag /
lifecycle records show whether it escaped the parent Job Object — a
job-teardown kill is no longer indistinguishable from any other silent
death.
3. Post-update resume verifies liveness before vouching
(hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
runs the same provisional-hit + 2s-confirmation liveness poll every
other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
with all_profiles= for the fleet) before printing ✓, writes the #91675
start attestation for the verified PIDs, and fails the resume with a
"restart could not be verified" warning + recovery hint when no stable
gateway appears. Suggestion (2) of the 4th repro; closes the last
silent-success hole in the family (#84185 fixed the cold-start leg,
#91675 the direct-start leg; this is the relaunch leg).
Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.
Fixes the Bug-1 relaunch-trust leg of #48820.
Curated picker lists (OPENROUTER_MODELS + _PROVIDER_MODELS['nous']) gain
claude-fable-5.1 above claude-fable-5 per newest-first ordering; manifest
regenerated via scripts/build_model_catalog.py.
Provider-agnostic metadata verified as already resolving for the 5.1 slug
(no new entries needed): DEFAULT_CONTEXT_LENGTHS fuzzy-matches the
claude-fable-5 prefix (1,000,000), reasoning stale-timeout floor fires
(600s), and both routes bill via official_models_api (live pricing, no
snapshot entry required).
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).
The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).
Fixes#98028Fixes#100325