Commit Graph

35501 Commits

Author SHA1 Message Date
teknium1 a1cbed3592 chore: map atmaksri co-author email for contributor audit
Co-authored-by trailer in the STT salvage uses a non-noreply address.
2026-09-15 18:25:17 -07:00
Sahil Vishnalya afe9e25c57 fix(stt): consume lazy whisper segments inside the CUDA→CPU retry guard
faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2
dlopens the CUDA runtime on the FIRST encode, which happens while the segments
are iterated in `_join_confident_segments()` — outside the try/except that
implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an
NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux
`libcublas.so.12`) the model loads fine, the error escapes the guard, and every
voice note fails with "Local transcription failed: Library cublas64_12.dll is
not found or cannot be loaded" until the user pins `stt.local.device: cpu`.

Materialize the segments inside the guarded block (first attempt and CPU retry)
so the dlopen failure reaches the existing evict-and-retry-on-CPU path.

Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class).
Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`,
`cudnn64_`, `cudart64_`) — the Windows message already matches the existing
"cannot be loaded" marker, proven by the live probe with the reporter's exact
string; the 6-test file was reduced to 2 invariant tests in the existing suite.

Fixes #111929
Fixes #105295
Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate)

Co-authored-by: atmaksri <sri.atmakur@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>
2026-09-15 18:25:17 -07:00
teknium1 0ff9941b76 test: one invariant for the importable-resource stub, drop the change-detector
Keep a single test on the seam that actually failed: an importable
``resource`` module without ``getrlimit`` (the third-party Windows stub
behind #111877/#111879) makes ``_fd_soft_limit`` return None and
``_fd_headroom_ok`` fail open, so session reads proceed.

Drop the ``resource_limits`` test from the salvaged PR: ``apply_nofile_soft_limit``
already wraps the rlimit calls in ``except Exception`` on main, so that test
was green on base and asserted nothing new.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:24:48 -07:00
KoNit-K 576e7c9cfc fix(state): tolerate Windows resource stubs 2026-09-15 18:24:48 -07:00
teknium1 d1bd778a5e fix(checkpoints): clear-legacy trims to two real-path tests, aligns failure line with the log
Replace the three monkeypatch-heavy tests from the salvaged commit (which
faked clear_legacy's return dict, so they could not catch the manager hunk
regressing) with two invariant tests that drive the real cmd_clear_legacy
against a temp checkpoint base: an undeletable legacy-* dir yields exit 2
plus the "Could not delete" line while the archive stays on disk; a clean
sweep keeps exit 0 and the unchanged success line. The green-path guard is
harvested from #111789.

Reword the CLI failure line to "Could not delete N archive(s) (see logs)."
so it matches the manager's WARNING wording and the text proposed in
#111776, and document the exit code in the CLI reference.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:24:22 -07:00
KoNit-K 065bc91846 fix(checkpoints): report legacy archive deletion failures 2026-09-15 18:24:22 -07:00
teknium1 dc7e52874e docs(agent): record why the Z.AI vision default is glm-5.3-flash
Pinned vision ids rot silently (glm-5v-turbo was retired from the Coding Plan endpoints
while still valid on pay-as-you-go), so name the data behind the new pin — the only
image-capable GLM id on every Z.AI surface — and why the pin cannot simply be dropped in
favour of ProviderProfile.default_vision_model() (ZaiProfile returns None, which would route
vision to a text-only chat model).

Co-authored-by: POWERFULMOVES <142271328+POWERFULMOVES@users.noreply.github.com>
2026-09-15 18:23:55 -07:00
KoNit-K c93f2e1d59 fix(agent): update zai vision fallback 2026-09-15 18:23:55 -07:00
teknium1 9aaa71367b fix(catalog): apply the opencode free exclusion on the live-first keyed Zen/Go picker 2026-09-15 18:23:29 -07:00
teknium1 bb5d745ab1 fix(catalog): keep the delisted deepseek-v4-flash-free out of the live OpenCode picker too
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).

Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.
2026-09-15 18:23:29 -07:00
KoNit-K 4766cb2c0c fix(catalog): remove delisted opencode free model 2026-09-15 18:23:29 -07:00
teknium1 cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
finn763 a7dde8a57d fix(sessions): keep a stream-interrupt recovery inside the original session
A stream that dies mid-answer could leave an orphan session behind: source='unknown', its
first message an assistant message and no user prompt anywhere — invisible to the startup
orphan sweep, unrepairable by the session's own creator. Three links made it permanent:

* the token-accounting guard (hermes_state_usage.update_token_counts, the only writer that
  mints source='unknown') mints whenever the row is missing — which is exactly the state a
  recovery dispatch resumed from: _run_prompt_submit (the crash auto-continue, the
  queued-prompt drain) went straight into the turn without persisting the session's own row,
  unlike the prompt.submit handler, so the first durable writer for that session was the
  accounting side effect, and the turn's prompt could not be written at all (the messages FK
  needs the row);
* _insert_session_row's upsert deliberately keeps what the first writer set, so the real
  creator could never repair that placeholder;
* _ORPHAN_SWEEP_SOURCES skipped 'unknown', so such a row stayed ended_at IS NULL forever.

Every dispatch now binds its own row (original session_key, real source) before the turn
writes anything; the upsert repairs the placeholder source when the session's real creator
arrives; the sweep collects a phantom an older build already left on disk. Regression tests
(red before, green after) in tests/tui_gateway/test_stream_interrupt_recovery_orphan.py.

Refs #111999
2026-09-15 18:23:07 -07:00
teknium1 b027a4658e fix: trim DeepInfra reasoning salvage to the invariant and document it
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.

- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
  config -> top-level field table, and the transport main-turn path with
  ``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
  two-directional reasoning control

Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.

Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
2026-09-15 18:22:40 -07:00
KoNit-K 4fe3d288eb fix(providers): send DeepInfra reasoning effort 2026-09-15 18:22:40 -07:00
teknium1 ee49b7d25d fix(agent): file-mutation footer states failed edits, not "files were NOT modified"
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.

- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
  and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
  (mtime_ns, size) snapshot; a later success clears every entry with the same
  identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
  entries whose file changed since the failed call, so a receipt-less
  mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
  way the file tools resolved them.

Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.

Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:22:12 -07:00
teknium1 0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1 66c9440826 fix: stream retry no longer replays the old model on the switched-to provider
After a mid-turn /model switch while a stream was stalled, the streaming
retry loop re-sent the request it had captured at construction time. That
payload still named the OLD model, but every stream (re)open builds its
request client from the LIVE agent, so the new provider's base_url received
a foreign model slug: 404 "Not found the model ...", then the turn sat in
the provider's rate-limit hold (#112121).

_StreamingCall now records the route (model, provider, base_url, api_mode)
its api_kwargs were built for. When a retry is about to be issued and the
live route differs, the streamer stops and hands the transient error back
to the turn loop instead. The turn loop already rebuilds the request per
attempt for the CURRENT route (turn_api_request.build_api_request: model,
wire shape, prompt-cache decoration, provider request overrides), so a
re-keyed model alone would still have shipped a payload shaped for the old
provider. Non-streaming requests have no in-process retry, and the
fallback / restore-primary paths go through the same turn-loop rebuild, so
this is the only site that replayed a captured route.

Fixes #112121

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:21:17 -07:00
teknium1 d54ae02606 fix(pricing): custom-provider /models prices already per-million are no longer inflated 1e6x
`_extract_pricing`'s generic path copied catalog values verbatim, while
usage_pricing unconditionally applies OpenRouter's per-token convention and
multiplies by 1e6. A provider quoting USD per 1M tokens (Neosantara 0.6/M,
Crof cost.input 0.04/M) or declaring `unit: per_1m_tokens` therefore priced
at $600,000/M and corrupted estimated_cost_usd in state.db and every cost
report summing across providers.

Normalize at the producer, where Novita/DeepInfra unit handling already
lives: an explicit `unit` beside the rates wins (per_token / per_1k_tokens /
per_1m_tokens); without one, a token rate at or above $0.001 per token
($1,000/MTok - no real model) can only be a per-million quote. Output keeps
the per-token-string contract, so the consumer is untouched; `request` fees
and per-token catalogs pass through unchanged.

The $0.001/token magnitude threshold is the one proposed in #34263 by
@Bartok9 (earliest fix); #112036 by @kvnloo proposed the same heuristic at
the consumer.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:20:51 -07:00
teknium1 5435ac8cc4 fix(aux): clamp extra_body.reasoning effort too so auxiliary.<task>.reasoning_effort ultra never reaches the wire 2026-09-15 18:20:13 -07:00
teknium1 9e45a90488 fix: clamp aux reasoning effort once before profile projection
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.

Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.

Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
2026-09-15 18:20:13 -07:00
KoNit-K c02db64077 fix(agent): clamp auxiliary ultra reasoning effort 2026-09-15 18:20:13 -07:00
teknium1 f55d4f6747 fix(aux): bare-custom AuthError yields no custom endpoint instead of a stale env OPENAI_BASE_URL 2026-09-15 18:19:50 -07:00
teknium1 a9fabe43c4 fix(runtime): keep the bare-custom fail-fast off local aliases
Key the post-ladder AuthError on the literal `custom` request again, in
addition to the dead runtime shape (provider=custom, empty api_key). The
previous commit widened it to every alias that resolves to custom (ollama,
vllm), which broke `/model <direct-alias>` switching: _creds_for_switched_provider
resolves the alias provider tolerantly and _apply_direct_alias_endpoint
supplies the alias endpoint AFTER that call, so the early raise turned a
working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery, red in CI).

The issue's own scope (#111741) is the bare non-routable placeholder; a
credential-less local alias is a legitimate intermediate state there.
2026-09-15 18:19:50 -07:00
teknium1 e9a54c48f2 fix(runtime): bare-custom fail-fast keys on the dead runtime shape, covers custom aliases
Tighten the post-ladder guard from #111744 to the exact failing shape: a
resolved ``custom`` runtime with an EMPTY api_key. Every other custom rung
(named entry, direct alias, local bypass, pool, key_cmd) yields a key, a
callable or the ``no-key-required`` placeholder, so the loopback heuristic and
the has_usable_secret() re-check were dead branches — and keying on the
requested name alone missed aliases that resolve to custom (``ollama``,
``vllm`` with nothing configured), which still returned the credential-less
OpenRouter fallback and died at agent construction as "No LLM provider
configured". The AuthError now names the requested provider, so cron's
runner/preflight, the CLI, the gateway and the TUI gateway (all of which
format AuthError or walk their fallback chain on it) show the culprit.

Tests folded to two invariants: raise-and-name (bare + alias, plus the
OpenRouter-key control from #111750) and the loopback no-auth control.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-15 18:19:50 -07:00
KoNit-K 879d65ec78 fix(runtime): fail fast for bare custom credentials 2026-09-15 18:19:50 -07:00
teknium1 c1bbcf9712 fix(agent): hint-preview truncation log names the real remedy; invariant tests
Follow-up to the two salvaged commits (#111777, #111781 by @KoNit-K):

- agent/prompt_builder.py::_truncate_content — with queue_warning=False the
  logged line no longer tells the operator to "pin a larger
  context_file_max_chars, or use a larger-context model": the subdirectory
  hint cap is a constant neither knob raises. It now points at the read_file
  recovery the marker already discloses.
- tests/gateway/test_startup_environment_probe.py — replace the
  call-detection test with the behavioural invariant: an oversized SOUL.md in
  HERMES_HOME and a warm-up leave the truncation-warning queue empty for the
  next default-executor task (the api_server turn path runs on that executor
  without copy_context, which is how the boot warning reached a foreign
  session).
- tests/agent/test_subdirectory_hints.py — fold the new drain assertion into
  the existing oversized-hint test (same fixture) and pin that the log carries
  no context_file_max_chars advice.
- agent/AGENTS.md, website/docs/.../context-files.md — the hint cap is 32,000
  (docs said 8,000) and is fixed; document that it is logged, not surfaced as a
  chat warning.
2026-09-15 18:19:26 -07:00
KoNit-K fc7cbc7d8e fix(agent): suppress subdirectory hint truncation warning 2026-09-15 18:19:26 -07:00
KoNit-K 87661da314 fix(gateway): defer context files until active turn 2026-09-15 18:19:26 -07:00
teknium1 9b3042902e fix: trim hygiene bound salvage to two invariant tests and document the knob
Salvage follow-up to the cherry-picked #112055 hunk (@Finn763):

- tests: fold the landed-commit control into the turn-hold-miss test and drop the
  two extra end-to-end tests (sub-limit identity is already pinned by the pure
  contract test) so the fix carries exactly two invariant tests.
- gateway/run_turn.py::bound_model_input_without_hygiene: drop the isinstance /
  limit<=0 defences — the caller always passes the transcript list and a
  _knob()-validated positive int, so the checks could never fire.
- docs: configuration.md now says hygiene_hard_message_limit also bounds the
  model payload on any turn where hygiene did not land (payload-only, disk
  untouched, landed summaries adopted as-is).

Co-authored-by: Noor-Sol <187156657+Noor-Sol@users.noreply.github.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:18:58 -07:00
finn763 5ae06d38d3 fix(gateway): bound the model payload when hygiene has not landed (#111988)
A long-lived gateway DM rots because a turn where hygiene did not commit
(turn-hold expiry, idle/ceiling timeout, cooldown, or one already in flight)
fed the model the FULL uncompressed transcript — watermark-fenced summaries
can land late, but a turn with none had no fail-closed bound at all.

bound_model_input_without_hygiene() keeps the leading system/session_meta
setup rows plus the newest tail, total <= hygiene_hard_message_limit, and
never starts the kept tail on an orphaned tool result. It is deterministic
and payload-only: the stored transcript is untouched, so the agent's durable
prefix slice (history_offset) and the disk rows are unchanged, and the
hygiene compressor still sees the full transcript.

Applied at the three exits of _hmwa_run_session_hygiene where hygiene did
not land. A landed compression publishes a NEW list on attempt.history, so
object identity (`attempt.history is history`) decides — an adopted summary
is passed through byte-identical, and below the limit the bound returns the
same list object (no copy, no behaviour change).
2026-09-15 18:18:58 -07:00
teknium1 c5c71ea1ad fix(docs): document the 10 s SSE keepalive comment for custom parsers 2026-09-15 18:18:30 -07:00
teknium1 8e00832d9b fix(api-server): /v1/runs events stream shares the SSE keepalive cadence; trim tests
The durable-run events stream (`GET /v1/runs/{id}/events`) hardcoded its own 30 s idle
keepalive, so it still fell outside the ~20 s idle deadline common to remote
OpenAI-compatible clients even after CHAT_COMPLETIONS_SSE_KEEPALIVE_SECONDS dropped to
10 s. Point it at the shared constant so every API-server SSE writer (chat completions,
Responses, native session stream, run events) keeps the wire alive on the same cadence.

Rewrite the salvaged regression test as two invariants: an idle OpenAI stream writes a
`: keepalive` comment inside the remote-client window (module-local fake clock, no global
`time.monotonic`/`asyncio.wait_for` patching, which would also steer asyncio), and the run
events stream follows the shared constant end-to-end through aiohttp. Both are red on
origin/main.
2026-09-15 18:18:30 -07:00
KoNit-K 67a959ac20 fix(gateway): keep SSE active during preflight 2026-09-15 18:18:30 -07:00
teknium1 eba9b5551c docs: note the automatic non-streaming fallback for contentless SSE frames
The streaming section documented only the manual model.streaming escape
hatch; users hitting a degraded gateway need to know the session flips to
non-streaming on its own and why the warning appears.
2026-09-15 18:18:08 -07:00
teknium1 857133babd refactor: trim empty-frame salvage to the translated error only
Drop the raw JSONDecodeError belt from _is_provider_stream_empty_frame_error:
every chat stream iteration goes through _iter_provider_stream_chunks, which
already translates the decode failure, so the belt guarded a path that does
not exist. Drop the explicit classifier entry for the new code: unknown codes
already resolve to the same retryable unknown verdict (verified live: identical
ClassifiedError with and without the entry).
2026-09-15 18:18:08 -07:00
moxian 71281fb266 fix(agent): contentless SSE keepalive frames retry without streaming
An empty `data:` frame (or a frame carrying only `event:` / `id:`) is a legal SSE
no-op, but the OpenAI SDK still hands it to `json.loads`, which raises
`JSONDecodeError` with an empty document. The streaming helper translated that
into `ProviderStreamError(provider_stream_non_json_data)`, nothing recognised it
as recoverable, and the main loop retried streaming -- identically -- until the
retry budget ran out: "API call failed after 3 retries: Provider stream returned
non-JSON SSE data". A degraded gateway answers EVERY streaming request that way,
so the retries only repeated the failure while the same request sent
non-streaming succeeded seconds later.

An empty document means no payload, which is a different fact from a malformed
payload: give it its own code, and when it appears before any delta switch the
session to non-streaming (the existing `_disable_streaming` mechanism, already
used for "stream not supported", Bedrock IAM denials and adapter-returned final
responses) so the retry goes out on a channel the degraded gateway can answer.
Non-empty payloads keep today's fatal semantics; failures after deltas keep the
existing stream-drop handling.

Tests: the turn-level recovery test drives the real SDK decoder over a real
httpx response and is red on base (three identical streaming attempts, turn
lost); the boundary test pins the malformed-payload path so a future "ignore bad
frames" change cannot swallow genuine provider errors.
2026-09-15 18:18:08 -07:00
teknium1 416a8177c2 perf(plugins): scan each plugin's source for removed imports once per process, not once per profile
A multiplex gateway runs plugin discovery for every served profile, and each
discovery ran plugin_compat.scan_plugin (an ast.parse + two ast.walk passes per
file) over every external plugin: on an 11-profile host that was ~0.4s per
profile of identical work on the gateway boot path, ahead of adapter connect.
The scan result is now cached process-wide on the plugin dir's (relpath,
mtime_ns, size) signature whenever the loaded manifest is used; a changed file
rescans, a caller-supplied manifest bypasses the cache.

Measured over the 11-profile MCP discovery loop on the reporting host:
3.78s -> 3.01s; per empty profile 0.25-0.30s -> 0.11s.
2026-09-15 15:11:33 -07:00
teknium1 f81c33cb6a fix(update): stop the fleet settle poll early when the restarted unit is dead
Salvage of #111385 (@JoaoMarcos44): the 30s -> 120s settle window is kept so a
slow host's gateway can publish its state stamp. This commit bounds the other
side of that trade: when every restarted systemd unit reports neither active
nor activating, the successor has died and nothing will ever publish, so the
poll fails closed at once instead of spending the full 120s. Unknown states
(no units, systemctl missing or slow) keep waiting. Also drops the test
assertion pinning the 2s poll cadence.
2026-09-15 15:10:41 -07:00
joaomarcos c99fee4e0a fix(update): wait for fleet state publication after supervised restart
A systemd unit can be active before the replacement gateway finishes
bootstrap and publishes gateway_state.json. The 30s poll in
_collect_fleet_snapshot() then returned [] with rows_expected=true,
causing _verify_fleet_after_update() to mark the restart incomplete and
exit(1) before clearing fleet_restart_pending. Every later CLI
startup/doctor therefore printed a false "did not restart" warning even
though the live gateway already served expected_sha.

Allow the default systemd startup budget plus publication slack by
extending the bounded settle window to 120s. Keeps fail-closed on
stale/down/empty after the deadline, only widens the window for slow
bootstraps (e.g. Raspberry Pi).

Fixes #111272
2026-09-15 15:10:41 -07:00
teknium1 c91770b268 fix(desktop): route tile title heals when its plugin route registers late
A route tile restored (or opened) before its plugin route registered kept the
humanized-path fallback as its tab title forever: paneMirror recomputes titles
only when one of its atoms changes, and watchRouteTiles listened to $routeTiles
alone. #112147 made the pane CONTENT heal on late registration; the title
still lagged.

routes.ts exposes $routesVersion, an atom bumped from registry.subscribeArea
(ROUTES_AREA) only while listened to, and watchRouteTiles passes it as `also`
so the sync re-runs and the pane re-registers with the contribution's title.
One test, red on origin/main.
2026-09-15 15:07:24 -07:00
teknium1 212f661048 fix(desktop): remembered plugin page survives boot when the session list beats disk plugins
A remembered plugin-page route is session-shaped until its route registers,
and disk plugins load through an async IPC chain. When the backend was
already running (remote/URL connection, macOS close-reopen with the process
alive) the session list could arrive first; the restore effect then read
'/html-gallery' as a session id nobody owned, navigated to the last chat and
erased the remembered route, so the page was lost on every later boot too.

The disk door now publishes $diskPluginsScanPending for the duration of its
first scan, and the restore latch holds a session-shaped remembered route
while that scan is pending, mirroring the existing "sessions not loaded yet"
latch. Once the scan settles the route is either a registered page (restored)
or a genuinely stale session (dropped as before). Two tests: the late
registration restores; the stale session still drops.
2026-09-15 15:07:03 -07:00
KoNit-K 5910de20bc fix(gateway): setup.status / setup.runtime_check scope the launch profile under multiplex
`_readiness_check` bound a secret scope only for a named non-launch
profile and used nullcontext for the launch profile. Once the process
multiplexes (`set_multiplex_active(True)`) `get_secret` fails closed, so
the launch profile's `setup.runtime_check` died on the first profile-scoped
read inside `resolve_runtime_provider` (`HERMES_CODEX_BASE_URL` in
`_pool_entry_mode_and_url`, added by b62bb2a3d5) and the Desktop showed
onboarding while sessions — which resolve under
`_session_profile_runtime_scope` — worked fine.

Route the launch profile through the same helper: `profile_home=None`
binds the launch profile's frozen `.env` scope only when multiplexing is
active (`_profile_runtime_scope_tokens` returns None otherwise, keeping
the single-profile `os.environ` fallthrough for systemd / `op run`
injection). The unknown-profile `ok:False` answer is untouched — no
`@_profile_scoped`, whose `_profile_home` raise would turn it into an
error.

Slimmer shape than the PR's `scope_launch_profile` flag: the flag guarded
nothing the helper does not already decide, and setup.status reads the
same `.env`-derived state.

Fixes #112061
2026-09-15 14:24:49 -07:00
teknium1 57c9e92ae3 test(tui): setup.runtime_check agrees with the session fallback chain
Invariant for #111775 on the registered RPC: primary AuthError + one complete
fallback entry -> ok:True with the fallback provider and model (what
_make_agent builds); explicit `provider` still answers the primary's strict
failure. Red on origin/main (ok:False, "No Anthropic credentials found.").
2026-09-15 14:24:49 -07:00
KoNit-K 6f975b768e fix(tui): setup.runtime_check resolves like session creation and reports the model
Without an explicit `provider`, `setup.runtime_check` called the strict
`resolve_runtime_provider(requested=None)` while `_make_agent` goes through
`_resolve_agent_model_runtime` (startup model + provider pin, then the
configured fallback chain). With the primary blocked and a working fallback
entry the probe answered ok:False (the primary's auth error) and the Desktop
showed onboarding for a backend whose sessions built fine. The probe also
reported `model: null` because the resolver never populates that key.

The default case now runs the session builder's resolver; an explicit
`provider` stays a strict single-provider check (onboarding verifies the
provider just connected — another provider's fallback must not mask a failed
connection) and resolves against the startup model like the session would.
Both branches report the selected model.

Fixes #111775
2026-09-15 14:24:49 -07:00
Austin Pickett 05e82b741d fix(desktop): Tab Strip → Auto copy names the other-zone exception 2026-09-15 16:07:39 -05:00
Austin Pickett 36010d0bba fix(desktop): a lone chat keeps its tab strip while another chat zone is open
Dragging a session tab out of the main strip into its own zone left the
workspace alone in main. Auto treated a lone pane as "not a tab", and the
uncloseable workspace is not stranded, so main lost its tab and its "+" —
while the tile it had just split from kept a strip of its own. Two chats
side by side, one with tabs and one without, reads as "the tabs
disappeared"; ⌥⌘T only helped because it wrote an explicit `always`.

The resolver gains one input, `siblingMainZone`: on auto, a lone main tile
keeps its strip whenever another zone in the layout also hosts a main tile.
A chat that is the whole window stays chromeless, side chrome is untouched,
and an explicit `never` still wins so Hide tabs keeps working.

Both callers read it from `$mainTileZoneCount`, a computed over the tree,
the hidden set and the registry version. It is a number, so zones re-render
only when a main zone appears or goes — never per sash-drag frame — and the
store's answer (what the toggle command flips against) and the strip on
screen cannot disagree.
2026-09-15 16:07:39 -05:00
teknium1 5d59366010 feat(plugin-catalog): add 66 community plugins from the Discord #plugins-skills-and-skins sweep
Every entry is pinned to the commit reviewed on 2026-09-15 and passed, at that
commit: `hermes plugins validate` (declared tools/hooks/middleware match
registrations, no built-in tool collisions) in both the main venv and a bare
`pip install -e .` venv, `tools.plugin_guard.scan_plugin` with a non-dangerous
verdict, a self-updater grep over every .js file, and a vendor-credential-store
write grep over every .py file. Capabilities and requires_env are copied from
each plugin.yaml; category was assigned by hand.
2026-09-15 12:55:47 -07:00
teknium1 eb562b10ad docs(plugin-catalog): allow maintainer-curated sweep entries alongside owner submissions
Teknium ruled that maintainers may add batches of community plugins from a
reviewed sweep instead of waiting for each owner to submit. Rule 5 and the
user-guide checklist now say so, and give authors the explicit right to adjust
or remove a swept-in entry via their own PR.
2026-09-15 12:55:47 -07:00
teknium1 eaef52371f fix(desktop): late-contributed keybind actions reach the Keybinds settings map
Same class as the plugin-route bug: KeybindSettings subscribed to
useContributions(KEYBINDS_AREA) but discarded the snapshot and called the
impure allKeybindActions() in render, so the React Compiler memoized the
action list without the subscription as a visible input. A plugin keybind
registered after the tab mounted never appeared, despite the comment
promising "appear/disappear live".

allKeybindActions()/contributedKeybinds() take an optional contribution
snapshot (default: registry read, so the store/imperative callers are
unchanged) and the settings tab passes its subscription through. One test,
red on origin/main.
2026-09-15 12:08:52 -07:00