Commit Graph

16383 Commits

Author SHA1 Message Date
ywatanabe 419a050427 fix(agent): preserve tool cache in iteration summary 2026-09-14 20:35:28 +05:30
teknium1 4748caff76 fix(gateway): explicit tool_progress new/all keeps text progress in un-cardable Slack chats
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.

Test proven red on the salvaged head (adapter.sent == [] with `all`).
2026-09-14 07:46:51 -07:00
Victor Kyriazakos a4e2a82a6d fix(gateway): preflight task-card destination before transport fallback 2026-09-14 07:46:51 -07:00
Victor Kyriazakos dab7eebf47 fix(gateway): resolve progress mode and intent from the same source 2026-09-14 07:46:51 -07:00
Victor Kyriazakos bc125d59d5 fix(gateway): null tool_progress inherits; name the task-card suppression latch
Review findings (Salt, adversarial pass on the two preceding commits):

- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
  counted as an explicit mode because the gate tested key presence, while
  the display resolver skips None and inherits. Null resolved to Slack's
  tier default `off` and disabled cards, which is the default-off trap the
  change exists to avoid. Explicit intent is now a non-None value (or the
  env bridge). Tests cover null at each level plus null-over-global-all;
  mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
  destinations, so the name no longer described the field. Renamed to
  `publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
  described the opt-in as independent of tool_progress. Rewritten: cards
  follow an operator-written off (including /verbose), null inherits, an
  un-threaded chat with the card lane active shows no tool progress, other
  native failures keep the editable fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos 3412490ad1 fix(gateway): no text tool progress when a Slack chat cannot host a task card
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.

Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos ed25a40917 fix(gateway): explicit tool_progress off disables Slack task cards
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.

Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.

Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
2026-09-14 07:46:51 -07:00
kshitijk4poor 62e5f46656 test(streaming): one Anthropic event-stream fake for the three parse-error tests
Three inline context-manager classes shared the same __enter__/__exit__
boilerplate and differed only in the events yielded before the raise.
Also drop the unreachable 'or agent.base_url' fallback: every
anthropic_messages init path sets _anthropic_base_url, and the two
sibling call sites read it bare.
2026-09-14 20:01:19 +05:30
kshitijk4poor 982e504262 fix(anthropic): partial tool names are reset per stream attempt
A tool_use block name recorded by a stream attempt that died before any
visible text survived into the next attempt: only the deltas_were_sent
mid-tool branch cleared result["partial_tool_names"]. When the retry then
streamed plain text and dropped, the partial stub blamed the stale tool
("Stream stalled mid tool-call (old_tool)") and the stale name could make
a later attempt look mid-tool-call when deciding whether the drop is
retryable. Reset it in _start_stream_attempt alongside
provider_tool_in_flight, which already has attempt-local semantics.

Regression: two-attempt stream (tool_use start + parse error, then text +
drop) — stub content and emitted deltas carry no stale tool name.
2026-09-14 20:01:19 +05:30
kshitijk4poor 3d88259483 fix(anthropic): retry a malformed tool-JSON stream with buffered tool input
Follow-up to the cherry-picked fix: keep the classifier widening
("expected value at line" is now a transient stream parse error on the
main turn) but replace the messages.create() fallback with a retry on the
same stream wire.

Why not create(): the fallback ran outside Relay (lost request rewrites),
outside _handle_stream_error (could replace text already shown to the
user with a different generation), ticked no liveness events for the
whole buffered payload, and a bare identical retry still re-emits the same
malformed JSON.

Why not drop the fine-grained-tool-streaming beta (#108583/#109056): live
probe on claude-sonnet-4-5, ~500-line tool call - beta on: max inter-event
gap 1.6 s; beta off: 139 s zero-event gap while Anthropic buffers the
args, which the 180/240 s stale-stream detector kills on larger payloads
(the regression 80a899a8e2 fixed).

Instead, on a parse error the retry sets `eager_input_streaming: false`
on every tool for that request only (the SDK/API per-tool field overrides
the legacy beta header), so Anthropic returns buffered, server-validated
args while the happy path keeps fine-grained streaming. A tool_use that
started streaming is registered in partial_tool_names so the mid-tool
transient retry fires the same way it does on the chat_completions wire.
2026-09-14 20:01:19 +05:30
joaomarcos dfb4caf4b7 fix(anthropic): recover malformed streamed tool JSON 2026-09-14 20:01:19 +05:30
teknium1 d57c28a554 fix(tests): compression stall-fallback tests stop racing the 0.2s ceiling
Under a loaded runner the primary stall plus the fallback retry overran the
0.2s total ceiling, so the retry never started and attempts==1 failed
intermittently (seen once in a 40-worker tests/agent run). Idle stays 0.05s;
the ceiling moves to 2s per the >=2s wall-clock rule in AGENTS.md.
2026-09-14 07:21:18 -07:00
teknium1 d28938d3da fix(computer-use): screenshot dedup forgets its last frame at a compaction boundary
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
2026-09-14 07:21:18 -07:00
teknium1 31964ff4c6 fix(computer-use): dedup keyed by the scoped session, cleared on release
Key the screenshot-dedup state by the same profile-scoped session id the
backend cache uses, so two multiplexed profiles sharing a session id (or a
DISPLAY) never dedup against each other's frames, and forget the state in
release_computer_use_session so a re-created session's first capture
always delivers pixels. Reword the unchanged note to cover the aux-vision
path (where the prior result was an analysis, not an image). Tests: the
dispatch path (explicit capture + capture_after) honours the streak cap;
release forgets state.
2026-09-14 07:21:18 -07:00
Teknium 682b973b32 feat(computer-use): stop resending unchanged screenshots
Port from openclaw/openclaw#129924: a capture whose pixels are
byte-identical to the previous capture of the same target in the same
session returns its full text metadata (element index included) plus an
explicit 'screen unchanged' note instead of the multimodal image block.

Adapted for Hermes: openclaw gates dedup on per-frame context-epoch
tracking; Hermes bounds staleness with a consecutive-omission streak cap
(2) so full pixels are re-delivered before compaction could evict the
referenced image. Dedup state is per-session (no cross-session leaks),
append-only (no history rewrites — prompt cache prefixes untouched),
and skipped entirely when no session_id is present.
2026-09-14 07:21:18 -07:00
kshitijk4poor 66ddd5f83c fix(auth): only a Codex token refresh writes through to root
Following the grant's source on every save made a fresh device-code
login (or `hermes auth import`) under a profile that had been borrowing
root's Codex grant overwrite root's account instead of creating the
profile's own. Redirecting a save into another file is the exception, so
it is opt-in: the refresh path passes write_through=True; login, import
and recovery keep saving locally. The two save branches collapse into one
(store, path, set_active) triple.

Test: root discovery on Windows comes from LOCALAPPDATA — set it so the
fixture's root is the resolved root on every host.
2026-09-14 19:49:36 +05:30
kshitijk4poor 0ff20dc98a test: trim Codex write-through tests to two invariants on the real profile layout
The picked tests monkeypatched _auth_file_path/_global_auth_file_path
directly and leaned on a HOME override to dodge the pytest seat belt.
Isolate the way the rest of tests/hermes_cli does instead: Path.home ->
tmp_path and HERMES_HOME -> <root>/profiles/<name>, so the fixture drives
the same get_default_hermes_root() resolution production uses. Drop the
classic-mode test (no new behaviour: source == active store is the
pre-existing save path). Two invariants remain: root-borrowed refresh
lands in root (singleton + pool) with no profile shadow; profile-owned
grant stays local with root untouched.
2026-09-14 19:49:36 +05:30
liuhao1024 6bd29f26f6 fix(auth): write profile-refreshed Codex tokens through to the global store
Codex refresh tokens are single-use with rotation-family reuse
detection. _save_codex_tokens resolved the state via the profile's
root fallback but always persisted into the ACTIVE (profile) store, so
a profile-scoped refresh left the global store holding the consumed
refresh token — the next process to read it replayed it and OpenAI
revoked the whole rotation family, forcing a manual device-code
re-auth (#87503; observed four times on one multi-profile deployment).

Mirror the xAI source-aware save (#43589/#74339): resolve the state
with _load_provider_state_with_source; when the grant came from the
global root, write the rotated chain back to root only — singleton AND
credential_pool entries, under the root store's own lock, without
creating a shadowing profile key. Best-effort, with the same pytest
seat belt as the xAI path.
Fixes #87503
2026-09-14 19:49:36 +05:30
kshitijk4poor 329257c060 fix(cli): keep venv-pinned console scripts exec-able on relaunch
The launcher guard rejected every python shebang, so a pip/uv console
script pinned to the running venv (#!<venv>/bin/python) was also
discarded in favour of `python -m hermes_cli.main`. Reuse
linux_desktop_entry._shebang_escapes_running_env, which already knows
that `env` shebangs escape and a shebang inside the running
interpreter's directory does not; only the escaping launcher loses the
venv. Also drops the second shebang classifier the fix had introduced.
2026-09-14 19:49:12 +05:30
frozen 5c2ddeb55a fix(cli): preserve venv across self-relaunch 2026-09-14 19:49:12 +05:30
kshitijk4poor 7e7641561b refactor(gateway-windows): one death predicate, no second process scan
attested_gateway_died() re-ran find_gateway_pids() (current profile only)
although both callers had just proven the process table empty with
all_profiles=True, and it re-implemented check_start_attestation's
liveness rule. Callers now pass the liveness they hold (current_pids=[])
and both probes share _attested_dead(), so the consuming and read-only
twins cannot drift.
2026-09-14 19:47:09 +05:30
kshitijk4poor da382a413e test(update): trim #109538 coverage to two invariant tests
Four new tests overlapped: plan-time and spawn-time attested-death overrides
both exercised the same predicate via monkeypatched lambdas. Collapse to:
- one end-to-end test using a real attestation marker in a tmp home: dead
  attested gateway keeps the plan under Desktop ownership, survives the
  spawn-time re-check, and the marker is consumed by the spawn;
- one probe test: no marker / null or non-list pids / non-dict / non-JSON all
  read False (fail closed), alive and clean-exit read False, read-only when
  it does read True.
Existing #76129 tests keep their attested_gateway_died=False pins unchanged.
2026-09-14 19:47:09 +05:30
ennheng 830c8f443d fix(update): keep the Windows cold-start plan for a dead attested gateway
A Desktop self-update hand-off exits the app before the updater runs and can
kill the messaging gateway in those same seconds (#109538), so the updater's
discovery finds no live PID while the one-shot start attestation still
vouches for the dead one. Both Desktop-ownership checks then read "nothing
running" as "nothing to restore" and the bot stayed down until a manual
start.

Consult the attestation non-destructively before Desktop-owned lifecycle
suppresses a cold-start: a vouched-for PID gone without a clean ledger exit
keeps the plan and is restored; no attested death preserves the #76129 skip
unchanged.
2026-09-14 19:47:09 +05:30
joaomarcos 6bc0e9e6df fix(agent): same-model review fork keeps the parent's affinity header and Portal conversation root (#109964)
Trimmed salvage of #110045 (deltas 1 + 2 only), stacked on the #110009 scope inheritance:

- `declared_conversation_scope` treats an inherited value as a DECLARED scope only when it
  carries the `gwk_` prefix. A rotated CLI parent publishes no affinity scope (None → sticky
  key falls back to the conversation root); the fork now publishes exactly the same instead
  of an explicit physical lineage root. `resolve_prompt_cache_scope` honors any inherited
  value directly, so the body `prompt_cache_key` still matches.
- `build_cache_parity_fork` snapshots `parent._conversation_root_id()` as
  `_cached_conversation_root`; with `_session_db=None` the fork's own walk fell back to the
  parent's PHYSICAL id, so after a compression rotation the review's Portal
  `conversation=` tag fragmented usage attribution across one logical conversation.

Dropped from the original: copying `_gateway_session_key` onto the persistence-detached
fork (no cache-identity consumer reads it there; the compression-boundary hooks were
deliberately severed by `_detach_fork_compression`), and the defensive
hasattr/callable/try wrapper around `_conversation_root_id()`.
2026-09-14 06:55:54 -07:00
salch-cred a4b620f17c fix(agent): same-model review fork inherits the parent's resolved cache scope (#109964)
build_cache_parity_fork gives the same-model fork the parent's session_id,
cached system prompt, tools[] and session_start — but with
_persist_disabled=True and _session_db=None, BOTH cache-identity resolvers
diverged from the parent on their own: declared_conversation_scope failed
closed on _persist_disabled, and the lineage walk skipped on the missing
DB. The fork's affinity header (set_affinity_scope) and body
prompt_cache_key (cache_scope_id on the OpenAI-wire transports) therefore
keyed a different bucket than the gateway parent, costing one cold
~full-context request per review. Not gateway-only: any parent whose
lineage root != current physical id diverges too (teknium1's triage table).

Fix, per the triage's suggested direction: on the not-routed branch only,
the fork stamps _inherited_cache_scope = resolve_prompt_cache_scope_safe
(parent) — the parent's ALREADY-RESOLVED scope, no DB access from the fork,
persistence fully detached. Both declared_conversation_scope and
resolve_prompt_cache_scope return the inherited scope first when set, so
the header path and the body path are fixed together (fixing only one
leaves the other divergent — Vivamisu's header/body split observation).
Routed (different-model) forks, /branch children, delegate/tool children
and fresh sessions set nothing; the fail-closed default stands untouched.
/btw shares build_cache_parity_fork and gets the repair for free.
2026-09-14 06:55:54 -07:00
teknium1 274fd56dca fix(state): WAL lock guard follows the handle's lifecycle
Three gaps in the #110544 guard, all reported in its review and reproduced:

- A writer reopened by _reopen_after_close_locked (teardown/worker race,
  #94736) came back with no guard: the next stray close + foreign close
  deleted its WAL again.
- _try_wal_checkpoint refreshed the guard outside self._lock; landing after
  close() it pinned an OFD lock with no connection behind it, so a foreign
  `PRAGMA journal_mode=DELETE` saw `database is locked` forever.
- Refcounts keyed on (fd, inode) treated a recycled fd number as a surviving
  lock: A+B live, close A, C reuses A's fd, close B left C recorded as guarded
  while a foreign EXCLUSIVE succeeded.

The guard now counts handles per inode, re-locks every matching descriptor on
each hold (OFD re-lock is idempotent), and unlocks on the last handle only;
the reopen path holds it; the checkpoint refresh runs under self._lock and
skips a closed handle. The macOS holder scan folds case so a case-only alias
of the sidecar path on APFS still matches.
2026-09-14 06:54:07 -07:00
teknium 743140cd82 feat(gemini): send full JSON Schema tool parameters via parametersJsonSchema
Clean-room port of the approach in zed-industries/zed#63342. The native
Gemini adapter previously down-translated every tool schema into the
restricted FunctionDeclaration.parameters subset, which was lossy: anyOf
unions without an outer type, bare arrays, $ref/$defs indirection and
additionalProperties had to be stripped or repaired, and one
unrepresentable construct could 400 the entire request (live repro:
INVALID_ARGUMENT ...properties[bare_array].items: missing field).

Google now accepts plain JSON Schema in parametersJsonSchema on all
current models. The adapter sends full schemas through that field; the
old subset translator is replaced by a light normalizer that deep-copies,
strips root $schema, inlines same-document $refs (MCP pydantic / zod
emit them; unresolvable or circular refs pass through untouched with the
reason logged), and guarantees an object root.

Live-verified against the real API: the union+bare-array+$ref schema
that 400s through the legacy parameters field is accepted with 200 via
parametersJsonSchema on gemini-3.7-flash and gemini-2.5-flash, and
gemini-2.5-flash returns a correct functionCall against it.
2026-09-14 06:44:54 -07:00
teknium1 49c6d4a9e0 test(contracts): tests mirror tui_gateway/; the runtime-artifact spoof test asserts the new 4000
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
2026-09-14 06:12:19 -07:00
teknium1 f6306d1920 feat(contracts): TypeScript consumes the generated contract; hand-typed wire shapes deleted
apps/shared/src/gateway-events.ts is now a thin layer over
gateway-contract.generated.ts (client-local synthetic events + the
GatewayEvent envelope); gateway-events.json, its two rendezvous tests and
the duplicated BillingBlock / SessionInfo / ProjectInfo hand copies are
gone. Desktop, TUI, web and shared typecheck against the generated
RpcMethods / ServerRequestMap / BackendGatewayEventMap.

What tsc found once the types were honest: three phantom fields the
backend never sent (tool.start.todos, error.reason,
voice.transcript.voice_stopped) - the TUI todo tests were driving the
list through the phantom and are retargeted to tool.complete, where the
wire actually carries it; nullable fields (`None` on the wire) were typed
as plain optionals in eight places and now coerce at the boundary;
SessionResumeResult had a stale generic.

Contract fixes from the consumer pass: TranscriptMessage is the gateway
projection (text/row_id/context/args), not the stored row; SkinPayload
matches HermesSkin (empty-string defaults, never null); SessionLiveInfo
model/tools/skills are required (always emitted); BillingBlock.billing_url
is required-nullable (dataclass asdict).

tui_gateway/AGENTS.md documents the declare -> regenerate -> tsc loop.
2026-09-14 06:12:19 -07:00
teknium1 24ffc8d23c fix(contracts): params validation rejects only unknown keys; accepted params + results are checked after the handler
Handlers own their documented domain codes (4006 missing session_id, 4015 bad
url, 4009 orphan claim); the contract's job on the way in is the one check no
handler performs — an unknown key (4000 with the key path). Missing/mistyped
fields are re-checked AFTER a successful handler answer under the strict
test policy, so a contract narrower than the wire still fails the suite.
Two models widened from the suite: SeedMessage (clients forward stored rows
verbatim), tool.complete.args (mirrored child rows omit it). Tests that
drove session.activate with prompt params (and vice versa) or stubbed
_live_session_payload with a bare {session_id} now send the real shapes.
2026-09-14 06:12:19 -07:00
teknium1 0250c8bcae feat(contracts): declare every gateway method, server request and event; commit the generated TS + OpenRPC (#110522, part 2)
215 methods, 13 server→client requests and 67 notifications now have Pydantic
contracts under tui_gateway/contracts/<topic>.py, rendered to
apps/shared/src/gateway-contract.generated.ts (616 types) and
gateway-contract.openrpc.json. tests/contracts/test_generated.py pins both
files to an in-memory regeneration and asserts catalog completeness from the
CODE side (every registered handler / emitted event / sent request has a
contract, nothing orphaned). scripts/ci/classify_changes.py runs the Python
lane when either generated file changes.

Phantom fields the hand-typed TS carried and no emitter ever set:
tool.start.todos, error.reason, voice.transcript.voice_stopped.
2026-09-14 06:12:19 -07:00
teknium1 eabf2e46f9 test(tui_gateway): hosted-room suites stop waiting on 0.5-2s thread joins under a loaded runner
Under the 40-worker file runner, service.stop(timeout=1.0) / runtime.stop(timeout=0.5)
and the 2s _wait_for lost the race to scheduler latency (a different test each run,
green on retry). Bounds move to the 5s the other 25 sites already use; the bounded-stop
invariant keeps a 2s ceiling, still far inside the join timeout.
2026-09-14 06:05:53 -07:00
teknium1 9f7f2f28c0 feat(gateway): server→client JSON-RPC requests replace the *.request/*.respond event pairs (#110521)
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.

`tui_gateway/server_requests.py` owns the one mechanism:

  send()          block the agent thread until the response frame
                  (`srq-<n>` ids; ints belong to the client)
  send_async()    fire-and-callback variant (bot relay)
  cancel*()       withdraw with ONE `request.cancel {id, method, reason}`
                  event (timeout / interrupt / process exit /
                  answered elsewhere) instead of per-kind *.expire
  open_requests() the still-open frames, replayed by session.resume,
                  session.activate and session.events.since so a
                  reconnecting client re-renders every kind, not two
  clarify.lock    stays a real client→server RPC (locks one batch
                  answer early); locked answers merge into the final
                  set even when the closing response carries only the
                  tail the user answered last

A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.

Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.

Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
2026-09-14 06:02:05 -07:00
teknium1 ebe8cda8ea feat(tui_gateway): real JSON-RPC server→client requests replace the *.request / *.respond notification pair
The backend never sent a JSON-RPC request; when it needed an answer from the
renderer it hand-correlated a `*.request` notification with a later `*.respond`
method through four module-level dicts, a timeout thread and 13 derived
`*.expire` names, plus a separate reconnect snapshot per prompt kind. That is a
second request/response layer built on a protocol that already has one.

`tui_gateway/server_requests.py` sends `{id: "srq-…", method, params}` and
blocks on the response frame with that id (string ids never collide with the
clients' integer ids). One `request.cancel {id, method, reason}` notification
withdraws a request on timeout / interrupt / session close. `open_requests` on
`session.resume` / `session.activate` / `session.events.since` re-delivers
unanswered requests after a reconnect; the shared TypeScript channel does that
itself before the caller sees the result. Batch clarify keeps its per-question
locks as a normal `clarify.lock` RPC (the last lock resolves the request).
Approvals stay queue-backed (`tools.approval` owns the timeout, `/approve all`,
coalescing): the request resolves the queue entry and the entry's own
resolution withdraws the request through `register_gateway_settle`.

Deleted: `_block`, `_respond`, `_pending`, `_answers`,
`_pending_prompt_payloads`, `_batch_clarify`, `_EXPIRING_REQUESTS`, the
`*.respond` methods, every `*.request` / `*.expire` event, `pending_clarify`.
Compute-host (turn isolation) mirrors the child's open request and relays the
response frame / lock to it. Desktop, TUI and shared clients register
`onRequest` handlers where they used to switch on `*.request` events; answers
are response frames over the socket the request arrived on, so #91684's
owner-routing class cannot recur for prompts.
2026-09-14 06:02:05 -07:00
teknium1 500e133bee fix(config): Desktop fallback editor keeps per-entry routing; MoA save writes only the moa section
Two remaining halves of #89184 (Desktop Settings saves rewriting unrelated
config):

- The `fallback_providers` structured editor normalized every entry down to
  `{provider, model}`, so any edit (remove a row, pick a model) re-emitted a
  hand-written local-gateway chain without its `base_url` / `api_key` /
  `key_env` / `api_mode` — the next autosave persisted bare pairs and the
  fallbacks silently routed to the public provider. Entries now carry every
  key through; the editor only owns the two selects.

- `PUT /api/model/moa` did `cfg = load_config(); cfg["moa"].update(...);
  save_config(cfg)`: the whole default-expanded snapshot went back to disk,
  so a Desktop MoA autosave re-persisted every other section too (the
  2026-09-10 repro: `fallback_providers: []` written alongside the MoA block
  the user had just edited). It now saves `{"moa": ...}` with
  `merge_existing=True`, the same section-scoped write every other sparse
  writer uses since #110535. Hand-edited moa keys (#58819) still survive.

The `model.default not persisted / base_url cleared` symptom from the 0.20.4
report no longer reproduces on main through the real REST path (Config page
diffs against a baseline since 5361867c6d32; `_denormalize_config_from_web`
keeps the on-disk `model:` block).
2026-09-14 05:56:29 -07:00
yoniebans 9d75f20630 fix(insights): short-circuit before opening when state.db is absent
SessionDB(read_only=True) cannot create a missing store, so a fresh install
running hermes insights / /insights errored instead of reporting no data
(reported by @ehz0ah on #110718; guard shape from @kshitijk4poor's #110026).

Co-authored-by: kshitijk4poor <kshitijk4poor@users.noreply.github.com>
2026-09-14 05:28:46 -07:00
yoniebans a30ccdbeab fix(tui_gateway): keep opportunistic writes off read-only foreign handles
Two write paths were still reachable through the now read-only foreign-profile
handle (reported by @ehz0ah on #110718): the repo-root backfill in
_discover_repos_payload raised and was swallowed per RPC, silently dropping the
persistence; the Bot Chat unarchive in session.list's exact-title lookup failed
the RPC with error 5006. Backfill now skips on read-only handles (that
profile's own gateway backfills on its refreshes); the unarchive escalates to a
short-lived registry writer for the rare recoverable-archive case.

Reported-by: ehz0ah
2026-09-14 05:28:46 -07:00
yoniebans 31e0300b52 fix(tui_gateway): foreign-profile RPCs open state.db read-only
_profile_db acquired a registry WRITER on another profile's state.db for
every RPC about it and closed it in the handler's finally — schema init
plus the write lock on a store owned by that profile's own gateway, once
per sidebar refresh under desktop app-global remote mode. Reads now go
through the dashboard sidebar's read-only open helper; the four handlers
that mutate (session.move_cwd, session.delete, session.set_hidden,
session.foreign.import) opt in with writer=True.

Field trace and read-only approach by @samalone on #109737.

Co-authored-by: samalone <samalone@users.noreply.github.com>
2026-09-14 05:28:46 -07:00
B0on d680d66998 fix(insights): open the session store read-only
hermes insights and /insights only read; taking the default writer runs
schema init and contends the write lock on a live store.

Salvaged from #109737.
2026-09-14 05:28:46 -07:00
teknium1 3e43cee505 fix(state): refcount guard locks shared by several handles in one process
Two SessionDB handles on one state.db in one process see the same
descriptors; the first release must not drop the range the second still
needs. Test covers the same-process sibling and the foreign-process close.
2026-09-14 05:28:22 -07:00
赵桂雄 68b10bbbf9 fix(state): enumerate deleted-WAL sidecar holders on macOS via libproc
iter_deleted_sqlite_sidecar_holders() returned [] on every non-Linux
platform, so refuse_deleted_wal_generation() -- the pre-connect refusal
that stops a second opener from minting a replacement WAL under a live
writer -- was a permanent no-op on macOS. The reporter of #109641 hit
exactly that: after an update/restart took the sidecars away, a fresh
opener minted a new generation at the path, the still-live writer's next
write raised DeletedWalGenerationError, and each event copied the whole
database (16 halts / 13 minutes / 4.2 GB of captures).

macOS has no /proc and never reports a " (deleted)" suffix, which is why
the scan was restricted to Linux, but libproc does describe other
processes' descriptors: proc_pidinfo(PROC_PIDLISTFDS) lists a process's
fds and proc_pidfdinfo(PROC_PIDFDVNODEPATHINFO) returns each vnode fd's
(st_dev, st_ino) plus the vnode's last pathname -- for same-user
processes, without elevation. Both survive unlink, which is also why
psutil.Process.open_files() cannot stand in for it (it hides unlinked
descriptors, so the retired generation is structurally invisible).

The judgement itself is unchanged and now shared: a descriptor counts
only when it names a watched sidecar path while its identity no longer
matches what that path holds, i.e. _fd_is_truly_unlinked()'s identity
test (#108082), never a path suffix or a link count. Only the source of
that identity differs per platform -- readlink on /proc for Linux,
libproc for macOS -- and the darwin side resolves symlinks before
comparing paths because libproc reports the kernel's path
(/private/var/... where the caller opened /var/...).

Scope is this one function: the enumeration legs, the gate (Windows
still returns [] -- it cannot unlink a held sidecar) and the stale
docstring reason. Enumeration failures keep the existing fail-open
behaviour (logged at debug, no holders), and the new tests are marked
macos_only so the existing Linux-only ones stay untouched.

Cost, measured on macOS 26.4 (darwin 25.4.0) with 721 processes /
4383 vnode descriptors: ~20 ms per full enumeration, versus the Linux
leg's ~11 ms / ~4.4k syscalls measured in #108910 -- the same order,
paid once per open, on the platform where the guard previously did
nothing at all.

(cherry picked from commit f1501dfe7141e3c2521c5192857c37d7b60a922b)
2026-09-14 05:28:22 -07:00
teknium1 75e155ab09 fix(state): a live writer's WAL generation survives lock cancellation and sibling closes
SQLite protects a WAL generation with per-PROCESS POSIX locks (SHARED on
state.db, DMS byte on -shm). Any in-process open()/close() of either file
cancels both (sqlite.org/howtocorrupt.html §2.2); the next last-connection
close in ANY process then checkpoints and unlinks -wal/-shm, and the holder
sticky-halts with DeletedWalGenerationError. #109841 removed one such
close (mode tightening) but the class is open-ended: raw header probes,
plugins, tool reads of ~/.hermes, any library that touches the files.

hermes_state_lockguard re-holds the same two ranges as OFD locks
(F_OFD_SETLK) on private descriptors for as long as a writer handle is
open. OFD locks belong to the open file description, so a stray close()
cannot cancel them, and they conflict with the EXCLUSIVE a sibling needs
for the close-time reset exactly like SQLite's own. Released before the
handle's own close so a true last close still ends the generation; the
descriptors are closed only once no connection to the path remains, so a
holder scan from another process never counts them. Works on Python 3.11
(where sqlite3 cannot arm SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE) and on macOS
(F_OFD_SETLK=90 per XNU bsd/sys/fcntl.h); no-op on Windows.

Live repro (Linux, Python 3.11.15, SQLite 3.53.1): holder = SessionDB
writer; in-process os.open/os.close of state.db and -shm; then a foreign
sqlite3.connect()+close(). Before: -wal unlinked, holder write raises
DeletedWalGenerationError. After: -wal keeps its inode, holder writes.
2026-09-14 05:28:22 -07:00
liuhao1024 47714f9402 fix(agent): routed background reviews honor auxiliary.background_review.reasoning_effort
The review fork is a full AIAgent, not an auxiliary_client call, and its
routed branch deliberately skips the parent's reasoning_config (the parent's
effort vocabulary may be invalid for the routed provider). It also never read
the per-task key, so an explicit `auxiliary.background_review.reasoning_effort`
was silently ignored and the routed fork ran at the provider default (#94825).

Routed forks now parse the task key through the shared parse_reasoning_effort
(same levels and `none` alias as every other aux task); unset keeps the
provider default, an unknown level warns and falls through. The same-model
path is untouched: it still inherits the parent's reasoning_config verbatim
for prompt-cache parity.

Salvaged from #94832 (liuhao1024), re-applied on the decomposed
_fork_init_kwargs seam.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-14 05:25:01 -07:00
Finn763 6da540d3ec fix(background_review): surface the ignored reasoning_effort on same-model forks (#104116)
auxiliary.background_review.reasoning_effort was a silent no-op on the
same-model path: the fork inherits the parent's reasoning_config verbatim to
keep prompt-cache parity (#30532), and nothing told the user. Emit a one-time
user-visible warning (parent-scoped, so a nudge-per-turn session warns once,
not per fork) when the key is actually set, document the no-op in the config
reference, and leave the fork-birth request bytes unchanged.
2026-09-14 05:25:01 -07:00
teknium1 01338e88ec test(video-gen): keep the i2v-only guard real once every family is dual-modality
The rebased synthetic guard built a bare dict that main's _build_payload
no longer accepts (it indexes aspect_ratios/resolutions directly) and never
exercised generate(); it now builds the family via _family() and asserts
generate() returns modality_unsupported without submitting. The surface
matrix's i2v-only parametrization collected zero cases after Gemini Omni
Flash 1.1 gained t2v, so it is removed along with the dead branch in the
t2v matrix.
2026-09-13 21:40:26 -07:00
Teknium aecee6f66a feat(video): Gemini Omni Flash 1.1 — text-to-video now works (was image-only)
FAL shipped google/gemini-omni-flash/v1.1/* (Aug 2026): the family gains a
text-to-video endpoint, 360p/720p/1080p/4k resolution enum, and keeps
3-10s integer durations with always-on native audio.

- plugins/video_gen/fal: bump gemini-omni-flash to the v1.1 endpoints,
  declare the resolution enum, refresh display/strengths copy
- tests: family test asserts the versioned dual-modality endpoints; the
  i2v-only clean-error guard survives via a synthetic family; catalog
  invariant now requires both endpoints on every family

Schema verified against FAL OpenAPI (t2v: prompt required, 16:9/9:16,
360p-4k, duration 3-10 int; i2v adds image_url + optional end_image_url).
Pricing: $0.03/s 360p, $0.10/s 720p, $0.15/s 1080p, $0.30/s 4K.
2026-09-13 21:40:26 -07:00
teknium1 60559d4e0e fix(delegation): late-attached children take the parent's soft/hard stop kind; dedupe the fallback replay
_attach_child now mirrors a pending parent stop with the same split
interrupt() uses for its own fan-out (hard -> hard_interrupt, soft ->
interrupt), so a redirect is not turned into a cancel on a child that was
attached late. _restore_parent_cancellation collapses to re-attaching the
rejected unit's children: the replay is the attach step's job now.

Test fixture: _Batch gained origin_session_history_delivery on main after
the salvaged PR was written.

Co-authored-by: illidan <noequal666@gmail.com>
2026-09-13 21:31:57 -07:00
illidan 86258b9912 test(delegation): enable independent units in partial admission coverage 2026-09-13 21:31:57 -07:00
illidan 946111cb11 fix(delegation): restore parent cancellation for rejected async units 2026-09-13 21:31:57 -07:00
teknium1 2178b3ebe8 test: trim guard-cache race tests to the invariants
Fold the whitespace-only case into the empty-file miss test and drop
test_overwrite_is_all_or_nothing: it wrote sequentially, so it exercised
nothing about os.replace atomicity that the publish test does not.
2026-09-13 21:31:18 -07:00