Commit Graph

56 Commits

Author SHA1 Message Date
Teknium d44228c609 refactor(tui_gateway): _TurnScopes/_TurnRun as dataclasses, tighter module docstring 2026-09-02 23:44:12 -07:00
Teknium cb0352d6db refactor(tui_gateway): AST-neutral layout compaction + docstring tightening for prompt modules 2026-09-02 23:21:59 -07:00
Teknium f80e2c77c1 refactor(tui_gateway): split prompt turn/submit god functions into phase helpers, unify side-agent + followup dispatch 2026-09-02 23:14:19 -07:00
Teknium 99433742dc refactor(tui): extract the prompt turn into prompt_turn.py; compact methods_*/session_*/agent_callbacks; split _start_agent_build into scope/wiring helpers
- server.py 7663 -> 5319: _run_prompt_submit and its goal/loop/voice/scope
  phases move to tui_gateway/prompt_turn.py (bound via method_ctx.bind_module);
  _start_agent_build split into _bind/_release_build_profile_scopes,
  _deferred_build_agent_kwargs, _wire_session_agent, _start_session_services;
  _load_enabled_toolsets split (_enabled_mcp_server_names, _resolve_explicit_toolsets).
- methods_slash: _LIVE_SLASH_OUTPUT dispatch table; methods_tools: _SLASH_BUILTINS,
  _guarded; tool_progress: _PROGRESS_HANDLERS; methods_config_set:
  _REASONING_DISPLAY_WORDS; methods_voice: _VOICE_TOGGLE_ACTIONS.
- Unified: _denied_source (methods_session) replaces _WORKER_SOURCES in
  methods_profiles; _compress_live_with_feedback / _compute_host_slash shared by
  the slash mirror and /compress; _end_voice_chat shared by stop-phrase paths;
  _watcher_mtime_ns; _reaper_session_is_detached_idle; _notif_* helpers.
- Dead: _profile_dir_or_err, _resume_info, _slash_builtin_table (refs.py: 0 hits).
- acp_adapter/server.py: docstring compaction only (AST-identical).
- Every file semantically reviewed hunk-by-hunk for wire/log/lock/order parity.
2026-09-02 14:09:13 -07:00
Teknium 0a2fd94515 refactor(tui): table-drive the 8 identical *.respond handlers; drop dead mcp_rpc_helpers.resolve_profile + server shims; bind_module publishes cross-module aliases 2026-09-02 14:09:12 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Futahua a5f0fbb262 fix(sessions): per-session exclusivity is correctness, not a capacity policy
Cherry-picked from PR #94595 (author: Futahua) onto current main, with the
maintainer-review revision points folded in during the rebase:

- the lease engages UNCONDITIONALLY: try_acquire_active_session no longer
  returns a disabled no-op lease when max_concurrent_sessions is unset;
  the concurrency cap stays an orthogonal, optional policy checked second
- ownership uncertainty fails CLOSED (SESSION_COORDINATION_UNAVAILABLE)
  instead of degrading to an untracked go-ahead: a corrupt/unreadable
  registry must not be collapsed into 'no owner exists' (review blocker 2)
- the ownership admission sits at the _run_prompt_submit chokepoint that
  EVERY fresh turn source crosses, and crash auto-continue acquires (or
  bails) BEFORE emitting message.start — closing the #94778 bypass where
  backend B's auto-continue ran a duplicate turn while backend A was live
  (review blocker 1)
- the TUI gateway claim helper fails closed on claim exceptions for every
  surface, not just desktop
- CLI and messaging-gateway call sites pass live_session_id metadata so
  the (pid, live id) re-entrancy identity protects them from self-fencing
  on a leaked lease

Co-authored-by: teknium1 <teknium1@users.noreply.github.com>
2026-08-31 12:36:33 -07:00
konsisumer 3ae74119c9 fix(gateway): relay compute-host clarify state 2026-08-31 12:18:00 -07:00
Alvin T. Veroy e17fd0a708 fix(state): decode errors now reach the heal path and fail loud in TUI (residual #98924 surfaces)
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:

- web_server._open_session_db_at_path: the one-writable-open heal only
  caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
  to decode SQLite's own error message over corrupt file bytes) bypassed
  it, so the heal documented for malformed schema never fired (#98924
  Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
  vtables whose probe raised UnicodeDecodeError, the same too-narrow
  catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
  could not open, so prompt.submit streamed the turn while persisting
  nothing (#98924 Failure 2). It now returns False and prompt.submit
  fails the RPC with code 5072 so desktop maps it to a toast, mirroring
  the disk-full/5070 convention. session.create stays silent per its
  pinned degraded-mode contract.
2026-08-31 09:56:43 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
David Dudok de Wit 93c7089f70 feat(bot-mode): run same-gateway Group Chats without Desktop 2026-08-30 22:19:06 -07:00
Teknium 578f85cfb0 feat: /btw rides the background-review cache-parity fork for full-context answers
The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.

- agent/background_review.py: extract the review-fork construction into
  build_cache_parity_fork() — same runtime/credentials as the parent,
  byte-identical system prompt / tools[] / reasoning config on the
  same-model path, shared session_id for prefix warmth, full persistence
  detachment (no state.db writes, no rotation, no external memory,
  in-place-only compaction). The review thread now calls the helper;
  behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
  AIAgent exists — replays the untruncated snapshot as warm cache reads,
  denies every tool at dispatch via an empty thread whitelist (tools[]
  stays byte-identical for cache parity), attributes usage to the parent,
  and trims a mid-turn snapshot tail so role alternation holds. The
  one-shot digest remains as fallback (no live agent = cold cache anyway,
  and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
  the chat's cached agent (parity with how turns reuse it).

Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
2026-08-29 08:23:47 -07:00
Teknium 74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Teknium 9b3f60c029 fix(gateway): resolve approval.respond by durable identity before failing 4001
Server half of #91684: the desktop can answer an approval prompt with a
stale live sid — its runtime record was re-minted after a reconnect while
the prompt stayed on screen. approval.respond now falls back, on 4001
only, to resolving the target session (1) by the unique approval
request_id across every live session's pending gateway approvals, then
(2) by treating session_id as a STORED session id mapped to its live
runtime record. Only when neither resolves does it return 4001.

Tests: request_id fallback, stored-id fallback, and 4001 when nothing
resolves.
2026-08-23 17:43:39 -07:00
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
Teknium 98f6fc549a feat(desktop): failed turns name the failing layer with recovery actions
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.

Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
2026-08-21 15:24:03 -07:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
brooklyn! 02e270a47e fix: editing a message in an old session fails (profile DB + window-relative ordinal) (#91302)
* fix(gateway): persist prompt.submit truncation to the session's own profile DB

`_get_db()` returns the LAUNCH profile's SessionDB handle. App-global
remote mode gives a session its own profile (`session["profile_home"]`)
whose transcript lives in that profile's `state.db`, so a write keyed on
`session_key` that goes through `_get_db()` addresses the wrong database.

In the `prompt.submit` truncate branch that has two consequences. The
edit/resend never sticks — `session.resume` reopens the profile db and
resurrects the undone turns — and when the launch profile happens to hold
a row under the same session id, the truncated transcript is inserted
into a profile the session does not belong to.

It also silently voids the branch's own fail-closed contract. The handler
persists before it rewrites `session["history"]` precisely so that a
failed write refuses the turn and leaves memory and DB aligned; that only
holds if the handle it checks is the one that owns the row.

`_session_db(session)` is the profile-aware resolver that already exists
for this: the profile's `state.db` when `profile_home` is set, otherwise
the shared launch handle. Non-profile sessions are unaffected —
`_session_db` borrows the same shared handle and leaves it open.

`active_only=True` and `archive_dropped=True` are carried through
unchanged; only the handle the call is made against changes.

* fix(gateway): resolve the /undo command against the session's own profile DB

`command.dispatch`'s `/undo` branch opened the launch profile's handle via
`_get_db()`, but every read and write under it is scoped by session id:
`list_recent_user_messages`, `rewind_to_message` and the
`get_messages_as_conversation` reload all key on `session_key`.

For a session with its own profile (`session["profile_home"]`) the rows
live in that profile's `state.db`, so against the launch handle
`list_recent_user_messages` returns nothing and the command fails closed
with `4018 "no user messages to undo"` — for the entire session, on every
invocation, even though the transcript is right there in the profile db.

Route the whole branch through `_session_db(session)`, which yields the db
that owns the session's row and closes a profile handle on exit. Sessions
without a profile keep borrowing the shared launch handle exactly as
before, so this is behaviourally identical for them.

* fix(gateway): read /history and /context from the session's own profile DB

`_format_live_history_output` and `_format_live_context_output` rebuild the
transcript from the database rather than from `session["history"]`, because
the in-memory list is empty for a session this process did not run itself.
Both reads are scoped by session id but were issued against `_get_db()`,
the launch profile's handle.

A session with its own profile (`session["profile_home"]`) keeps its rows
in that profile's `state.db`, so both reads come back empty and the
commands under-report: `/history` renders "No conversation history yet."
and `/context` falls back to the empty in-memory list and reports a
conversation of zero messages. Both swallow their exceptions, so there is
no error either — just a wrong answer about the user's own transcript.

Resolve both through `_session_db(session)`, the profile-aware resolver
used by the rest of the session-scoped paths.

* test(gateway): cover session-scoped transcript ops against a profile DB

Regression coverage for the three session-scoped sites that resolved
against the launch profile's handle instead of the db owning the
session's row. Each test drives the real JSON-RPC entry point with a
session carrying `profile_home`, seeds the transcript into the profile's
own `state.db`, and asserts against both databases.

Per site, with the production change reverted to its pre-fix form:

- `prompt.submit` truncation — `test_truncation_persists_to_the_profile_db`
  and `test_truncation_does_not_copy_rows_into_the_launch_profile` fail.
  The second seeds a row under the same session id in the launch db so the
  foreign write succeeds instead of failing a key check, which is the case
  that copies a transcript into a profile it does not belong to.
- `/undo` — `test_undo_rewinds_the_profile_transcript` fails with
  `4018 "no user messages to undo"`.
- `/history` and `/context` — `test_history_reads_the_profile_transcript`
  and `test_context_reads_the_profile_transcript` fail, reporting an empty
  conversation.

`test_undo_still_uses_the_shared_handle_without_a_profile` and
`test_truncation_without_a_profile_uses_the_shared_handle` pin the
unchanged path: with no `profile_home` the resolver must borrow the shared
launch handle and leave it open. Both stay green in every direction, so a
future change cannot satisfy the profile cases by abandoning the shared
one.

* fix(desktop): aim truncations by durable id alone on tail-only transcripts

The cold-open transcript is a newest-first prefetch page
(LATEST_SESSION_MESSAGES_LIMIT = 120) with the resume RPC sent
omit_messages — older rows only arrive via "Show earlier" backfill.
planEdit/planReload/planRestore still counted truncate ordinals over
that windowed list, so every edit/reload/restore in a session longer
than the prefetch page sent a window-relative ordinal alongside the
durable row/message id. The gateway's #82959 cross-check resolved the
durable id to its full-history ordinal, read the offset as drift, and
refused with 4030 — making the Edit affordance permanently dead in
long sessions.

When the transcript may be tail-only (the transcript-tail
bookkeeping's possiblyTruncated), drop the client ordinal and address
the truncation by durable id alone — the same rule runRewindSubmit
already applies to content-resolved row ids (#87059). The ordinal
tripwire stays on whenever the transcript is complete.

Closes #88082

* fix(desktop): drop client rewind ordinal whenever a durable id is present

#88092 gated the drop on tail-only prefetch. After in-place compact the
live scrollback is treated as complete, so Restore still sent a
display-lineage ordinal next to a resolved row id and the gateway
refused with 4030 (#89244). prefix_user_count is structurally 0 on
in-place because get_ancestor_display_prefix is cross-session.

Same choke point: if a durable truncate_before_row_id or a real
truncate_before_message_id is present, omit the client ordinal.
confirm_empty_truncate is still carried from a caller ordinal of 0.
Unknown ids still fail closed at 4018.

Closes #89244

---------

Co-authored-by: briandevans <252620095+briandevans@users.noreply.github.com>
Co-authored-by: zengzheqing <yuntianqing@yahoo.com>
2026-08-21 06:02:19 +00:00
Brooklyn Nicholson c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Brooklyn Nicholson 23d88c2b0e feat(tools): tour — let the agent walk a user through the UI
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.

Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
2026-08-19 00:52:54 -05:00
Brooklyn Nicholson 3ead0f8dc1 feat(desktop): widget clicks reach the agent as hidden user turns — the widget updating IS the response
An inline ::preview widget could render and be clicked, but the click went
nowhere: the sandbox has no channel to the agent, so an interactive chart
was a dead end. Now the frame injects a second script beside the measurer
that gives the page one voice:

  window.hermes.send('get-price eth')
  <button data-hermes-send="get-price eth">ETH</button>  (zero-script form)

The prompt rides postMessage up tagged with the mount token, then goes
through the composer's own send path (requestComposerSubmit -> prompt.submit)
flagged display_kind=hidden — the same row-typing auto-continue and internal
notifications already use. The agent wakes and takes a real turn; the
durable row persists (context, resume, DB audit); but NO bubble renders,
live or on reload. The user clicks ETH and the chart just changes — the
off-screen loop is click -> hidden turn -> agent rewrites the widget file ->
frame hot-swaps.

Trust boundary matches size reports and is tighter where it matters: mount
token required (frames can't forge each other's intents), string-only,
trimmed, capped at 500 chars, throttled to one intent per second per frame.
The gateway whitelists display_kind to "hidden" — the RPC can't mint
arbitrary row types — and the flag threads through both turn paths (inline
and compute-host isolation) so isolated sessions don't resurrect bubbles on
resume.

The desktop platform hint teaches the model to wire interactive widgets
with data-hermes-send and to answer clicks by updating the widget's file
rather than with prose; the SDK doc documents the contract.
2026-08-17 16:08:18 -05:00
poisdahl d67583ac66 fix(tui): classify repaired rows in rewind rebinds 2026-08-16 12:27:15 +02:00
poisdahl 7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
rainbowgits e14b095f20 fix(tui): map lineage edit ordinals past compression prefix
Desktop/TUI count full displayed lineage after compression, but
prompt.submit validated truncate ordinals against tip-only history.
Translate via display_history_prefix and recover stale 4018s on Desktop.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 02:24:40 -07:00
poisdahl fcea2175e2 Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tests/test_tui_gateway_server.py
#	tui_gateway/methods_prompt.py
#	tui_gateway/methods_tools.py
2026-08-15 23:02:54 +02:00
fangliquanflq 3863de3155 revert(gateway): keep profile truncation routing out of scope 2026-08-15 13:36:25 -07:00
fangliquanflq 0640fe7119 fix(gateway): route truncation writes to profile database 2026-08-15 13:36:25 -07:00
fangliquanflq 79b7d969d3 fix(gateway): reject unstamped durable ordinal rewinds 2026-08-15 13:36:25 -07:00
poisdahl 3f075d41dd Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tui_gateway/methods_prompt.py
#	tui_gateway/server.py
2026-08-15 12:05:14 +02:00
jdgg777 6da30f72a2 fix(tui_gateway): fall back to session id when session_key is NULL in truncation persist
CLI-origin sessions have no session_key; the Desktop history-truncation
path called replace_messages(session["session_key"], ...) with None,
whose reinsert violated the messages.session_id FK -> "FOREIGN KEY
constraint failed" -> "Restore failed" on resume. Key the persist off
the durable session id instead.

Extracted from PR #81904 (the scope=compacted API half was superseded by
include_compacted, #86595). The PR's companion change defaulting
session_key to the session id at insert time is deliberately NOT taken:
main treats a non-NULL session_key as "this is a gateway session"
(list_gateway_sessions, orphan gateway-session repair), so the default
would misclassify every CLI session.

(extracted from PR #81904, commit cef9b9b27d)
2026-08-15 00:35:16 -07:00
fangliquanflq 4642f9630d fix(gateway): reject unsafe ordinal-only truncation 2026-08-14 21:33:31 -07:00
VooDoo Pixels f703e70618 fix: make desktop approval routing reliable
Correlate approval requests, reject stale responses, replay pending approvals after reconnect or session resume, and preserve fail-closed timeout behavior.
2026-08-14 20:24:32 -07:00
David Gutowsky 3abe6cc503 fix(gateway): bind profile HERMES_HOME override in ephemeral agent threads (#50233)
Normal prompt turns bind session['profile_home'] via set_hermes_home_override
before run_conversation, but the two ephemeral RPC paths (prompt.background,
preview.restart) spawn a fresh AIAgent on a new thread where the HERMES_HOME
ContextVar does not propagate — so a background/preview turn under a
non-default profile ran against the wrong home. Re-bind for the duration of
the ephemeral turn and restore in finally, mirroring the normal prompt turn.

Surgically reapplied from PR #50777 (handlers moved to methods_prompt.py
since the PR was authored; handler bodies rebind onto server.py globals, so
the original pattern transplants verbatim). Includes the contributor's
regression tests unchanged.
2026-08-14 15:04:46 -07:00
Brooklyn Nicholson 8b06f7df8e fix: attach RPCs no longer wait on the deferred agent build
image.attach, image.attach_bytes, file.attach, pdf.attach, clipboard.paste
and image.detach resolved their session through _sess(), which blocks on
_wait_agent(). None of them needs the agent — they read cwd/profile_home and
mutate attached_images, all populated when the session record is created.

None of these methods is in _LONG_HANDLERS either, so the wait ran inline on
the socket reader thread. Attach runs before prompt.submit, so pasting an
image into a session whose deferred build was still warming (MCP discovery,
model metadata, skills scan) stalled the send and every RPC queued behind it
on the same socket, with no spinner to explain it. prompt.submit already
resolves via _sess_nowait and waits later, off the reader thread — which is
why the symptom reads as "text is instant, images hang".

_sess_building() resolves the session and still kicks off the build (so the
following prompt.submit finds a warm agent), it just doesn't block on it.
_sess() is now expressed in terms of it, so the two differ in exactly one
way: the wait.
2026-08-14 13:24:40 -07:00
poisdahl 3e5e4c5d20 fix(agent): preserve live turns in compaction carriers 2026-08-13 21:38:30 +02:00
kshitij a4f468e832 refactor(gateway/desktop): consent-first truncation precedence + dedup (simplify pass)
Final-diff simplify/review pass findings on #83785:

- Consent gate (confirm_truncate -> 4029) now checked BEFORE target
  resolution, restoring the pre-PR precedence: an unconfirmed submit
  carrying truncation params refuses without paying the durable-transcript
  read or heal-stamping live history dicts, and an unconfirmed out-of-range
  ordinal returns 4029 (not 4018). Malformed params still refuse first
  with 4004. Regression test added (spy DB asserts zero reads pre-consent;
  mutation-checked against the previous commit).
- _coerce_truncate_ordinal generalized to _coerce_truncate_int(param_name):
  the row_id branch was inlining the exact bool-guard + int() -> 4004
  pattern the helper had just extracted.
- Deleted the dead user_indices re-read after _resolve_truncate_row_id
  (heal mutates dicts in place; the filter output is identical) and the
  duplicate range check that had deadened the pre-existing guard.
- Desktop: exported isVisibleUserMessage from use-prompt-actions/utils and
  used it in visibleUserOrdinal / visibleUserIndexAtOrdinal /
  rebindSurvivorRowIds — one predicate for the ordinal parity all three
  depend on instead of three verbatim copies.
- Docs: programmatic-integration.md documents survivor_user_row_ids.
2026-08-13 13:35:55 +05:30
kshitij 42eec4ab38 fix: return survivor row ids after rewind so clients can rebind stale rowIds
Review follow-up (StanleyStetson + egilewski on #83785/#83202): a successful
rewind's replace_messages(archive_dropped=True) re-inserts the surviving
prefix as NEW SQLite rows. Gateway memory picks up the fresh _row_id stamps
via lastrowid, but the Desktop's surviving bubbles kept their pre-rewind
ChatMessage.rowId — so a second rewind/edit/regenerate of an older surviving
turn sent a stale truncate_before_row_id and was (correctly) refused with
4018 until a transcript reload. Fail-closed stays untouched, per both
reviews; the fix is rebinding, not ordinal fallback.

Server: prompt.submit now returns survivor_user_row_ids (fresh post-rewrite
ids of surviving visible user turns, in visible-user-ordinal order) on both
the inline and compute-host paths whenever a durable truncation committed.

Desktop: runRewindSubmit surfaces the field; restore/edit/reload on both the
primary chat and session tiles rebind surviving user bubbles positionally
(same visible-user filter the ordinal math uses) and clear any rowId they
cannot rebind — a cleared id degrades to the ordinal path instead of a 4018.
Absent field (older gateway) leaves state untouched.

Tests: consecutive-rewind regression on a real SessionDB (stale id 4018s,
returned id succeeds; mutation-checked) + vitest for survivorRowIdsFrom /
rebindSurvivorRowIds (rebind, null-clear, past-end clear, hidden skip,
identity preservation).
2026-08-13 13:35:55 +05:30
kshitij 040420bd11 refactor(gateway): dedupe truncation-target validation; drop dead state and redundant test
Review cleanup on the #83202 salvage (findings from the 4-angle + 3-reviewer
passes, all verified against the diff):

- Extract _coerce_truncate_ordinal() and _reconcile_client_ordinal(): the
  bool-check/int-coercion block was duplicated verbatim 3x and the 4030
  ordinal-mismatch block 2x across the row-id/message-id/ordinal branches
  (~90 lines of copy-paste with drift risk between the two durable branches).
- Delete target_idx (4 assignments, 0 reads — the cut uses
  user_indices[ordinal]) and replace the stale inline user-indices
  comprehension with the _history_user_indices helper it duplicated.
- Drop test_reproduce_row_id_truncation: a strict subset of
  test_prompt_submit_truncates_by_row_id +
  test_prompt_submit_refuses_ordinal_and_row_id_mismatch with weaker asserts.
- Collapse PR-introduced blank-line runs in the test file.

Behavior-preserving: error codes, messages, and log fields unchanged
(4004/4018/4029/4030 wording identical); full test_tui_gateway_server.py
suite green (549 passed).
2026-08-13 13:35:55 +05:30
kshitij 16de3c3f1b fix(gateway): verify memory/durable alignment before trusting position in row-id resolve
The #83202 heal path zip-stamped _row_id onto live-memory dicts purely by
position whenever the durable and live lists had equal length, and the DB
fallback mapped durable user-ordinals onto live indices with only a bounds
check. Equal length is not proof of alignment: the durable copy is loaded
with repair_alternation=True (merges user;user pairs, collapses consecutive
assistants, drops orphan tool rows) while live memory is unrepaired and can
carry optimistic/marker rows — the two can coincide in length while
position-shifted. A misaligned stamp is sticky: it permanently attaches the
wrong durable id to a live dict and re-aims every later rewind (E2E probes
showed a wrong-content cut and a persisted alternation break).

_mem_db_pair_agrees() now gates both paths: the heal loop stamps only when
EVERY zip pair agrees on role, display-marker status, and (for addressable
user turns) content; the ordinal fallback verifies the mapped live turn
shows the durable target's content, else refuses via the existing
fail-closed 4018. Regression tests derived from the review probes (content
swap, role shift, repaired-merge ordinal shift); the misalignment guards
fail on the pre-fix code.

Surfaced during review of PR #83202 for #82959.
2026-08-13 13:35:55 +05:30
StanleyStetson 23da6d6fe2 fix(gateway/desktop): durable row-id addressing for rewind truncation
Address rewinds/edits via SQLite messages.id (truncate_before_row_id)
instead of shifting user ordinals. Resolve against in-memory stamps,
then durable session history when live turns drop _row_id; refuse
unknown durable targets with 4018 (no ordinal fallback) and 4030 on
ordinal/row_id mismatch. Stamp _row_id on insert, load row ids on
resume paths, send rowId from Desktop, filter renderer-synthetic ids,
and stop silently resending failed targeted edits without truncation.
Add production-shaped SessionDB tests for resolve and fail-closed paths.

Fixes #82959
2026-08-13 13:35:55 +05:30
Brooklyn Nicholson adbc77eb50 feat(desktop): setup_mcp tool — inline MCP consent card over the clarify-style blocking bridge
New desktop_ui tool: the agent proposes an MCP server (install/enable/
authorize + a one-line reason) and blocks on mcp.setup.request until the
renderer's consent card answers mcp.setup.respond with the outcome
(installed/enabled/authorized/declined/unanswered/error). Same lifecycle
as clarify: 10-min timeout, allow_expired late answers, tool lifecycle
events forced on so the card mounts even with tool progress off. Desktop
prompt hint steers the model to the tool instead of hand-editing config;
every other surface keeps the schema out and is pointed at hermes mcp
install.
2026-08-13 01:06:51 -05:00
StanleyStetson 4d79bd3d02 fix(gateway): reject boolean ordinals and bare confirm_truncate on prompt.submit
Two hardening guards extracted from #82766 by @StanleyStetson:

- bool is an int subclass, so a JSON `true` in truncate_before_user_ordinal
  coerced via int() to ordinal 1 and aimed a CONFIRMED rewind at the second
  user turn — the same silent-loss class as #82756. Reject with 4004.
- confirm_truncate with no truncation target is leaked client rewind state
  on an ordinary submit; fail fast with 4004 instead of silently ignoring
  the flag, so the corrupted client state is surfaced.

Part of the composite fix for #82756.
2026-08-10 11:01:15 +05:30
joaomarcos 60645f8a53 fix(state): make a rewind truncation recoverable instead of a hard DELETE (#82756)
Guarding the *aim* of a rewind still leaves every other way of aiming it
wrong terminal. All three reported incidents (#70516, #80763, #82756) ended
at the same write — `replace_messages()` in the `prompt.submit` truncation
path — and all three were unrecoverable for the same reason: the rows are
DELETEd, which also evicts them from the FTS index, so there is no `active=0`
archive and nothing to restore from.

The codebase already draws this distinction and already has the safe half of
it. `archive_and_compact` is documented as "the durability-preserving
alternative to replace_messages"; `rewind_to_message` — the `/undo` path —
soft-deletes to `active=0, compacted=0` and keeps the rows "on disk for audit
/ forensic inspection". The desktop rewind is the same user-facing operation
as `/undo` and was the one taking the destructive branch.

`replace_messages(..., archive_dropped=True)` flips the DELETE to a
content-preserving `UPDATE messages SET active = 0`, reusing the existing
transaction and the existing `active=0, compacted=0` marking so the dropped
turns stay readable via `get_messages(..., include_inactive=True)` and stay
out of session search (`compacted=0` = "the user took it back", vs
compaction's `compacted=1` = "summarized away, still discoverable").

The live transcript is byte-identical either way — only the durability of the
dropped turns changes. The parameter defaults to False, so the fork handler,
the ACP adapter and `gateway/session.py` keep their current semantics
untouched; a test pins that.

`active_only=True` stays on the call: #80216 still applies, and archiving must
not disturb rows an earlier compaction deliberately archived.

Test doubles for `replace_messages` in the gateway suite are widened to the
real signature — they are stand-ins for SessionDB, and a double that does not
accept what production passes silently converts this write into a 5008.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:01:15 +05:30
Brooklyn Nicholson 0665cd4b5b style(hud): tighten the surface-note comments and test helper
Comment wording only, plus the desktop test's boolean parameter becomes
an 'app' | 'hud' union so the call site says which window it means.
2026-08-08 16:21:15 -05:00
Brooklyn Nicholson e24bac49fa feat(desktop): tell the agent when it is floating in HUD mode
In HUD mode Hermes is a strip over the app the user is actually working
in, so "what's under you?" or "look up the weather" is almost always
about that app — but the agent had no way to know it was floating, and
answered from its own browser and panes instead.

The desktop tags a HUD submit with `surface: 'hud'` and the gateway turns
that into a per-turn note pointing at read_window_below, and at carrying
the work out in the app underneath. It rides the model-bound message
beside the reaction and speech-interrupted notes rather than the system
prompt: one session can be driven from the app window on one turn and the
HUD on the next, and the system prompt has to stay byte-stable.

Every tool the note names is checked against the agent's own schema
first, so a session without computer_use or read_window_below is never
pointed at a tool it cannot call.
2026-08-08 15:37:41 -05:00
Brooklyn Nicholson 406501fd97 feat(agent): read_window_below tool — which OS window is underneath the desktop app
Desktop-gated (desktop_ui toolset) metadata-only window awareness: the agent
can ask which application window sits directly behind the Hermes window
(app, title, bounds — never pixels). Rides the same blocking bridge as
read_terminal: the gateway emits window.read.request and the renderer
answers window.read.respond.
2026-08-08 12:17:50 -05:00
HexLab98 c24ff38c51 fix(gateway): make a history-dropping submit prove it meant to
prompt.submit honored truncate_before_user_ordinal on every request. A
client that carried a leftover ordinal into an ordinary send therefore
issued something the gateway could not tell apart from a real rewind —
same method, same shape, an in-range target — and the cut was applied
with replace_messages(), which DELETEs the durable rows. One report lost
244 messages (296 -> 52) with no prompt and nothing to restore from.

The existing guard only covered ordinal 0, where the cut empties the
transcript; a mid-session ordinal sailed straight through. Only the
client knows whether a submit is a rewind, an edit, or a regenerate, so
require it to say so: an ordinal without confirm_truncate is refused on
4029 and neither memory nor the DB is touched. Desktop sends the flag
from the one place that builds these params, so every rewind path is
covered and a stale build fails closed with an actionable error instead
of quietly deleting a conversation.
2026-08-07 18:20:19 +05:30
kshitij ee6d79648a fix(state): finish the #80216 bug class — archive-preserving rewrites at the two remaining sibling sites
#80216 fixed /retry (and a follow-up fixed yuanbao recall) destroying
soft-archived active=0/compacted=1 in-place-compaction rows via the
destructive replace_messages default. Two sibling sites still carried the
same class:

- acp_adapter/session.py _persist (non-owned-agent branch): probed
  has_archived_messages and FAILED OPEN into the destructive full replace
  on any probe error; the probe can also race a concurrent
  archive_and_compact. Now passes active_only=True unconditionally — on a
  fresh create/fork every row is active=1 so behavior is identical, and
  the probe (its only production caller) is deleted.
- tui_gateway/methods_prompt.py edit/regenerate truncation: bare
  replace_messages() deleted the archived transcript of a compacted
  session on every edit/regenerate. Now active_only=True.

hermes_state.has_archived_messages docstring updated (probe is now
test/diagnostic-only). Test stubs in test_tui_gateway_server.py accept the
new kwarg. New regression tests: real-SQLite archive-survival for both
write shapes, fresh-session equivalence (the claim the unconditional
switch rests on), and source-level guards pinning that neither site
re-grows the fail-open probe (both mutation-checked: revert either fix and
its guard fails).
2026-08-07 14:42:32 +05:30
brooklyn! 64646dda56 Hermes can read the in-app browser (#79482)
* feat(agent): read_preview — the desktop-gated tool that reads the in-app browser

The agent could open the preview pane (open_preview) and read the embedded
terminal (read_terminal), but the browser it had just opened was a black box —
'what does this page say?' had no answer. read_preview mirrors read_terminal
end to end: HERMES_DESKTOP-gated via check_fn (zero schema footprint outside
the GUI), dispatched through the same agent callback pattern, windowed with
start/count so a long page pages instead of flooding context.

* feat(gateway): preview.read blocking bridge

Same lifecycle as terminal.read: the tool blocks on preview.read.request, the
renderer answers preview.read.respond (allow_expired — a slow page extraction
losing the 45s race must not surface a raw 4009), and a timeout emits
preview.read.expire so late answers resolve quietly.

* feat(desktop): the renderer serializes the active preview tab for the agent

preview-reader.ts is the preview analog of the terminal's buffer registry: the
URL pane registers a page reader (webview executeJavaScript → title + visible
innerText) keyed by tab id; readActivePreview resolves the ACTIVE tab, windows
the text (24k cap per read), and answers file/artifact tabs with identity plus
a note pointing at the tool that reads that content directly. The gateway
event handler answers preview.read.request beside terminal.read.request.
2026-08-05 16:35:00 +00:00