Commit Graph

392 Commits

Author SHA1 Message Date
Teknium fda54a613e chore(tests): remove two flaky test files that tax CI
tests/tools/test_website_policy.py and tests/cli/test_surrogate_sanitization.py
repeatedly failed under the parallel CI runner (process-teardown timeouts /
async-timeout flakes) across multiple unrelated PRs this cycle, while passing
locally. Removed per maintainer direction to stop the flake taxing every PR.
2026-08-20 04:54:54 -07:00
Teknium c32119b12c feat(config): default agent.max_turns to unlimited; accept inf/infinity/null spellings
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
2026-08-20 04:50:39 -07:00
Teknium 05b4ab0ceb feat(cli): declutter /help + Ctrl+P command palette (C-04/C-05) 2026-08-20 04:30:33 -07:00
Teknium 9bdff6ab68 feat(cli): /status shows reasoning, approval mode, and context usage (C-02) 2026-08-20 04:14:07 -07:00
Teknium e931081a44 feat(cli): type-to-fuzzy-filter the /model picker model list 2026-08-20 03:59:16 -07:00
Teknium 3f67921fca fix(cli): /undo typo no longer quits the CLI; +3 slash-command papercuts 2026-08-20 01:46:08 -07:00
Teknium fc9b4186a0 fix(worktree): pruner reaps rebase-merged trees via PR state; worktree add survives disk contention
Two gaps behind the recurring 'hermes -w timed out after 30 seconds':

1. Rebase-merge leak: git cherry only catches patch-identical commits.
   Salvage flows routinely change the diff (conflict resolution, follow-up
   commits), so 12 of 22 'unpushed' trees on the incident box had MERGED
   PRs and were preserved forever. The pruner now falls back to
   'gh pr list --head <branch> --state merged' — authoritative, memoized
   on (branch, head_sha) with True-only caching, fail-safe to preserve.

2. Creation timeout 30s -> 120s: the ~10k-file checkout measured 113s at
   near-zero CPU under multi-agent disk contention vs 1.2s idle. 30s
   killed legitimate creates and threw away completed work.
2026-08-20 00:00:28 -07:00
kshitijk4poor 446da3ef56 fix(cli): widen lock-bit coverage to alias installers, legacy nav keys, and PUA functional keys
Follow-up to the salvaged #89676 + #90291 lock-bit fixes: extract a shared _lock_variants() helper and cover the sites both PRs missed - install_shift_enter_alias / install_ctrl_enter_alias / install_cmd_backspace_alias CSI-u spellings, legacy CSI-letter and CSI-tilde navigation twins derived from the existing table for ALL modifiers 1-16 (not just plain/shift), plain F1-F4 SS3 fallback, unmodified CSI-u keys (Tab/Enter/Space/Backspace), and kitty PUA functional keys (keypad, F13-F24, Ignore range) under lock bits. 8 new tests.
2026-08-20 12:16:20 +05:30
liuhao1024 118dbe871f fix(cli): map kitty CSI-u lock-bit variants so key combos survive NumLock
kitty and ghostty OR the CapsLock (64) / NumLock (128) state into the
CSI-u modifier parameter. With NumLock on, Ctrl+C arrives as
ESC[99;133u (5 + 128) instead of ESC[99;5u; the alias table had no
entry for it, so every key combo leaked as literal text like
[127;133u (#89651). Install every CSI-u alias with the lock-bit
variants (+64/+128/+192); the xterm modifyOtherKeys encoding never
carries lock bits, so the ESC[27;N;CP~ form is left untouched. The
Esc-key registration now covers modifier 1 as well (1+128=129 is a
lone Esc with NumLock on).
2026-08-20 12:16:20 +05:30
kshitijk4poor 45f11263bd fix(tui): skip the kitty protocol push for Ghostty in the Ink TUI too
Widen the cli.py Ghostty exception to the sibling sites the review found: the Ink TUI pushes CSI >1u at raw-mode entry (App.tsx), on alt-screen exit, and on the extended-keys re-assert path (ink.tsx) for every EXTENDED_KEYS_TERMINALS entry including ghostty - same Alt-stripping bug. New skipKittyKeyboardProtocol() helper in terminal.ts gates the ENABLE push at all 3 sites; the DISABLE (pop) stays unconditional since popping an empty stack is a spec no-op. Also fix the cli.py comment citing the modifyOtherKeys encoding where the kitty CSI-u form (ESC[127;3u) is what the broken path expected, dedupe the quadruplicated Ghostty comment, and update the stale 'mirroring the Ink TUI' docstring. 7 new vitest cases.
2026-08-20 11:39:03 +05:30
kshitij 1a8fea3ce2 fix(cli): skip Kitty keyboard protocol push for Ghostty, use modifyOtherKeys only
Ghostty's Kitty disambiguate-mode implementation strips the Alt modifier
from the Backspace key — Option+Backspace arrives as bare \x7f instead of
the expected \x1b[27;3;127~, breaking backward-kill-word.  This was a
regression introduced when PR #87630 re-added the CSI >1u Kitty protocol
push for all allowlisted terminals including Ghostty.

Under modifyOtherKeys mode (CSI >4;2m), Ghostty correctly sends
\x1b[27;3;127~ for Option+Backspace, which the alias table in
pt_input_extras already maps to (Escape, ControlH) = backward-kill-word.

Fix: for Ghostty only, push just modifyOtherKeys and skip the Kitty
protocol push.  All other terminals (iTerm2, WezTerm, kitty, tmux, VS Code)
still get the full dual-protocol push.

Ghostty upstream tracking: discussion #9560, issue #9895 (cmd+backspace
variant of the same root cause).
2026-08-20 11:39:03 +05:30
Teknium f4a866b484 fix(cli): /config displays the live agent credential, not the env-var constructor seed 2026-08-19 19:42:07 -07:00
Teknium b035036582 fix(cli): /yolo reports locked-ON under process-frozen YOLO instead of a false OFF 2026-08-19 19:41:59 -07:00
kronexoi 4f12cfd98d fix(cli): wire /whoami slash command in classic CLI 2026-08-19 19:41:32 -07:00
ethernet 100380bbf5 feat(cli): batch clarify back-navigation and answer visibility
Shift-Tab walks backwards through the questions, with wrap, the same
way Tab walks forward. A locked answer now renders on its own indented
line in a distinct color under its question, instead of an arrow
suffix on the status line, so the current answers stay readable while
the cursor moves.

A re-visited question restores its earlier state: a choice answer puts
the cursor back on that choice, a typed answer highlights the Other
row and shows the typed text next to it. Enter on an answered Other
switches to freetext with the composer prefilled with the earlier
text, so the user edits instead of retyping. Answer metadata records
how each answer was produced to drive the restore.
2026-08-18 21:28:53 -04:00
ethernet b4526d46d0 feat(cli): compact multi-question clarify panel
The clarify callback accepts a questions list and renders a batch
panel. The batch panel shows all questions as a status list with one
expanded active question. Enter locks the active answer and moves to
the next unanswered question. Tab cycles questions for any-order
answering. Locked answers stay editable until the batch completes. A
timeout returns the locked partial answers with a timed_out flag. The
single-question panel is unchanged.
2026-08-18 21:28:53 -04:00
Teknium 56de3c1428 fix(cli): persist one-shot resumed-session turns (Bot Chat bot-to-bot messages)
Bot Mode's bot-to-bot send (`hermes -p <bot> chat --in ~ -c "Bot Chat"
--create-if-missing -Q -q "..."`) runs one turn and exits. When the turn's
in-loop transcript flush failed transiently (state.db write-lock contention
with a multiplex gateway), the one-shot path had no end-of-run durable
retry: the reply reached stdout and agent.log while the resumed titled
session's stored history never changed (#88583). The interactive CLI is
immune — it retries the flush on the next persist point and finalizes the
row on quit — but every one-shot exit path lacked both.

Fix the whole class with cli._flush_one_shot_session_store():

- final _persist_session retry at one-shot exit (idempotent — per-message
  persisted-marker stamps mean already-written turns are not re-written)
- drain queued async token-accounting deltas
- end_session(..., "cli_close") so resumed/created titled session rows no
  longer dangle open forever after one-shot runs

Wired into _finalize_single_query (quiet -Q -q AND human -q paths, ahead
of memory-provider shutdown so nothing later can lose the turn) and into
the kanban SIGTERM handler before os._exit(0), which skips atexit and the
SessionDB token-drain hook entirely (same gap class as PR #50881).

Handed-off sessions (#88234) and persistence-isolated forks
(_persist_disabled) are skipped.

Fixes #88583

🤖 Generated with Hermes Agent
2026-08-17 16:55:07 -07:00
kshitij 979ca57a50 Merge pull request #88244 from kshitijk4poor/fix/handoff-cleanup-race
fix: prevent handoff leg data loss + surface state.db corruption to users
2026-08-17 17:41:44 +05:30
Teknium 5975aff8a1 test: assert sprawl as a behavior contract, not a pack-count snapshot
CI git consolidates during incremental pack creation differently per
build (4 packs from 6 attempts on ubuntu-latest, 6 locally, 3 on the
previous run) — even pack-objects counts drift with auto-maintenance.
The fixture now only guarantees strictly-more-packs-than-threshold and
the test asserts consolidation strictly decreases the count.
2026-08-17 03:03:47 -07:00
Teknium 31fe024e93 test: make pack fixture deterministic via git pack-objects
Incremental 'git repack' consolidates small packs on newer git builds
(CI produced 3 packs from 6 commits), making the sprawl fixture count
nondeterministic. pack-objects with an explicit sha per commit creates
exactly one pack each on every git version.
2026-08-17 03:03:47 -07:00
Teknium 25c051d480 fix(cli): self-heal git health around worktree creation
Two gaps from the Aug 2026 'hermes -w timed out after 30s' incident:

1. Atomic failure cleanup: a timed-out/failed `git worktree add` left a
   partially-materialized directory plus a LOCKED admin entry under
   .git/worktrees/ (lock pid = the live hermes process that timed out),
   which the startup pruner's dead-pid unlock never reaps — retries of
   the same name fail forever. _cleanup_failed_worktree_add sweeps dir,
   admin entry, and orphaned branch on every failure path (timeout,
   nonzero exit, remote-base retry).

2. Pack maintenance: nothing consolidated the object store; on a
   multi-agent box packs sprawl (39 packs / 638MB at the incident) and
   every object lookup scans all pack indexes until worktree creation
   blows its timeout. _maintain_pack_health repacks (niced, background,
   fail-soft) when *.pack count reaches 15, wired into the existing
   startup maintenance thread on both the CLI (-w) and TUI paths.
   gc --auto doesn't cover this: its threshold is 50 packs.

Both sabotage-verified; full repack on the incident box: 39 packs ->
2, 638MB -> 287MB, worktree add 30s-timeout -> 0.5s.
2026-08-17 03:03:47 -07:00
kshitij 3d6cfdd83b fix: close sibling finalize path for handed-off sessions (#88234)
/simplify-code review found _notify_single_query_session_finalize was
missing the _handed_off_session_ids guard that _should_emit_cleanup_session_finalize
and _emit_interrupted_session_end already had. One-shot CLI queries that
somehow handed off would still finalize the session via this path.

Added guard + test.
2026-08-17 13:26:29 +05:30
kshitij 59f302fef9 fix: prevent handoff leg data loss + surface state.db corruption to users
Two data-loss bugs reported by users:

1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
   cleanup called finalize_session on the session the gateway just
   reopened. This set end_reason on a row the gateway was actively
   writing to, causing the handoff leg to vanish from session history
   and breaking session_search recall. Fix: add _handed_off_session_ids
   module-level set (mirrors _single_query_finalize_attempted_session_ids
   pattern). _handle_handoff_command registers the session_id on
   completion; _should_emit_cleanup_session_finalize and
   _emit_interrupted_session_end check it before firing.

2. state.db corruption silent failure (#88235): When SessionDB init
   failed at gateway startup, the error stayed in logs — messages
   flowed but nothing was persisted, with no user-visible indication.
   Fix: store _session_db_init_error on GatewayRunner, broadcast a
   recovery-guidance message to all home channels via
   _send_session_db_warning_notifications() after the gateway connects.
   Also improved the 'corrupt' persistence cause wording in
   _format_turn_completion_explanation to include the full recovery
   path (hermes doctor --fix, sqlite3 .recover, backups).

Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.
2026-08-17 13:19:58 +05:30
kshitij 56526bc0d3 Merge pull request #87630 from kshitijk4poor/fix/kitty-extended-keys-proper
fix(cli): restore Kitty keyboard protocol push and complete the extended-key alias table
2026-08-16 15:57:33 +05:30
kshitij 03bf85d83b fix(cli): restore Kitty keyboard protocol push and complete the extended-key alias table
Commit 4c34eeb416 fixed dead Ctrl+C by removing the Kitty protocol push
(CSI >1u) from _EXTENDED_ENTER_KEYS_SEQ, keeping only modifyOtherKeys
level 2. That regressed kitty-the-terminal completely: kitty removed
xterm modifyOtherKeys support (kovidgoyal/kitty#4075) and only speaks
its own protocol, so after the removal kitty users lost Shift+Enter and
every other extended key — the CSI >4;2m we still pushed is a no-op
there (kitty even logs a PARSE ERROR for it).

The original reason for removing the push is obsolete: #87511 mapped
CSI-u control sequences, so Ctrl+C as ESC[99;5u now parses to
Keys.ControlC and fires the existing c-c binding. (The kernel-INTR
concern in that commit was moot — prompt_toolkit's raw mode clears
ISIG, so Ctrl+C is always handled by the binding, never the kernel.)

Restore the dual push (CSI >1u + CSI >4;2m), exactly mirroring the Ink
TUI, and complete the alias table for what the kitty disambiguate flag
actually emits — #87511 left real gaps, some of which its PR body
wrongly claimed were covered:

- Esc key: ESC[27u (+ modifiers) — previously leaked '[27u' as text
- Ctrl+Backspace -> backward-kill-word (#78285 was closed on the wrong
  claim that codepoint-127 mapping existed; it did not)
- Shift+Space -> space (#86866's second symptom; the Ctrl+Space
  mapping never covered modifier 2)
- Alt+Enter -> newline tuple; Shift+Tab -> BackTab; Ctrl+Tab -> Tab;
  Alt/Shift+Backspace
- Multi-modifier letters (Shift+Alt 4, Ctrl+Shift 6, Ctrl+Alt 7,
  Ctrl+Alt+Shift 8) normalized onto their Ctrl/Escape-prefix targets,
  both unshifted (kitty) and shifted (mok emitters) codepoints
- Kitty PUA functional keys: keypad -> non-keypad equivalents,
  F13-F24, and Ignore for lock/media/modifier-event keys so they are
  consumed instead of leaking (kitty emits these even in legacy mode)

Also: clear the VT100 parser's prefix cache after installing (stale
answers could misparse), and re-push extended keys after
_recover_terminal_input_modes' reset — the recovery previously popped
both modes mid-session and never re-enabled them, silently killing
Shift+Enter until restart.

Refs #87511, #87074, #56684, #56645, #78285, #86866, #87390.
2026-08-16 15:52:24 +05:30
zhao 0a8092ac32 fix: guard exit watchdog against mid-cleanup overlap 2026-08-16 01:56:43 -07:00
Christopher 3cbd86aac8 fix(cli): warn when --provider default_model cannot be resolved
A named --provider without -m used to swallow lookup failures and
silently keep the global model.default. Log the resolution error so
the fallback is visible.
2026-08-16 01:51:35 -07:00
Christopher d6688adcc2 fix(cli): use custom provider default_model when --provider is set
hermes chat --provider <name> without -m sent the global model.default
to the custom endpoint. Named custom entries already expose
default_model via _get_named_custom_provider(); honor that when the
user selected the provider and did not pass an explicit model.

Fixes #86978
2026-08-16 01:51:35 -07:00
kshitij 2be183142c refactor: extract _install_paired helper, fix misleading comment, isolate test fixture
Apply findings from /simplify-code 3-agent review:

1. Extract _install_paired() inner helper — the Ctrl, Alt, and Shift
   sections all repeated the same mok+csiu sequence generation pattern
   (~30 lines of duplication). Now each section builds a dict and
   delegates to _install_paired(modifier, mapping).

2. Replace 10 hardcoded Ctrl+digit lines with a loop matching the
   Ctrl+letter pattern above it.

3. Fix misleading comment: claimed 'Ctrl+0 doesn't produce a control
   byte' but chr(ord('0') & 0x1F) = 0x10 = ControlP. The code was
   correct (maps directly to Keys.Control0..9); only the comment was
   wrong.

4. Add comment explaining why Shift+letter maps both lowercase and
   uppercase codepoints (some terminals send the already-shifted
   codepoint with modifier=2).

5. Test fixture: snapshot/restore ANSI_SEQUENCES in teardown so 294
   mappings don't leak into sibling test files (global mutable state).
2026-08-16 13:02:27 +05:30
kshitij b353ac39a2 fix(cli): map all Ctrl/Alt/Shift+key combos under modifyOtherKeys level 2
Commit 4c34eeb416 stopped pushing the Kitty keyboard protocol (CSI >1u)
because Ctrl+C arrived as ESC[99;5u instead of \x03, breaking SIGINT.
But modifyOtherKeys level 2 (CSI >4;2m) was kept so Shift+Enter stays
distinguishable from Enter.

Under modifyOtherKeys=2, terminals re-encode EVERY Ctrl+key combo as
ESC[27;5;<codepoint>~ instead of the raw control byte. prompt_toolkit
3.x only maps ESC[27;5;13~ (Ctrl+Enter = Ctrl+M); all other Ctrl+letter
combos are unmapped and leak as literal text or get swallowed — breaking
Ctrl+A, Ctrl+C, Ctrl+D, Ctrl+E, Ctrl+K, Ctrl+R, Ctrl+U, Ctrl+W, Ctrl+Z,
etc. Shift+letter combos (ESC[27;2;<codepoint>~) have the same problem,
causing the 'caps locked sessions' symptom where typed text appears
corrupted or stuck.

Fix: add install_modify_other_keys_aliases() to pt_input_extras.py that
populates prompt_toolkit's ANSI_SEQUENCES dict with 294 mappings covering:
- Ctrl+letter (a-z): ESC[27;5;<code>~ and ESC[<code>;5u -> Keys.ControlA..Z
- Ctrl+digit (0-9): same formats -> Keys.Control0..9
- Ctrl+symbol ([ \ ] ^ _ @ Space): same formats -> matching Keys.Control*
- Alt+letter (a-z, A-Z): both formats -> (Escape, <letter>) tuple
- Shift+letter (a-z, A-Z): both formats -> uppercase character

Uses setdefault semantics — never clobbers existing mappings from
install_shift_enter_alias or install_ctrl_enter_alias. The Ink TUI
(Node.js) already handles this via a regex parser; prompt_toolkit 3.x
uses dict lookup only, so we populate the dict.

Refs #56684, #87711.
2026-08-16 13:02:27 +05:30
Teknium bfd9cef389 fix(worktree): deepen shallow clones so worktree cleanup can verify push state
The installer clones with --depth 1, so every default install is shallow.
In a shallow repo, an older worktree HEAD (a past snapshot of main) is
disconnected from current origin/main by the shallow boundary, so
'git log HEAD --not --remotes' misreports thousands of already-public
commits as unpushed. The fail-safe unpushed guard then preserves every
aged 'hermes -w' worktree forever, and the git-cherry squash-merge
escape hatch never rescues them (22k 'ahead' >> max_ahead=20).
Real incident: 21 of 25 hermes-* worktrees stuck on one install.

Fix at the root, one owner:
- _deepen_shallow_repo(): one-time blobless unshallow
  (fetch --unshallow --filter=blob:none; plain --unshallow fallback)
  run from the background startup pruner thread before classification,
  so history verdicts become correct and the backlog self-clears on the
  next 'hermes -w' startup. Fail-soft offline: keep preserving.
- _cleanup_worktree(): when the unpushed verdict comes from a shallow
  clone, say 'Shallow clone — cannot verify push state' instead of the
  misleading 'has unpushed commits' message.
- Document the shallow caveat on _worktree_has_unpushed_commits (the
  primitive stays conservative on purpose).

Tests: real shallow clone over file:// reproducing the disconnect shape,
covering detection, deepen+verdict flip, pruner E2E reap, offline
fail-soft preserve, full-clone noop, and genuine-unpushed-work survival.
Sabotage-verified: the E2E test fails with the deepen call disabled.
2026-08-15 04:02:07 -07:00
Teknium ca5c04818b feat: bare /save shows a usage card instead of exporting
Per review: /save with no arguments prints usage (formats, filename,
redact, examples) and writes nothing; the export requires an explicit
format. Unknown formats print the error plus the same usage card. One
shared SAVE_USAGE string in hermes_cli/session_export.py serves both the
CLI and gateway handlers. args_hint updated to <json|md|html> to reflect
the now-required format. Tests: bare-save and bad-format usage cases
added; location tests pass an explicit format.
2026-08-15 02:04:49 -07:00
Teknium e5a2c80c0e fix: harden /save for test doubles, async session-store boundary, snapshot shape pin
Three CI-red follow-ups on the /save rework:

- cli.py save_conversation: getattr-guard _session_db/session_id so
  SimpleNamespace/object.__new__ test doubles (and any embedder passing a
  minimal stub) don't AttributeError (pitfall #17 pattern).
- gateway/slash_commands.py: route through the awaited
  async_session_store.get_or_create_session boundary — the architecture
  test forbids raw session_store calls in async gateway code.
- tests/cli/test_save_conversation_location.py: /save now emits the
  canonical export_session payload shape; the session id key is "id"
  (was "session_id" in the legacy snapshot format) — update the pin.

The slice-8 lost-and-found failure was pre-existing on main and is fixed
there by f2a30fa400 (test: derive lost-and-found synthetic width from the
live schema); picked up via rebase.
2026-08-15 02:04:49 -07:00
John Lussier b3447c2129 test: align submit bindings with multiline default 2026-08-15 01:05:39 -07:00
John Lussier 2ae7884ffa fix: make CLI multiline shortcuts work by default 2026-08-15 01:05:39 -07:00
coe0718 7f7aefe5cb fix: restore complete message timestamp coverage 2026-08-15 01:04:19 -07:00
mariobgsp 998329a621 fix: flatten dict-valued model.default at config load boundary
Extends the fix to the config-load chokepoint so every reader sees plain
strings, not just the interactive CLI paths. _normalize_root_model_keys
now flattens a dict-valued model.default/model.model into a string default
plus the nested provider (promoted to model.provider when no explicit
outer provider or "auto" is set), covering the residual readers the
review flagged: doctor, status/dump, fallback picker, prompt-size, and
the context-switch guard — all of which called .strip()/flowed the raw
value and would crash or misroute on a nested dict.

Adds _normalize_root_model_keys regression coverage for the flatten
(precedence, auto-override, explicit-provider-win, alias shape, flat
strings untouched).
2026-08-15 10:29:05 +05:30
mariobgsp 86054eff62 fix: keep nested model.default provider paired through HermesCLI
The prior dict coercion converted a dict-valued model.default to a plain
model string but dropped the nested provider. On the interactive CLI path
requested_provider then fell back to the outer merged model.provider
(typically "auto", authoritative at runtime resolution), so the model
could be routed through the wrong active provider.

Canonicalize both halves at the shared boundary: _split_model_config_default
flattens a dict-valued default into (model, provider) and HermesCLI.__init__
feeds the nested provider into the requested_provider chain (still below an
explicit --provider argument). new_session reuses the same helper.

Adds regression coverage asserting the nested provider stays paired with the
model, that flat string defaults keep the outer provider behavior, and that
an explicit provider argument still wins.
2026-08-15 10:29:05 +05:30
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
Teknium 3bd98ec4bf fix(cli): redraw on terminal focus regain + docs for scrollback rebuild config
Completes the duplicated-chrome class fix:
- Focus-in (CSI I) now routes through the same rate-limited full-redraw
  recovery as Ctrl+L//redraw, clearing ghost prompt/composer copies after
  Alt+Tab / tab switches (focus-regain variant reported on #60920, #25337)
- Document display.cli_rebuild_scrollback_on_redraw in configuration.md
- Register the new default in hermes_cli/config_defaults.py (moved from
  the pre-refactor config.py location the salvaged commit targeted)
2026-08-14 13:09:37 -07:00
HunterSThompson c3d7f6eefd fix(cli): recover prompt_toolkit paint after tmux attach
Same-width SIGWINCH (typical tmux attach) skipped screen clear and
left previous_screen inconsistent, crashing redraw with
'cell' object has no attribute 'char'. Always clear on resize
recovery and retry _output_screen_diff with previous_screen=None
on AttributeError/TypeError.
2026-08-14 13:09:37 -07:00
angeon 88a1a9fd95 fix(cli): let redraw recovery rebuild scrollback 2026-08-14 13:09:37 -07:00
angeon 0e1cba326b fix(cli): honor persisted status bar visibility 2026-08-14 13:09:37 -07:00
halaprix a22fbba340 fix(cli): don't replay transcript on the session's first benign SIGWINCH
The resize recovery treated the first SIGWINCH of a session as a width
change (no prior width to compare against), running the Ctrl+L-style
viewport clear + _OUTPUT_HISTORY replay. The 2J clear preserves
scrollback, so everything in the deque printed a second copy below the
still-visible original. After --continue/--resume the deque holds the
whole "Previous Conversation" recap plus the first live exchange, so a
benign resize signal (GNOME Terminal tab bar appearing, monitor-scale
change, focus events) duplicated the entire conversation.

Seed the width baseline when the resize hook is installed, and replay
only on an observed width change. The baseline is read from app.output
— get_app() at install time is still the DummyApplication whose
DummyOutput reports a fake 80 columns, which would turn the first real
signal back into a phantom width change. A real initial maximize or
restore still differs from the seeded width and is still recovered
(#49120 behavior preserved; verified in a pty harness both ways).

Fixes #65293
2026-08-14 13:09:37 -07:00
Alli 6625c72c92 fix(cli): use _suspend_output_history for interrupt marker instead of clearing _OUTPUT_HISTORY
The original fix for #60920 cleared _OUTPUT_HISTORY in _recover_terminal_after_interrupt
to prevent the interrupt marker from being replayed on redraw. This approach:
- Discarded legitimate scrollback content unnecessarily
- Broke /redraw and Ctrl+L replay for any content after an interrupt

Instead, the interrupt marker is now printed via _cprint inside a
_suspend_output_history() context so it never enters _OUTPUT_HISTORY.
_recover_terminal_after_interrupt no longer needs to clear history — the
marker was never recorded, so _replay_output_history cannot duplicate it.

Also adds:
- _show_interrupt_marker flag to cleanly separate marker rendering from
  response construction
- Focused regression tests covering the marker recording suppression,
  history preservation after recovery, replay cleanliness, and flag logic

Fixes: #60941
2026-08-14 13:09:37 -07:00
webtecnica 56a41715dc fix: persist provider on model switch and add billing_provider fallback
Salvage of #79604 (webtecnica) + #85721 (pierrenode), combined and
rebased onto current main with simplify-code findings folded in.

#79604: update_session_model() wrote the model name to sessions.model
but never persisted the provider into model_config. On resume, the
runtime recombined the persisted model with the config.yaml primary
provider (which may not serve that model), producing auth errors.
Fix: add optional provider parameter to update_session_model, merged
into model_config via the shared _merge_model_config_json helper (not
hand-rolled SQL). Wire both gateway /model call sites to pass
result.target_provider.

#85721: session_gateway_runtime() had no billing_provider fallback.
A CLI session that never ran /model has no gateway_runtime or
top-level provider in model_config — billing_provider (written on
every session's first accounted API call) is the only durable record.
Fix: add billing_provider as the last-resort fallback in
session_gateway_runtime(), filtering bare billing buckets (auto/custom)
that are not routable identities.

Simplify-code findings addressed:
- Use _merge_model_config_json instead of 40 lines of branched SQL
- Share _BARE_BILLING_PROVIDERS from hermes_state.py (was duplicated
  as a set in tui_gateway/server.py)
- Merge None-filtering from #85920 with the billing_provider fallback
  into one coherent return path

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-08-14 14:39:08 +05:30
kshitij 9cb456a9b9 fix: unify route dict or-None discipline in /model persist
The route dict in _persist_model_switch_to_session used  filtering
(omits falsy values) while the top-level keys used  (writes
explicit None to trigger deletion in _merge_model_config_json). This
asymmetry meant stale keys from a previous /model switch survived in
the nested gateway_runtime dict even after the fix in #85261 that
properly deleted them from the top-level keys.

Fix: build the route dict with  and derive the top-level keys
from **route so both shapes always use identical deletion semantics.
Also filter None values in session_gateway_runtime's reader since
gateway_runtime is replaced as a whole dict (not deep-merged), so None
values written by the persist path survive in the nested dict.

Found by /simplify-code 3-reviewer review on #85261 (all 3 reviewers
converged on the route dict asymmetry as the verdict-relevant finding).
2026-08-14 13:18:53 +05:30
Teknium 2ae96939f5 fix(cli): self-heal cooked-mode termios drift that freezes CLI input
When a prompt_toolkit run_in_terminal cooked->raw restore is lost (cancelled
coroutine, racing chained cross-thread windows from background-review
summaries / process-notification prints), the tty stays in cooked mode while
the Application still expects raw. The kernel line-buffers keystrokes and the
CLI appears to stop taking input even though the event loop is healthy.

Observed live 2026-08-13: interactive session left in 'icanon echo' after a
background skill-review fork + notify_on_complete turn; only an external
stty rescue restored input.

Fix: _heal_cooked_mode_drift() re-applies prompt_toolkit's own raw-mode flag
surgery when stdin's lflag has drifted cooked, and process_loop's idle branch
runs a rate-limited _check_termios_drift() watchdog that skips legitimate
cooked windows (app._running_in_terminal), agent-running phases, non-tty
stdin, and Windows.
2026-08-13 14:22:14 -07:00
kshitij d002167390 fix: delete stale top-level route keys on /model persist
patch_session_model_config merges key-level and only deletes on explicit
None. Dropping falsy values from the top-level patch let a previous
switch's api_mode/base_url survive the next switch — TUI/desktop resume
then restored e.g. openrouter with anthropic_messages wire mode, and a
failed bare-custom heal produced a stale-provider/new-endpoint route.
Write absent top-level values as explicit None so each switch fully
replaces the persisted route. Regression test against a real SessionDB;
mutation-checked. Also correct the heal comment (CLI is deliberately
stricter than the TUI recovery, which keeps bare custom with a base_url).
2026-08-14 02:15:22 +05:30
kshitij dbe24dfc12 fix: heal bare-custom provider at persist AND restore; persist --global switches to the row
- Bare 'custom' from ModelSwitchResult.target_provider is the resolved
  billing class, not a routable identity; persisting it verbatim made a
  later --resume hard-fail once the config default moved off the custom
  endpoint. Heal to custom:<name> via canonical_custom_identity at
  persist time, and again on restore for rows written by older builds
  (mirrors tui_gateway's _stored_session_runtime_overrides recovery).
- --global switches now also update the session row: the row records
  what THIS session runs, otherwise resume restored the stale
  creation-time model over the user's new global choice.
- Only adopt resolved credential_pool alongside its api_key (don't null
  the ambient pool when resolution returns no credentials).
- 3 new tests; healing path mutation-checked.
2026-08-14 02:15:22 +05:30