Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium 03662dca77 chore(release): map KHALIDagara contributor email 2026-08-23 21:12:26 -07:00
Teknium 859010b94e fix(desktop): gate transport retries to idempotent or provably-unsent requests
Follow-up hardening on #92977 (issue #92976). The cherry-picked retry
wrapped every verb, so an ECONNRESET arriving after the backend had
already processed a POST (prompt submitted, session created) would
silently double-submit on retry.

- Extract the transport policy into electron/api-transport.ts so it is
  unit-testable without Electron: keep-alive agent pools, transient
  error classification, and a verb-gated withRetry.
- Retry rule: GET/HEAD/OPTIONS retry on any transient transport error;
  POST/PUT/PATCH/DELETE retry only when the request provably never
  reached the server (connect-phase failures like ECONNREFUSED /
  ENOTFOUND, or an error thrown before the body was flushed —
  requestState.bodySent === false). Ambiguous resets after the body
  went out surface to the caller; when in doubt, don't retry.
- Separate keep-alive pools for JSON calls vs streaming downloads so
  long downloads can't starve latency-sensitive JSON calls.
- Destroy pooled agents on app will-quit.
- Tests: shouldRetryRequest truth table, withRetry behavior, plus LIVE
  transport tests against real misbehaving node HTTP servers: a GET
  burst where the server resets keep-alive sockets (bare attempt fails,
  retried succeeds) and a POST whose socket is RST after server-side
  processing (hit counter stays 1 — no double submit).
2026-08-23 21:12:26 -07:00
KHALIDagara 1cc76fce3a fix(desktop): harden Hermes API transport 2026-08-23 21:12:26 -07:00
Teknium 95af45419f chore: map beplee contributor email for attribution gate 2026-08-23 21:12:20 -07:00
Teknium 03c3554fc2 fixup(curator): align #93002 test stubs with #93149 set_pinned bool contract
Combining both PRs for issue #92993: #93149 makes set_pinned() return a
bool and _cmd_pin/_cmd_unpin exit 1 on a no-op write; #93002's tests
stubbed set_pinned with a None-returning lambda, which the combined
_cmd_pin now reads as failure. The stub reports True (write landed) so
#93002's messaging assertions exercise the intended success path.
2026-08-23 21:12:20 -07:00
liuhao1024 ef882a5595 fix(curator): say what pin actually does on an unmanaged skill
`hermes curator pin` guarded on is_agent_created (a filesystem-shape
check), but the flag only matters when the skill carries the
curator-management marker: curated_report() walks marker-carrying skills
only, so auto-transitions never consider an unmanaged (pre-marker)
skill at all. Pinning one recorded the flag and then printed
"will bypass auto-transitions" — an effect that does not exist.

Keep the write (the flag becomes meaningful after `hermes curator
adopt`) and branch the message on is_curator_managed: unmanaged pins
now say the skill is unmanaged and point at adopt. Unpin gets the
symmetric wording.
2026-08-23 21:12:20 -07:00
beplee dd20c30dec fix(curator): check unpin result, guard status ghost rows, tighten test
Review feedback on #93149:
- _cmd_unpin now checks set_pinned's return (same false-success defect
  existed symmetrically on the unpin path)
- curated_report() pinned-visibility branch requires a local skill dir,
  so stale records for deleted dirs don't render as ghost rows
- test 2 asserts rc==0 unconditionally instead of vacuous-passing
- error message points to list-unmanaged (status doesn't render reasons)
2026-08-23 21:12:20 -07:00
beplee 7caa731e80 fix(curator): report pin failures instead of false success and surface pinned unmanaged skills
`hermes curator pin <skill>` printed success even when the underlying
write never landed. set_pinned() routes through _mutate() with
require_curation_eligible=True, which silently returns None for skills
that pass is_agent_created() but fail is_curation_eligible() — e.g. a
user-created skill named "plan", which PROTECTED_BUILTIN_SKILLS blocks
by name. The CLI then announced a pin that does not exist (#92993).

Also, a pin that DID land on an eligible-but-unmanaged skill (no
created_by marker) was invisible: curated_report() only iterated
list_agent_created_skill_names(), which requires the management marker,
so the skill showed up under 'unmanaged' with no trace of its pin.

- set_pinned() now returns bool write success; _cmd_pin() checks it,
  exits nonzero and explains the refusal when the write did not land
- curated_report() additionally includes curation-eligible skills whose
  usage record carries pinned=true, so their pins are visible in status

Fixes #92993
2026-08-23 21:12:20 -07:00
hermes-seaeye[bot] 6ed8bcee8d fmt(js): npm run fix on merge (#93503)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-24 03:39:06 +00:00
Brooklyn Nicholson 3c19644419 feat(desktop): HUD game-overlay mode
While a fullscreen app owns the screen, the HUD becomes an in-game chat frame:
the idle bar steps back to a glanceable opacity, and the transcript is held
open for as long as the game is there rather than fading on a timer — you look
back at a chat log during a lull, not while the text happens to be fresh.

Detection is a pure pass over the same front-to-back window enumeration
read_window_below uses (electron/hud-game-overlay.ts); main polls it while the
HUD is open and pushes changes to the renderer, which owns the treatment. Two
details the enumeration forced:

- Hysteresis. Entering needs the game to be what the user is actually looking
  at, so a windowed app on top vetoes it. Staying only needs the game to still
  exist: clicking the HUD to type de-foregrounds the game and floats every
  other open window above it, which otherwise dropped overlay mode at the
  moment the user engaged with it.
- The last state is replayed on did-finish-load. The watch pushes only on
  change and its first tick fires at window creation, before the renderer has
  mounted its listener, so a HUD opened over an already-fullscreen game
  consumed its only message and sat at 'no game' forever.

The band itself is reworked for living over someone else's window:

- Light-on-dark unconditionally. The theme's near-black body ink is unreadable
  over a dark game, and every attempt to gate the light ink on some
  condition — focus, then the game flag — produced a state where it evaluated
  false and the words went black on black. The sheet is a dark scrim in every
  theme so white is always right; anything that paints its own light surface
  (a clarify question, an approval card, a code block, a form control) opts
  back into theme ink by re-pointing the ink variable, matched on the fill it
  paints rather than the feature it belongs to.
- Your own lines are gold rather than bubbled. With no card the log otherwise
  reads as one voice; blue and purple are what most game UIs use for their own
  text, so they disappear into the background.
- The scrollback ramps out at the top instead of being cut off, masked on the
  scroller (the band is a static box — its rows overflow the thread viewport
  nested inside it, so a mask on the band ramps over empty space).
- The sheet is inset under the bar, so its square top corners no longer poke
  out past the bar's rounded ones.
2026-08-23 22:34:38 -05:00
Brooklyn Nicholson 433f518cad fix(desktop): Windows HUD paints opaque white
setBackgroundMaterial on a transparent window permanently kills per-pixel
alpha on Win11 — every transparent pixel composites as opaque white, so the
HUD showed a white slab instead of the desktop behind it. Verified against a
minimal repro on Electron 40.10.2: the break happens with ANY material value
including 'none', which is exactly what the idle HUD asks for, and neither
'auto' nor a follow-up setBackgroundColor('#00000000') restores it.

The DWM backdrop and window transparency are mutually exclusive, so the
Windows HUD keeps the CSS tint its sheet already paints and skips the native
frost. macOS is untouched: setVibrancy composites correctly.
2026-08-23 22:34:38 -05:00
Brooklyn Nicholson 1c75e05982 fix(desktop): resolve get-windows from the staged copy first
window-below asked node_modules for get-windows, whose lib/windows.js locates
its native binding through preGyp.find() — by HOST platform. When the tree was
installed on one OS and Electron is running on another (a WSL-hosted dev run
driving a win32 Electron), pre-gyp picks the host's slot, ignores the correct
binding sitting beside it, and upstream's fail-soft path returns no-op stubs.
Enumeration then reports 'unavailable' on a machine that answers perfectly
well, which silently disables read_window_below.

scripts/stage-native-deps.mjs already writes a staged lib/windows.js that
requires its binding directly, so prefer it and keep the bare import as the
fallback.
2026-08-23 22:34:38 -05:00
Teknium c584d15cdc feat(bots): typed failure reasons reach the sending agent on A2A calls (#93091)
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:

- Desktop relay drain forwards bot_relay.deliver's error.data.reason
  into bot_relay.reply (and prefers it for the attention badge over
  free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
  text, so the completion notification the sending agent receives is
  machine-branchable.

Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
2026-08-23 20:07:21 -07:00
Teknium 0c3a507535 fix(classifier): 429 quota walls route to billing across providers; reset signals stay rate-limited
Consolidates the 429-quota-classifier cluster on top of the merged #93419
Anthropic core. Three independent contributor findings salvaged into one
coherent change to the single 429 branch:

- Broaden the 429 usage-limit check from the narrow 'usage limit' string to
  the full _USAGE_LIMIT_PATTERNS ('quota', 'limit exceeded', 'key limit
  exceeded') and add _BILLING_PATTERNS detection on 429 ('insufficient
  credits' wrapped in a 429 instead of 402), guarded by a _RATE_LIMIT_PATTERNS
  exclusion so an explicit 'Rate limit exceeded' never promotes to
  non-retryable billing. (credit @Pluviobyte, #39441 — earliest submitter)
- Add 'resets in' to the transient signals: Codex's 'Weekly usage limit
  reached. Resets in 6hr 29min.' wrongly read as terminal billing because
  main only had 'reset in' (no substring match). (credit @LeonSGP43, #63021)
- Add 'reset after' / 'available in' / 'per minute' / 'per second' transient
  signals. (credit @jtstothard, #74785)

Supersedes #65633 (defective branch placement, no tests). The aux-client
path already covers these shapes (_is_payment_error catches weekly/quota
walls; _is_rate_limit_error treats 'resets in' as transient), so no change
there.

Tests: 6 new cases (generic quota wall, insufficient-credits 429, rate-limit
guard, Codex resets-in, extra transient phrases). Guard sabotage-verified.

Co-authored-by: Pluviobyte <Pluviobyte@users.noreply.github.com>
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
Co-authored-by: jtstothard <jtstothard@users.noreply.github.com>
2026-08-23 20:02:07 -07:00
fangliquanflq fe24525605 fix(state): reap only proven database holders 2026-08-23 20:01:41 -07:00
fangliquanflq 8f3a82f96a fix(state): recover FTS after orphan holder deferrals 2026-08-23 20:01:41 -07:00
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
fangliquanflq 030edf9774 fix(auth): canonicalize configured provider display names 2026-08-23 20:00:53 -07:00
fangliquanflq 3a7c094582 fix(auth): preserve configured provider compatibility 2026-08-23 20:00:53 -07:00
fangliquanflq c527b2c0a4 fix(auth): normalize configured provider pool keys 2026-08-23 20:00:53 -07:00
Adolanium 2912c36aa4 fix(gateway): stop multiplex allowlist leak and bot-relay python -c injection
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.

bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
2026-08-23 20:00:30 -07:00
Teknium 57649294be test(bots): turn-lock fake Proc gains stdout/stderr attrs
_run_delivery now captures output to drive the retry policy; the
turn-lock test's minimal _P fake predates that contract. Sibling-test
blast radius fix, no behavior change.
2026-08-23 20:00:18 -07:00
Teknium b274b346d8 feat(bots): retry session policy — resume transient turns, compress-and-resume on context overflow (#93091 item 5)
Maintainer ruling (2026-08-23): a retried bot turn never mints a fresh
session. retry_action() maps the #93091 item-1 reason enum to one of
resume / compress_then_resume / none:

- transient classes (runtime_offline, delivery_timeout, rate limit,
  server error) re-run the same Bot Chat session once;
- context_overflow also re-runs the same session — the retried turn
goes through the pre-API compaction pass in conversation_loop.py,
  which compacts the over-threshold transcript first (the one
  sanctioned context mutation); no fresh-session escape hatch exists;
- auth/quota/config/model classes never auto-retry.

Wired at both delivery surfaces (fix the class, not one site):
bot_relay.deliver (relay handler) and _run_delivery (local
message_agent runner). Failed deliveries now carry the classified
reason in the structured error payload (error.data.reason).

Sabotage-verified: with the retry blocks removed, 3 consumer tests
fail; with them present, 22/22 pass.
2026-08-23 20:00:18 -07:00
joaomarcos f5a9ba9ee6 perf(bluebubbles): move attachment reads off the event loop 2026-08-23 20:00:07 -07:00
Teknium fdff700dc2 chore: map e-macgregor contributor email 2026-08-23 19:55:17 -07:00
Teknium b03b8ac51d fix(dashboard): name the exact gate trigger in fail-closed refusals
When the bind is loopback and the only gate trigger is
dashboard.public_url, the startup refusal now says so explicitly and
gives both exits (configure a dashboard auth provider, or remove
dashboard.public_url if the proxy no longer exists). Prevents the
stale-public_url mystery-locked-dashboard upgrade trap.

Adds a truth-table regression suite for should_require_auth and the
fail-closed message shape.
2026-08-23 19:55:17 -07:00
e-macgregor d3df14a7e3 fix(dashboard): secure loopback public URL proxy mode 2026-08-23 19:55:17 -07:00
briandevans 608a56ed7f fix(state): stop rebuilding the whole FTS index on every open when the trigram tokenizer is missing
`_init_schema` decided whether the FTS triggers needed repair by comparing
the live trigger count against `len(_FTS_TRIGGERS)`, the full six-name set.
Three of those six are the `messages_fts_trigram_*` triggers, and they are
declared only inside `FTS_TRIGRAM_SQL` / `LEGACY_FTS_TRIGRAM_SQL`, whose
`CREATE VIRTUAL TABLE ... tokenize='trigram'` needs a tokenizer SQLite only
gained in 3.34.

On an older build `_ensure_fts_schema` soft-fails that DDL by design (via
`_is_trigram_unavailable_error`) and returns False, so those three triggers
can never be created. The count is therefore pinned at 3, `3 < 6` is
permanently true, and the repair path ran on every single `SessionDB` open,
forever, while holding the SQLite write lock. It never converged: every
`hermes` command, gateway start, dashboard request and cron tick paid a full
re-index of the message corpus. That is ordinary LTS territory — Ubuntu
20.04 ships 3.31, RHEL/CentOS 8 and Alibaba Cloud Linux ship 3.26, and
Hermes has no minimum-SQLite gate precisely because it is supposed to
degrade gracefully here.

The v23 repair also ends by clearing `fts_rebuild_high_water` and
`fts_rebuild_progress`, which is correct after a genuine full rebuild but
means an interrupted `hermes sessions optimize-storage` silently lost its
resume point on the next open, restarting the chunked backfill from zero
every time.

Fix: keep `_FTS_TRIGGERS` as the single source of truth and derive two
subsets from it, then measure each half against the DDL that can actually
create it. `_fts_trigger_count` takes an optional `names` sequence
(defaulting to the full set, so no caller changes), and both branches gate
on `base_triggers_missing or (trigram_enabled and trigram_triggers_missing)`.
The counts are still taken before the DDL runs so they describe the
pre-repair state, while `trigram_enabled` is only known afterwards — hence
the combination at the `if` rather than at the assignment.

Behaviour is unchanged wherever the tokenizer exists: a genuinely missing
trigram trigger on a capable host still triggers the rebuild. Only the
permanently unsatisfiable comparison changes.
2026-08-23 19:31:35 -07:00
Finn763 bf15b050b1 fix(telegram): watchdog silent long-poll death via last getUpdates progress (#92991) 2026-08-23 19:26:41 -07:00
Teknium 3f5d37568e fix: managed-runtime guard no longer trips on sdist/build copies in the workspace
The bare-which() scanner rglobs the repo root; a CI job that builds the
wheel leaves an sdist extraction (hermes_agent-<version>/) in the
workspace, and the scanner re-found every already-exempted call site
under that versioned prefix — which can never match an _ALLOWED key —
failing the guard on untouched code (flaked PR #93420's Python-tests
job). _source_files now skips build/, dist/, *.egg-info, and any
top-level dir carrying PKG-INFO.

A/B: planted a fake hermes_agent-9.9.9/ sdist with a which('node')
site — old scanner 1 failed, fixed scanner 7 passed, clean tree
unchanged.
2026-08-23 19:15:28 -07:00
Teknium 7526bd39a8 feat: every subagent's prompt embeds the workspace's project context files
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.

The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.

Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 23fb949f2c feat: /review briefing embeds the workspace's project context files
load_workspace_context() resolves the parent's workspace via the same
_resolve_workspace_hint used for child prompts (explicit sources only —
TERMINAL_CWD / agent cwd hints, never a bare getcwd fallback, so the
#64590 install-tree-leak guard concern doesn't apply) and runs it
through agent.prompt_builder.build_context_files_prompt — the exact
discovery/priority/cap logic the main system prompt uses (.hermes.md >
AGENTS.md chain > CLAUDE.md > .cursorrules; SOUL.md skipped). The
result is embedded in the reviewer briefing as binding review
standards. Subagents are built with skip_context_files=True, so without
this the reviewer judged repo work without the repo's own conventions.

5 new tests incl. real-filesystem AGENTS.md discovery through the real
loader. Docs updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 22381edc11 feat: /review briefing carries the parent's loaded skills
The reviewer subagent now inherits the primary agent's working skill
context: collect_parent_loaded_skills() gathers launch-preloaded skills
(from the activation notes in ephemeral_system_prompt) and mid-session
skill_view loads (from assistant tool_calls in history), deduped and
capped at 8, and the briefing instructs the reviewer to skill_view each
and treat their conventions as binding for the assessment.

Reference-file reads (file_path=...) don't count as loads; full-skill
injection was rejected as too costly (a single dev skill can be 40KB+).

Docs: delegation.md /review flow updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium ea25bf204d chore: map contributor email for cxxCoolStar 2026-08-23 19:02:30 -07:00
Teknium 580060ffd8 fix: reuse first-observed sequence when announced items land via output_item.done
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.

Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
2026-08-23 19:02:30 -07:00
cxxCoolStar 4f3ae189a3 fix(agent): harden pending Responses tool call settlement 2026-08-23 19:02:30 -07:00
cxxCoolStar 720344cfba fix(codex): settle pending Responses tool calls when output_item.done is omitted
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
2026-08-23 19:02:30 -07:00
Teknium 081cdd9911 fix(terminal): subagents no longer hijack the tty with an interactive sudo prompt
delegate_task children run on worker threads of the parent process and
inherit the process-wide HERMES_INTERACTIVE=1 the CLI sets at startup.
_transform_sudo_command's interactive gate therefore fired inside
children with no sudo callback registered, falling through to the raw
/dev/tty password prompt: a password box printed mid-TUI from a
background thread, parallel children racing for the tty, and each child
blocked for the full 45s timeout.

Gate the prompt (and the sibling 'you will be prompted again' message
after an auth failure) on agent.delegation_context.is_delegated_child_context(),
the ContextVar set around every child run and propagated through
contextvars.copy_context onto the executor thread. Children now behave
as headless for sudo: configured SUDO_PASSWORD, the session cache, and
the NOPASSWD probe still work; otherwise the command fails gracefully
with a subagent-specific tip.

A/B verified: 3 regression tests fail on merge-base, 7/7 pass at head.
2026-08-23 19:00:47 -07:00
Teknium 74e483d326 chore: add contributor email mapping for jackijianxa 2026-08-23 19:00:36 -07:00
Teknium 9d0727d49b fix(state): single fail-closed cross-process authority for all full FTS rebuilds
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:

- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
  30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
  _rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()

Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.

Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
2026-08-23 19:00:36 -07:00
jackijianxa 0f33c207e6 fix(state): serialize cross-process FTS rebuild with file lock
When two Hermes processes (e.g. gateway + serve) detect FTS corruption
simultaneously, both run rebuild_fts() on the same database file in
parallel. rebuild_fts() only holds an in-instance threading lock, so the
concurrent rebuilds collide on write and structurally corrupt the
database ('file is not a database' / 'database disk image is malformed').

This happened twice in production (2026-08-15 and 2026-08-23), each time
requiring a full page-level salvage of state.db: sessions b-tree
clobbered, 5507 messages recovered row-by-row.

Fix: acquire an exclusive fcntl.flock on <db_path>.fts_rebuild.lock
before rebuilding, with a bounded 30s wait. The SQLite writer lock
remains the final backstop. POSIX-only; no-op elsewhere.
2026-08-23 19:00:36 -07:00
Teknium ca226b5e06 chore: map contributor emails for ring-2 salvage (A2chitect, c-pompa) 2026-08-23 18:58:40 -07:00
Christian Pompa d6bc3f2bca fix(tui_gateway): enable TCP keepalive on websocket sockets (dead-peer detection)
Without SO_KEEPALIVE a silently-dropped client (SSH tunnel reset, laptop
sleep, NAT timeout) leaves the TCP leg half-open forever: receive_text()
blocks indefinitely and the disconnect teardown (detach, orphan reap,
resume replay) never runs. The server then leaks the session and never
reclaims its orphans.

_disable_nagle already reaches the raw socket, so enable keepalive there:
SO_KEEPALIVE on, plus TCP_KEEPIDLE=30s / TCP_KEEPINTVL=10s /
TCP_KEEPCNT=3 on Linux and TCP_KEEPALIVE=30s on macOS. A dead peer is now
detected in ~60s instead of never. Best-effort like the Nagle tuning —
any failure to reach the socket is logged at debug and skipped.

Tests: new tests/tui_gateway/test_ws_keepalive.py fakes the socket and
pins SO_KEEPALIVE + the platform-specific idle tuning, plus the
no-transport no-raise path. tests/tui_gateway: 336 passed.
2026-08-23 18:58:40 -07:00
Teknium a7977771a6 fix(tui-gateway): revalidate transport ownership before sentinel-parking on WS disconnect
Reimplements the concept from #77129 on the current structure (viewer
rebinding from #83716 and _client_gone_interrupt_requested clearing are
preserved).

_close_sessions_for_transport snapshots owned sessions under
_sessions_lock, then wrote session['transport'] = _detached_ws_transport
WITHOUT re-checking that the session still pointed at the disconnecting
transport. A session.resume that rebinds the session to a new live
transport between the snapshot and the stomp got knocked back onto the
drop sentinel with an orphan-reap Timer armed against a client that is
attached right now.

The park now happens under _sessions_lock and first revalidates
ownership: if the session already moved to a different live transport,
the disconnect has nothing to tear down — skip the sentinel park AND the
reap scheduling. Regression test simulates the rebind landing between
snapshot and stomp.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 18:58:40 -07:00
Teknium 47d6ce78a2 fix(tui): make startup_orphan_reap recoverable and move its config onto dashboard.*
Follow-up to the #65422 salvage:

- startup_orphan_reap joins _RECOVERABLE_END_REASONS (kept distinct from
  ws_orphan_reap for forensics): every recovery fence
  (find_latest_gateway_session_for_peer, unarchive_recoverable_session,
  promote_to_session_reset) now treats a startup-swept row as an
  accidental end, so a sweep never makes a session unresumable.
- Config key moves from sessions.orphan_reaper to
  dashboard.startup_orphan_sweep in DEFAULT_CONFIG, next to its siblings
  ws_ping_interval / ws_ping_timeout / ws_orphan_reap_grace_s; the raw
  loader in tui_gateway.server reads the new key (fail-open on missing).
  cli-config.yaml.example and website/docs/user-guide/configuration.md
  follow the dashboard.* documentation pattern.
- New regression test: a stranded 'active' row (ended_at NULL, no live
  runtime) is swept AND still recoverable via peer-keyed lookup and fully
  revivable via reopen_session afterward.
2026-08-23 18:58:40 -07:00
halaprix d3e4b50e68 fix(tui): sweep orphaned tui/desktop/subagent session rows at gateway startup
Close session rows left ended_at IS NULL when the in-process websocket
orphan timer dies with the process (#65194). Dual-clock staleness
(started_at AND newest message), desktop included, live in-memory
sessions excluded, scheduled once from both entry.main and the WS
sidecar so desktop/dashboard boots also run the sweep.
2026-08-23 18:58:40 -07:00
A2chitect c305839442 fix(tui): log 4001 session-not-found rejections for diagnosability
Messages sent into a session whose in-memory runtime was detached on WS
disconnect and orphan-reaped vanished silently: _sess_nowait returned
4001 with no log line, so 'request arrived and was rejected' was
indistinguishable from 'request never arrived' in a 'message vanished'
report. Log a WARNING with the session id and request id on every
session-scoped RPC rejected against an unknown runtime id.

Adds a regression test asserting the 4001 response and the warning.

Closes #90428
2026-08-23 18:58:40 -07:00
pierrenode 525597c9c3 fix(bot-mode): fail closed on transient group-session resume failures
ensureGroupChatSession's resume loop caught ANY session.resume error
(stored sid, then title lookup) identically and fell through to
session.create — the same bug findExistingCanonicalChat was fixed for
hours earlier (87b645f52c) in the same file: a transient failure (the
backend still warming up after a restart, a network blip on a
cross-connection lookup, an oversized-resume refusal) read as "no
session, mint a new one". That forks the member's real session AND
silently overwrites room.sessions[key], making the original
unreachable from the room. ensureGroupChatSession is actually more
exposed than the 1:1 case: it runs every group turn
(runGroupChatMemberTurn), with two independent swallow points.

Distinguish "genuinely doesn't exist" from "transient failure" the
same way the gateway itself does: session.resume's own handler
(tui_gateway/methods_session.py) returns JSON-RPC code 4007 only when
the target truly isn't found; every other failure (including 4130,
"session too large to resume" — a session that DOES exist) now
surfaces instead of being silently swallowed. The existing outer
try/catch at the call site already treats a thrown error as "this
member passes the round" (recordGroupActivity kind: 'failed'), so
nothing new needs to catch it — a transient hiccup now costs one
skipped round instead of a permanent fork.
2026-08-23 18:58:40 -07:00
kshitijk4poor 80b202f53a harden(adoption): review findings — exact-id donors only, divergence guard, honest donor_retired
Review batch (3 reviewers) on the final diff surfaced:
- H1: title-based donor matching could adopt AND non-recoverably retire
  an UNRELATED default-store conversation (bot titles collide by design;
  get_session_by_title has no archived filter/ordering). Donor probe is
  now exact-id only — the stranded repro always has the id.
- H2: re-adoption after a partial run could retire a donor that had
  accumulated NEWER messages than the profile copy (skip-based
  idempotency never merges). New divergence guard compares message
  counts and refuses retirement when the donor is ahead (still adopts).
- M1: donor_retired reported True even when every retirement step
  failed under suppress. Now per-segment tracked + warn-logged;
  True only when all applied.
- M3: adopted=False (e.g. import validation limits) was silent — now
  warn-logged with import errors.
- M4: archived donors are never re-adopted (no cross-profile cloning).
- Dead 'from pathlib import Path' dropped; contextlib no longer needed.

5 new red-first-verified regressions (title-collision immunity,
archived-donor immunity, non-vacuous owns_db gating with a real donor
seeded, divergent-donor retirement refusal, donor_retired truthfulness).
tests/tui_gateway: 578 passed. ruff clean.
2026-08-23 18:58:40 -07:00
kshitijk4poor 26a4f89ada fix(gateway): adopt stranded bot sessions from the default store on profile resume
Pre-#93296, the desktop routed session RPCs by the focused tile, so a
profile bot's turns executed on the default backend and its canonical
session accumulated in the DEFAULT profile's state.db. Post-fix, the
profile backend correctly receives the resume — but its store has never
seen the session, so the same chat 4001s forever (unreachable instead
of misrouted). Live repro: Teknium's Developer bot, session c93770.

- hermes_state_portability: SessionDB.adopt_session_lineage_from() —
  composes the existing export_session_lineage()/import_sessions()
  primitives; donor rows are archived (never deleted) with
  end_reason=adopted_by_profile, which is deliberately NOT in
  RECOVERABLE_END_REASONS so canonical-lookup resurrection cannot undo
  an adoption. Idempotent (already-present ids skip).
- tui_gateway/methods_session: profile-scoped session.resume falls back
  to adoption from the default store right before the 4007; ids unknown
  to BOTH stores still 4007 exactly as before, and launch-profile
  resumes never consult the fallback.
- tests: 10 new (7 unit on the primitive incl. compression-lineage
  unit adoption + non-resurrectable archive; 3 handler-level through
  server.handle_request incl. the live repro shape); db-ownership
  leak test taught that the shared launch handle probe is by design.

Follow-up to #93296/#93311; part of #93091.
2026-08-23 18:58:40 -07:00