Commit Graph

28271 Commits

Author SHA1 Message Date
Brooklyn Nicholson d0e41cda69 fix(gateway): let config.set write the desktop's Appearance switches
config.set matches an explicit key list and answers 4002 for anything
else, so a renderer mirroring an unlisted key wrote nothing at all. The
reactions toggle shipped that way: every write was rejected into a
swallowed .catch(), and react_to_message stayed dark no matter what the
user picked.

Adds the display booleans as a recognized group so the toggle reaches
the config of whichever gateway the app is actually talking to, which is
the only place a check_fn can read it.
2026-08-31 21:11:11 -05:00
hermes-seaeye[bot] de590cc3cb fmt(js): npm run fix on merge (#99917)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 02:09:50 +00:00
Brooklyn Nicholson 0b247df94f fix(desktop): dock the sidebar rails down to 640px
The rails left the grid for the hover-reveal overlay at 768px, which cost
the docked sidebar on any half-screen split (1512 -> 756, 1440 -> 720).

Derive the collapse point from the layout instead: a rail costs 237px
(SIDEBAR_DEFAULT_WIDTH) and the chat beside it wants roughly the 420px a
popped-out session window enforces on itself. Express it as a dock floor
rather than a collapse ceiling so an exactly-640px window still docks.
2026-08-31 21:08:57 -05:00
Brooklyn Nicholson b0bcb3640c fix(desktop): drop redundant onReorderSessions check after sessionsDraggable 2026-08-31 21:03:25 -05:00
Brooklyn Nicholson 9052ce76c4 feat(desktop): collapse Yesterday / Last week groups in the sessions sidebar
Date and status labels were separators only. Clicking one now hides the
sessions under it, same as project and profile rows, and Collapse all
covers those buckets too.
2026-08-31 21:03:25 -05:00
Brooklyn Nicholson 1b47277067 test(tui_gateway): cover stop-vs-marker recovery races
Pin that session.interrupt retires the original and rotated marker keys
on ACK, and that a Stop arriving before the disk write cannot leave a
resume-able marker behind.

Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
2026-08-31 20:50:39 -05:00
Brooklyn Nicholson 5c3bfb6da8 fix(tui_gateway): retire recovery marker on local stop
A confirmed Desktop Stop left the crash-recovery marker on disk until the
run thread finished. If the backend exited in that window, resume treated
the leftover as a crash and auto-continued the turn the user had stopped.

Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
2026-08-31 20:50:39 -05:00
Gille e07172319d fix(desktop): make /stop interrupt active turns
Preserve the existing background-process cleanup after interrupting the targeted Desktop session. This salvages the narrow /stop behavior from the broader, conflicting PR #45030 on the current slash-command architecture.

Co-authored-by: AlvaroBiano <alvarobiano@users.noreply.github.com>
2026-08-31 20:50:10 -05:00
Brooklyn Nicholson adb23c13cb fix(bot-mode): keep Bot Chat resume on a proven compression tip
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-08-31 20:47:53 -05:00
Gille 4abf5c0790 test(bot-mode): scope title lookup regressions 2026-08-31 20:47:53 -05:00
Gille bb28056efd fix(bot-mode): keep canonical lookup on compression lineage 2026-08-31 20:47:53 -05:00
Ben Barclay 56916841b5 refactor(dashboard-auth): replace PKCE cookie payload with base64url(JSON) codec (#99210)
The PKCE cookie's payload has now needed three serialization fixes at
the same spot: the original flat 'k=v;k=v' string tripped http.cookies'
\073 quoted form (dropped whole by strict cookie parsers like Go's
net/http — #83832 field case), and #99176 URL-encoded the whole flat
payload to stay inside the RFC 6265 cookie-octet set. The stacked
layers (single-encoded next=, ';' joins, whole-payload encoding, legacy
discriminator) were the recurring defect source.

Kill the bug class instead of patching it again: the payload is a dict
end-to-end and goes on the wire as base64url(JSON) — the urlsafe
alphabet is a strict subset of cookie-octets, and JSON framing means no
segment value can ever collide with a delimiter. parse_pkce_payload
keeps a three-rung compatibility ladder (base64url(JSON) -> oldest flat
form split-as-is -> #99176 unquote-then-split) for in-flight cookies
during a rolling upgrade (10-minute TTL); a new cookie hitting an old
server fails the OAuth state check and the user just retries.

The 'next' segment is stored as its plain validated path — no extra
encoding layer, so the post-login redirect Location is byte-for-byte
the original target.

Refs #99176, #84065.
2026-09-01 11:47:38 +10:00
xxxigm bcecd675f7 fix(desktop): unwrap Models-page code-skew 503 and recycle the owned backend (#97046)
Show a Restart backend action that kills the SSH serve before the local child
so reconnect cannot reuse a stale lockfile, instead of dumping raw IPC JSON.
2026-08-31 20:43:36 -05:00
xxxigm c99a3919a8 fix(dashboard): drop systemd-only advice from the model-picker code-skew 503 (#97046)
Desktop-owned serve and macOS hosts have no hermes-dashboard unit, so the
restart hint now follows HERMES_SERVE_HEADLESS instead of hardcoding systemctl.
2026-08-31 20:43:36 -05:00
Brooklyn Nicholson 42c415820c test(desktop): assert regenerate rejection restores the full history
Drive the real reloadFromMessage catch, not a copied merge object.
Same coverage on the session-tile sibling.

Co-authored-by: ygd58 <buraysandro9@gmail.com>
2026-08-31 20:43:06 -05:00
Brooklyn Nicholson f615f3796d fix(desktop): restore the transcript when regenerate is rejected
reloadFromMessage hid/truncated messages for the optimistic UI, then on
a failed submit (provider error, compressed-away turn) reset only the
busy flags. restoreToMessage already rolled messages back; do the same
on the primary chat and the session-tile path.

Co-authored-by: ygd58 <buraysandro9@gmail.com>
2026-08-31 20:43:06 -05:00
Brooklyn Nicholson 5c6dbe22c3 fix(compression): take reasoning_content when the summarizer leaves content empty
Local and thinking backends (DeepSeek, Qwen, Kimi) often return a usable
summary in reasoning fields. Treat that as the summary instead of burning
another 100s+ empty-content retry. Leave the wire max_tokens omit intact.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
Brooklyn Nicholson 02458e67ad fix(auxiliary): accept dict and object messages in extract_content_or_reasoning
Compression and some OpenAI-compatible proxies hand us a dict-shaped
response or a bare message, not a ChatCompletion. Reuse the existing
helper instead of a second extractor, and bound an optional reasoning
fallback so a chain-of-thought dump cannot become the summary.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
Brooklyn Nicholson 6eddeca6ef fix(desktop): stop hidden composers from stealing the caret
Keep-alive tabs remount their composer on transcript/status backstops
and were calling focus() while the user typed in the front tab. Gate
autofocus on pane visibility, and refuse to steal the caret from
another visible composer. A hidden tab that still holds DOM focus
does not block the pane the user just switched to.

Co-authored-by: Dan Bennett <dan@danbennett.me>
Co-authored-by: mor44-AI <andrefmontemor@gmail.com>
Co-authored-by: d4rk pr10r <darkpriorlabs@gmail.com>
2026-08-31 20:41:08 -05:00
Brooklyn Nicholson 60ba6024b5 fix(desktop): isolate keep-alive panes from the visible transcript
Hidden session panes published stick-to-bottom into window-global atoms
and every mounted list subscribed to the same jump broadcast, so a buried
tab yanked the reader and flashed the composer. Only the visible pane
may publish, scroll requests are keyed by session, and a run start or
same-session refresh leaves a scrolled-up reader where they were.

Co-authored-by: gamewocao <gamewocao@users.noreply.github.com>
Co-authored-by: mor44-AI <andrefmontemor@gmail.com>
Co-authored-by: d4rk pr10r <darkpriorlabs@gmail.com>
Co-authored-by: Jackal991 <lawrence@hydra-flow.co.uk>
2026-08-31 20:41:08 -05:00
Brooklyn Nicholson d72ce5e434 fix(desktop): rebind the visible chat after a reaped runtime
session.reclaimed now heals through markRuntimeGone before dropping
cache. A 4001 on the visible session's dispatcher request asks for a
durable resume. prompt.submit uses the window dispatcher so the turn
lease survives the ACK, and a primary route no longer registers a
phantom turn lease.

Co-authored-by: Lester Liang <153183032+lesterlxt@users.noreply.github.com>
Co-authored-by: wz-heng <68931789+wz-heng@users.noreply.github.com>
2026-08-31 20:40:31 -05:00
Brooklyn Nicholson 36bea50139 fix(desktop): latch approval and goal polls off a dead runtime
process.list already stopped hammering a reaped id. approval.pending and
goal status did not, and session.info heartbeats republished an equivalent
state object so every tab rerendered. Share the gone-latch from
runtime-gone and keep heartbeat identity when nothing changed.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: Dolverin <5910064+Dolverin@users.noreply.github.com>
2026-08-31 20:40:31 -05:00
chelsealong cbe007a559 fix(desktop): stop status-stack remounts from re-arming a dead-runtime poll storm (#98434)
A boot-restored chat can stay bound to a dead runtime id and remount its
composer status stack repeatedly with no genuine rebind ever occurring.
The stack's mount effect cleared the gone-polling latch on every mount, so
each remount re-armed process.list + slash.exec('goal status') against
the same phantom id forever, churning the composer every ~5s.

Real rebinds already reset the latch at the runtime-mint seams
(use-gateway-boot.ts, store/gateway.ts). Drop the redundant per-mount
reset so the latch actually holds across a remount.
2026-08-31 20:40:31 -05:00
Brooklyn Nicholson 308454b372 fix(desktop): route model catalogs and picker reseeds through the focused owner
requestModelOptions now sends the owner profile on the RPC, forced reseeds call getGlobalModelInfo(profile), and picker/cache keys include the registry connection so a tile cannot fall back through the ambient socket.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-08-31 20:40:09 -05:00
Brooklyn Nicholson f08666d42a fix(desktop): persist composer model/provider per remote connection and profile
Sticky composer keys were global, so a provider picked on one remote profile could ride into session.create on another. Persistence now follows an explicit (connectionId, profile) owner published before the active-profile reseed; unresolved legacy owners fail closed.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-08-31 20:40:09 -05:00
Brooklyn Nicholson 0e7eebc266 fix(desktop): scope remote model catalog and primary-label REST to the focused profile
A shared dashboard's launch HERMES_HOME is not the selected profile. model.options now runs under @_profile_scoped, and global-remote REST keeps ?profile= even for the primary label.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-08-31 20:40:09 -05:00
Brooklyn Nicholson 8b28bdceb5 fix(desktop): stop a stale composer model pinning every new chat
Settings → Model while a chat is open flips the composer source to
'default' but leaves the live session's model painted. Sending that
value on session.create pinned every new chat and skipped model.default.

Only a manual composer pick is a per-session override.

Co-authored-by: Tharanee <tharanee@tharanee.net>
2026-08-31 20:39:06 -05:00
Gille f98f5e74e0 fix(desktop): preserve terminal startup output 2026-08-31 19:43:31 -05:00
hermes-seaeye[bot] 66b844c967 fmt(js): npm run fix on merge (#99871)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 00:41:26 +00:00
Brooklyn Nicholson e600507a8f fix(desktop): drop the stray paren that broke the /btw event test typecheck 2026-08-31 19:36:03 -05:00
Brooklyn Nicholson 38e4de48d8 test(desktop): cover /btw prompt.btw dispatch and btw.complete
Pin the RPC route (not slash.exec), bare-/btw usage, older-gateway
fallback, and originating-session answer rendering.

Co-authored-by: SsSs <w-kwan@hotmail.com>
Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-31 19:36:03 -05:00
Brooklyn Nicholson 5994b833b0 fix(desktop): route /btw through prompt.btw so the answer reaches the chat
Desktop sent /btw through the slash worker, which printed the answer after
process_command returned, so only the acknowledgement ever showed. Use the
TUI's prompt.btw RPC and persist btw.complete on the originating session.

Co-authored-by: SsSs <w-kwan@hotmail.com>
Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-31 19:36:03 -05:00
Brooklyn Nicholson a2907a8bcd fix(desktop): skip cold process probes for dead backend-ownership PIDs
Windows Get-Process and macOS ps exit 1 on a missing PID, so reapOrphans
kept stale records and the next launch paid another 2-8s spawn each.
Throw ESRCH from the existing isPidAlive helper before any shell-out.

Closes #92875

Co-authored-by: Jackal991 <139240222+Jackal991@users.noreply.github.com>
Co-authored-by: jonotonfoto <126111813+jonotonfoto@users.noreply.github.com>
Co-authored-by: foras910521-lab <268267187+foras910521-lab@users.noreply.github.com>
2026-08-31 19:32:14 -05:00
hermes-seaeye[bot] e22f8a7fbd fmt(js): npm run fix on merge (#99862)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 00:28:40 +00:00
Gille 9dffcc431a fix(desktop): keep project previews visible on drill-in 2026-08-31 19:21:16 -05:00
briandevans cebe2168ab fix(desktop): resolve the clarify card by zone, not document order
visibleClarifyCard() bottomed out in queryVisible(), which only drops
[data-pane-hidden] and then returns the first remaining DOM match. That is
enough to tell a foreground tab from a background one, but a split layout
has two chat surfaces on screen at once: both cards survive the hidden-pane
filter, so the winner is decided by document order. The earlier zone owns
Enter and 1..N permanently and the other visible card can never receive its
own shortcut — worse than the mount-order behaviour it replaced, because
mount order at least changed as panes came and went.

Break the tie on the ladder the app already has instead of inventing a
second notion of which surface is "the" one: tree/store's tabTargetGroup
walks hovered zone, then focused zone, and every tab verb (Cmd+1..9,
Ctrl+Tab, the Cmd+W family) targets through it; composerTargetInHoveredZone
(#74447) mirrors it for the model hotkey. visibleClarifyCard now does the
same, matching each visible card's closest [data-tree-group] against those
rungs. The single-card path short-circuits before any store read, so the
common case costs nothing extra.

The final fallback stays document order rather than null on purpose. When
neither rung names a zone holding a card — pointer off every zone, nothing
interacted with yet — returning null would leave Enter doing nothing at all,
a strictly worse regression than answering the first visible card.

Regression coverage pins both halves: the resolver unit tests assert the
active zone selects either card (not just the later one), that hover
overrides focus, that a hovered zone with no card falls through to the
focused one, and that neither rung resolving still yields a card; the
ClarifyTool integration tests render two visible cards in two zones and
assert exactly one clarify.respond fires, carrying whichever zone is
focused. All four order-sensitive assertions fail against the previous
resolver.
2026-08-31 19:20:32 -05:00
briandevans a3662bb32c fix(desktop): answer the visible clarify card, not the first-mounted one
A pending clarify card binds Enter / 1-9 / A-Z / arrows on `window`, and
inactive tabs stay MOUNTED, so every parked clarify keeps a live listener.
Nothing in the handler asked whether the card was on screen: the first
listener registered won, called preventDefault(), and the rest bailed on
defaultPrevented. Registration order is mount order, not visibility.

With two chats waiting on a question, answering the one in front of you
sent `clarify.respond` for a background session's request instead —
silently answering a question the user never saw and resuming that turn.

The invariant is already stated one file over, in composer-focus-keys.ts:
"a clarify card waiting in a background thread must not take the
foreground composer's letter keys". `clarifyCardOwnsKey` honours it via
queryVisible(); the card's own listener did not. That asymmetry is the
bug — the key-ownership resolver and the key handler disagreed about
which card is live.

Export that lookup as `visibleClarifyCard()` and have both sides use it,
so they cannot drift apart again. The card bails unless it IS the visible
card, which also leaves the keystroke unprevented for the composer when
the only pending card is hidden.

Scoped by visible-card identity rather than `usePaneVisible()`: split
zones each render their own active pane, so two cards can be visible at
once and a per-pane visible flag would not disambiguate them.
2026-08-31 19:20:32 -05:00
kshitijk4poor b20cc5f787 docs(agent): explain intentional preflight vs in-loop message divergence
The preflight handler surfaces the boundary exception's per-request text
(token count, 'provider call was not sent') rather than the in-loop
_COMPRESSION_TIMEOUT_FINAL_RESPONSE constant, which describes a
different state (compression ran and could not reduce). Document the
divergence so it is not 'fixed' into a single message later.
2026-09-01 03:48:40 +05:30
kshitijk4poor cc0931d235 test(agent): loosen brittle error/final_response equality to substring
result['error'] and result['final_response'] are independently settable
keys that only coincidentally share _COMPRESSION_TIMEOUT_FINAL_RESPONSE
today; assert the actionable substring instead so a benign prefix or
rewording does not break the terminal-contract test.
2026-09-01 03:48:40 +05:30
kshitijk4poor 1e8f6a0491 fix(agent): surface preflight compression timeout as typed result, not generic error
When the turn-start fail-closed boundary (#98424) raises
PreflightCompressionTimedOut, the exception escaped run_conversation to
the surfaces' generic exception handlers. The gateway deliberately never
exposes raw exception text, so users saw 'Sorry, I encountered an
unexpected error... Try again or use /reset' instead of the boundary's
actionable guidance, and the compression_exhausted clean-session
recovery contract (#9893/#35809) never engaged.

Catch it at the build_turn_context callsite and convert it into the
same typed recovery dict the in-loop timeout consumers return
(salvaged #98741 / PR #99710): failed=True, partial=True,
compression_exhausted=True, turn_exit_reason=context_compression_timeout,
with the actionable message in final_response and error.

Regression test proves the exception no longer escapes and the typed
contract fields survive to the caller (mutation-checked: test fails on
main without the handler).
2026-09-01 03:48:40 +05:30
kshitijk4poor 1eaa2b0a9b chore: map everest.kill1@gmail.com -> anhtahaylove in contributor emails
PR #98424 merged with commits authored as everest.kill1@gmail.com
(anhtahaylove) but the email was never added to contributors/emails/,
so the next release's contributor_audit would fail on the unmapped
address. One-file mapping, same mechanism as every other entry.
2026-09-01 03:48:40 +05:30
Teknium 02ecc4be27 fix(gateway): gate stream commentary reply-to on platforms that need it
Discord interim commentary was replying to the user's trigger message on
every update (66fa6e41c4 added reply_to unconditionally for Buzz). Discord
uses native thread_id; only Buzz/Slack/Mattermost/Feishu need reply-to
anchoring for threading.

- Stream consumer _send_commentary checks adapter.platform.value
- Passes reply_to only for buzz/slack/mattermost/feishu
- Discord/Telegram commentary posts flat or threads via metadata
2026-08-31 15:03:05 -07:00
Teknium cf21e28d39 fix(gateway): _routing_db tolerates bare test instances (object.__new__)
Bare SessionStore instances built without __init__ lack _db_pinned,
_routing_home, and the handle cache behind the _db property. Restore
main's old getattr contract for them: report no DB and fall through to
the sessions.json path instead of raising AttributeError.
2026-08-31 14:54:18 -07:00
Teknium 5a407a0c52 test(gateway): pin routing-index harness to the tmp HERMES_HOME store
The #66887 fix pins the routing index to HERMES_HOME state.db; the
fast-path harness still read entries through the ambient store. Point
get_hermes_home at the test tmp so both are the same file, matching
the new single-store contract.
2026-08-31 14:54:18 -07:00
caya8205-2 8d74cb52da fix(gateway): give the routing index one store instead of the ambient one
Second half of #66887. _entries is a single flat dict holding every
profile's keys, so the index it persists to has to be a single file — but it
was read and written through _db, which resolves whichever profile scope is
active. A whole-index rewrite during one profile's turn copied every other
profile's routing rows into that profile's store, and startup, which runs
unscoped, then loaded a different copy than the last writer produced.

That is why the startup recovery pass never sees a secondary profile's crash
marker, which is the half this issue's title names. mark_turn_active()
persists through the single-entry fast path (state.db only, no sessions.json
mirror), so a marker written during a profile's turn landed in that
profile's store and _recover_unclean_sessions(), running with no scope, read
a store that had never heard of it. The turn was silently never promoted to
resume_pending.

Capture the gateway's own home at construction — the store is built at
startup before any profile scope exists — and route the index through it:
_ensure_loaded_locked, _reconcile_recovered_routing_locked,
_persist_routing_data and _save_entry now use _routing_db. A pinned handle
still wins, so suites that install a fake or disable the DB are unaffected.

_prune_stale_sessions_locked is the mixed case and is split accordingly: it
now asks _db_for_key(key) whether each session ended, because that is a
per-session question, while the index write stays on the single store. One
ambient handle previously answered it for every profile at once, which could
prune a live secondary-profile route on the strength of the root store's
copy of that session.

Regression as requested on the issue: mark a turn active under a secondary
profile's scope, then build a fresh store with no scope and run
recover_interrupted_turns(). It promotes exactly one turn to resume_pending
here and promotes zero against the previous behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 a60e04a32c fix(gateway): prove the compression child's owner before writing to it
Review P1 #2. _append_to_transcript_serialized() writes the compression
continuation to child_id BEFORE publishing either _transcript_reroutes or
the _entries update — that ordering is load-bearing for backlog order, so it
must not move. At that moment nothing in the routing index points at the
child, so _db_for_session_id(child_id) missed its scan and fell through to
_db_for_key(None), i.e. the ambient store. The fail-closed guard did not fire
because root is a live handle.

The row therefore targeted root rather than the already-proven parent owner.
With no child row there the append is rejected by the FOREIGN KEY constraint,
the pending queue never drains and the reroute cannot advance; against a
split-brain root the message would instead be written cross-profile.

Record ownership before the mutation instead of moving the publication: a
private _session_owner_hints map carries session_id -> owning key for ids
whose owner is proven but not yet published, consulted by the new
_owner_key_for_session_id() after the index scan misses, and dropped as soon
as routing publishes. Signatures are unchanged, so the existing suites that
stub _append_transcript_message keep working untouched; the map is read
through getattr for stores built via object.__new__.

The regression is physical rather than mocked: an ended compression parent
and a live child that exist only in profiles/fitness/state.db, no active
profile scope, append to the parent, then assert all four effects — the row
lands on the child in the profile store, the pending queue drains, the
reroute and the routing entry advance, and root state.db stays untouched.
Without the hint it fails exactly as the review predicted, on
"FOREIGN KEY constraint failed" against root.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 49fb58523f fix(gateway): fail closed when a named profile has no resolvable home
Review P1. _profile_home_for_key() returned the same None for three
different states — multiplexing off / legacy agent:main namespace, a named
profile whose directory does not exist yet, and a resolution error — and
_db_for_key() collapsed all of them to the ambient store.

That recreated the very split this change removes. The enrollment bridge
provisions profiles/<name>/ at runtime, so a key such as
agent:fitness:telegram:dm:1 can legitimately be seen first: the first lookup
landed in root state.db, and the next one, after provisioning, in
profiles/fitness/state.db. One qualified session identity, two physical
stores. The resolver-exception path fell open the same way.

Ownership is now tri-state:
  - no named owner            -> ambient DB (single-profile behavior intact)
  - named owner + home        -> that profile's DB
  - named owner, unresolvable -> None, and a warning; never root

Callers already treat a missing DB as "skip the mutation", which is the
defer-don't-misroute behavior wanted here. _append_transcript_message is the
one path reached with an id the entry-point guard did not check (the
compression-child id), so it now raises explicitly and lets the caller's
retry queue hold the row instead of relying on an AttributeError.

Tests exercise the effect boundary rather than cache state: a named key
before its profile exists leaves root untouched and lands only in the
profile store once provisioned, and a resolver exception fails closed too.
Both fail against the previous two-state behavior by returning a live
SessionDB where None is required.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 58dcc10049 fix(gateway): memoize only resolved profile homes, never the miss
Review feedback on the memo introduced with _profile_home_for_key. The
problem is sharper than "no invalidation on profile deletion": caching the
miss pinned a profile that appears AFTER the gateway started to the ambient
store for the life of the process, which is the exact failure this helper
exists to prevent.

That is not hypothetical — an enrollment bridge can provision
profiles/<name>/ at runtime, so a key is legitimately seen before its
directory exists.

Memoize hits only. A miss costs one profile_exists() stat and recurs only
for profiles that genuinely do not exist, so the hot path for real profiles
is still a dict hit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 5ffaed6e45 fix(gateway): resolve session storage from the key's profile, not ambient scope
#88734 made SessionStore._db follow the ambient HERMES_HOME so a multiplexed
profile's rows reach its own state.db. That is correct for the inbound message
path, which installs the scope via _profile_runtime_scope. Nothing else does.

_session_expiry_watcher (gateway/run.py) walks the single process-wide
_entries dict — every profile's keys — and finalizes expired sessions with no
scope installed, so _db resolved the ROOT store for rows that live under
profiles/<name>/state.db. The scoped inbound path and the unscoped background
path then maintained two copies of the same logical session whose end_reason
drifted apart independently. Once they disagreed, the #54878 stale-routing
guard read one copy while the routing index pointed at the other, and a live
conversation was dropped and recreated — silently, since that branch only sets
was_auto_reset when a reset policy also fired.

Field evidence from a live two-profile install: session 20260814_234313 was
end_reason=None in the root store but agent_close in the profile store, while
20260822_225807 was inverted. Both directions, which rules out a single
mis-scoped writer.

The owning profile is already encoded in the session key, so derive the store
from it: _profile_home_for_key / _db_for_key, plus _db_for_session_id for the
entry points addressed by session id. 40 self._db uses across 14 methods now
resolve that way. No signature changed and no existing test was modified.

_profile_home_for_key returns None when multiplexing is off, when the key
carries the legacy agent:main namespace, or when the profile has no live
directory, so single-profile installs resolve exactly where they always did.
The explicit-path branch still goes through SessionDB.__init__ ->
_ensure_test_isolation, keeping the live-DB guard over per-profile paths.

Part of #66887. The routing-index half — _routing_scope() and the sessions.json
mirror still pinned to one frozen sessions_dir while the handle moves — is left
for a follow-up rather than mixed in here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
kshitijk4poor f20bbfa40d fix: a declined liveness abort must not cancel a pending compression
Closes the #99758 review P1 (andrexibiza): with a generation claim in
play, `interrupt()` called `_admit_hard_cancel()` BEFORE the claim was
validated at the final mutation edge, and the production
`CompressionCommitFence.cancel_before_commit()` irreversibly sets
`_cancelled = True` whenever no commit has started. So a watchdog abort
that ultimately DECLINED (real progress landed in the window, claim went
stale) had already killed the recovered turn's legitimate pending
compression commit — `begin_commit()` refuses a cancelled fence forever.
Generation authority covered interrupt publication but not the
compression-fence mutation that preceded it.

Split hard-cancel admission into two halves:

- `_wait_for_compression_commit()` runs pre-claim and is NON-mutating:
  it only blocks when `commit_in_flight` is true (the started-commit
  branch of the production fence waits for `finish_commit` without
  cancelling), so the interrupt still publishes only after an in-flight
  SessionDB mutation has finished — exactly as before.
- `_cancel_pending_compression_commit()` runs AFTER
  `_consume_claim_and_publish_first_state()` survives, so the
  destructive pending-commit cancellation can never outlive a stale
  claim. If a commit crossed its boundary in between, it is no longer
  fence-cancellable and completes on its own.

Regression coverage (both use the real `CompressionCommitFence`):

- `test_declined_abort_does_not_cancel_pending_compression_commit`:
  parks the interrupt at the claim-reservation release, lands real
  progress (G+1), lets the interrupt decline, then proves
  `fence.begin_commit()` still admits. Red on the pre-fix tree
  (mutation-checked: the fence was left cancelled).
- `test_declined_abort_parks_and_leaves_fence_operational`: the
  in-flight-commit window variant — activity lands while the interrupt
  waits on a started commit; the interrupt declines and a fresh
  `begin_commit()` still admits afterwards.
- The round-6 witness (`...resumes_inside_interrupt_publication`) now
  models an in-flight commit (`commit_in_flight = True`) so its park
  point stays inside the pre-claim wait, matching the new admission
  shape.

Also updates the `interrupt()` docstring for the deferred destructive
cancellation.
2026-09-01 03:19:59 +05:30