Six commits merged via #95366/#95367 carry an incorrect author email
(kshitijkapoorr@gmail.com — not an address the contributor owns; it was
set by tooling error during salvage). The address maps to no GitHub
account (verified via the users search API). Canonicalize it to the real
identity so shortlog/contributor tooling attributes correctly; the
contributors/emails mapping file from the same PRs already covers
release attribution.
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."
Fixes#84802
- _target_mirror_eligible accepts a precomputed origin_match so the sole
production caller stops re-resolving origin + re-running the origin
match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
precedence (standalone correctness: per-job False must beat raw global
True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
'not inchannel_continuable' and reword it for the widened seed scope.
The feature page still said 'only the origin chat is ever touched', which
this change makes stale. Documents the three attach-eligible shapes and
that the global flag never activates explicit targets (review nit).
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
tests/cron/test_cron_relay_delivery_guards.py landed on main after #89329
branched; its exact-dict assertion needs the new _resolved_from field the
salvaged commit adds to origin-resolved targets.
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).
The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.
Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
gate: origin unchanged; origin_fallback eligible under the same flags
as origin (per-job attach_to_session wins, else global
cron.mirror_delivery); explicit platform:chat targets eligible ONLY
under per-job attach_to_session — the global flag never activates
them, so it cannot start writing transcript entries into arbitrary
explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
session keys are user-isolated, so a seed without a user_id (origin-
less job into a shared channel) would create an orphan session no
reply resolves to — those targets fall back to the plain mirror. DM
targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.
Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.
15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.
Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).
Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
active non-local gateway are stamped with its registry connection_id
(explicit owners from multi-source responses preserved), so a later
resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
switch leaves the first read RPC on a newly-opened socket without a
response (request timed out / gateway connection closed), only while
the same gateway/profile is still foreground. Component fix for the
live-confirmed #92352 sidebar-never-paints gap.
Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).
Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)
Follow-up to the #95007 partial cherry: waitForInitialConnection() was
an unbounded listen on $connection — a primary that never publishes its
descriptor (spawn failure, dead SSH target) would strand
initializeConnectionsRegistry() forever and the last-used source would
never be restored. Bound it with the codebase's withTimeout helper
(same pattern as the sibling SWITCH_* call sites in this file): after
45s (the primary spawn budget) the restore proceeds exactly as it did
before the wait existed, and the listener is torn down either way.
Regression test: boot restore proceeds after the deadline with the
descriptor never arriving.
Original-PR: #95007
Partial cherry-pick of PR #95007 (weismanfamily). Surviving scope:
- electron connection-registry: registrySourceOwnsPrimaryBackend() —
descriptor-level proof that a registry-scoped request names the
already-running primary backend, wired into ensureRegistryBackend as
the generic (non-SSH-fingerprint) primary-owns short-circuit so a
cloud/url registry primary cannot spawn a second isolated server.
- store/connections: waitForInitialConnection() before the boot-time
source restore, so the sidebar registry cannot dial the preferred
source a second time while the identical primary backend is still
publishing its connection identity.
Dropped scope (superseded on main): the renderer routing half —
primaryConnectionId plumbing, primaryOwnsAgent short-circuits in
requestGatewayForAgent/openGatewayForAgent/ensureGatewayForAgent
(main has isPrimaryRegistryRoute via 1ec32e738), the 3-arg
setPrimaryGateway boot wiring (main has setPrimaryGatewayConnection),
the use-session-list-actions stampConnectionOwner (row stamping lands
via #94656/#94901), and the wiring.tsx bare-profile promotion commit
65d106e41 (main's knownSessionOwner covers it).
Original-PR: #95007
Dropped-scope: renderer routing half of d1c0fb093; all of 65d106e41
selectProfile(name) / newSessionInProfile(name) keep only $newChatProfile and
clear $newChatRoute. #94147 taught the send path to capture the (registry
source, profile) pair as the draft's exact owner, but three gaps still let a
session created on the composite gateway conn:local::omar degrade to the bare
string "omar" — and requestGatewayForProfile("omar") is a DIFFERENT socket
than the one that minted the runtime, so the next session-scoped RPC 4001'd
"session not found" while the runtime was ws-orphan-reaped:
1. Ownership persistence. The exact owner lived only in the bounded,
in-memory owner-hint map. The primary aggregate serves a `local` registry
source's rows WITHOUT connection_id (the unified-list splice tags non-local
sources only), and mergeSessionPage replaced the optimistic row with that
untagged row on the first sidebar refresh. After a hint eviction or a
relaunch nothing exact was left.
- One canonical SessionOwnerRoute type (store/session-request-router);
AgentProfileRoute / SessionProfileRoute / SessionRpcOwnerRoute alias it.
- Every owner ladder gains the connection-tagged ROW rung
(knownSessionOwner / sessionOwnerRouteFromRow): the session-RPC
dispatcher, knownOwnerForSession, the tile delegate, foregroundSessionScopes
and the async probe (resolveSessionOwner) all yield the exact route when
the row carries its connection.
- mergeSessionPage carries connection_id onto a row that comes back
untagged for the same profile (merge, don't clobber).
- Owner hints are persisted (bounded LRU, hermes.desktop.sessionOwnerHints.v1)
and rehydrated in LRU order; a removed registry connection drops its
hints (use-gateway-boot onChanged) so fail-closed can't pin sessions to a
dead source.
2. Fail closed. A request carrying session_id whose owner no rung could name
silently fell to the ambient presentation gateway, turning missing metadata
into a misleading backend "session not found". createSessionRpcDispatcher
and requestForOwnedSession now reject with an explicit
SessionOwnerResolutionError (store/session-owner-resolution). The ONE case
where ambient is the owner by construction stays ambient: no registry
source live AND at most one profile (legacy single-backend Desktop, whose
older backends omit `profile` on rows). Main-pane runtime ids (native
approval.respond, queued sends) now translate to their stored id through
the per-runtime state mirror (storedSessionIdForRuntimeId), so they resolve
an owner instead of tripping the gate.
3. Lifecycle. Between session.create returning and the foreground publication
($selectedStoredSessionId via navigate → route effect, or $sessionTiles),
the owner entry had no active request and was not yet foreground-pinned:
a prune recompute or a refcount-0 lease release could close the socket
holding the just-minted runtime before the first prompt.submit. Both create
paths now hold retainGatewayForAgent across the create RPC and hand off to
holdSessionOwnerUntilForeground (session-states), which names the owner in
foregroundSessionScopes — every registry dispose path honors it — until the
session becomes selected/tiled, the caller releases it (failed create,
mid-create drift close), or a 60s TTL expires. Nothing latches.
Regressions:
- profile-rail-fresh-chat-owner.test.tsx: new case evicts the hint AND
merges an untagged refresh row after turn one; turn two still rides the
same conn:local::omar socket, no probe, no session.close, no v1 socket;
the existing case also asserts the owner is foreground-pinned from create.
- session-rpc-dispatcher.test.ts (new): fail-closed error + ambient never
called; legacy single-backend stays ambient; tagged-row rung routes by
runtime id; hint outranks untagged row; probe result routes exactly.
- session.test.ts: persisted hints survive a simulated relaunch in LRU order,
malformed storage is ignored, per-connection forget; knownSessionOwner;
mergeSessionPage carries / drops / preserves identity on connection_id.
- session-states-foreground-scopes.test.ts: tagged-row scope for the selected
thread; hold pins from create, retires on selection / tile mount, explicit
release, TTL expiry.
- session-states-runtime-map.test.ts: state-mirror rung; knownOwnerForSession
through mirror + hint / tagged row; requestForOwnedSession fail-closed vs
legacy ambient.
- wiring-routing.test.ts: connection-tagged row rung; hint still outranks it.
Stacked on #94145, #94147 and #94178 (merged as the integration base);
review from this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Sessions / profile-rail path (selectProfile, newSessionInProfile, a
connection switch, `/profile`) sets $newChatProfile and deliberately clears
$newChatRoute, so a fresh chat had no explicit owner. The session was created
on the active registry gateway (conn:local::omar) but its durable owner
degraded to the bare string "omar": follow-up RPCs dialed
requestGatewayForProfile("omar") — a different socket than the one that
minted the WebSocket-scoped runtime — and 4001'd "session not found" while
the runtime was ws-orphan-reaped.
- store/profile: capture the active registry source together with the
new-chat profile intent ($newChatConnectionId / captureNewChatSource) in
selectProfile, newSessionInProfile, newSessionInAgent, connection switches
and `/profile`; resolveNewChatOwnerRoute() derives the exact
{ connectionId, profile } route whenever a registry source is live, even
with $newChatRoute null (legacy v1 primary still yields null).
- use-session-actions: session.create, the owner hint, the optimistic row's
profile + connection_id, and the failed-create cleanup all use that
effective owner (main chat and tile paths).
- use-prompt-actions/submit: re-pin targetStoredSessionId after a fresh
create. It was captured before the create (null) and seedOptimistic handed
it to updateSessionState, which the state cache read as a DETACH — the
fresh stored↔runtime binding was severed the moment the chat existed, so
every later session-scoped RPC failed to translate the runtime id, never
saw the tile route / owner hint / row, probed REST by runtime id and fell
to the ambient socket.
- contrib: the session-RPC dispatcher is factored out of wiring.tsx
(createSessionRpcDispatcher) so the exact production routing is what the
integration test drives.
Regression (profile-rail-fresh-chat-owner.test.tsx) drives the real path:
mocked sockets under the real registry store, primary = remote default,
active source = local, selectProfile("omar") ($newChatProfile = "omar",
$newChatRoute = null), real useSessionStateCache / useSessionActions /
usePromptActions and the production dispatcher; asserts session.create and
BOTH prompt.submit calls hit the same conn:local::omar gateway object, no
session-scoped RPC reached the primary or a v1 "omar" socket, no
session.close, the binding survives both turns, no REST probe.
Verified with the packaged Linux Desktop against the real ~/.hermes
(primary = remote OAuth gateway, "This device" as registry source, omar via
the profile rail, two prompts): both prompts persisted on one session in
profiles/omar/state.db, no ws_orphan_reap.
Refs #94071
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A fresh chat created through $newChatRoute lost its owner the moment
session.create returned. The create RPC rode the captured route
(requestGatewayForAgent), but the optimistic row was stamped from
$activeGatewayProfile — still `default` in All-profiles / Bot routing —
and no owner hint was recorded. The first turn ran on the routed
backend (e.g. local::omar); every later session-scoped RPC resolved the
row as `default` and 4001'd "session not found", leaving the routed
runtime to be ws-orphan-reaped.
Make the ownership transition atomic with the create:
- record capturedRoute as the stored session's exact owner hint the
moment a routed create returns a stored id (main chat and tile paths);
- upsertOptimisticSession accepts an explicit owner and stamps
profile = targetProfile || profile plus the owning connection_id,
falling back to the ambient profile only for an unrouted create;
- contrib/wiring resolves session RPC owners as: persisted tile owner
route → exact unique owner hint → session-row profile → cross-profile
probe (resolveSessionRpcOwner, pure + unit-tested), so prompt.submit,
session.resume, attachments, interrupt, redirect and recovery all use
the exact owner; knownOwnerForSession follows the same ladder;
- the mid-create drift-abort session.close rides capturedRoute too.
Regressions: an integration test (ambient default, route local::omar,
two turns → both prompt.submit hit local::omar, no session-not-found,
no session.close) and a unit test asserting a routed fresh create never
yields { followupOwner: 'default', foregroundScope: 'conn:local::omar' }.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 95e2ef66becb641d3bdedf016dab6c104a777ae7)
Bot Mode resolves the forever-chat by exact-title lookup on
(profile, 'Bot Chat'); no session-id pointer exists. A user rename
therefore orphaned the whole conversation: resolution missed, the next
click minted an empty replacement, and UNIQUE(title) then blocked ever
renaming back. Refuse the rename at SessionDB._set_session_title — the
single write path every surface funnels through (gateway session.title,
/title, CLI rename, REST). Hidden discriminates the registry row, so a
normal visible session a user happens to call 'Bot Chat' stays freely
renameable; re-asserting the same canonical title stays a no-op.
Hardening on top of the TCC daemon-identity salvage:
- _validate_cua_driver_app_signature: codesign -dv gate requiring EXACT
Identifier=com.trycua.driver and the official team (4YEC26S9KF) before
/usr/bin/open hands the bundle to LaunchServices — the identity fix must
not double as a launcher for arbitrary/impostor bundles (suffixed
identifiers and wrong teams rejected; unsigned dev builds only via
computer_use.allow_unsigned_driver: true in config.yaml).
- _resolve_cua_driver_app_path: derive the bundle ONLY from the resolved
driver binary — the /Applications fallback could launch a DIFFERENT
install than the manifest resolved.
- open -n -g: don't activate/steal focus when launching the daemon.
- 7 new tests incl. sabotage-verified exact-match assertions.
Grafted from #76433's review direction (@Chadmc9889's original fail-closed
validation requirement).
Co-authored-by: Chadmc9889 <Chadmc9889@users.noreply.github.com>
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.
Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).
Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
The embedded spawn now consults the overlay policy (capability probe via
subprocess.run) when _cua_no_overlay() is true — which it is on headless
CI since the Linux X11 default flip. The fixed two-entry run side_effect
in this test didn't budget for the probe call; pin the policy off since
this test pins the socket/ack contract, not overlay behavior.
The intro kickoff ('Hey, tell me about yourself!') now fires ONLY from
genuine New Agent creation. The bot-click canonical resolution path mints
silently: the eager session.title write already persists the lazy row on
modern gateways, so the kickoff's session-persistence job is obsolete
there. A resolution miss (retitled row, hidden-listing gap, post-update
skew) previously re-fired the kickoff on EVERY click — a burned model
turn plus a user-attributed prompt the user never typed (ScottFive
report). Older gateways that reject the eager title keep a narrow compat
kickoff, else the pruner reaps the empty lazy session.
teardownSshConnection closed the tunnel and SSH transport but never
killed the detached serve --isolated process. Spawn uses setsid/nohup,
so the backend reparents to pid 1, keeps state.db open, and accumulates
across Cmd+Q. Reuse cleanupStale via disconnect while SSH can still
exec, sequence remote kill before close, and seal the bootstrap
coordinator so reconnect during a prevented first quit cannot respawn.
The quit race is 6s to cover cleanupStale's 5s wait-for-exit loop.
Avoid opening a second SSH lifecycle when a migrated registry request targets the same primary/default backend already booted through the legacy route. Compare effective SSH configuration for representation-only drift, treat empty and default as the same root profile, and keep named profiles isolated.
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
'invalid response' errors (#7264) into the same empty-content abort
carve-out so those shapes also preserve the session (#94459's wider
classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
_COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.
- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py
Fixes#94448
The reuse reviewer found a third copy of the fc_->call_ synthesis in
agent/transports/codex.py _pair_ids (item_id[3:] spelling, which is why
the len('fc_') grep missed it). All three sites now share
_canonical_call_id_from_fc(), keeping the pairing invariant in one place.