- async timeout marks once with a label-carrying rounded-timeout reason
- completion never marks
- sync flavor marks on timeout
- non-adopting backends keep the real timeout result
- a raising mark_suspect cannot eat the timeout or the label
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
The sidebar's home-dir repo crawl descends into ~/Pictures, ~/Music,
~/Movies and ~/Public — and into Photos/Music library packages — which
triggers Photos, Media Library and Files & Folders permission prompts
attributed to Hermes.app. Because the app is ad-hoc signed and re-signed
on every self-update (#49110), macOS drops all TCC grants after each
update, so these prompts re-fire every time.
Skip the media folders as direct children of a search root (nested dirs
like ~/dev/Music are ordinary and still scanned; an explicitly passed
root is still walked), and skip Apple library packages
(*.photoslibrary/*.musiclibrary/*.tvlibrary/*.aplibrary) at any depth.
Partial mitigation for #49110 / #52010: removes the Photos and Media
Library prompts entirely; the identity reset itself needs Developer ID
signed release artifacts (tracked in #49110).
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.
_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.
Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.
Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).
Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.
The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
After macOS sleep/resume, pooled remote SSH descriptors kept serving dead
tunnels: a remote entry has no child 'exit' to clear it, the renderer
keepalive spares it from the idle reaper, and the wake-path nudges from
b90289b04/febed060a only re-drive the PRIMARY renderer socket. The
background failure-streak policy needs several probe rounds before it
drops a descriptor, so the Bots pane showed 'Gateway offline' long after
the network was back.
New revalidateSuspectPooledRemoteBackends(): on resume every pooled
remote is suspect — probe each once (bounded by
REMOTE_LIVENESS_TIMEOUT_MS), retire the dead ones immediately (pool
entry + SSH bootstrap + tunnel/master teardown) and rebuild them through
the caller's dial path, while healthy descriptors are left untouched. A
failed retire skips the rebuild (never dial on top of an installed
descriptor); a failed rebuild is logged and left to the renderer's
normal reconnect — the sweep never throws.
attachPowerResumeRemoteRevalidation() wires the sweep to the Electron
powerMonitor 'resume'/'unlock-screen' seam with a 15s holdoff so the
near-simultaneous macOS wake signals coalesce into one sweep and can
never form a hot loop; overlapping kicks additionally join the one
in-flight sweep via the existing RemoteRevalidationCoordinator.
composer-status imports $gateway from this module (cycle forces the
dynamic import), and awaiting the module load inside openSecondary sat on
the timed redial path — under CI load that pushed cold-start redials past
waitFor budgets in the lifecycle suite. The reset needs no ordering
guarantee relative to the dial; detach it.
The combined dynamic import meant a failed composer-status import (mocked
test graphs) silently skipped resetTileRuntimeBindings too — the exact
lifecycle regression CI caught. Separate best-effort trys per module.
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).
Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.
Tests:
- electron/hardening.test.ts pins both broadcast branches in
saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
(/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
'saved' push arrives, menu after — without reload — shows it.
A gateway connection that dies mid-turn leaves cached session snapshots
whose busy/awaitingResponse flags can never settle: the respawned
backend re-mints runtime ids, so no terminal publish ever reaches the
orphaned snapshot again. #isWarmSettled treated those frozen flags as
live work, so every orphan pinned its full warm transcript until app
restart — roughly 5MB per reconnect cycle, which turned the restart
loop in #95189 into renderer OOM.
SessionStateCache now accepts an optional isAuthoritativelyActive
probe. When wired, in-flight flags only block eviction while the
authoritative $sessionStates record still claims work for the same
runtime id; without the probe the legacy always-block behavior is
preserved byte-for-byte. Eviction remains gated on needsInput, pending
drafts, and active references, so a genuinely running turn (which
re-asserts busy on every publish) is never a casualty.
The useSessionStateCache hook wires the probe to the store it already
imports. Reconnect reconciliation (reconcileBusyStatesOnReconnect)
settles the authoritative record, and the next prune drains the
orphaned cache entry through the normal LRU path, ownership included.
Two review-thread deltas on the salvaged #94950 latch:
- Match the gateway's structured 4001 code, not a message substring, when
the rejection carries one (JsonRpcGatewayError). A coded error that
merely mentions 'session not found' in wrapped text (e.g. a 5007 tool
failure) must not latch the guard and freeze the status stack on a
healthy session. The substring fallback survives only for codeless
legacy errors.
- Reset the latch on runtime re-mint, not only on status-stack rebind:
wire resetBackgroundPollingGuard() at both reconnect seams that already
drop stale runtime bindings (use-gateway-boot's post-reconnect
resetTileRuntimeBindings and gateway.ts's reopening path), so ids the
dead runtime 4001'd resume polling once a respawned backend re-mints
them.
Tests: code-specific match both directions; full-reset resumes every
latched session.
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".
`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.
The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.
Distinguish the two failure classes:
- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
session may well still be alive. Misclassifying that direction would
silently freeze the status stack on a healthy session.
The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.
Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.
Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
Follow-up to the #95039 salvage: the cherry-picked bound used the 20s
RECONNECT_ATTEMPT_TIMEOUT_MS on boot()/softSwitch() getConnection(), but a
reviewer note (and the Phase A registry-restore work) established that
boot-class awaits must ride out a full backend cold spawn — main's spawn
budget is 45s (DEFAULT_BACKEND_READY_TIMEOUT_MS). A 20s renderer bound would
latch boot errors on healthy-but-slow cold boots.
Introduce BACKEND_BOOT_WAIT_TIMEOUT_MS (45s) in lib/with-timeout.ts as the
single shared boot-class budget, point boot()/softSwitch() getConnection()
and connections.ts BOOT_DESCRIPTOR_WAIT_TIMEOUT_MS at it, and keep the 20s
reconnect budget only for reconnect-class awaits against an already-spawned
backend. No magic-number drift: 45_000 now appears once in renderer code.
resolveGatewayWsUrl() in attemptReconnect() because a wedged IPC
round-trip into the main process (e.g. a stuck revalidation after a
liveness-probe trip) can hang these awaits forever. A later fix
(e8d5660bae) extended the bound to resolveGatewayWsUrl() in boot() and
softSwitch() too, but left the getConnection() call immediately above
it in both functions unbounded.
If that call wedges: during initial boot the 'Starting Hermes...'
screen never resolves (bootCompleted never flips, nothing hits catch),
and during a soft gateway/profile switch latches
true forever since the try block's finally never runs.
Wrap both with the same withTimeout()/RECONNECT_ATTEMPT_TIMEOUT_MS
pattern already used for the sibling calls.
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.
The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.
Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".
- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.
Closes#82614
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.
Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.
Two fixes:
1. _check_via_local_git now detects fetch failure (returncode != 0 or
exception) and returns None instead of falling through to stale
refs. The caller treats None as 'check could not run' rather than
'up to date'.
2. check_for_updates no longer caches None results. Previously, a
None from a failed check was cached for 6 hours, suppressing
retries until the cache expired. Now only conclusive results
(0 or >=1) are cached, so the next check attempt runs immediately
on the next call.
Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
Follow-up to the salvaged #95207 fix, completing the misreport bug class:
- restore() now also drops delete_targets whose unlink failed (OSError
swallowed) from restored_files — the sibling of the kept-oversize
misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
(slash_commands + gateway.rollback.kept_oversize locale key in all 17
catalogs) now tells the user which files were kept because the size
cap excluded them from every checkpoint; previously the file was
correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.
The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.
Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.
No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.
Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.
Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.
Reproduced on main with shipped defaults:
corpus.jsonl (2 MB), agent appends to it, then /rollback 1
safe_restore_plan restore=['corpus.jsonl', 'notes.py']
restore ok=True restored_files=['corpus.jsonl', 'notes.py']
notes.py exists=True content="v1 = 'original source'"
corpus.jsonl exists=False <- deleted, was in no checkpoint
Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.
The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.
The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.
The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.
Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.
Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
Integration fixups so #82529's generation signing and #95131's interpreter
anchor cover each other's gaps (without these, each fix leaves the other's
rotation path broken):
- macos_tcc_anchor store detection now recognizes .hermes-runtime/python/
generation-* stores: repair_vulnerable_runtime() rebuilds the venv against
a generation interpreter, replacing the anchored bin/python with a fresh
symlink — previously the anchor then read 'not uv-managed' and NEVER
re-anchored, so every SQLite CVE repair silently orphaned terminal TCC
grants (the exact #82427 scenario, path-keyed).
- _install_anchor signs the anchor copy with the same identifier-pinned DR
(via managed_uv._macos_sign_managed_python) before it goes live: copy2
carries the source build's cdhash-based signature, so an unsigned refresh
would still change the stored csreq on every patch bump/repair despite
the stable path. Best-effort, never blocks the anchor.
- Tests: generation-store recognition + repair-generation anchoring +
sign-on-install call (sabotage-verified: dropping the generation root
marker fails both new tests).
Between the registry miss and the eager session.title write, another
writer can take the canonical title (peer dm minting server-side, a
second machine, cross-connection sync). UNIQUE(title) rejects our write
with 'already in use' — which the compat path previously read as 'old
gateway' and prompted into OUR stray lazy session, forking the forever
chat. A uniqueness rejection now re-consults the registry and adopts the
winner; the zero-message stray is abandoned to the gateway pruner.
Genuine old-gateway failures (unknown method) keep the compat kickoff.
Follow-up on the cherry-picked #90030: candidate.resolve() breaks the fix
on the standard Linux venv layout, where venv/bin/python is a symlink to
the base interpreter. Resolving makes the venv python compare equal to
the dependency-less base (so the swap never happens), and returning the
resolved target would spawn the bare base interpreter, bypassing
pyvenv.cfg — the fix would silently not fix#90026 on the exact platform
it was reported from. Compare and return normalized UNRESOLVED paths:
the venv path IS the interpreter's identity. Adds the symlink-layout
regression test; live-E2E'd with a real dependency-less base runtime.
Under an SSH remote backend the web server is launched by running the uv
BASE interpreter with the venv's site-packages injected into sys.path at
startup, so sys.executable is a dependency-less python. Detached
dashboard actions spawned from it (Update now, restart, anything routed
through _spawn_hermes_action) inherited neither the injected path nor a
PYTHONPATH and died on the first third-party import — 'Update now'
always failed instantly with ModuleNotFoundError: No module named 'yaml'
while 'hermes update' from the venv worked (#90026).
_dashboard_spawn_executable now prefers the install's own venv
interpreter (venv/bin/python, venv/Scripts/python.exe) when it differs
from sys.executable, resolving the same dependency set the venv launcher
provides. Same-interpreter launches return sys.executable unchanged,
preserving the Windows console-ownership behavior verbatim, and layouts
without an install venv keep the old fallback.
Six commits merged via #95366/#95367 carry an incorrect author email
(kshitijkapoorr@gmail.com — not an address the contributor owns; it was
set by tooling error during salvage). The address maps to no GitHub
account (verified via the users search API). Canonicalize it to the real
identity so shortlog/contributor tooling attributes correctly; the
contributors/emails mapping file from the same PRs already covers
release attribution.
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."
Fixes#84802
- _target_mirror_eligible accepts a precomputed origin_match so the sole
production caller stops re-resolving origin + re-running the origin
match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
precedence (standalone correctness: per-job False must beat raw global
True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
'not inchannel_continuable' and reword it for the widened seed scope.
The feature page still said 'only the origin chat is ever touched', which
this change makes stale. Documents the three attach-eligible shapes and
that the global flag never activates explicit targets (review nit).
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
tests/cron/test_cron_relay_delivery_guards.py landed on main after #89329
branched; its exact-dict assertion needs the new _resolved_from field the
salvaged commit adds to origin-resolved targets.
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).
The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.
Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
gate: origin unchanged; origin_fallback eligible under the same flags
as origin (per-job attach_to_session wins, else global
cron.mirror_delivery); explicit platform:chat targets eligible ONLY
under per-job attach_to_session — the global flag never activates
them, so it cannot start writing transcript entries into arbitrary
explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
session keys are user-isolated, so a seed without a user_id (origin-
less job into a shared channel) would create an orphan session no
reply resolves to — those targets fall back to the plain mirror. DM
targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.
Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.
15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.
Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).
Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
active non-local gateway are stamped with its registry connection_id
(explicit owners from multi-source responses preserved), so a later
resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
switch leaves the first read RPC on a newly-opened socket without a
response (request timed out / gateway connection closed), only while
the same gateway/profile is still foreground. Component fix for the
live-confirmed #92352 sidebar-never-paints gap.
Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).
Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)