Commit Graph

25477 Commits

Author SHA1 Message Date
Ayush Nangia 60b93eb521 test(deadline): Phase 3a poisoned-state coverage
- async timeout marks once with a label-carrying rounded-timeout reason
- completion never marks
- sync flavor marks on timeout
- non-adopting backends keep the real timeout result
- a raising mark_suspect cannot eat the timeout or the label
2026-08-26 17:47:22 +05:30
Ayush Nangia 7a3aaf0143 feat(deadline): SuspectableBackend protocol — mark timed-out backends suspect
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
2026-08-26 17:47:22 +05:30
hermes-seaeye[bot] 86ae906e88 fmt(js): npm run fix on merge (#95511)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 11:57:06 +00:00
miha 1bd5da3ac6 fix(desktop): skip macOS TCC-protected media dirs in git repo scan
The sidebar's home-dir repo crawl descends into ~/Pictures, ~/Music,
~/Movies and ~/Public — and into Photos/Music library packages — which
triggers Photos, Media Library and Files & Folders permission prompts
attributed to Hermes.app. Because the app is ad-hoc signed and re-signed
on every self-update (#49110), macOS drops all TCC grants after each
update, so these prompts re-fire every time.

Skip the media folders as direct children of a search root (nested dirs
like ~/dev/Music are ordinary and still scanned; an explicitly passed
root is still walked), and skip Apple library packages
(*.photoslibrary/*.musiclibrary/*.tvlibrary/*.aplibrary) at any depth.

Partial mitigation for #49110 / #52010: removes the Photos and Media
Library prompts entirely; the identity reset itself needs Developer ID
signed release artifacts (tracked in #49110).
2026-08-26 04:51:48 -07:00
Teknium 279726cc2f chore: map contributor email for wiseconnex 2026-08-26 04:51:41 -07:00
Dimar Anez 9eb13d07b6 fix(terminal): tolerate macOS TCC PermissionError in _safe_getcwd
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.

_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.

Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.

Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
2026-08-26 04:51:41 -07:00
Teknium bd134d0f30 test: loosen frozen bare-verdict dict in cua_0_9 sibling test to decision contract
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
2026-08-26 04:50:13 -07:00
Teknium 3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
Teknium 65605d4a7a fix(computer_use): pitch background-FIRST (not background-only) in schema, prompt block, and skill 2026-08-26 04:50:13 -07:00
Teknium bb3421bf25 fix(desktop): single-owner backend dial claim in Electron main (#90812)
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).

Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.

The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
2026-08-26 04:49:56 -07:00
Teknium 3123624c07 fix(desktop): revalidate pooled remote/SSH backends on power resume (#93910)
After macOS sleep/resume, pooled remote SSH descriptors kept serving dead
tunnels: a remote entry has no child 'exit' to clear it, the renderer
keepalive spares it from the idle reaper, and the wake-path nudges from
b90289b04/febed060a only re-drive the PRIMARY renderer socket. The
background failure-streak policy needs several probe rounds before it
drops a descriptor, so the Bots pane showed 'Gateway offline' long after
the network was back.

New revalidateSuspectPooledRemoteBackends(): on resume every pooled
remote is suspect — probe each once (bounded by
REMOTE_LIVENESS_TIMEOUT_MS), retire the dead ones immediately (pool
entry + SSH bootstrap + tunnel/master teardown) and rebuild them through
the caller's dial path, while healthy descriptors are left untouched. A
failed retire skips the rebuild (never dial on top of an installed
descriptor); a failed rebuild is logged and left to the renderer's
normal reconnect — the sweep never throws.

attachPowerResumeRemoteRevalidation() wires the sweep to the Electron
powerMonitor 'resume'/'unlock-screen' seam with a 15s holdoff so the
near-simultaneous macOS wake signals coalesce into one sweep and can
never form a hot loop; overlapping kicks additionally join the one
in-flight sweep via the existing RemoteRevalidationCoordinator.
2026-08-26 04:49:56 -07:00
Teknium 600d5166f0 test(gateway): prove delivered rows are never reclaimed by the reconnect sweep 2026-08-26 04:49:39 -07:00
milnerrad 8e1db41041 fix(gateway): redeliver transient failures after reconnect 2026-08-26 04:49:39 -07:00
Teknium b455abe0b3 fix(desktop): poll-guard reset is fire-and-forget off the redial path
composer-status imports $gateway from this module (cycle forces the
dynamic import), and awaiting the module load inside openSecondary sat on
the timed redial path — under CI load that pushed cold-start redials past
waitFor budgets in the lifecycle suite. The reset needs no ordering
guarantee relative to the dial; detach it.
2026-08-26 04:49:22 -07:00
Teknium 62534e2b5a fix(desktop): isolate the poll-guard reset import + sort-imports lint
The combined dynamic import meant a failed composer-status import (mocked
test graphs) silently skipped resetTileRuntimeBindings too — the exact
lifecycle regression CI caught. Separate best-effort trys per module.
2026-08-26 04:49:22 -07:00
Teknium 6bbae974d5 chore: map justinjohnson25600 and BrunoBza contributor emails 2026-08-26 04:49:22 -07:00
Teknium fe615a0099 fix(desktop): republish the connections registry to renderers after every successful save (#95393)
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).

Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.

Tests:
- electron/hardening.test.ts pins both broadcast branches in
  saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
  (/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
  'saved' push arrives, menu after — without reload — shows it.
2026-08-26 04:49:22 -07:00
Bruno Bza 06be6cffbc fix(desktop): release reconnect-orphaned warm transcripts once their authoritative state settles
A gateway connection that dies mid-turn leaves cached session snapshots
whose busy/awaitingResponse flags can never settle: the respawned
backend re-mints runtime ids, so no terminal publish ever reaches the
orphaned snapshot again. #isWarmSettled treated those frozen flags as
live work, so every orphan pinned its full warm transcript until app
restart — roughly 5MB per reconnect cycle, which turned the restart
loop in #95189 into renderer OOM.

SessionStateCache now accepts an optional isAuthoritativelyActive
probe. When wired, in-flight flags only block eviction while the
authoritative $sessionStates record still claims work for the same
runtime id; without the probe the legacy always-block behavior is
preserved byte-for-byte. Eviction remains gated on needsInput, pending
drafts, and active references, so a genuinely running turn (which
re-asserts busy on every publish) is never a casualty.

The useSessionStateCache hook wires the probe to the store it already
imports. Reconnect reconciliation (reconcileBusyStatesOnReconnect)
settles the authoritative record, and the next prune drains the
orphaned cache entry through the normal LRU path, ownership included.
2026-08-26 04:49:22 -07:00
Teknium a7ea156470 fix(desktop): harden the dead-session poll guard per #94950 review
Two review-thread deltas on the salvaged #94950 latch:

- Match the gateway's structured 4001 code, not a message substring, when
  the rejection carries one (JsonRpcGatewayError). A coded error that
  merely mentions 'session not found' in wrapped text (e.g. a 5007 tool
  failure) must not latch the guard and freeze the status stack on a
  healthy session. The substring fallback survives only for codeless
  legacy errors.

- Reset the latch on runtime re-mint, not only on status-stack rebind:
  wire resetBackgroundPollingGuard() at both reconnect seams that already
  drop stale runtime bindings (use-gateway-boot's post-reconnect
  resetTileRuntimeBindings and gateway.ts's reopening path), so ids the
  dead runtime 4001'd resume polling once a respawned backend re-mints
  them.

Tests: code-specific match both directions; full-reset resumes every
latched session.
2026-08-26 04:49:22 -07:00
Justin Johnson c19849cd02 fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".

`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.

The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.

Distinguish the two failure classes:

- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
  and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
  session may well still be alive. Misclassifying that direction would
  silently freeze the status stack on a healthy session.

The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.

Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.

Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
  never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
2026-08-26 04:49:22 -07:00
Teknium 574bd7175c fix(desktop): unify boot-class getConnection() budgets on one shared 45s constant
Follow-up to the #95039 salvage: the cherry-picked bound used the 20s
RECONNECT_ATTEMPT_TIMEOUT_MS on boot()/softSwitch() getConnection(), but a
reviewer note (and the Phase A registry-restore work) established that
boot-class awaits must ride out a full backend cold spawn — main's spawn
budget is 45s (DEFAULT_BACKEND_READY_TIMEOUT_MS). A 20s renderer bound would
latch boot errors on healthy-but-slow cold boots.

Introduce BACKEND_BOOT_WAIT_TIMEOUT_MS (45s) in lib/with-timeout.ts as the
single shared boot-class budget, point boot()/softSwitch() getConnection()
and connections.ts BOOT_DESCRIPTOR_WAIT_TIMEOUT_MS at it, and keep the 20s
reconnect budget only for reconnect-class awaits against an already-spawned
backend. No magic-number drift: 45_000 now appears once in renderer code.
2026-08-26 04:49:22 -07:00
nftpoetrist 31f3de1f06 fix(desktop): bound getConnection() on the boot and soft-switch paths (#93454)
resolveGatewayWsUrl() in attemptReconnect() because a wedged IPC
round-trip into the main process (e.g. a stuck revalidation after a
liveness-probe trip) can hang these awaits forever. A later fix
(e8d5660bae) extended the bound to resolveGatewayWsUrl() in boot() and
softSwitch() too, but left the getConnection() call immediately above
it in both functions unbounded.

If that call wedges: during initial boot the 'Starting Hermes...'
screen never resolves (bootCompleted never flips, nothing hits catch),
and during a soft gateway/profile switch  latches
true forever since the try block's finally never runs.

Wrap both with the same withTimeout()/RECONNECT_ATTEMPT_TIMEOUT_MS
pattern already used for the sibling calls.
2026-08-26 04:49:22 -07:00
Teknium 20d33e385a fix(web_server): a dashboard started without a build recovers the moment one appears (#82614)
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.

The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.

Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
2026-08-26 04:48:58 -07:00
codexbt 3dea11d703 fix(web_server): recheck WEB_DIST existence dynamically in mount_spa
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".

- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.

Closes #82614
2026-08-26 04:48:58 -07:00
Shakti Prasad Mohapatra 98e87ac886 fix(cli): preserve stale positive behind-count on fetch failure (#92578)
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.

Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
2026-08-26 04:17:39 -07:00
Shakti Prasad Mohapatra 55d50d5c92 fix(cli): don't serve stale update-check results after fetch failure (#82166)
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.

Two fixes:

1. _check_via_local_git now detects fetch failure (returncode != 0 or
   exception) and returns None instead of falling through to stale
   refs. The caller treats None as 'check could not run' rather than
   'up to date'.

2. check_for_updates no longer caches None results. Previously, a
   None from a failed check was cached for 6 hours, suppressing
   retries until the cache expired. Now only conclusive results
   (0 or >=1) are cached, so the next check attempt runs immediately
   on the next call.

Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
2026-08-26 04:17:39 -07:00
kshitijk4poor d62a05e94c fix(checkpoints): surface skipped_oversize to users and stop misreporting failed deletes as restored
Follow-up to the salvaged #95207 fix, completing the misreport bug class:

- restore() now also drops delete_targets whose unlink failed (OSError
  swallowed) from restored_files — the sibling of the kept-oversize
  misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
  (slash_commands + gateway.rollback.kept_oversize locale key in all 17
  catalogs) now tells the user which files were kept because the size
  cap excluded them from every checkpoint; previously the file was
  correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
2026-08-26 16:44:43 +05:30
RickyYii 595b5ce68a refactor(checkpoints): call the size-cap predicate instead of restating it
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.

The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.

Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.

No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.

Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
2026-08-26 16:44:43 +05:30
RickyYii d28bf79927 fix(checkpoints): stop safe restore deleting files the size cap excluded
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.

Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.

Reproduced on main with shipped defaults:

    corpus.jsonl (2 MB), agent appends to it, then /rollback 1
    safe_restore_plan restore=['corpus.jsonl', 'notes.py']
    restore ok=True restored_files=['corpus.jsonl', 'notes.py']
    notes.py     exists=True   content="v1 = 'original source'"
    corpus.jsonl exists=False  <- deleted, was in no checkpoint

Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.

The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.

The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.

The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.

Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.

Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
2026-08-26 16:44:43 +05:30
Teknium 7a10d91b29 chore: map contributor email for notkisk 2026-08-26 04:14:16 -07:00
Teknium 9f8cdf89d6 fix(macos): keep the TCC anchor alive across CVE-repair rotations + sign anchor copies
Integration fixups so #82529's generation signing and #95131's interpreter
anchor cover each other's gaps (without these, each fix leaves the other's
rotation path broken):
- macos_tcc_anchor store detection now recognizes .hermes-runtime/python/
  generation-* stores: repair_vulnerable_runtime() rebuilds the venv against
  a generation interpreter, replacing the anchored bin/python with a fresh
  symlink — previously the anchor then read 'not uv-managed' and NEVER
  re-anchored, so every SQLite CVE repair silently orphaned terminal TCC
  grants (the exact #82427 scenario, path-keyed).
- _install_anchor signs the anchor copy with the same identifier-pinned DR
  (via managed_uv._macos_sign_managed_python) before it goes live: copy2
  carries the source build's cdhash-based signature, so an unsigned refresh
  would still change the stored csreq on every patch bump/repair despite
  the stable path. Best-effort, never blocks the anchor.
- Tests: generation-store recognition + repair-generation anchoring +
  sign-on-install call (sabotage-verified: dropping the generation root
  marker fails both new tests).
2026-08-26 04:14:16 -07:00
notkisk 8d6c0a3098 fix: preserve macOS TCC identity for managed Python 2026-08-26 04:14:16 -07:00
Teknium d22e2b9f6e fix(desktop): a canonical-title race adopts the winner instead of forking the forever chat (#92473, part 2)
Between the registry miss and the eager session.title write, another
writer can take the canonical title (peer dm minting server-side, a
second machine, cross-connection sync). UNIQUE(title) rejects our write
with 'already in use' — which the compat path previously read as 'old
gateway' and prompted into OUR stray lazy session, forking the forever
chat. A uniqueness rejection now re-consults the registry and adopts the
winner; the zero-message stray is abandoned to the gateway pruner.
Genuine old-gateway failures (unknown method) keep the compat kickoff.
2026-08-26 04:03:35 -07:00
Teknium 19fde8a450 fix(dashboard): compare and spawn the venv interpreter by UNRESOLVED path
Follow-up on the cherry-picked #90030: candidate.resolve() breaks the fix
on the standard Linux venv layout, where venv/bin/python is a symlink to
the base interpreter. Resolving makes the venv python compare equal to
the dependency-less base (so the swap never happens), and returning the
resolved target would spawn the bare base interpreter, bypassing
pyvenv.cfg — the fix would silently not fix #90026 on the exact platform
it was reported from. Compare and return normalized UNRESOLVED paths:
the venv path IS the interpreter's identity. Adds the symlink-layout
regression test; live-E2E'd with a real dependency-less base runtime.
2026-08-26 04:03:21 -07:00
liuhao1024 ad8f995bcf fix(dashboard): spawn detached actions from the install's venv interpreter
Under an SSH remote backend the web server is launched by running the uv
BASE interpreter with the venv's site-packages injected into sys.path at
startup, so sys.executable is a dependency-less python. Detached
dashboard actions spawned from it (Update now, restart, anything routed
through _spawn_hermes_action) inherited neither the injected path nor a
PYTHONPATH and died on the first third-party import — 'Update now'
always failed instantly with ModuleNotFoundError: No module named 'yaml'
while 'hermes update' from the venv worked (#90026).

_dashboard_spawn_executable now prefers the install's own venv
interpreter (venv/bin/python, venv/Scripts/python.exe) when it differs
from sys.executable, resolving the same dependency set the venv launcher
provides. Same-interpreter launches return sys.executable unchanged,
preserving the Windows console-ownership behavior verbatim, and layouts
without an install venv keep the old fallback.
2026-08-26 04:03:21 -07:00
kshitij 16c67248d4 Merge pull request #95477 from kshitijk4poor/chore/mailmap-kshitijk4poor
chore: mailmap kshitijkapoorr@gmail.com to kshitijk4poor
2026-08-26 16:22:14 +05:30
kshitijk4poor d873ee8e25 chore: mailmap kshitijkapoorr@gmail.com to kshitijk4poor's canonical noreply
Six commits merged via #95366/#95367 carry an incorrect author email
(kshitijkapoorr@gmail.com — not an address the contributor owns; it was
set by tooling error during salvage). The address maps to no GitHub
account (verified via the users search API). Canonicalize it to the real
identity so shortlog/contributor tooling attributes correctly; the
contributors/emails mapping file from the same PRs already covers
release attribution.
2026-08-26 16:17:34 +05:30
kshitijk4poor 365cbc242b test: assert omitted attach_to_session stays absent from formatted list output
Closes the gap flagged in review: the raw store was checked but not the
_format_job surface.
2026-08-26 16:06:41 +05:30
StanleyStetson 5e9adc9e4d fix(cron): forward attach_to_session through cronjob handler
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."

Fixes #84802
2026-08-26 16:06:41 +05:30
kshitijk4poor e513f3fb40 chore: map kshitijkapoorr@gmail.com to kshitijk4poor in contributors/emails 2026-08-26 16:06:09 +05:30
kshitijk4poor ded9470990 refactor(cron): fold simplify-review findings into mirror eligibility
- _target_mirror_eligible accepts a precomputed origin_match so the sole
  production caller stops re-resolving origin + re-running the origin
  match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
  precedence (standalone correctness: per-job False must beat raw global
  True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
  'not inchannel_continuable' and reword it for the widened seed scope.
2026-08-26 16:06:09 +05:30
kshitijk4poor b5e7842442 docs(cron): describe widened continuable scope (origin fallback + explicit opt-in)
The feature page still said 'only the origin chat is ever touched', which
this change makes stale. Documents the three attach-eligible shapes and
that the global flag never activates explicit targets (review nit).
2026-08-26 16:06:09 +05:30
kshitijk4poor 8313185449 fix(cron): unify in_channel flatten and seed behind one continuable gate
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
2026-08-26 16:06:09 +05:30
kshitijk4poor 2f425872ef test: extend exact-shape assertion in relay-delivery guard for provenance tag
tests/cron/test_cron_relay_delivery_guards.py landed on main after #89329
branched; its exact-dict assertion needs the new _resolved_from field the
salvaged commit adds to origin-resolved targets.
2026-08-26 16:06:09 +05:30
Victor Kyriazakos 580daa7b96 fix(cron): mirror continuable-cron briefs for origin-fallback and opted-in explicit targets
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).

The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.

Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
  origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
  gate: origin unchanged; origin_fallback eligible under the same flags
  as origin (per-job attach_to_session wins, else global
  cron.mirror_delivery); explicit platform:chat targets eligible ONLY
  under per-job attach_to_session — the global flag never activates
  them, so it cannot start writing transcript entries into arbitrary
  explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
  keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
  session keys are user-isolated, so a seed without a user_id (origin-
  less job into a shared channel) would create an orphan session no
  reply resolves to — those targets fall back to the plain mirror. DM
  targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.

Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.

15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
2026-08-26 16:06:09 +05:30
hermes-seaeye[bot] c427367938 fmt(js): npm run fix on merge (#95457)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 10:27:25 +00:00
Teknium 8e9459c97f style(desktop): sort registrySourceOwnsPrimaryBackend imports (lint) 2026-08-26 03:22:37 -07:00
Teknium 213f46a4e3 chore: map contributor email for attribution gate 2026-08-26 03:22:37 -07:00
Teknium 2947272233 refactor(desktop): one canonical write shape for connection_id row stamping
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.

Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).

Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
2026-08-26 03:22:37 -07:00
joe-rodgers fb393ee08b fix(desktop): stamp remote list rows with their owning connection; retry one transient projects.tree loss
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
  active non-local gateway are stamped with its registry connection_id
  (explicit owners from multi-source responses preserved), so a later
  resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
  switch leaves the first read RPC on a newly-opened socket without a
  response (request timed out / gateway connection closed), only while
  the same gateway/profile is still foreground. Component fix for the
  live-confirmed #92352 sidebar-never-paints gap.

Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).

Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)
2026-08-26 03:22:37 -07:00