Commit Graph

35797 Commits

Author SHA1 Message Date
teknium1 911620ed9f docs(tui): describe what happens when the gateway dies or the socket drops
User-visible copy for the spawned-gateway restart and the attached-mode
reconnect landed with the #111594 fix; document both paths.
2026-09-15 18:53:05 -07:00
teknium1 4ccbc2936d test(tui): events keep flowing and backoff grows across gateway reconnects
Invariant for #111594: after drain() on mount, every later transport
generation still emits gateway.ready live, and the reconnect delay grows
across consecutive failures instead of restarting at the base delay.
2026-09-15 18:53:05 -07:00
gustavosmendes 583dbb534c fix(tui): preserve sessions across gateway reconnects
An attached (dashboard-embedded) Ink TUI whose WebSocket dropped never
recovered even though the backend stayed alive, and a spawned gateway that
crashed resumed the wrong session.

- GatewayClient no longer resets `subscribed` on each transport generation:
  the renderer drain()s once on mount, so every post-reconnect event
  (gateway.ready included) stayed buffered forever.
- clearReconnect() keeps the attempt counter; it is reset on gateway.ready
  (and kill()), so backoff actually grows across failed reconnects.
- useMainApp's exit handler no longer calls start() for an attached socket
  closure — GatewayClient owns that reconnect; it only respawns a still-owned
  child, and plans the resume with the durable stored_session_id (what
  session.resume takes) instead of the process-local runtime sid.
- session.create's stored_session_id is carried into ui state / the active
  session file so recovery and the exit epilogue target the durable id.
- The recovery target is cleared only after resumeById resolves into a live
  sid, so a second disconnect during setup/history loading keeps it.
- Stale-socket identity guards on 'open'/'message'.

Salvaged from #111599 (@gustavosmendes) with trims: kept the client-side
backoff reconnect for spawned children too (the "keeps trying to reconnect
in the background" copy depends on it), kept the RPC-triggered reconnect
during backoff, kept the spawn-mode "reply in progress was lost" wording
(true for a dead child) and added attached-mode copy in userMessages.ts,
dropped the config_warning contract regen and the SessionCreateResponse
re-export (local type gains stored_session_id instead).

Fixes #111594
2026-09-15 18:53:05 -07:00
teknium1 6f0be923bc fix(tui_gateway): refuse a turn whose session has no agent instead of crashing the turn thread
`_admit_prompt_turn` is the gate every turn source crosses (prompt.submit,
auto-continue, heartbeat/loop ticks, bot mailbox, compute host). When the
record's agent is None (a deferred build that attached nothing, see previous
commit) it used to hand the None agent through; `_invoke_agent` raised
AttributeError, `_recover_turn_exception` emitted a frame, and then the
`finally` dereferenced `st.agent.interim_assistant_callback` again — the
turn thread died before `running = False`, so the prompt vanished and the
session stayed "busy" for every later prompt.

Now the gate releases `running`, emits the same retryable
`agent_init_failed` terminal frame `_run_after_agent_ready` uses (carrying
the recorded build reason when there is one), logs the refusal, and returns
None so `_run_prompt_submit` reports the turn as not started. The turn body
therefore never sees a None agent and needs no guard in its `finally`.

Live probe (real `_run_prompt_submit`, inline thread, agent None):
before: `CRASH out of _run_prompt_submit: AttributeError ... running: True`
after:  `returned: False ... error_surface {'layer': 'runtime', 'code':
'agent_init_failed', 'retryable': True} running: False`

Fixes #111531
Co-authored-by: zqy1-1 <185413950+zqy1-1@users.noreply.github.com>
2026-09-15 18:52:43 -07:00
zqy1-1 a944a4d8ae fix(tui_gateway): record why a deferred agent build attached nothing
A deferred build whose session record was closed or replaced while it ran
leaves through `_build` early, yet the `finally` still sets `agent_ready`
with `agent` None and `agent_error` None. Record the cause on the record so
the turn that was waiting on it can refuse with a real reason instead of
running against a missing agent (#111531).

Trimmed from PR #111534: the `_wait_agent_for_prompt` readiness hunk and the
turn-body guard are superseded by a single refusal at `_admit_prompt_turn`
(next commit), which every turn source crosses.
2026-09-15 18:52:43 -07:00
teknium1 88593f0f0d chore: map 0xmosta contributor email
Attribution for the cherry-picked boot-recovery focus-trap commit (#111857).
2026-09-15 18:52:15 -07:00
0xmosta 758638c64c fix(desktop): trap focus in boot recovery 2026-09-15 18:52:15 -07:00
teknium1 d747a71df2 test(desktop): type the fresh-module MCP health import without an import() annotation
The repo's eslint config forbids `typeof import(...)` type annotations
(@typescript-eslint/consistent-type-imports); derive the module type from the
dynamic-import thunk instead so `npm run check:lint` stays at 0 errors.
2026-09-15 18:51:49 -07:00
KoNit-K c9277d25b9 fix(desktop): honor MCP health snooze after restart 2026-09-15 18:51:49 -07:00
teknium1 b469be8cc3 fix(desktop): a push that lands mid-delivery claims its envelope now — lanes outlive the drain
The previous commit claims every outbox before delivering and runs
deliveries per target profile, but `drainBusy` still spanned the delivery
phase: a `bot_relay.outbox.pending` push that arrived while one lane ran a
long turn (up to RELAY_DELIVER_TIMEOUT_MS) only set `drainRerun`, and the new
envelope was claimed after that turn — the reporter's step 3 (bot C mails D
while A→B runs) still ended in `queued_expired`, because the gateway checks
the TTL at the claim.

Scope `drainBusy` to the claim phase and make the delivery lanes module
state: `relayLanes` maps `target_connection::target_profile` to the tail of
that target's in-flight deliveries, so a later drain appends to the running
lane (same target stays ordered, one turn at a time) or starts a new one
(other targets run now). Lane entries drop once idle; stopBotRelay clears
them so a restart begins fresh.

Tests: keep the contributor's red-on-base test (claims every outbox first,
delivers to different targets concurrently) and replace the ordering-only
test — green on base — with one that pins the mid-delivery claim plus the
same-target ordering (red on base AND on the previous commit alone).

Docs: bot-mode.md states the delivery concurrency contract.

Part of #111587 (with the previous commit: Fixes #111587)
2026-09-15 18:51:23 -07:00
John Paul Soliva 1eb771e2ff fix(desktop): the bot relay claims every outbox first and delivers per target, not one envelope at a time
drainRelayOutboxes drained one gateway's outbox and delivered its envelopes
before draining the next gateway, and delivered every envelope one after
another. One long turn — bot_relay.deliver may take up to
RELAY_DELIVER_TIMEOUT_MS, 25 minutes — therefore held every other bot's mail:
a sibling gateway's envelope sat unclaimed in its outbox, and since the
gateway checks the envelope's age against bot_mode.envelope_ttl_seconds
(15 minutes) at the claim, the sender's waiter received queued_expired for a
message nothing was wrong with; envelopes that did get claimed still waited
their turn behind unrelated deliveries, against a finite waiter.

Claim every gateway's outbox first, then deliver in lanes keyed by target
connection and profile: a lane runs its envelopes in order, one turn at a
time (the target gateway serialises that profile's turns behind its turn
lock anyway), and lanes run concurrently.
2026-09-15 18:51:23 -07:00
teknium1 81805a97ef fix(tui): drop a pending key-burst flush when an external value replaces the draft
The composer hands key bursts to the parent on a 16ms timer. When
history navigation, a slash completion or a submit clears/replaces the
draft while such a flush is still armed, the [value] effect reset the
local buffer correctly but left the timer running, so 16ms later the
stale burst was handed to the parent and overwrote the external value
(second hole in #111934).

Disarm the pending flush in the external branch of the [value] effect:
once the parent has replaced the draft, a burst typed against the old
draft can never be the newer value. Second invariant test covers it
alongside the stale own-echo case.
2026-09-15 18:50:50 -07:00
liuhao1024 c9fe50f39d fix(tui): keep stale own-echo flushes from rewinding composer keystrokes
A deferred key-burst flush can still be in flight when the parent's
re-render lands: the echoed value is the one we emitted, older than
vRef because the user typed past it. The [value] effect treated any
non-equal incoming value as an external assignment and rewound local
state — the cursor jumped backward and freshly typed letters were
overwritten (#111934).

Track the last value handed to onChange; an echo matching it stays on
the own-change path, and the pending flush for the newer local value
converges the parent on its next timer.
2026-09-15 18:50:50 -07:00
teknium1 672fd6b96f chore(desktop): drop a restating comment from the one-walk boot path
The WHY (one resolution instead of two) already lives on
resolveRendererIndexWithMissing; the call-site copy only repeated the code.
2026-09-15 18:50:24 -07:00
John Paul Soliva ab4bfda360 perf(desktop): resolve the renderer bundle once per window, not twice
createMainWindow walked the entire renderer generation twice before it
could call loadWindowUrl:

    const rendererIndex = DEV_SERVER ? null : resolveRendererIndex()
    const tornAssets = rendererIndex ? missingRendererAssets(rendererIndex) : []

resolveRendererIndex already computes exactly that list while choosing the
copy — it needs it to decide whether a copy is torn — and then throws it
away. missingRendererAssets is a BFS that readFileSync's every present
chunk whole and regex-scans it for the inline __vite__mapDeps table, so on
a release tree it is not a stat walk: measured against the real
apps/desktop/dist (252 chunks, 28.7 MiB of JS), one walk is 162
readFileSync calls reading 28.23 MiB, 576 existsSync calls, and 56.6 ms
median (min 55.8, 9 reps, warm page cache, darwin-arm64). Both walks run
synchronously on the main thread before the window gets its URL.

Return the list alongside the index. resolveRendererIndexWithMissing()
carries the existing body and hands back { index, missing }; the
path-only resolveRendererIndex() stays as a one-line wrapper so the nine
other call sites are untouched. The primary-window path takes one
resolution.

Semantics are unchanged in every branch: the same candidate is chosen, the
same log lines are emitted, and the missing list always describes the copy
actually returned. The all-copies-torn branch now reuses the first
candidate's list, captured on the first loop iteration, rather than
recomputing it for present[0] — recomputing there would have reintroduced
the second walk in exactly the case that matters most, and using the
loop's last value would have described a bundle we do not load.

Net effect on every primary-window boot: one fewer full walk, so 162 fewer
readFileSync calls, 28.23 MiB less synchronous reading, 288 fewer
existsSync calls, and ~57 ms of main-thread blocking removed before
loadURL. The win lands on packaged and --prod launches; DEV_SERVER skips
the walk entirely, so `vite dev` is unaffected.
2026-09-15 18:50:24 -07:00
teknium1 a22c731744 fix(desktop): refresh stale 4px scrollbar-width comments to 8px 2026-09-15 18:49:57 -07:00
teknium1 00c66225d4 fix(desktop): widen the portaled-menu scrollbar to match the app theme
`.dt-portal-scrollbar` is the same themed bar as `.scrollbar-dt`, applied
to overlays that portal under document.body (dropdown/context menus, the
command palette, the session and connection switchers). Widening only the
#root theme (#111634) would have left those lists on the 4px bar that was
too thin to grab; keep the two variants on one width (0.5rem = 8px).
2026-09-15 18:49:57 -07:00
kvnloo 22dc293f2a fix(desktop): widen themed scrollbars from 0.25rem to 0.5rem
The app-wide .scrollbar-dt theme (on #root) rendered 4px scrollbars
everywhere, including the conversation window, making them nearly
impossible to see or grab (#111634). Bump the webkit track size to 8px
so the thumb is actually hittable while staying a slim themed bar.

Fixes #111634
2026-09-15 18:49:57 -07:00
teknium1 be67e1d31e fix(tools): tolerate an untraversable HOME when probing ~/.local/bin
CI runs the suite as an unprivileged user with HOME=/root in one fixture;
Path.is_dir() raised PermissionError from _user_local_bin_entries and the
run-env builder crashed. An unreadable home has no usable ~/.local/bin, so
treat the OSError as absent.
2026-09-15 18:49:29 -07:00
teknium1 43e7e830fd fix(tools): fold ~/.local/bin into the POSIX PATH completion siblings, tests + docs
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.

Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.

Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.

Fixes #111778
2026-09-15 18:49:29 -07:00
KoNit-K c68e306ea4 fix(tools): include user local bin in POSIX PATH 2026-09-15 18:49:29 -07:00
teknium1 b74f158b0d fix(desktop): do not prune the desktop half of a package folder the app cannot read
The ghost-prune loop in reconcileUnifiedDesktopHalves used existsSync on
the marker's source, which answers false for EACCES/EPERM as well as
ENOENT, so a mode-000 / ACL-denied package folder was treated as an
uninstall and its materialized half rm -rf'd with no warning.
Distinguish a genuinely missing source (ENOENT/ENOTDIR) from one the app
is not allowed to stat: keep the half and warn. materializeDesktopHalf
now also warns on a non-ENOENT stat failure instead of swallowing it.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1 e3a90079b4 fix(plugins): one unreadable plugin child no longer aborts iter_plugin_dirs or memory-provider discovery
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1 b915405f92 fix(desktop): reconcile of unified plugin halves survives one unreadable package
`reconcileUnifiedDesktopHalves` let a stat/copy failure on one package
(`EPERM: lstat` on Windows in #111804; EACCES on a mode-000 file here)
reject the whole pass. The `hermes:fs:desktopPluginsRoot` IPC runs that
reconcile before returning the root, so the renderer never got a root and
every disk desktop plugin silently stopped loading. Warn about the one
package and keep materializing the siblings.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1 22716fab2a fix(plugins): one unreadable plugin dir no longer aborts loader or list discovery
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.

Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.

Part of #111804
2026-09-15 18:48:59 -07:00
KoNit-K 9af62e7ef0 fix(dashboard): skip unreadable plugin manifests 2026-09-15 18:48:59 -07:00
teknium1 0b48ae8d96 test(cli): split the incomplete-fleet-restart hint tests per host
Once the launchd file carried the macos_only marker, the first real macOS
lane run showed TestIncompleteWarningMentionsLaunchctl asserting the
systemd-host branch (kickstart hint, systemctl line) that never runs on
macOS. The two Linux-arm tests move to their own linux_only file; the
launchd file keeps the macOS arm (bootstrap/list hint, no systemctl).
2026-09-15 18:48:35 -07:00
teknium1 10faa3d7c7 test(cli): launchd fleet tests run on the real host, hermetic to a live fleet
With the macos_only marker the win32 skipif is redundant (the marker already
skips every non-darwin host) and the is_macos() fake in _wire contradicts the
repo rule that host-specific behaviour is tested on that host, not by making
the interpreter believe it is elsewhere. Drop both.

_get_service_pids(all_profiles=True) also runs a bare `launchctl list` prefix
scan, so on a macOS box with a live ai.hermes.gateway* fleet the scoping
assertions picked up real PIDs (the one failure the reporter could clear by
booting the gateway out). Stub subprocess.run in _wire so the tests assert on
routing, not on the developer machine.

The cherry-picked tests/ci change-detector (asserting one specific file is in
the macos_only list) is dropped: the selector test suite already covers the
mechanism, and the marker is now the file's declaration.
2026-09-15 18:48:35 -07:00
KoNit-K 8624e632fd fix(tests): select launchd restart tests on macOS 2026-09-15 18:48:35 -07:00
KoNit-K 6c05cbf4bd test(agent): make aux timeout FD test deterministic 2026-09-15 18:48:35 -07:00
teknium1 4179c5a99c fix: describe protected_instruction_files as an always-ask approval gate
The example comment said the guard refuses writes to instruction files;
the code (tools/file_tools_write_guards.py) always prompts a human, even
under --yolo, and refuses only when nobody can answer.
2026-09-15 18:46:54 -07:00
teknium1 9c0633b631 docs(cron): name idle-vs-wall-clock and the separate script timeout
The corrected invariant still read as if "inactive loops" were the thing
being bounded. Say what the watchdog measures (idle time, not wall-clock),
what that means for operators sizing work (an active job is never cut off),
and that attached scripts have their own bound (_DEFAULT_SCRIPT_TIMEOUT),
which is the second misreading #111633 calls out.
2026-09-15 18:46:54 -07:00
teknium1 462ca7c112 docs(config): explicit ollama_num_ctx is never capped by context_length
agent/agent_init.py::_configure_ollama_num_ctx caps only the auto-detected
value (the cap is skipped when an explicit override is set). The salvaged
example wording said context_length caps the explicit value too, which
would send readers chasing a cap that does not apply.
2026-09-15 18:46:54 -07:00
KoNit-K 6c9a665de2 docs(cron): clarify inactivity watchdog 2026-09-15 18:46:54 -07:00
DavidMetcalfe 3e833fd56f docs(browser): cover logins across scheduled and unattended runs
Real-profile browsing and cron's per-job toolsets are both documented, but
nothing connected them: browser.md never mentioned scheduled runs and cron.md
never mentioned login state, so the constraints of driving a login-gated site
from a cron job were only discoverable in source.

Adds a Scheduled and unattended runs subsection to the real-profile section:
the browser.use_real_profile prerequisite (off by default), the credentials a
login form or a fresh 2FA challenge needs saved ahead of time, the auth re-sync
behaviour and its Windows consequence. Plus a pointer from cron.md's toolset
section.
2026-09-15 18:46:54 -07:00
KoNit-K 30c0dce20e docs: add fail-loud boundary convention 2026-09-15 18:46:54 -07:00
Konstantin Khlopkov 0e7ffc641d docs(contributing): name the cause and the remediation in user-facing error messages 2026-09-15 18:46:54 -07:00
John Paul Soliva f20acf5d9f docs(config): align cli-config.yaml.example with the code defaults and document six missing keys
Two default claims in the example contradict the code:

- agent.verify_on_stop: the example says the default is "auto"; the default
  in config_defaults.py is False, agent/verification_stop.py treats "auto"
  as the explicit opt-in for the surface-aware mode, and the website
  already says off. The comment and the sample value now say so.
- display.show_reasoning: the example marks false as the default and ships
  show_reasoning: false; DEFAULT_CONFIG and cli.py's own defaults are True
  ("Default ON ... with this off the user stares at a spinner"). Anyone who
  copied the example silently turned reasoning display off. The marker and
  the sample value now match the code.

Six keys the runtime reads but the example never mentioned:

- security.protected_instruction_files / protected_instruction_extra_patterns
  (tools/file_tools.py) — a write guard for AGENTS.md/CLAUDE.md/SOUL.md and
  friends with no documentation anywhere.
- browser.restrict_evaluate / browser.allow_unsafe_evaluate — both in
  DEFAULT_CONFIG, only the former mentioned in the website browser page.
- tts.delivery_profiles.<platform>.{max_file_bytes,safety_ratio}
  (tools/tts_tool.py) — per-platform audio upload limits over the built-in
  discord/telegram/default table.
- gateway restart_after_turn_timeout (config_defaults.py, 1800) — the
  in-band restart knob its own sibling comment tells readers to prefer.
- model.ollama_num_ctx (agent/agent_init.py) — the documented-in-code VRAM
  cap for Ollama's num_ctx.

display.streaming was left alone on purpose: the CLI builds its config from
its own defaults (streaming: True) rather than DEFAULT_CONFIG (False), so
the example's "(default)" marker is correct for the surface it documents.
2026-09-15 18:46:54 -07:00
teknium1 1cfa892db1 fix(desktop): make the zsh probe test legs visible and run them on CI
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.

- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
  passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
  probe's grandchildren) in the SSH connection guide.
2026-09-15 18:46:32 -07:00
teknium1 2388d401c9 test(desktop): capability probe through a real zsh login shell (#111949)
The user-visible failure was the ownership capability probe reporting a
current remote as unsupported when sshd ran it under a non-interactive zsh.
Pin the probe itself, not only the wrapper: run remoteSupportsSshOwnership
through `zsh -c` against a fixture CLI that advertises both flags; skipped
where no zsh is installed. Fails on the pre-fix wrapper (empty capture).

Co-authored-by: the-repeter <50600051+the-repeter@users.noreply.github.com>
2026-09-15 18:46:32 -07:00
KoNit-K 1267d7fdd0 fix(desktop): support zsh SSH probe watchdogs 2026-09-15 18:46:32 -07:00
teknium1 d30c05bebf fix(desktop): blank-line padding in the foreground-dial test 2026-09-15 18:46:08 -07:00
teknium1 5eb0ed4531 fix(desktop): Vault settings tab dials its scoped profile foreground
requestGatewayForProfile always dialed with the 'background' spawn priority,
so the Vault tab under the Settings 'Applies to' selector kept the #111651
infinite-spinner path on a cold profile even after the REST side moved to
scopedDialPriority. Let requestGatewayForProfile take a spawnPriority option
and pass 'foreground' from the Vault panel; ambient callers are unchanged.

Part of #111651
2026-09-15 18:46:08 -07:00
teknium1 d3ba5b307a fix(desktop): one scoped-dial priority helper; Capabilities selector dials foreground too
Fold the salvaged per-call ternaries in api/config.ts into a single
scopedDialPriority(scope) helper on api/client.ts and apply it to the
Capabilities scope selector's cold-start reads (getSkills / getToolsets /
getMcpCatalog), which hit the same background-capped pool queue when the
selector targets a stopped profile.

Drop the salvaged main.ts source-text change-detector test; keep only the
string update the existing #90812 wiring test needs. The renderer seam test
(hermes-capability-scope) is the invariant: an explicit scope carries
priority 'foreground', the ambient path stays untagged.

Part of #111651 (salvage #111672)
2026-09-15 18:46:08 -07:00
KoNit-K d6ff2777b7 fix(desktop): prioritize scoped settings backend dials 2026-09-15 18:46:08 -07:00
teknium1 2134950069 fix(desktop): keep the original boot error when post-spawn cleanup cannot prove ownership
In connect()'s post-spawn catch, a cleanupStale rejection (indeterminate
ownership probe) replaced the boot failure's message and kind. Catch it,
attach it as error.cleanupCause, and rethrow the original error.
2026-09-15 18:45:47 -07:00
teknium1 e368c01e28 fix(desktop): a lost SSH probe answer no longer kills or orphans a live remote backend
The Desktop SSH bootstrap proves the remote `hermes serve --isolated` alive
(`kill -0 … && echo ALIVE || echo DEAD`) and owned (argv probe printing
OWNED/FOREIGN) over the same SSH channel that is often mid-teardown right
after the served token was resolved. Both probes treated ANY answer that was
not the positive sentinel as the negative verdict, so an exec that resolved
with empty output was read as death: the boot failed with "remote dashboard
exited while its served token was being resolved", the post-spawn cleanup ran
the ownership probe on the same channel, read the lost answer as FOREIGN,
skipped the kill and removed the lockfile — one orphaned ~144 MB backend per
failed attempt, ten in one session on a 1 GB host (#111810).

`execProbeVerdict` now runs both probes: an answer that is neither sentinel is
indeterminate and is retried over a short bounded window (3 attempts, 500 ms);
with no definite answer it fails closed with a `transient-transport-error`, so
`cleanupStale` keeps the ownership record for the next connect to reap instead
of leaving a lockless orphan. The post-spawn catch probes liveness first so a
child that genuinely died at startup skips the ownership proof and the
original error ("exited before announcing") stays visible.

Slimmer redo of #111828 by @kokhlo: same mechanism, without the parallel
`verifyRemotePidAlive`/`verifyPidOwnership`/`readOwnershipVerdict` layer that
left `remotePidAlive` and the original probe as dead duplicates.

Fixes #111810

Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:45:47 -07:00
teknium1 d4d9cf16d9 docs(desktop): record the partition-name invariant and the one-time re-sign-in
The salvaged comment promised the path component stays within
[A-Za-z0-9._-], but sanitizePartitionComponent() also emits the other
characters encodeURIComponent leaves alone (!~*'()). State the real
invariant — nothing Electron percent-escapes, pinned by the test as
"no ':' and no '%'" — and record the migration decision: one partition
name on every platform, so a non-primary cookie-auth remote signed in
on macOS/Linux under the old `:conn:` name is re-prompted once, and the
old `%3Aconn%3A` folder stays on disk, inert. apps/desktop/AGENTS.md gets
the rule so the next partition does not repeat #92183's folder name.
2026-09-15 18:45:13 -07:00
Peyton Lu 401a7d087a fix(desktop): keep cookie partitions colon-free so Windows jars work
Electron escapes ':' in a session partition name as '%3A' for the on-disk
profile folder. On Windows, a profile folder whose name contains '%3A' gets a
cookie store the network stack can neither read nor write: a jar placed there
reads back zero cookies, `cookies.set()` never reaches disk, and every request
goes out with no cookies at all — the gateway answers 401 `no_cookie` and the
desktop concludes the user is signed out.

Per-connection partitions (#92183) are the only ones carrying a colon beyond
the 'persist:' prefix, so every NON-primary cookie-auth remote hits this: the
dial mints no ticket, the reauth copy fires ("Remote Hermes gateway uses
OAuth, but you are not signed in…"), and the login window — riding the same
partition — writes into the same invisible jar. Nothing ever persists, so the
prompt returns on every connect. The registry primary and the v1 remote keep
the legacy `persist:hermes-remote-oauth` jar, whose folder has no '%3A', which
is why a single-gateway setup never shows the symptom.

Verified against the app's own Electron runtime (40.10.2, Windows): identical
jar bytes seeded into two partition folders read 3 cookies and mint a
ws-ticket (200) from `…-conn-…`, and 0 cookies / 401 from `…%3Aconn%3A…`.

Keep the partition path component inside [A-Za-z0-9._-]; a regression test
pins that it never contains ':' or '%'.
2026-09-15 18:45:13 -07:00
teknium1 1e2cb57973 fix(approval): session teardown and interrupted leaders withdraw the prompt instead of denying it
clear_session (/new, /reset, auto-reset boundary) stamped entry.result="deny"
before waking the wait, and an interrupted coalesced leader published the
same deny to its followers, so both still rendered outcome="denied" /
"denied by user". Carry the cause on the entry (entry.cancelled) and let
_cancel_cause map a result-less wake to a withdrawn prompt; the wait still
unwinds fail-closed and the leader's own decision is unchanged.
2026-09-15 18:44:46 -07:00