Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:
- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
poisoned state never does I/O; the NEXT caller pays once for a health
probe that clears the suspicion or forces a reconnect. A single
teardown-vs-keepalive race or auth-lock corruption can no longer park
a connection permanently — park stays reserved for genuinely
exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
in-flight RPC races a child-watcher task, so a dead subprocess fails
the call immediately with a retryable timeout instead of riding out
the full 300s. Deliberate teardown/reconnect also fails in-flight
calls now instead of leaving them attached to a dying transport.
Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.
Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.
Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).
Cover the null-path draft path, Home-scope createBackendSessionForSend,
and openNewSessionTile({ cwd: null }) so the last project folder cannot
leak back into a Home chat.
Home's "+" passes path/cwd null on purpose, but null was falsy and fell
through into resolveNewSessionCwd(), so "New session in Home" (especially
the openTab path while main chat is occupied) still created under the
previous project folder and showed its branch.
Browser tabs took the lowest free slot, so an id was reused once its tab
closed. That is only safe while every store keyed by the id is wiped on
close — true today, but a discipline rather than a guarantee, and stale
state would resurface under an unrelated tab the day it lapses.
Mint like a terminal does instead: no id is ever handed out twice.
Committing an address dropped the field back to the url of the page you
were leaving, so typing baby.com over google.com flashed google.com back
before baby.com arrived — and nothing said a load was underway.
The address you asked for now stays in the field until the page actually
lands somewhere (a redirect supersedes it, as it should), and progress
spins inside the field beside it. The pane owns that loading state from
the moment it accepts the address, because the reach probe it runs first
delays did-start-loading.
A URL tab used to be a singleton — every link navigated the one Browser,
so there was no way to keep a page open beside another and the strip's
"+" never appeared next to it.
A Browser tab is now a vessel with its own id: links still land in the
browser you are looking at (an agent opening five pages must not leave
five tabs behind), while the "+" mints another one on request. Tabs name
themselves after the page they are showing, since three tabs reading
"Browser" name nothing.
The "+" itself is now a pane capability rather than a session-only
button, so any pane kind that can make more of itself contributes one.
Move, ignore-mouse, placement, resize edges, snap, cursor feed, and
overlay promote all read the same Ozone-normalized capabilities instead
of re-deriving linux/Wayland/X11 at each call site.
The chord previously only acted when a terminal or preview selection
existed. A new bubble-phase window fallback now moves focus to the
composer on an unclaimed press, like the address-bar chord in a browser.
Existing owners keep priority: selection handlers claim the press on the
capture phase, a user-rebound action marks the event handled, and a
focused terminal with no selection keeps Ctrl+L as clear-screen via
composerFocusBlockedBySurface().
The chord matcher moves from the terminal feature to
src/lib/keybinds/chords.ts as isComposerChord: it now has three
consumers and the old name (isAddSelectionShortcut) was wrong at the
composer call site. The fixed panel row view.terminalSelection becomes
view.selectionToComposer because it covers preview selections too.
Omarchy tiles the HUD like any other toplevel, so always-on-top and
xdg_toplevel.move never apply. Ask Hyprland to float+pin after map,
trying classic dispatch then Lua for 0.55+ configs.
A start-marker probe failure (Get-Process timing out on a PowerShell 5.1
cold start, #87169) in claimBackendChild used to stop the freshly spawned
backend and rethrow — killing a healthy backend, triggering the renderer's
repair respawn, and looping. And because stderr piping only attached after
the claim, every before-ready failure surfaced as a bare exit code.
- extract probe + claim policy into electron/backend-claim.ts:
processStartMarker/execText (moved verbatim from main.ts), probeStartMarker,
and a pure claimDecision(childAlive, probe) a Windows CI lane can drive
with real PowerShell
- probe failure + LIVE child now degrades to PID-only identity
(pid-only:<pid> marker, WARNING logged), matching the existing
createParentStartMarkerResolver degrade pattern; processIdentityMatches
verifies degraded identities by PID liveness (command check still layers
on top in backendIdentityMatches)
- probe failure + DEAD child keeps the fail-closed throw, now carrying the
child's buffered stderr/stdout tail
- ring-buffered ~8KB output tail attached at spawn time in BOTH spawn paths
(pool + primary); tail appended to claim errors, before-ready exit
messages, and backend-ready's exited-before-port-announcement errors so
the real exit reason reaches desktop.log and the boot UI
- tests: claimDecision matrix (degrade test fails against the old
stop+throw behavior), real processStartMarker probe, ring-buffer caps,
and output-tail suffixes on backend-ready exit errors
A held port made 'hermes serve' print only uvicorn's bare
'ERROR: [Errno 98/10048] error while attempting to bind on address'
and exit 1 — indistinguishable from a broken backend for the desktop
spawn and wrapping scripts.
- Preflight bind probe (matching uvicorn's SO_REUSEADDR bind flags)
before uvicorn.Server; on conflict print machine-readable
'BACKEND_PORT_IN_USE port=<port>' + a human hint naming likely
holders, exit 75 (EX_TEMPFAIL — existing repo convention, see
gateway/restart.py, kanban_db.py).
- Probe-to-bind race covered: SystemExit(1) from uvicorn's own bind
failure is re-checked and translated on both POSIX and Windows
runner paths.
- --port 0 (ephemeral) short-circuits the probe: unchanged behavior.
- HERMES_BACKEND_READY contract untouched.
- Tests: real held-socket repro (sentinel + exit 75, sabotage-proven
to fail as bare exit 1 without the fix), free-port boot regression,
ephemeral-port regression, probe/classification units.
- Docs: port-conflict paragraph under 'hermes serve' in
reference/cli-commands.md.
A persisted tall/narrow size has no way out on Linux. Put a reset next
to Exit HUD so the default size (and position, where the compositor
allows it) is one click away.
Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>
Wayland clients cannot place themselves, so the JS setBounds drag is a
no-op there. Make the composer bar a -webkit-app-region drag handle on
Linux (input carved out with no-drag) and let the compositor move it.
Co-authored-by: Tony Simons <214744153+asimons81@users.noreply.github.com>
Both readers of the per-server MCP tool timeout (the connection's run()
and the cache-path registration) read config.get("timeout", 300) as
their own private resolution. Route them through _resolve_tool_timeout:
per-server mcp_servers.<name>.timeout still ALWAYS wins (most specific),
then timeouts.mcp.tool_call from the unified timeouts: section, then
the unchanged 300s default. Values pass through resolve_timeout's
platform clamp; resolution failure falls back to the historical default.
Default-behavior invariance pinned by contract tests (nothing
configured -> exactly 300, per-server beats section, section beats
default, invalid/failed resolution falls back).
The adapter's private thread-deadline helper was the ancestor of the
unified deadline layer's run_bounded_async (#85147 was extracted from
it, plus the caller-cancellation leak fix the original still lacked).
Consolidate: the helper body becomes a thin wrapper mapping
BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call
sites (the PTB retry ladder) expect. ~90 duplicated lines die, along
with the adapter-local copies of the abandon-cleanup runner and the
blocked-loop faulthandler diagnostics (both live in agent/deadline.py).
Everything the call sites rely on is preserved by the unified layer:
- thread-timer deadline that survives a blocked event loop (#63309)
- abandonment of cancellation-shielded tasks (PTB/httpcore anyio init)
- detached best-effort on_abandon cleanup (no httpx pool leak per retry)
- off-loop stack dump when the loop never processes the expiry
Plus one behavior IMPROVEMENT inherited from the shared copy: a caller
cancelling the wrapper no longer leaks the inner task unobserved (the
telegram original had that leak; the extraction fixed it).
test_telegram_init_deadline.py: the #63309 diagnostics probe now pins
the shared layer's dump hook (label "telegram-init") — same contract,
new seam. Wedge + cleanup-crash tests pass unchanged.
Simplify-pass follow-ups on the #87033 fix:
- _gateway_liveness_notice(plural=) authors both wording variants at one
site; removes the exact-substring .replace() that would silently no-op
if the create-path text is ever edited.
- Collapse the operator-precedence-trap conditional in list to a plain
'if jobs' — an empty list has nothing inert and now skips the probe.
- Fix docstring/code mismatch (builder returns gateway_running: True on
the happy path) and drop the dead try/except in
_warn_if_gateway_not_running (the helper never raises).
Follow-ups for the salvaged #93098:
- Move the tri-state liveness heuristic into hermes_cli.cron
(_builtin_gateway_liveness) so the CLI warning and the cronjob tool
share one implementation instead of two drifting copies.
- Surface gateway_running/warning on the list action too — an agent
inspecting jobs in a gateway-less environment has the same silent-
inert-job failure mode (#87033) as create. Empty lists stay quiet.
The builtin cron ticker only runs inside the gateway process. The CLI
surfaces this ('hermes cron list' / 'hermes cron status' both warn when
no gateway is running), but the model-facing cronjob tool returned a
clean success on create even with no gateway running - so the agent
confidently told the user a recurring task was scheduled while the job
could never fire.
Mirror the CLI's liveness heuristic in the tool's create path and attach
a tri-state gateway_running field to the result:
- true -> gateway running (or a non-builtin scheduler provider owns
firing, e.g. Chronos, which is exempt by design)
- false -> explicit warning telling the model the job is saved but will
NOT fire until the gateway starts, so it can relay that to
the user instead of reporting unqualified success
- null -> probe failed; claim neither way
Fixes#87033
the client half of the gateway.ping heartbeat contract (#89958); detects a silently-dropped socket via missed ping-acks and reconnects with bounded backoff; part of the #83166 recovery series.
The shared-client half of the gateway.ping heartbeat contract (#89958);
tracks lastInboundAt, sends pings, invalidates a silently-dead socket.
Part of #83166.
Additive WebSocket wire contract for a client-driven heartbeat.
The gateway.ready payload now advertises "heartbeat": True so clients can
discover the capability, and the WS read loop answers a gateway.ping request
with a {"ok": True} pong short-circuited BEFORE method dispatch (no method is
invoked). The WSTransport gains closed / last_inbound_at properties and a
mark_inbound() hook, updated on every inbound frame, for later liveness checks.
Backward-compatible in both directions: old clients never send gateway.ping,
and old servers simply never advertise the heartbeat flag. This is the first
slice of a WebSocket-recovery series; the follow-on slices consume this
contract (server-side transport rebind, TUI/desktop clients).
Receipts:
bash scripts/run_tests.sh tests/test_tui_gateway_ws.py -q
=> 1 file, 7 tests passed, 0 failed (100%) in 0.8s; exit 0
The Windows update hand-off shim was hardcoded to Microsoft Edge
(Find-EdgeExe), so machines whose default browser is Chrome still got
an Edge --app progress window, and every run leaked a throwaway
browser profile (hermes-update-ui-<pid>) under %TEMP% that was never
removed.
- Get-DefaultBrowserExe replaces Find-EdgeExe: resolves the OS default
browser from the UserChoice ProgId (https first, http fallback).
ChromeHTML -> Chrome, MSEdgeHTM -> Edge; any other ProgId returns
$null and degrades to the existing WinForms card.
- The dedicated --user-data-dir profile is now removed when the shim
closes, and stale hermes-update-ui-* leftovers from interrupted
runs are swept from %TEMP% in the same pass.
- --app + --user-data-dir is Chromium-only, so the whitelist is
intentionally limited to chrome/msedge; Edge keeps
--disable-features=msImplicitSignin to suppress the implicit MSA
sign-in that leaks into shim windows (#88410).
Review round (#86412):
- the except-import fallback returned the RAW oversized value, re-opening
the exact macOS time_t overflow this fix prevents; it now fails closed
to a finite ~1-year cap matching agent.deadline.MAX_SAFE_TIMEOUT_S
- clamp engagement logs a WARNING so operators see the semantic change
- new tests: float-form oversized value (YAML 1e18), warning emission,
and the import-failure fail-closed path (blocked-import probe proving
the result stays Lock.acquire-safe)
A group member turn is a session-scoped RPC sequence (resume → attach →
prompt.submit → poll) issued with the runtime id its first RPC minted, but
requestForBot routes every RPC through its own request-scoped socket lease
(retained:false secondaries in store/gateway). Between two RPCs the refcount
hits 0, the leased socket closes, the gateway detaches the runtime session on
WS disconnect, the orphan reaper frees it after grace, and the next RPC —
prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts
and the sub-profile bot goes silent in the room.
- store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted
hold on the pooled socket with an idempotent release, mirroring the
existing request-lease machinery.
- sdk: host.retainProfile(route) exposes the retain to plugins
(feature-detected by consumers; older hosts keep working).
- hermes-bots plugin: runGroupChatMemberTurn acquires the lease before
ensureGroupChatSession's first RPC and releases in finally, so the socket
that minted the runtime id stays open across attach+submit+poll; and
prompt.submit gets a one-shot catch-and-retry that re-resumes via the
STORED session id on 4001-class failures (belt-and-braces for routes the
lease can't cover). 4007 'never existed' keeps flowing to session.create.
Tests: simulated 4001 on first submit recovers via re-resume and delivers;
lease held across attach+submit (mock refcount never hits 0 mid-turn); lease
released after success AND failure; no-retainProfile host feature detection;
store-level retain/release + idempotent double-release + the unretained
disposal race.
Widen the salvaged gate-site clamp to the bug class. The gate min() from
the contributor PR capped the gate bound at 360s, which would break the
#79719 contract (gate must extend while a legitimate >360s approval
prompt is answerable) and left the sibling overflow sites live: the CLI
prompt thread.join, the gateway poll deadline, and human_wait_ceiling
all consume the same config value.
Clamp once in _get_approval_timeout() via agent.deadline.MAX_SAFE_TIMEOUT_S
(1 year - semantically unbounded, platform-safe). The gate keeps its
approvals.timeout-tracking behavior above 360s; 7 regression tests pin
lock-acquire/thread-join safety and the gate-extension contract.
Bilateral E2E: with approvals.timeout=1e20 in a real config.yaml, main
crashes every consumer with OverflowError; this branch survives all 5
probes.
The _ConcurrentToolAuthorizationGate uses threading.Lock.acquire(timeout=...)
where timeout comes from human_wait_ceiling() (approvals.timeout + 60s).
When approvals.timeout is set very large (e.g. 999999999999 to effectively
disable timeouts), this overflows macOS timespec and raises:
OverflowError: timestamp out of range for platform time_t
The gate only needs to serialize parallel dispatch — it should never wait
longer than the wedged-holder bound (_AUTHORIZATION_GATE_LOCK_TIMEOUT_S =
360s). Clamp the return value with min() so unbounded approvals.timeout
cannot overflow the lock acquire.
Fixes#83220
#93515 reports auto-speak reading each reply twice when the Edge TTS
streaming attempt falls back to the POST endpoint and the reply's
renderer id gets rewritten to its durable id mid-flight. That was true
before 63565fa26b, but resolveSpokenReply()'s ordinal-anchored dedupe
(landed 2026-08-19, five days before this issue was filed) already
follows the rewrite. No source change — this pins the behavior with a
regression test at the hook/store integration level, one layer above
the existing spoken-reply.ts unit tests.
forkBranch was unconditionally routing every branched session into the
main pane via resumeSession, including sidebar/background branches of
a session the user isn't currently viewing. That reintroduces the
#69750 focus-stealing bug for that path: branching a different session
from the sidebar yanked the active view away from whatever was open.
Only take over the main pane when the branch's parent is the session
already selected; otherwise keep opening it as its own tile.
forkBranch ended by opening the branch as a session-tile and leaving the
primary selection on the parent (#69750). In the default layout there is
no visible tile pane, so branching only added a sidebar row with no
feedback in the main area — and openSessionTile no-ops when the target
is already the selected session, the common case of branching the chat
you're viewing.
Load the branch as the primary session via resumeSession instead, which
reuses the runtime already warm-cached by forkBranch's
ensureSessionState/updateSessionState calls, so it doesn't cost an extra
resume RPC.
Fixes#93444
Salvage round on the Phase 1 skeleton (three findings from review):
1. The DELETE leg was silently upgraded to WAL on healthy SQLite:
pre-seeding the file via PRAGMA was undone by SessionDB.__init__'s
apply_wal_with_fallback(), which upgrades any non-WAL file whenever
the configured mode (default wal) says so — only WAL-reset-vulnerable
interpreters preserved DELETE, i.e. the leg tested the advertised
mode only where CI wasn't running. Each matrix leg now pins
database.journal_mode in an isolated HERMES_HOME for the child and
audits the ON-DISK mode after the run (effective_mode_or_skip):
a leg that ran in a different mode skips instead of double-counting.
2. Cell 2 exit-code conflation: a claimant crashing with an unhandled
exception exits 1 — indistinguishable from the clean "lost the
claim" exit(1), so one winner + seven crashes passed as consume-once
proof. Codes are now disjoint (0=won, 10=lost, anything else=crash).
3. Crashed-writer diagnostics: wait_for now fails immediately with the
child's stderr when the writer dies before reaching the kill window
(was: 60s opaque deadline, stderr discarded). spawn_child prepends
to an inherited PYTHONPATH instead of clobbering it.
Plus: dead `if False` scaffolding removed from cell 1's writer; README
matrix section corrected (cell 2 is default-mode-only by design — the
consume-once property rests on a single predicated UPDATE).
The unlayered *:focus-visible reset in styles.css intentionally zeroes
--tw-ring-shadow ('No focus rings, anywhere'), so any control that relied
solely on focus-visible:ring-* had no visible keyboard focus state at all.
Mirror each control's hover treatment as a focus-visible background/text
affordance instead, keeping the global reset intact:
- ui/sidebar.tsx: group label, group action, menu button, menu action,
menu sub-button get focus-visible:bg-sidebar-accent + accent foreground
- ui/tabs.tsx: TabsTrigger gets focus-visible:bg-background + text-foreground
- ui/text-tab.tsx: focus-visible:text-foreground (matches its hover)
- chat/composer/micro-actions.tsx: pill gets focus-visible chrome-action-hover
- right-sidebar/index.tsx HEADER_ACTION_CLASS: focus-visible sidebar-accent
- right-sidebar/terminal/rail.tsx RAIL_ACTION: focus-visible chrome-action-hover
- chat/sidebar/cron-jobs-section.tsx (row body + run rows): focus-visible
chrome-action-hover
- chat/sidebar/session-row.tsx <time>: focus-visible:text-foreground
Sweep verified: remaining focus-visible:ring-* usages under apps/desktop/src
already pair with a border/bg/text companion (button/checkbox/switch/input,
starmap share-controls) or are covered by PR #93460's row-hover work
(cron/index.tsx run rows).
Fixes#93462. Reported by @fred0m.
Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the
hindsight_retain tool schema, threaded into the retain item's timestamp
field. When absent, the item timestamp defaults to the configured event
clock (base from PR #82928 by @ragingbulld, authorship preserved) so the
Hindsight server can resolve relative time phrases; previously no item
timestamp was ever sent and temporal memories landed with null
occurred_start/occurred_end.
Fixes#93568. Salvages #82928.
Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.