Commit Graph

28271 Commits

Author SHA1 Message Date
Will Lynas b91845c2a0 fix(slack): honor unfurl controls for media captions 2026-08-27 15:33:32 +10:00
Will Lynas f7c87d9536 test(slack): cover default and split unfurl behavior 2026-08-27 15:33:32 +10:00
Andrew Bennett aeaa0c784e feat(slack): add link unfurl controls 2026-08-27 15:33:32 +10:00
Ben Barclay 4bdabb21ed fix(telemetry): renew the claim atomically before every POST
Seventh review found the claim-token fix incomplete, and its
reproduction is exact: the pre-POST check was READ-ONLY. A claimant
whose lease expired while suspended still passes it when it wakes
BEFORE anyone reclaims - its token is still in the row - and then a
second process legitimately reclaims while the first one's POST is in
flight. Both send. Reproduced at 60addb16e2: posts ['B', 'A'], both
reporting 'sent'. This is the check-to-POST expiry race, not the
documented mid-POST residual: A's lease was already dead before its
authority check passed.

The check is now an atomic RENEWAL (single CAS UPDATE): it requires the
token to match, the row to be pending, AND the current lease to be
unexpired, and only then extends next_attempt_at a fresh lease into the
future. rowcount == 1 is the only grant. A claimant that wakes past its
own lease fails the unexpired condition and yields even though its
token was never replaced - expiry alone means another process may
claim at any moment, so waking stale is disqualifying regardless of
whether anyone has taken the row yet. The renewed lease (300s) covers
the POST (30s timeout) with margin, and renewal runs before every
retry, not just the first attempt.

Regressions: the reviewer's exact ordering (expired wake before any
reclaim -> zero POSTs, row stays claimable), plus a healthy-claimant
renewal test. Mutation-checked: dropping the lease-unexpired condition
or the token condition each fails the suite.

The at-least-once scope note on _send_one stands: a suspension landing
mid-POST remains client-unfixable; the fixable window is now closed on
both sides (before the check, and between check and POST).

277 tests pass; ruff + footguns clean; staging E2E 202.
2026-08-27 15:21:30 +10:00
Ben Barclay 29033a3fd5 fix(relay): restore voice-note STT — wire media[] MIMEs, message_type "voice", and a User-Agent for CDN downloads (#95274)
* fix(relay): map wire media[] → event.media_types; accept message_type voice

A relayed voice note arrived as MessageType.AUDIO with media_types=[] —
the STT gate (_event_media_is_stt_input) excludes AUDIO unconditionally
and its per-attachment MIME rescue was unreachable, so STT never fired
and the agent fell back to the "user sent an audio file attachment"
context note (live-verified on staging 2026-08-26, Discord + Telegram).

Two wire-boundary fixes, both additive within contract_version 1:

- "voice" parses to MessageType.VOICE: the enum already had it — pinned
  by test so a future refactor can't collapse the two.
- media[] is now mapped into event.media_types (positional alignment
  with media_urls; mime-less entries keep their slot as ""). This is
  what run.py's per-attachment classifiers key off, so EVERY relayed
  attachment — image vs document, audio vs voice — now routes like its
  native-adapter equivalent, not just voice notes.

Behaviour pinned: new-connector voice → STT-eligible; legacy
audio-typed events unchanged (no STT); music uploads never STT-eligible
(direct _event_media_is_stt_input assertions on real wire-parsed
events, not mocks).

Pairs with the gateway-gateway PR that puts "voice" on the wire.

* review: pin the STT gate by test; fail safe on media/media_urls mismatch

Addresses independent review of #95274.

1. The PR's acceptance criterion is STT ROUTING, but no committed test
   called _event_media_is_stt_input — it was only asserted ad-hoc. Adds
   TestSttGate: voice→eligible, voice-without-media_types→eligible
   (the new-connector/old-gateway shape), legacy audio-typed voice
   note→not eligible, music→not eligible. Mutation-verified: removing
   the VOICE branch from the gate turns these RED.

2. media_urls and media[] are INDEPENDENT wire fields that consumers
   index by the same i. Mapping MIMEs positionally without checking
   agreement means a disagreeing producer misassociates a MIME with the
   wrong URL and mis-routes that attachment — strictly worse than no
   MIME, which degrades safely to message-level classification.
   _media_types_from_wire() now maps only when the lengths agree, warns
   and returns [] otherwise.

Note for the record: MessageType.VOICE predates this PR and the gate's
VOICE branch ignores media_types, so a NEW connector against an OLD
gateway ALREADY fires STT. That is desirable, but it is not "unchanged"
— the PR body's rollout matrix said otherwise and is corrected.

* fix(relay): send a User-Agent on relay media requests (Discord CDN 403)

Discord's CDN rejects urllib's default "Python-urllib/x.y" User-Agent
with HTTP 403, and RelayMediaClient never set one. Every Discord CDN
pass-through download therefore failed; _localize_inbound_media then
kept the raw URL (its "a public URL still has value" branch), and the
consumer tried to open a URL as a FILE PATH:

  WARNING gateway.relay.media: relay media download failed for
    https://cdn.discordapp.com/...voice-message.ogg: HTTP Error 403
  INFO gateway.run: Voice transcription failed for https://cdn.discord...
    : Audio file not found: https://cdn.discordapp.com/...

This killed ALL Discord relay media inbound — voice notes, images and
documents alike — not just the voice lane. Telegram/WhatsApp were
unaffected because their media is connector-re-hosted (/relay/media/{id},
fetched from our own host) and localizes to real /tmp paths.

Reproduced from a clean shell against a live CDN URL:
  curl (own UA)                -> 200
  urllib, no UA                -> 403 Forbidden
  urllib + descriptive UA      -> 200, 14583 bytes, OggS magic

Fix: a module-level _MEDIA_USER_AGENT sent on both download() and
upload(). upload() only ever targets our own connector so it was not
broken, but a single client should identify itself consistently.

Validated on staging: hot-patched hermes-agent-stg-test-6698, restarted
the gateway service, and Ben's Discord voice note transcribed
successfully — zero new 403s and zero new transcription failures after
the patch (last 403 predates it).

Test is mutation-verified: removing the UA from download() turns it RED
while the other five media tests stay green.

* fix(relay): keep url↔mime pairing through media localization

Addresses a blocking review finding on my own change: mapping media[]
into media_types created a POSITIONAL contract that the rest of the
inbound path then broke.

1. _localize_inbound_media (adapter.py) filtered media_urls without
   filtering media_types. Dropping a dead connector re-host is a NORMAL
   best-effort path, so every surviving attachment inherited its
   neighbour's mime. Reproduced through the real functions:

     before urls  [.../relay/media/dead, .../kept.png]
            types [application/pdf, image/png]
     after  urls  [.../kept.png]
            types [application/pdf, image/png]   <-- PNG reads as PDF
     _event_media_is_image(ev, 0) -> False

   The loop now carries (url, mime) as PAIRS, so a dropped URL drops its
   mime with it.

2. _media_types_from_wire compared LENGTHS only, which is not alignment:
   equal-length-but-reordered wire fields were accepted and paired
   wrongly, and an absent media_urls skipped the check entirely while
   still emitting types. Resolution is now BY URL (url -> mime lookup
   over media_urls); an unmatched URL degrades to "" and falls back to
   message-level classification.

Tests: 4 new cases driving the real chain (wire parse -> localization ->
run.py classifier), incl. the dropped-first-attachment case the existing
localization test could not catch (it builds events without
media_types). The obsolete length-mismatch test now asserts the stronger
by-url guarantee. Both fixes mutation-verified: reinstating the URL-only
filter fails 1 test, reverting to positional resolution fails 3.

Relay suite 258 passed; media/voice/stt selection 685 passed; ruff clean;
cross-repo integration payload re-verified.

* fix(relay): media_types is always one slot per media_url

Self-review after two review rounds flagged this bug class in adjacent
seams: I checked the function I edited, not every consumer of the
parallel arrays I created. Grepping ALL writers found a third instance.

merge_pending_message_event (gateway/platforms/base.py:2725-2735)
EXTENDS media_urls and media_types together when a second media message
merges into a pending one. My mapping could emit a POPULATED media_urls
with an EMPTY media_types (an older connector sends media_urls but no
media[]), so extend() concatenated lists of different lengths:

  A urls [old1.png, old2.png]  types []
  B urls [new.pdf]             types [application/pdf]
  merged urls  [old1.png, old2.png, new.pdf]
         types [application/pdf]
    -> old1.png reads as application/pdf; the real PDF gets ''

Fix: media_types is now ALWAYS len(media_urls), padded with '' — the
url-keyed lookup runs even when media[] is absent, and the localizer
rewrites the list unconditionally (no  short-circuit that
could leave a stale/short list behind).

Tests: 4 new cases — padding with no media[], the merge shift above
driven through the real merge_pending_message_event, localization
preserving the invariant while dropping an entry, and normalization of
a short/empty media_types arriving from a non-wire source. All
mutation-verified: removing the padding fails 4; restoring the
 guard fails 1.

Relay 262 passed; media/voice/stt selection 689 passed; ruff clean;
cross-repo integration payload re-verified.
2026-08-27 15:07:38 +10:00
Ben Barclay 60addb16e2 fix(telemetry): fence send authority on a per-claim token
Responds to the independent PR review (andrexibiza). Both P1s were
checked against current HEAD rather than taken on authority - the
review was written against 613849c190, before the interval-model
consent replacement landed.

P1-1 (same-UTC-day revoke/re-enable releases refused data): already
fixed by the interval model. The reviewer's exact reproduction - opt in
06:00, revoke 12:00, package collected 18:00, re-enable 20:00 same day
- was re-run at HEAD: the off-window package stays local, and a full-day
aggregate straddling the revocation boundary also stays local (period
containment, timestamp precision). The consent-windows harness already
pins both. The reviewer's related ask that consent-ledger persistence
failures fail closed also holds structurally now: reconciliation derives
state rather than recording transitions, so a lost write means a shorter
confirmed horizon - less is released, never more.

P1-2 (lease has no owner) was VALID at head. Reproduced exactly as
described: A claims, is suspended past the 300s lease, B reclaims and
POSTs, A resumes and POSTs again - and the ingest key is minute-
prefixed, so the duplicate lands as a DISTINCT stored object, making
this worse than a benign idempotent overwrite.

Fix: every claim now mints a claim_token (additive nullable column,
schema version unchanged). Ownership is revalidated immediately before
every external POST, and every settlement, rejection, and backoff write
is compare-and-set on (package_id, claim_token, pending). A lapsed
claimant that resumes yields without transmitting, and its stale
backoff cannot move next_attempt_at under the live claim's lease.

Two deterministic regressions ship with it: expiry -> reclaim -> resume
(the reviewer's schedule), and the subtler stale-backoff-clobber case.

Honest scope, documented on _send_one: delivery remains at-least-once.
The token closes the claim->POST gap; a suspension landing mid-POST
(bytes already on the wire) is not client-revocable. The residual
duplicate is byte-identical content; collapsing it fully needs
package_id-keyed dedupe at the ingest service.

275 tests pass; ruff + footguns clean; staging E2E 202.
2026-08-27 15:04:02 +10:00
hermes-seaeye[bot] 36b0a96dcb fmt(js): npm run fix on merge (#96076)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-27 04:43:58 +00:00
Ben Barclay 151632c333 Merge branch 'main' into feat/telemetry-exporter
387 commits from main; no conflicts (verified with merge-tree before
merging). Overlap limited to hermes_cli/config_defaults.py and
hermes_cli/setup.py, both auto-merged; all shared-metrics surfaces
untouched by main.
2026-08-27 14:42:41 +10:00
Teknium ff03d46eef test(desktop): method-aware gateway mock for the refresh-reconcile confirm interaction test
The reconcile-to-guarded-model interaction test's requestGateway mock
must only answer config.set with the confirm handshake — the panel's
model.options read rides the same dispatcher and was eating the
first mocked response.
2026-08-26 21:38:41 -07:00
墨綠BG 477053538c 🐛 fix(desktop): include route scope in gateway dial errors 2026-08-26 21:38:41 -07:00
1052326311 a8169f3e65 fix(desktop): preserve group turn reason codes 2026-08-26 21:38:41 -07:00
Teknium 61b6788dd4 fix(desktop): Bots-mode picker routes guarded model switches through the shared confirm handler
The Bots editor's model write (profiles.configure) was the one switch
surface that bypassed the data-policy / expensive-model selection guard:
a guarded pick (e.g. muse-spark contributor tier) was applied silently,
with no confirm flow anywhere — the #95293 remainder after the core
picker's confirm handshake landed in use-model-controls.

Gateway: profiles.configure now answers confirm_required +
confirm_message for a guarded model (same handshake as config.set
model) and writes NOTHING until the client resends with
confirm_expensive_model: true. Other sections still apply; the pending
model section is not reported as failed.

Desktop: the confirm flow is extracted out of use-model-controls into
one shared applier (lib/guarded-model-switch.ts, exported through the
plugin SDK) — warning toast, staleness-guarded Confirm, single
confirmed resend, never a retry loop. The core picker and the Bots
editor now consume the SAME handler; the Bots editor's Confirm resends
only the model section with confirm_expensive_model: true.

Fixes #95293 (Bots surface remainder).
2026-08-26 21:38:41 -07:00
rainbowgits 0fa98d6a1b fix(desktop): switch model after refresh when it leaves the catalog
Refresh Models only updated the catalog cache, so the composer kept showing a model that was no longer in the new group list.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-26 21:38:41 -07:00
Finn763 911978da62 fix(desktop): Bot Mode model picker always settles and stops remount churn (#95279)
The Bots model picker's catalog read rode the bot's own socket with no
deadline: a wedged dial left the query pending forever, so the picker
spun indefinitely. On top of that every fetch forced refresh:true,
bypassing the staleTime cache, so each Bots view remount (tab re-front,
dialog reopen, pane visibility flip) knocked the picker back into its
loading state and wiped the staged provider/model pick mid-edit.

Bound every attempt to 20s (rejection falls through to the picker's
existing free-text fallback), drop the forced refresh so the read
participates in the cache like every other surface's catalog, and pin
the contract with a red->green regression suite.
2026-08-26 21:38:41 -07:00
Teknium bc4eea7773 feat(desktop): Managed updates section drives per-connection SSH updates
Slim renderer UI for the managed SSH remote update engine (#95942),
adapted from #93042's renderer unit with the deferred canary/rollout
scope stripped. Adds a per-connection store (idle/updating/terminal
states, receipt, managed-update-in-progress busy envelope) and a
'Managed updates' section on the Gateways settings page with an Update
button, progress line, and correlated receipt per registered
Desktop-managed SSH connection. Fails closed when the Electron main
lacks connections.updateManaged.
2026-08-26 21:38:21 -07:00
Finn763 00988f3b89 fix(desktop): widen pool keepalive-fresh window to absorb WSL2 IPC stalls (#95189)
The renderer pings each pool backend every 60s (`hermes:backend:touch` →
`touchPoolBackend` → updates `lastActiveAt`). The LRU eviction cap used a
keepalive-fresh window of 90s — only 1.5× the ping cadence — to decide
whether a backend was "plausibly still alive". One missed or delayed ping
pushed a live backend past the threshold and the cap-driven eviction killed
the active profile's backend mid-session, restarting the gateway and
re-minting runtime ids. On WSL2, where the renderer→Electron IPC roundtrips
through 9p, brief 9p hiccups commonly stretch a ping to seconds of observed
silence, producing the ~80–90s exit / ~2 min cycle reported in #95189
(122 gateway starts on 2026-08-26 alone, driving renderer OOM via reconnect
churn at ~5GB/day).

Widen POOL_KEEPALIVE_FRESH_MS to 4 minutes (3× ping cadence + IPC stall
headroom, still bounded well below POOL_IDLE_MS=10min). Backends with one or
even two missed pings are now spared; truly idle backends (multiple lapses,
minutes idle) remain eligible for eviction by the cap and the idle reaper.
The constant is also overridable via HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS
to make this tunable without a rebuild.
2026-08-26 21:38:20 -07:00
Cursor Agent 8246c4f92a fix(cli): repair interrupted update fleet restart
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.

Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
2026-08-26 21:38:20 -07:00
lesseradmin 1341dfbd12 fix(compaction): exclude operational notifications from tail anchor and auto-focus (#92703)
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:

- _is_actionable_user_turn (tail anchor) only checked role/content, so a
  notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
  operational notices leaked into the compact focus hint.

Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.

Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.

Fixes #92703
2026-08-26 21:38:20 -07:00
rainbowgits 11cf59d9a5 fix(tui): adopt live compression config on the next Desktop/TUI turn
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-26 21:38:20 -07:00
fangliquanflq cb54576b1a fix(gateway): isolate control routes from default executor 2026-08-26 21:38:20 -07:00
Teknium 26e48dd8c3 chore: map contributor email for kvnloo (#95173) 2026-08-26 21:38:20 -07:00
Teknium 2c495acb83 test: flush the prune-lease dial by condition, not by hop count
#95343's withTimeout wrapper added an await hop to openSecondary's dial,
breaking the sibling test's exactly-two-Promise.resolve flush. Condition-
bounded flushing survives future hop-count changes.
2026-08-26 21:38:00 -07:00
686f6c61 21054b751e fix(desktop): hide unreachable same-name Bot Mode roster twins
Group rooms persist source-qualified members. After Desktop switches to
the built-in This-device source, a dead loopback row still listed next
to the live profile and looked like a second agent. Collapse only the
sidebar tiles; $lastRoster, group seats, and mentions keep every
(connectionId, profile) identity.
2026-08-26 21:38:00 -07:00
kim-miram ac2ec56833 fix(desktop): retain local startup profile across SSH switches 2026-08-26 21:38:00 -07:00
Teknium 0cefc491c8 test: widen the update-all scan window for the salvaged #95606 wiring pin
main's hermes:connections:update-all handler grew (renderer-side exclusions +
the managed-SSH dispatch branch) after #95606 branched, pushing the claim-
guarded dial past the 2,000-char slice.
2026-08-26 21:38:00 -07:00
nftpoetrist 1e73281614 fix(desktop): claim-guard every remaining ensureRegistryBackend()/ensureBackend() call in Electron main
(connectionId, profile) scope, but it was only wired into hermes:connection,
hermes:connection:for, and the power-resume pool rebuild. Five other real
call sites still invoke ensureRegistryBackend()/ensureBackend() directly:
the media-protocol connection resolver, the terminal-pane backend resolver,
the ~5s roster-enumeration probe, the connections update-all dispatch, and
dispatchRegistryApiRequest (every registry-scoped hermes:api REST call).

ensureRegistryBackend() has a genuine await-then-check race
(reuseMatchingPrimarySshBackend before the pool entry check-and-set), so any
of these five racing a guarded dial for the same scope can each bootstrap
their own SSH tunnel / remote dashboard for the same connection — exactly
the symptom class #90812 was written to prevent, just reached through an
unguarded path instead of two renderer windows.

Left the ensureRegistryBackend() self-call inside its own dispatch-time
health-probe reconnect branch (registryDispatchRevalidation) unguarded —
that recursive path has its own coordination semantics and touching it
without modeling reentrancy against the newly-guarded entry points here is
a separate, riskier change better done on its own.
2026-08-26 21:38:00 -07:00
Finn763 b1f133d4ea fix(desktop): defer gateway liveness force-close while a turn is in flight (#95327)
A wake-path liveness-probe timeout force-closed the primary renderer
socket even when the backend was merely busy mid-tool-call; the gateway
then saw its client vanish, ws_orphan_reap expired, and the running turn
died as a bare "Operation interrupted." placeholder.

While any session still reports working, the first inconclusive probe
timeout now defers the teardown behind one bounded re-probe; only an
exhausted consecutive-failure streak (or no in-flight work at all)
rebuilds the transport. The streak resets on a successful ping, a
healthy -32601 answer, a clean open, a gateway switch, and unmount.
2026-08-26 21:38:00 -07:00
nftpoetrist fbe51fa734 fix(desktop): bound the onActiveConnectionInvalidated fallback getConnection() call
The registry's active-connection-invalidated fallback re-dial was the one
getConnection() await in this file the #93454 bound-every-IPC-round-trip
sweep never reached. A wedged main-process round-trip during an eviction
fallback (idle reap, connection removal, profile delete) left $connection
latched on a promise that never settles instead of rejecting into the
existing catch/publish(null) path.
2026-08-26 21:38:00 -07:00
Teknium 2a41f9b635 fix: restore the default resolving dial mock in the salvaged #95343 probe test
The #92434 mid-handshake pin (added on main after #95343 branched) latches
gatewayMocks.connect on a never-resolving mockImplementation; vi.clearAllMocks()
clears calls, not implementations, so the salvaged wedged-probe test inherited
a dial that never completes and timed out.
2026-08-26 21:38:00 -07:00
nftpoetrist d81b801fb4 fix(desktop): bound getConnection()/resolveGatewayWsUrl() on every remaining route
20s withTimeout() on the boot() and soft-switch paths (use-gateway-boot.ts),
since these are IPC round-trips into the main process with no timeout of
their own — a wedged main-process round-trip hangs the awaiting caller
forever instead of surfacing a failure.

Every other production call site of the same IPC pair was still unbounded:

- store/gateway.ts's openSecondary() and sharedPrimaryRoute() — the actual
  connection-establishment underneath requestGatewayForProfile/Agent,
  ensureGatewayForProfile/Agent, and every other exported routing entry
  point that opens a non-primary profile's socket.
- use-gateway-request.ts's on-demand reconnect (the primary gateway's
  "not connected" retry path hit by every RPC).
- voice-playback.ts's resolveSpeakStreamUrl().
- api/plugins.ts's activeConnection() (pluginSocket's connect()).

Extracted RECONNECT_ATTEMPT_TIMEOUT_MS into the shared lib/with-timeout.ts
(previously local to use-gateway-boot.ts) so every call site uses the same
budget instead of duplicating the constant.

Regression tests mirror the existing use-gateway-boot.test.tsx hang-repro
pattern: wedge getConnection()/getConnectionFor() with a never-resolving
promise, advance fake timers past the 20s bound, assert the caller settles
instead of hanging. Mutation-verified: reverted the production fix (kept
tests) and confirmed the 6 new tests fail — 4 by genuinely timing out at the
vitest level, 2 by TypeError on the not-yet-exported activeConnection —
restored the fix and confirmed all 70 tests across the gateway/voice/boot/
plugins suites pass, with tsc -p . --noEmit clean throughout.
2026-08-26 21:38:00 -07:00
Teknium 4896cab0eb fix(tui): _init_session cwd hydration must not fall back to the launch DB
Sibling site of the same class fixed in the previous commit: a failed
profile-store open during _init_session fell back to _get_db(), hydrating
and persisting a named-profile session's cwd row against the launch
state.db. Fail closed instead — skip the hydration (log a warning) so
nothing ever reads or writes the wrong profile's store. The other
profile-store open sites (_db_for_profile, _ensure_session_db_row,
_session_db) already degrade to None/skip and were left as-is.
2026-08-26 21:37:32 -07:00
EndeavorYen 0a4d3abaf3 fix(tui): fail closed when a named profile's state.db won't open
A deferred agent build for a named-profile session swallowed a failed
profile-store open (except Exception -> session_db = None), so _make_agent
silently bound the launch _get_db() handle and every turn bled into the
wrong profile's state.db exactly when the profile store was briefly
unopenable. Opening the named profile then looked blank.

Route the open through _open_profile_session_db, which raises a clear
'profile session store unavailable' error instead; the deferred build's
existing except path turns that into agent_error + an error event, so the
user gets a clear failure and no agent turn against the wrong store.

Salvaged from #90219 (hardening half), adapted to main's current
deferred-build/_transfer_db_to_agent structure.

Related to #87723 and #89789. #88532 covered SessionStore only.
2026-08-26 21:37:32 -07:00
Shaq Wu 77b5bb39e7 fix(desktop): preserve bounded bot hydration budget 2026-08-26 21:37:19 -07:00
Ben Barclay 67d152bc7e fix(telemetry): bound forward-clock damage to the consent horizon
Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.

The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.

Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
  call. Honest heartbeats never bind it; a machine off for months
  catches up in a few hook fires (fail-closed latency only); one insane
  sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
  Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
  raw stamp lets an honest clock at revoke time pull a poisoned horizon
  back to the true revoke moment. A rolled-back clock at close time
  only closes earlier - fail-closed.

Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
  (the harness re-implemented the insert; deleting the production line
  survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
  skip was dead code - the store constructor creates the directory
  before the exists() check ran. The probe now checks the default path
  without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
  upgrade (fail-closed; deliberate).

New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.

273 tests pass; ruff and windows-footguns clean; staging E2E 202.
2026-08-27 13:30:25 +10:00
hermes-seaeye[bot] 791e2ae325 fmt(js): npm run fix on merge (#96017)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-27 02:30:11 +00:00
Teknium 824f7e081a chore: remove Windows real-profile PROOF workflow + live test
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
2026-08-26 19:25:33 -07:00
Teknium e4451ec6e5 feat(browser): close-with-approval flow for Windows real-profile (toggle arms, agent asks, blocked if still locked) [proof do-not-merge]
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
   browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
   always returns the [profile-locked] signal and the copy is refused. A later
   attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
    (new CLI subcommand) runs
   close_browser_holding_profile only when the agent has the user's OK. The
   locked error tells the agent to ask first, then run it, then retry.

- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
  message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py:  subcommand (identity+binding-verified
  tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.

Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
2026-08-26 19:25:33 -07:00
Teknium 00d5632249 chore: remove Windows real-profile PROOF workflow + live test
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 75f402c302 test(ci): assert cookie DB at either location + dump Default contents on miss [do-not-merge]
Auto-close live test asserted the legacy Default/Cookies path, but modern Chrome
writes Default/Network/Cookies. Accept either; on miss, print the copy's Default
listing so a real copy gap (vs a path-assertion bug) is visible.
2026-08-26 19:25:33 -07:00
Teknium 9e9e1b2245 feat(browser): consented auto-close of a running browser for Windows real-profile [proof workflow do-not-merge]
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.

- close_browser_holding_profile: graceful terminate → kill → poll until the
  cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
  unrelated same-name process on a different dir).
- Config key + docs admonition.

Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.

Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
2026-08-26 19:25:33 -07:00
Teknium b73f78714a chore: remove Windows real-profile PROOF workflow + live tests
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 931bf613b1 fix(browser): Windows real-profile fails fast when the browser is running
Live windows-latest proof settled it: a running Chrome opens its cookie DB
deny-all (even CreateFile with FILE_SHARE_READ|WRITE|DELETE fails; sqlite
mode=ro/immutable/nolock all 'unable to open'), so copy-while-running is
impossible on Windows without VSS/admin — and the prior code HUNG ~24min on the
locked file.

Fix: a fast up-front lock probe (_profile_is_locked: one open() of the active
profile's cookie DB; PermissionError = locked) runs BEFORE any copy in
snapshot_real_profile. If locked, bail immediately with 'fully quit the browser
(incl. background/tray) and retry, or turn browser.use_real_profile off'. Never
hangs, never a silent signed-out copy. POSIX has no mandatory locking so the
probe never trips there — copy-while-running still works on macOS/Linux.

Docs: admonition stating Windows needs the browser fully closed (background
apps included); the live-drive-while-running path is #95669.

Tests: lock-probe unit coverage (readable/no-db/PermissionError), snapshot
fails-fast-no-copytree when locked. Windows live E2E asserts the fast-fail
contract (returns <30s with the quit message) + the read-strategy diagnostic.
2026-08-26 19:25:33 -07:00
Teknium 52772de402 test(ci): run Windows diagnostic first + hard-bound the live test [do-not-merge]
Prior run hung 24min in the product-path test (snapshot_real_profile against a
locked profile blocks on Windows — itself a finding). Diagnostic now runs FIRST
(each strategy internally bounded, reports fast), live test second under a
faulthandler 150s dump-and-die so a hang can't burn the job. Job timeout 12min.
2026-08-26 19:25:33 -07:00
Teknium 92b95f16ce test(ci): proof workflow uses per-sha concurrency, no cancel-in-progress [do-not-merge]
Previous runs were auto-cancelling each other (ref-scoped group + cancel-in-progress). Per-sha group lets each proof run finish so the diagnostic actually reports.
2026-08-26 19:25:33 -07:00
Teknium bd504bee6d test(browser): PROOF diag — probe which read strategy beats Chrome's Windows lock [do-not-merge]
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
2026-08-26 19:25:33 -07:00
Teknium 1740abbe64 test(browser): PROOF round 2 — Windows locked-profile fails CLOSED (live-corrected)
The first live Windows run DISPROVED the sqlite-online-backup claim: Chrome's
share lock on Windows is strong enough that even a read-only SQLite open is
refused by the OS (raw-copy precondition fired, _copy_auth_file still returned
False). Copy-while-Chrome-runs is impossible on Windows — the earlier fix was
theatre that only passed on Linux (no mandatory locking).

Corrected contract, now asserted live: with a running Chrome holding the cookie
DB, snapshot_real_profile FAILS CLOSED with 'could not read ... login data
(N locked). Close <browser> and retry' — never a silent signed-out/torn copy.
Second test proves the supported path (Chrome closed) copies cleanly. So
real-profile browsing on Windows requires the browser closed; Linux/macOS
unaffected; live-drive-the-real-profile is tracked in #95669.

PROOF branch evidence only — workflow + test reverted before merge.
2026-08-26 19:25:33 -07:00
Teknium 2ecb18c4a7 test(browser): PROOF — Windows live E2E for locked-DB real-profile copy [do-not-merge]
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.

This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
2026-08-26 19:25:33 -07:00
Teknium 42046b452b fix(browser): copy real-profile auth DBs lock-aware (Windows 'file in use')
On Windows a running Chrome holds Cookies / Login Data / Web Data with an
exclusive lock, so the raw file copy the snapshot used raised WinError 32
('being used by another process') and the best-effort skip left a signed-out
copy — the reported profile-cloning failure.

Fix: copy the SQLite auth DBs via SQLite's online-backup API (read-only
connection + Connection.backup()), which reads a consistent COMMITTED snapshot
while the writer holds the lock. Non-DB files (Preferences, Local State) stay a
plain copy. If even the online-backup can't read a DB, snapshot_real_profile now
FAILS CLOSED with an actionable 'close <browser> and retry' message instead of
launching a silently signed-out session.

- _copy_auth_file: sqlite-backup for Cookies/Login Data/Web Data, raw copy
  otherwise, raw-copy fallback if backup fails.
- Drop -journal/-wal/-shm sidecars from the auth set + snapshot ignore: the
  backed-up DB is self-contained; a stale sidecar next to it corrupts it.
- Fresh copytree excludes the auth DBs (raw copytree of a locked file raises on
  Windows); they're always mirrored lock-aware afterward.
- _mirror_profile_auth returns the count of DBs it could not copy so the caller
  can fail closed.

Tests: locked-DB copied-via-backup (open write txn = live-lock analog, 42
committed rows, uncommitted excluded, no journal sidecar), _copy_auth_file DB vs
plain, fail-closed when unreadable. 192 browser tests pass. Live: 68 real
cookies copied through the backup path and the session launches.
2026-08-26 19:25:33 -07:00
Teknium 6e854595e2 fix(browser): real-profile review round 3 — overlay-ordering race, torn-copy marker, consent cleanup, active-only copy
Addresses the round-3 findings from @Adolanium + @kshitijk4poor on #95620:

1. Overlay-before-reuse race (blocker): _real_profile_cdp ran snapshot_real_profile
   BEFORE the session-reuse check, so a cold resolve that ends in reuse rewrote
   Cookies/Login Data under a live Chromium holding the user-data-dir open (torn
   DBs, locked txns, phantom logouts). Now: resolve copy dir as a PATH, probe
   reuse first, return early on a hit; snapshot/overlay only on the relaunch
   path when no live browser owns the dir.
2. Torn first copy poisoned freshness forever: freshness keyed on isdir(Default),
   so a half-written copy (disk full / Ctrl+C) was treated as populated and only
   ever got auth overlays. Now gated on a .hermes-snapshot-complete marker
   written only after a full copy succeeds; a torn copy is rebuilt from scratch.
3. Consent revocation left copied credentials on disk: turning use_real_profile
   off now deletes ~/.hermes/browser-profile/ on next browser use
   (cleanup_real_profile_snapshots), so cookies/logins don't outlive consent.
4. Stale non-active profile copies: only the ACTIVE profile (last_used) is copied
   into the copy's Default now — other Chrome profiles are never snapshotted
   (smaller copy, no stale credential dirs lingering).
5. Docs/config/desktop wording aligned to actual behavior (active-profile only,
   refresh on fresh session, consent-off cleanup).

Tests: overlay-skipped-on-reuse + overlay-runs-on-relaunch, done-marker gating +
torn-copy rebuild, active-only copy, consent-off cleanup (removes store +
idempotent + triggered from _real_profile_cdp). 206 browser + 222
backup/file_safety pass. Live: reuse skips re-snapshot; direct launch on the
active-only copy loads the real signed-in Gmail inbox.
2026-08-26 19:25:33 -07:00
Teknium f780cb36d8 refactor(computer_use): drop the cua_browser_* route — browser work goes through browser_exec
Real-profile browsing routes all in-page browser work through the Browser Use
CLI (browser_exec), which obsoletes the cua-driver typed-browser surface baked
into computer_use. Remove it so computer_use is a pure DESKTOP-control tool
(screenshots / mouse / keyboard / window management) and every call's schema
drops ~24 browser-only params + 9 actions.

- schema.py: 9 cua_browser_* actions and the typed-browser param block removed;
  14 desktop actions + shared params kept; description drops the browser rung.
- tool.py: cua_browser entries out of _SAFE/_DESTRUCTIVE_ACTIONS; the whole
  cua_browser dispatch block deleted; {"type","cua_browser_type"} → "type"
  (desktop typing untouched); _config_preauthorized (browser-prepare-only, a
  no-op for every desktop action) and the browser-page escalation hint removed.
- browser_route.py deleted (no importers outside the package); cua_backend.py
  drops the import + typed_browser_* methods; backend.py drops the non-abstract
  defaults.
- tests: browser-route/contract suites removed; browser assertions trimmed.

Desktop control unchanged. 233 computer_use tests pass; the 1 remaining failure
(test_gateway_session_key_yolo_maps_to_unrestricted_mode) is a pre-existing
cross-test state leak — fails identically on origin/main, passes in isolation.
2026-08-26 19:25:33 -07:00