Apply-ready delta distilled by @andrexibiza: deterministic watcher-consumer
tests (watcher times out while a child is alive, resolves when all are dead),
psutil-unavailable fail-open pin, and probe-failure fail-open handling in
_stdio_children_dead (unknown is never proof that every child exited).
Local: 8 passed on tests/tools/test_mcp_stdio_children_dead.py
_stdio_children_dead returned True ('all children dead') on the first LIVE
pid — the intended False was dead code right below it. Every spawn path
that captures child PIDs (observed in hermes -z oneshots) then failed the
#81995 pre-call fast-fail with 'TimeoutError: MCP stdio subprocess ... has
exited' on every tools/call while the subprocess was demonstrably alive.
Long-lived gateway/dashboard sessions were unaffected only when
_stdio_child_pids was empty (the not-pids short-circuit).
Return False on the first live pid and drop the unreachable line.
TestGetHermesHome.test_default_path asserted ~/.hermes unconditionally,
but the native Windows default is %LOCALAPPDATA%\hermes (see
hermes_constants._get_platform_default_hermes_home). Branch the
assertion by platform so the test passes everywhere.
Salvaged from PR #96003 by @Aoshi-Dev (the parse-guard half of that PR
was superseded by #96169); authorship preserved.
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):
- config_command: catch RuntimeError for set/unset and print a clean
one-line error + exit(1) instead of a raw traceback on the primary
'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
ConsoleCommandError so 'hermes console' and the dashboard console
report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
branch — the old fallthrough claimed 'falling back to default config'
even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.
Fail closed when config.yaml is unparseable or non-mapping before set/unset writes, reuse the readable-config guard to return the parsed mapping, and cover refuse/empty-mapping paths with regression tests.
Eighth review round (the first against the atomic-renewal fix) verdict:
the production code holds - CAS exclusivity across real processes,
lease-extension schedules, clock skew both directions, renew-per-attempt
under 5xx backoff, defer accounting, and the author's mutants all
verified - but one shipped regression test could not fail against the
property it is named for.
test_renewal_extends_the_lease_across_the_post asserted
next_attempt_at >= lease_before under a frozen clock. A renewal that
matches the row but never extends the lease (M4: SET next_attempt_at =
next_attempt_at) satisfies >= trivially, and that mutant double-POSTs:
the un-extended lease expires mid-POST and a second process reclaims.
The reviewer demonstrated M4 surviving the whole suite while producing
a real duplicate send in a two-process schedule.
The test now renews 100s into the lease from an advanced clock and
requires the deadline to move strictly forward to exactly
renewal-clock + 300s. Verified: M4 now fails this test (61 others
unaffected); clean HEAD passes all 62.
No production code change. 277 tests; ruff + footguns clean.
The send_media lane (08b95c3) had no committed regression test: cover
explicit-bool stamping on media frames, the descriptor-platform
fallback when _platform_by_chat is empty (post-restart proactive
sends), and the omitted-key absence case.
The send and send_media lanes resolved the platform only from
_platform_by_chat, which is empty until an inbound frame arrives (e.g.
after a gateway restart). A proactive send to a Slack chat then missed
the unfurl stamp. Mirror the streaming gate and delivery resolver:
fall back to the negotiated descriptor's platform.
Relay-plane parity: hermes config set / Railway persist YAML booleans
as strings, and _slack_unfurl_kwargs silently dropped them — so
'unfurl_links: "false"' was a no-op on native while working on relay.
Coerce recognized string booleans exactly as _slack_unfurl_hints does;
unrecognized values still drop so junk config keeps Slack's default
instead of accidentally suppressing previews.
Replaces test_send_ignores_non_boolean_unfurl_options (which froze the
dropped-string behavior) with coercion + junk-drop tests.
Live staging (Coatue Slack):
- hermes config set / Railway knobs persist "true" as a string; bots that
omit unfurl_links do NOT inherit the human default, so dropping the
string looked like suppression.
- chat.startStream cannot carry unfurl_*. Native SlackAdapter already
falls back to chat.postMessage; the relay now matches.
Relay-fronted Slack reads platforms.relay.extra.slack.unfurl_links/unfurl_media
and stamps explicit booleans onto the frame metadata; the connector forwards
them to chat.postMessage with no config of its own (mirrors reply_in_thread).
Covers send, send_for_platform (cron/scheduled), and send_media lanes.
Seventh review found the claim-token fix incomplete, and its
reproduction is exact: the pre-POST check was READ-ONLY. A claimant
whose lease expired while suspended still passes it when it wakes
BEFORE anyone reclaims - its token is still in the row - and then a
second process legitimately reclaims while the first one's POST is in
flight. Both send. Reproduced at 60addb16e2: posts ['B', 'A'], both
reporting 'sent'. This is the check-to-POST expiry race, not the
documented mid-POST residual: A's lease was already dead before its
authority check passed.
The check is now an atomic RENEWAL (single CAS UPDATE): it requires the
token to match, the row to be pending, AND the current lease to be
unexpired, and only then extends next_attempt_at a fresh lease into the
future. rowcount == 1 is the only grant. A claimant that wakes past its
own lease fails the unexpired condition and yields even though its
token was never replaced - expiry alone means another process may
claim at any moment, so waking stale is disqualifying regardless of
whether anyone has taken the row yet. The renewed lease (300s) covers
the POST (30s timeout) with margin, and renewal runs before every
retry, not just the first attempt.
Regressions: the reviewer's exact ordering (expired wake before any
reclaim -> zero POSTs, row stays claimable), plus a healthy-claimant
renewal test. Mutation-checked: dropping the lease-unexpired condition
or the token condition each fails the suite.
The at-least-once scope note on _send_one stands: a suspension landing
mid-POST remains client-unfixable; the fixable window is now closed on
both sides (before the check, and between check and POST).
277 tests pass; ruff + footguns clean; staging E2E 202.
* fix(relay): map wire media[] → event.media_types; accept message_type voice
A relayed voice note arrived as MessageType.AUDIO with media_types=[] —
the STT gate (_event_media_is_stt_input) excludes AUDIO unconditionally
and its per-attachment MIME rescue was unreachable, so STT never fired
and the agent fell back to the "user sent an audio file attachment"
context note (live-verified on staging 2026-08-26, Discord + Telegram).
Two wire-boundary fixes, both additive within contract_version 1:
- "voice" parses to MessageType.VOICE: the enum already had it — pinned
by test so a future refactor can't collapse the two.
- media[] is now mapped into event.media_types (positional alignment
with media_urls; mime-less entries keep their slot as ""). This is
what run.py's per-attachment classifiers key off, so EVERY relayed
attachment — image vs document, audio vs voice — now routes like its
native-adapter equivalent, not just voice notes.
Behaviour pinned: new-connector voice → STT-eligible; legacy
audio-typed events unchanged (no STT); music uploads never STT-eligible
(direct _event_media_is_stt_input assertions on real wire-parsed
events, not mocks).
Pairs with the gateway-gateway PR that puts "voice" on the wire.
* review: pin the STT gate by test; fail safe on media/media_urls mismatch
Addresses independent review of #95274.
1. The PR's acceptance criterion is STT ROUTING, but no committed test
called _event_media_is_stt_input — it was only asserted ad-hoc. Adds
TestSttGate: voice→eligible, voice-without-media_types→eligible
(the new-connector/old-gateway shape), legacy audio-typed voice
note→not eligible, music→not eligible. Mutation-verified: removing
the VOICE branch from the gate turns these RED.
2. media_urls and media[] are INDEPENDENT wire fields that consumers
index by the same i. Mapping MIMEs positionally without checking
agreement means a disagreeing producer misassociates a MIME with the
wrong URL and mis-routes that attachment — strictly worse than no
MIME, which degrades safely to message-level classification.
_media_types_from_wire() now maps only when the lengths agree, warns
and returns [] otherwise.
Note for the record: MessageType.VOICE predates this PR and the gate's
VOICE branch ignores media_types, so a NEW connector against an OLD
gateway ALREADY fires STT. That is desirable, but it is not "unchanged"
— the PR body's rollout matrix said otherwise and is corrected.
* fix(relay): send a User-Agent on relay media requests (Discord CDN 403)
Discord's CDN rejects urllib's default "Python-urllib/x.y" User-Agent
with HTTP 403, and RelayMediaClient never set one. Every Discord CDN
pass-through download therefore failed; _localize_inbound_media then
kept the raw URL (its "a public URL still has value" branch), and the
consumer tried to open a URL as a FILE PATH:
WARNING gateway.relay.media: relay media download failed for
https://cdn.discordapp.com/...voice-message.ogg: HTTP Error 403
INFO gateway.run: Voice transcription failed for https://cdn.discord...
: Audio file not found: https://cdn.discordapp.com/...
This killed ALL Discord relay media inbound — voice notes, images and
documents alike — not just the voice lane. Telegram/WhatsApp were
unaffected because their media is connector-re-hosted (/relay/media/{id},
fetched from our own host) and localizes to real /tmp paths.
Reproduced from a clean shell against a live CDN URL:
curl (own UA) -> 200
urllib, no UA -> 403 Forbidden
urllib + descriptive UA -> 200, 14583 bytes, OggS magic
Fix: a module-level _MEDIA_USER_AGENT sent on both download() and
upload(). upload() only ever targets our own connector so it was not
broken, but a single client should identify itself consistently.
Validated on staging: hot-patched hermes-agent-stg-test-6698, restarted
the gateway service, and Ben's Discord voice note transcribed
successfully — zero new 403s and zero new transcription failures after
the patch (last 403 predates it).
Test is mutation-verified: removing the UA from download() turns it RED
while the other five media tests stay green.
* fix(relay): keep url↔mime pairing through media localization
Addresses a blocking review finding on my own change: mapping media[]
into media_types created a POSITIONAL contract that the rest of the
inbound path then broke.
1. _localize_inbound_media (adapter.py) filtered media_urls without
filtering media_types. Dropping a dead connector re-host is a NORMAL
best-effort path, so every surviving attachment inherited its
neighbour's mime. Reproduced through the real functions:
before urls [.../relay/media/dead, .../kept.png]
types [application/pdf, image/png]
after urls [.../kept.png]
types [application/pdf, image/png] <-- PNG reads as PDF
_event_media_is_image(ev, 0) -> False
The loop now carries (url, mime) as PAIRS, so a dropped URL drops its
mime with it.
2. _media_types_from_wire compared LENGTHS only, which is not alignment:
equal-length-but-reordered wire fields were accepted and paired
wrongly, and an absent media_urls skipped the check entirely while
still emitting types. Resolution is now BY URL (url -> mime lookup
over media_urls); an unmatched URL degrades to "" and falls back to
message-level classification.
Tests: 4 new cases driving the real chain (wire parse -> localization ->
run.py classifier), incl. the dropped-first-attachment case the existing
localization test could not catch (it builds events without
media_types). The obsolete length-mismatch test now asserts the stronger
by-url guarantee. Both fixes mutation-verified: reinstating the URL-only
filter fails 1 test, reverting to positional resolution fails 3.
Relay suite 258 passed; media/voice/stt selection 685 passed; ruff clean;
cross-repo integration payload re-verified.
* fix(relay): media_types is always one slot per media_url
Self-review after two review rounds flagged this bug class in adjacent
seams: I checked the function I edited, not every consumer of the
parallel arrays I created. Grepping ALL writers found a third instance.
merge_pending_message_event (gateway/platforms/base.py:2725-2735)
EXTENDS media_urls and media_types together when a second media message
merges into a pending one. My mapping could emit a POPULATED media_urls
with an EMPTY media_types (an older connector sends media_urls but no
media[]), so extend() concatenated lists of different lengths:
A urls [old1.png, old2.png] types []
B urls [new.pdf] types [application/pdf]
merged urls [old1.png, old2.png, new.pdf]
types [application/pdf]
-> old1.png reads as application/pdf; the real PDF gets ''
Fix: media_types is now ALWAYS len(media_urls), padded with '' — the
url-keyed lookup runs even when media[] is absent, and the localizer
rewrites the list unconditionally (no short-circuit that
could leave a stale/short list behind).
Tests: 4 new cases — padding with no media[], the merge shift above
driven through the real merge_pending_message_event, localization
preserving the invariant while dropping an entry, and normalization of
a short/empty media_types arriving from a non-wire source. All
mutation-verified: removing the padding fails 4; restoring the
guard fails 1.
Relay 262 passed; media/voice/stt selection 689 passed; ruff clean;
cross-repo integration payload re-verified.
Responds to the independent PR review (andrexibiza). Both P1s were
checked against current HEAD rather than taken on authority - the
review was written against 613849c190, before the interval-model
consent replacement landed.
P1-1 (same-UTC-day revoke/re-enable releases refused data): already
fixed by the interval model. The reviewer's exact reproduction - opt in
06:00, revoke 12:00, package collected 18:00, re-enable 20:00 same day
- was re-run at HEAD: the off-window package stays local, and a full-day
aggregate straddling the revocation boundary also stays local (period
containment, timestamp precision). The consent-windows harness already
pins both. The reviewer's related ask that consent-ledger persistence
failures fail closed also holds structurally now: reconciliation derives
state rather than recording transitions, so a lost write means a shorter
confirmed horizon - less is released, never more.
P1-2 (lease has no owner) was VALID at head. Reproduced exactly as
described: A claims, is suspended past the 300s lease, B reclaims and
POSTs, A resumes and POSTs again - and the ingest key is minute-
prefixed, so the duplicate lands as a DISTINCT stored object, making
this worse than a benign idempotent overwrite.
Fix: every claim now mints a claim_token (additive nullable column,
schema version unchanged). Ownership is revalidated immediately before
every external POST, and every settlement, rejection, and backoff write
is compare-and-set on (package_id, claim_token, pending). A lapsed
claimant that resumes yields without transmitting, and its stale
backoff cannot move next_attempt_at under the live claim's lease.
Two deterministic regressions ship with it: expiry -> reclaim -> resume
(the reviewer's schedule), and the subtler stale-backoff-clobber case.
Honest scope, documented on _send_one: delivery remains at-least-once.
The token closes the claim->POST gap; a suspension landing mid-POST
(bytes already on the wire) is not client-revocable. The residual
duplicate is byte-identical content; collapsing it fully needs
package_id-keyed dedupe at the ingest service.
275 tests pass; ruff + footguns clean; staging E2E 202.
387 commits from main; no conflicts (verified with merge-tree before
merging). Overlap limited to hermes_cli/config_defaults.py and
hermes_cli/setup.py, both auto-merged; all shared-metrics surfaces
untouched by main.
The Bots editor's model write (profiles.configure) was the one switch
surface that bypassed the data-policy / expensive-model selection guard:
a guarded pick (e.g. muse-spark contributor tier) was applied silently,
with no confirm flow anywhere — the #95293 remainder after the core
picker's confirm handshake landed in use-model-controls.
Gateway: profiles.configure now answers confirm_required +
confirm_message for a guarded model (same handshake as config.set
model) and writes NOTHING until the client resends with
confirm_expensive_model: true. Other sections still apply; the pending
model section is not reported as failed.
Desktop: the confirm flow is extracted out of use-model-controls into
one shared applier (lib/guarded-model-switch.ts, exported through the
plugin SDK) — warning toast, staleness-guarded Confirm, single
confirmed resend, never a retry loop. The core picker and the Bots
editor now consume the SAME handler; the Bots editor's Confirm resends
only the model section with confirm_expensive_model: true.
Fixes#95293 (Bots surface remainder).
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:
- _is_actionable_user_turn (tail anchor) only checked role/content, so a
notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
operational notices leaked into the compact focus hint.
Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.
Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.
Fixes#92703
Sibling site of the same class fixed in the previous commit: a failed
profile-store open during _init_session fell back to _get_db(), hydrating
and persisting a named-profile session's cwd row against the launch
state.db. Fail closed instead — skip the hydration (log a warning) so
nothing ever reads or writes the wrong profile's store. The other
profile-store open sites (_db_for_profile, _ensure_session_db_row,
_session_db) already degrade to None/skip and were left as-is.
A deferred agent build for a named-profile session swallowed a failed
profile-store open (except Exception -> session_db = None), so _make_agent
silently bound the launch _get_db() handle and every turn bled into the
wrong profile's state.db exactly when the profile store was briefly
unopenable. Opening the named profile then looked blank.
Route the open through _open_profile_session_db, which raises a clear
'profile session store unavailable' error instead; the deferred build's
existing except path turns that into agent_error + an error event, so the
user gets a clear failure and no agent turn against the wrong store.
Salvaged from #90219 (hardening half), adapted to main's current
deferred-build/_transfer_db_to_agent structure.
Related to #87723 and #89789. #88532 covered SessionStore only.
Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.
The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.
Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
call. Honest heartbeats never bind it; a machine off for months
catches up in a few hook fires (fail-closed latency only); one insane
sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
raw stamp lets an honest clock at revoke time pull a poisoned horizon
back to the true revoke moment. A rolled-back clock at close time
only closes earlier - fail-closed.
Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
(the harness re-implemented the insert; deleting the production line
survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
skip was dead code - the store constructor creates the directory
before the exists() check ran. The probe now checks the default path
without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
upgrade (fail-closed; deliberate).
New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.
273 tests pass; ruff and windows-footguns clean; staging E2E 202.
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
Auto-close live test asserted the legacy Default/Cookies path, but modern Chrome
writes Default/Network/Cookies. Accept either; on miss, print the copy's Default
listing so a real copy gap (vs a path-assertion bug) is visible.
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.
- close_browser_holding_profile: graceful terminate → kill → poll until the
cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
unrelated same-name process on a different dir).
- Config key + docs admonition.
Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.
Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
Live windows-latest proof settled it: a running Chrome opens its cookie DB
deny-all (even CreateFile with FILE_SHARE_READ|WRITE|DELETE fails; sqlite
mode=ro/immutable/nolock all 'unable to open'), so copy-while-running is
impossible on Windows without VSS/admin — and the prior code HUNG ~24min on the
locked file.
Fix: a fast up-front lock probe (_profile_is_locked: one open() of the active
profile's cookie DB; PermissionError = locked) runs BEFORE any copy in
snapshot_real_profile. If locked, bail immediately with 'fully quit the browser
(incl. background/tray) and retry, or turn browser.use_real_profile off'. Never
hangs, never a silent signed-out copy. POSIX has no mandatory locking so the
probe never trips there — copy-while-running still works on macOS/Linux.
Docs: admonition stating Windows needs the browser fully closed (background
apps included); the live-drive-while-running path is #95669.
Tests: lock-probe unit coverage (readable/no-db/PermissionError), snapshot
fails-fast-no-copytree when locked. Windows live E2E asserts the fast-fail
contract (returns <30s with the quit message) + the read-strategy diagnostic.
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
The first live Windows run DISPROVED the sqlite-online-backup claim: Chrome's
share lock on Windows is strong enough that even a read-only SQLite open is
refused by the OS (raw-copy precondition fired, _copy_auth_file still returned
False). Copy-while-Chrome-runs is impossible on Windows — the earlier fix was
theatre that only passed on Linux (no mandatory locking).
Corrected contract, now asserted live: with a running Chrome holding the cookie
DB, snapshot_real_profile FAILS CLOSED with 'could not read ... login data
(N locked). Close <browser> and retry' — never a silent signed-out/torn copy.
Second test proves the supported path (Chrome closed) copies cleanly. So
real-profile browsing on Windows requires the browser closed; Linux/macOS
unaffected; live-drive-the-real-profile is tracked in #95669.
PROOF branch evidence only — workflow + test reverted before merge.
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.
This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
On Windows a running Chrome holds Cookies / Login Data / Web Data with an
exclusive lock, so the raw file copy the snapshot used raised WinError 32
('being used by another process') and the best-effort skip left a signed-out
copy — the reported profile-cloning failure.
Fix: copy the SQLite auth DBs via SQLite's online-backup API (read-only
connection + Connection.backup()), which reads a consistent COMMITTED snapshot
while the writer holds the lock. Non-DB files (Preferences, Local State) stay a
plain copy. If even the online-backup can't read a DB, snapshot_real_profile now
FAILS CLOSED with an actionable 'close <browser> and retry' message instead of
launching a silently signed-out session.
- _copy_auth_file: sqlite-backup for Cookies/Login Data/Web Data, raw copy
otherwise, raw-copy fallback if backup fails.
- Drop -journal/-wal/-shm sidecars from the auth set + snapshot ignore: the
backed-up DB is self-contained; a stale sidecar next to it corrupts it.
- Fresh copytree excludes the auth DBs (raw copytree of a locked file raises on
Windows); they're always mirrored lock-aware afterward.
- _mirror_profile_auth returns the count of DBs it could not copy so the caller
can fail closed.
Tests: locked-DB copied-via-backup (open write txn = live-lock analog, 42
committed rows, uncommitted excluded, no journal sidecar), _copy_auth_file DB vs
plain, fail-closed when unreadable. 192 browser tests pass. Live: 68 real
cookies copied through the backup path and the session launches.
Addresses the round-3 findings from @Adolanium + @kshitijk4poor on #95620:
1. Overlay-before-reuse race (blocker): _real_profile_cdp ran snapshot_real_profile
BEFORE the session-reuse check, so a cold resolve that ends in reuse rewrote
Cookies/Login Data under a live Chromium holding the user-data-dir open (torn
DBs, locked txns, phantom logouts). Now: resolve copy dir as a PATH, probe
reuse first, return early on a hit; snapshot/overlay only on the relaunch
path when no live browser owns the dir.
2. Torn first copy poisoned freshness forever: freshness keyed on isdir(Default),
so a half-written copy (disk full / Ctrl+C) was treated as populated and only
ever got auth overlays. Now gated on a .hermes-snapshot-complete marker
written only after a full copy succeeds; a torn copy is rebuilt from scratch.
3. Consent revocation left copied credentials on disk: turning use_real_profile
off now deletes ~/.hermes/browser-profile/ on next browser use
(cleanup_real_profile_snapshots), so cookies/logins don't outlive consent.
4. Stale non-active profile copies: only the ACTIVE profile (last_used) is copied
into the copy's Default now — other Chrome profiles are never snapshotted
(smaller copy, no stale credential dirs lingering).
5. Docs/config/desktop wording aligned to actual behavior (active-profile only,
refresh on fresh session, consent-off cleanup).
Tests: overlay-skipped-on-reuse + overlay-runs-on-relaunch, done-marker gating +
torn-copy rebuild, active-only copy, consent-off cleanup (removes store +
idempotent + triggered from _real_profile_cdp). 206 browser + 222
backup/file_safety pass. Live: reuse skips re-snapshot; direct launch on the
active-only copy loads the real signed-in Gmail inbox.
Real-profile browsing routes all in-page browser work through the Browser Use
CLI (browser_exec), which obsoletes the cua-driver typed-browser surface baked
into computer_use. Remove it so computer_use is a pure DESKTOP-control tool
(screenshots / mouse / keyboard / window management) and every call's schema
drops ~24 browser-only params + 9 actions.
- schema.py: 9 cua_browser_* actions and the typed-browser param block removed;
14 desktop actions + shared params kept; description drops the browser rung.
- tool.py: cua_browser entries out of _SAFE/_DESTRUCTIVE_ACTIONS; the whole
cua_browser dispatch block deleted; {"type","cua_browser_type"} → "type"
(desktop typing untouched); _config_preauthorized (browser-prepare-only, a
no-op for every desktop action) and the browser-page escalation hint removed.
- browser_route.py deleted (no importers outside the package); cua_backend.py
drops the import + typed_browser_* methods; backend.py drops the non-abstract
defaults.
- tests: browser-route/contract suites removed; browser assertions trimmed.
Desktop control unchanged. 233 computer_use tests pass; the 1 remaining failure
(test_gateway_session_key_yolo_maps_to_unrestricted_mode) is a pre-existing
cross-test state leak — fails identically on origin/main, passes in isolation.
Addresses the five findings from @kshitijk4poor + @GottZ on #95620:
1. macOS 26 LSHandlers parser returned a version number ('7559.97') from the
nested LSHandlerPreferredVersions block instead of the bundle id — detection
returned None on a machine whose default IS Chrome. Strip the nested block
before the role regex.
2. Wrong profile launched (the LinkedIn/Gmail 'logged out' bug): Chrome opens
Default, but the session lives in Local State profile.last_used (e.g.
'Profile 6'). Resolve last_used and mirror its auth files into the copy's
Default on both fresh and refresh paths, so the launched browser is signed in.
3. Private-URL sidecar carried the real cookie jar to arbitrary LAN hosts:
_create_local_session gains allow_real_profile (default True); the
force_local sidecar passes False → always a throwaway profile, and a
real-profile resolve failure no longer breaks private-URL routing.
4. Snapshot permissions were set once (fresh only): now secure the snapshot dir
AND its browser-profile parent on every consented launch.
5. browser.engine=lightpanda + consent gave an unactionable error: guard with
_using_lightpanda_engine() before detection, naming the setting and the fix.
Tests: last_used mirroring (fresh+refresh+fallback), sidecar throwaway + error
isolation, macOS26 parser + detect, perms-on-refresh, lightpanda guard. 198
browser tests pass. Live: Profile-6 cookie DB lands in copy Default (file-level);
real Gmail (Default profile) still signed in.
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:
Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
(_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
(honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod
Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots
Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.