Follow-up to the #95605 salvage, closing the review findings:
- _copy_alias no longer swallows OSError silently: it warns (a leftover
alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
(update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
it now asserts the whole layout (anchor + aliases) is complete, so a
partially-materialized alias set can never read 'active' in doctor —
the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.
5 new regression tests.
Review fixes from kokhlo's live-hardware review:
- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
PYTHONHOME=<venv> boots a staged copy that would otherwise die with
"No module named 'encodings'", papering over the exact prefix
failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
images) still skip; EACCES after our own chmod now refuses the
install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
so the managed-runtime layout (cpython-3.11-macos-* symlinked to
cpython-3.11.15-macos-*) no longer reports stale on a fresh install.
Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
The gate is the never-brick guarantee, so each direction gets its own
test: nonzero exit (dyld/encodings crash), build-time prefix leak,
timeout, and the deliberate OSError skip (a binary that cannot execute
here means the symlinked venv was equally dead — installing cannot make
things worse).
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).
Re-land:
- Keep the signed real-file copy of bin/python (identifier-pinned
via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
symlinks. Copies boot on every build we could reproduce and keep
the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
plus the venv prefix, abort and leave the live venv untouched
on failure.
Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.
Closes#95596.
Inference is now any args_hint without subcommands → text. Mixed is the
only remaining hint-token path. desktop= and the few argument_mode
overrides live on the registry entry; the side tables are gone. Catalog
aliases get their own dict copy. Composer tests seed the catalog so
/goal stays mixed without an overlay row.
New commands and plugins declare argument_mode and desktop availability
on CommandDef / register_command. commands.catalog ships that map so
desktop does not need a second command list.
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.
Co-authored-by: metamindedu <metamind@kakao.com>
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
there, so the witness is permanently absent, the payload records
loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
WEDGED — deliberate fail-safe (graceful drain remains the backstop).
WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
stays inside the documented probe budget so retuning can't silently
blow past the 10s subprocess query tier.
/simplify-code follow-ups on the 90502 salvage:
- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
byte-identical — saw_node was effectively write-only. Collapsed to one
arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
from dead PIDs at arm time (POSIX-only liveness probe; Windows never
creates AF_UNIX nodes) so state/ does not accumulate nodes across
os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
was DISPROVED for this call site — asyncio's create_unix_server
os.remove()s an existing node before binding — but the contract is now
pinned by test_producer_rebinds_over_stale_socket_node (a live
producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
longer leaks a directory per test run.
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.
WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.
New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
the loop is frozen for longer than one tick timeout but shorter than
the wedge window -> probe is ALIVE and launchd_restart drains, never
escalates; loop frozen for longer than the window -> WEDGED.
The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.
The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).
The classifier is now two-witness:
- socket answers -> ALIVE (file age irrelevant: a stalled write
or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
agree the loop stopped scheduling)
- legacy payload (no flag) -> unchanged single-witness contract: the
legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity -> UNKNOWN, never escalate
Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):
- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
ignored mode'); with it, every ZIP update was refused as 'could not
check the working tree'. The mocked tests could not see this — a new
real-git test creates an actual repo + .gitignore and asserts the guard
runs clean, blocks on an ignored user file, and exempts ignored
preserved entries. --ignored=matching also reports an ignored dir as
one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
rename/copy status codes. Porcelain v1 does not quote plain filenames
with spaces, so an ignored file literally named 'venv -> node_modules'
parsed as two preserved tops and slipped past the guard into the
destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
instead of a comment-synced duplicate set (change-detector test added).
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.
Credit: @JoaoMarcos44, whose #87392 included this hardening.
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.
`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:
1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
`_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
used by `systemd_restart()` and the updater — had no launchd call site.
2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
systemd branch uses `_get_restart_exit_wait_budget()`
(drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
documents that callers falling back to a hard kill must cover both phases
or they reintroduce #77184.
The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.
Observed on macOS 27.0 / Hermes 0.20.4:
→ Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
⚠ Gateway PID 49787 still running after 0.0s — restart may fail
⚠ Gateway drain timed out after 0s — forcing launchd restart
Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.
The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.
Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.
Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.
Refs: #79096
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.
Salvaged from PR #96290 (kanban slice) with regression test.
- _write_machine_sentinel_line: wrap the print() fallback so a closed
redirected stream (ValueError, not OSError) can't propagate out of the
ready path and kill a healthy serve; document that pythonw port
discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
redirect active all server logging lands on stderr, and an unread
stderr pipe can fill and block the child before the sentinel, flaking
the test at the 120s timeout
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).
Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.
Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.
Refs #94724
TestGetHermesHome.test_default_path asserted ~/.hermes unconditionally,
but the native Windows default is %LOCALAPPDATA%\hermes (see
hermes_constants._get_platform_default_hermes_home). Branch the
assertion by platform so the test passes everywhere.
Salvaged from PR #96003 by @Aoshi-Dev (the parse-guard half of that PR
was superseded by #96169); authorship preserved.
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):
- config_command: catch RuntimeError for set/unset and print a clean
one-line error + exit(1) instead of a raw traceback on the primary
'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
ConsoleCommandError so 'hermes console' and the dashboard console
report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
branch — the old fallthrough claimed 'falling back to default config'
even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.
Fail closed when config.yaml is unparseable or non-mapping before set/unset writes, reuse the readable-config guard to return the parsed mapping, and cover refuse/empty-mapping paths with regression tests.
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.
Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
Auto-close live test asserted the legacy Default/Cookies path, but modern Chrome
writes Default/Network/Cookies. Accept either; on miss, print the copy's Default
listing so a real copy gap (vs a path-assertion bug) is visible.
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.
- close_browser_holding_profile: graceful terminate → kill → poll until the
cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
unrelated same-name process on a different dir).
- Config key + docs admonition.
Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.
Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
Live windows-latest proof settled it: a running Chrome opens its cookie DB
deny-all (even CreateFile with FILE_SHARE_READ|WRITE|DELETE fails; sqlite
mode=ro/immutable/nolock all 'unable to open'), so copy-while-running is
impossible on Windows without VSS/admin — and the prior code HUNG ~24min on the
locked file.
Fix: a fast up-front lock probe (_profile_is_locked: one open() of the active
profile's cookie DB; PermissionError = locked) runs BEFORE any copy in
snapshot_real_profile. If locked, bail immediately with 'fully quit the browser
(incl. background/tray) and retry, or turn browser.use_real_profile off'. Never
hangs, never a silent signed-out copy. POSIX has no mandatory locking so the
probe never trips there — copy-while-running still works on macOS/Linux.
Docs: admonition stating Windows needs the browser fully closed (background
apps included); the live-drive-while-running path is #95669.
Tests: lock-probe unit coverage (readable/no-db/PermissionError), snapshot
fails-fast-no-copytree when locked. Windows live E2E asserts the fast-fail
contract (returns <30s with the quit message) + the read-strategy diagnostic.
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
The first live Windows run DISPROVED the sqlite-online-backup claim: Chrome's
share lock on Windows is strong enough that even a read-only SQLite open is
refused by the OS (raw-copy precondition fired, _copy_auth_file still returned
False). Copy-while-Chrome-runs is impossible on Windows — the earlier fix was
theatre that only passed on Linux (no mandatory locking).
Corrected contract, now asserted live: with a running Chrome holding the cookie
DB, snapshot_real_profile FAILS CLOSED with 'could not read ... login data
(N locked). Close <browser> and retry' — never a silent signed-out/torn copy.
Second test proves the supported path (Chrome closed) copies cleanly. So
real-profile browsing on Windows requires the browser closed; Linux/macOS
unaffected; live-drive-the-real-profile is tracked in #95669.
PROOF branch evidence only — workflow + test reverted before merge.
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.
This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
real_profile_data_dir hard-wired Linux to $XDG_CONFIG_HOME/<name>, and the
xdg fragment map only knew the native package names. Ubuntu's default snap
Chromium (xdg reports chromium_chromium.desktop, profile under
~/snap/chromium/common/chromium) and Flatpak builds (~/.var/app/<id>/config/…)
therefore ended in 'profile directory was not found' for a browser the user
runs every day, and Flatpak Chrome (com.google.Chrome.desktop) was reported as
'not a supported Chromium browser'.
Try the native, snap and Flatpak locations and return the first that exists;
fall back to the native path so the error message still names a concrete
directory. Map the Flatpak application ids in the xdg lookup.
Tests cover the xdg names for all four browsers in native and Flatpak form,
and the directory preference order with a temp HOME.
_detect_default_darwin matched a Chromium bundle id and the literal 'https'
anywhere in the whole LSHandlers dump, so a browser registered for ftp or a
content type was reported as the https default, and map order decided ties.
When nothing matched it fell back to the first installed Chromium app — with
Safari or Firefox as the actual default that drove a browser the user never
consented to, contradicting the docstring, the config comment and the desktop
copy ('a non-Chromium default fails with a clear message').
Parse the dump entry by entry, take the LSHandlerRoleAll/Viewer of the entry
whose LSHandlerURLScheme is https, and fail closed on anything else — an empty
handler list is what macOS stores while Safari is still the implicit default.
Tests feed real 'defaults read' output shapes instead of patching the detector
(reviewer fixture from the PR discussion: Safari on https, Chrome on ftp).
Main's pin (added after this branch was cut) froze the #93406 bug as the
contract: resume-token services demanded probe rows that SCM-paused
services can never produce, stalling every healthy Windows desktop update
(#95589). The pin now asserts the exclusion; restart-phase and
pre-restart-pid signals keep failing closed.