Commit Graph

3331 Commits

Author SHA1 Message Date
kshitijk4poor 0976ceaa98 fix(macos): harden anchor alias failures — warn, unique staging, marker-last
Follow-up to the #95605 salvage, closing the review findings:

- _copy_alias no longer swallows OSError silently: it warns (a leftover
  alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
  (update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
  it now asserts the whole layout (anchor + aliases) is complete, so a
  partially-materialized alias set can never read 'active' in doctor —
  the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
  managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.

5 new regression tests.
2026-08-28 09:05:19 +05:30
Zheqing Zeng 37bccf343e fix(macos): scrub gate env, refuse EACCES, normalize marker paths
Review fixes from kokhlo's live-hardware review:

- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
  PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
  PYTHONHOME=<venv> boots a staged copy that would otherwise die with
  "No module named 'encodings'", papering over the exact prefix
  failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
  images) still skip; EACCES after our own chmod now refuses the
  install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
  so the managed-runtime layout (cpython-3.11-macos-* symlinked to
  cpython-3.11.15-macos-*) no longer reports stale on a fresh install.

Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
2026-08-28 09:05:19 +05:30
Zheqing Zeng 1c92b95a12 test(macos): cover every boot-gate refusal branch directly
The gate is the never-brick guarantee, so each direction gets its own
test: nonzero exit (dyld/encodings crash), build-time prefix leak,
timeout, and the deliberate OSError skip (a binary that cannot execute
here means the symlinked venv was equally dead — installing cannot make
things worse).
2026-08-28 09:05:19 +05:30
Zheqing Zeng aa72df4b42 fix(macos): re-land dylib-complete TCC interpreter anchor
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).

Re-land:

- Keep the signed real-file copy of bin/python (identifier-pinned
  via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
  symlinks. Copies boot on every build we could reproduce and keep
  the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
  devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
  plus the venv prefix, abort and leave the live venv untouched
  on failure.

Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.

Closes #95596.
2026-08-28 09:05:19 +05:30
Brooklyn Nicholson 595ee92289 fix(commands): put desktop slash metadata on the CommandDef
Inference is now any args_hint without subcommands → text. Mixed is the
only remaining hint-token path. desktop= and the few argument_mode
overrides live on the registry entry; the side tables are gone. Catalog
aliases get their own dict copy. Composer tests seed the catalog so
/goal stays mixed without an overlay row.
2026-08-27 22:05:40 -05:00
Brooklyn Nicholson b61408e95e feat(commands): attach desktop slash metadata to the registry
New commands and plugins declare argument_mode and desktop availability
on CommandDef / register_command. commands.catalog ships that map so
desktop does not need a second command list.
2026-08-27 22:05:40 -05:00
Gille 4956ff0cb9 fix(cli): keep journey labels readable 2026-08-27 17:45:43 -05:00
Gille 0dfba37b11 fix(dashboard): trust configured reverse proxies (#94126)
* fix(dashboard): trust configured reverse proxies

* fix(dashboard): trust IPv6 loopback proxies
2026-08-27 10:35:22 -07:00
Brooklyn Nicholson 51e67babca fix(cli): keep Desktop liveness leases when the session cap is off
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
kshitijk4poor e941be7a81 test(gateway): adapt witness-composition harness to the SIGUSR1 in-place drain path
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
2026-08-27 22:06:17 +05:30
kshitijk4poor 5abe2e1880 docs+test(gateway): pin Windows witness-absent behavior and the probe-budget math
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
  there, so the witness is permanently absent, the payload records
  loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
  WEDGED — deliberate fail-safe (graceful drain remains the backstop).
  WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
  stays inside the documented probe budget so retuning can't silently
  blow past the 10s subprocess query tier.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor b48540701c refactor(gateway): simplify witness-probe ambiguity arms; sweep stale tick-socket nodes; clean test tempdirs
/simplify-code follow-ups on the 90502 salvage:

- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
  byte-identical — saw_node was effectively write-only. Collapsed to one
  arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
  from dead PIDs at arm time (POSIX-only liveness probe; Windows never
  creates AF_UNIX nodes) so state/ does not accumulate nodes across
  os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
  was DISPROVED for this call site — asyncio's create_unix_server
  os.remove()s an existing node before binding — but the contract is now
  pinned by test_producer_rebinds_over_stale_socket_node (a live
  producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
  longer leaks a directory per test run.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor db154edbaf test: short tmp_path for loop-tick witness sockets (macOS AF_UNIX limit)
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
2026-08-27 22:06:17 +05:30
rodrigo ca4a9ec686 fix(gateway): a single tick-socket miss must not authorize the wedge kill
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.

WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.

New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
  required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
  the loop is frozen for longer than one tick timeout but shorter than
  the wedge window -> probe is ALIVE and launchd_restart drains, never
  escalates; loop frozen for longer than the window -> WEDGED.

The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
2026-08-27 22:06:17 +05:30
rodrigo a1c83ef901 fix(gateway): interlock the stale-heartbeat wedge verdict with a loop-scheduling witness
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.

The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).

The classifier is now two-witness:
- socket answers            -> ALIVE (file age irrelevant: a stalled write
  or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
  longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
  agree the loop stopped scheduling)
- legacy payload (no flag)  -> unchanged single-witness contract: the
  legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity    -> UNKNOWN, never escalate

Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
2026-08-27 22:06:17 +05:30
Brooklyn Nicholson 2119ed7b4a test(model): cover numeric YAML provider keys in picker and CRUD
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
kshitijk4poor f3cbb262c1 fix(update): valid --ignored=matching mode; rename-only path split; shared preserve constant
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):

- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
  ignored mode'); with it, every ZIP update was refused as 'could not
  check the working tree'. The mocked tests could not see this — a new
  real-git test creates an actual repo + .gitignore and asserts the guard
  runs clean, blocks on an ignored user file, and exempts ignored
  preserved entries. --ignored=matching also reports an ignored dir as
  one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
  rename/copy status codes. Porcelain v1 does not quote plain filenames
  with spaces, so an ignored file literally named 'venv -> node_modules'
  parsed as two preserved tops and slipped past the guard into the
  destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
  instead of a comment-synced duplicate set (change-detector test added).
2026-08-27 21:07:34 +05:30
joaomarcos e64db76982 fix(update): gitignored user files also block the ZIP overlay
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.

Credit: @JoaoMarcos44, whose #87392 included this hardening.
2026-08-27 21:07:34 +05:30
liuhao1024 b0a8d16c60 test(gateway): end-to-end rendered-unit coverage for the cron drain floor
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
2026-08-27 20:39:11 +05:30
HexLab98 1282362803 test(gateway): cover TimeoutStopSec including the cron drain floor 2026-08-27 20:39:11 +05:30
kshitijk4poor b26a359de8 polish: reuse launchd domain probe, fast-observe in wedged tests, blank-line cleanup
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
  up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
  observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
2026-08-27 20:06:04 +05:30
kshitijk4poor 8872cd137c fix: verify launchd replacement PID before trusting KeepAlive (review follow-up)
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
2026-08-27 20:06:04 +05:30
Joby Ellington 7a76046a86 fix(gateway): use SIGUSR1 graceful restart on launchd, not bare SIGTERM
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.

`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:

1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
   on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
   was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
   gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
   `_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
   used by `systemd_restart()` and the updater — had no launchd call site.

2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
   0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
   systemd branch uses `_get_restart_exit_wait_budget()`
   (drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
   documents that callers falling back to a hard kill must cover both phases
   or they reintroduce #77184.

The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.

Observed on macOS 27.0 / Hermes 0.20.4:

    → Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
    ⚠ Gateway PID 49787 still running after 0.0s — restart may fail
    ⚠ Gateway drain timed out after 0s — forcing launchd restart

Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.

The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.

Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 20:06:04 +05:30
x7peeps 1348e65e26 fix(gateway): rebase launchd --replace removal onto main
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.

Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.

Refs: #79096
2026-08-27 20:06:04 +05:30
Heath Harris a71be9852e fix(kanban): close half-open tracked connection when busy_timeout PRAGMA fails
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.

Salvaged from PR #96290 (kanban slice) with regression test.
2026-08-27 19:07:15 +05:30
kshitijk4poor 8d95ab1b37 fix(serve): review follow-ups — never-raise sentinel fallback, DEVNULL stderr in split-stream test
- _write_machine_sentinel_line: wrap the print() fallback so a closed
  redirected stream (ValueError, not OSError) can't propagate out of the
  ready path and kill a healthy serve; document that pythonw port
  discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
  redirect active all server logging lands on stderr, and an unread
  stderr pipe can fill and block the child before the sentinel, flaking
  the test at the 120s timeout
2026-08-27 17:11:56 +05:30
Kitson Kelly f2dd32d3e5 fix(serve): announce READY sentinel on fd 1, not the redirected sys.stdout
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).

Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.

Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
2026-08-27 17:11:56 +05:30
Adolanium a9611f3c6f feat(models): add GLM-5.3-Flash to z.ai and OpenCode Go pickers
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
2026-08-27 04:14:31 -07:00
fangliquanflq 091cc0e8be fix(hermes_cli): scope hook timeouts and fail closed on pre_tool_call
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
2026-08-27 16:13:45 +05:30
Teknium 01f7ce5b76 feat(sessions): one-shot single-match owner backfill for legacy NULL-profile rows (#94724)
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.

Refs #94724
2026-08-27 02:17:56 -07:00
Aoshi-Dev 5ab04c764c test(hermes_cli): make test_default_path pass on native Windows
TestGetHermesHome.test_default_path asserted ~/.hermes unconditionally,
but the native Windows default is %LOCALAPPDATA%\hermes (see
hermes_constants._get_platform_default_hermes_home). Branch the
assertion by platform so the test passes everywhere.

Salvaged from PR #96003 by @Aoshi-Dev (the parse-guard half of that PR
was superseded by #96169); authorship preserved.
2026-08-27 14:02:06 +05:30
kshitijk4poor 93a29d110d fix(hermes_cli): surface fail-closed config write refusals cleanly
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):

- config_command: catch RuntimeError for set/unset and print a clean
  one-line error + exit(1) instead of a raw traceback on the primary
  'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
  ConsoleCommandError so 'hermes console' and the dashboard console
  report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
  branch — the old fallthrough claimed 'falling back to default config'
  even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.
2026-08-27 13:27:24 +05:30
fangliquanflq df8b841c0c test(hermes_cli): align malformed YAML set/unset expectations with RuntimeError 2026-08-27 13:27:24 +05:30
fangliquanflq 77d4d23fbc fix(hermes_cli): stop config set/unset from wiping user overrides on invalid YAML
Fail closed when config.yaml is unparseable or non-mapping before set/unset writes, reuse the readable-config guard to return the parsed mapping, and cover refuse/empty-mapping paths with regression tests.
2026-08-27 13:27:24 +05:30
Gille 6defe7eb6c fix(config): preserve lossy decimal values as strings 2026-08-27 11:49:58 +05:30
Cursor Agent 8246c4f92a fix(cli): repair interrupted update fleet restart
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.

Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
2026-08-26 21:38:20 -07:00
fangliquanflq cb54576b1a fix(gateway): isolate control routes from default executor 2026-08-26 21:38:20 -07:00
Teknium 824f7e081a chore: remove Windows real-profile PROOF workflow + live test
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
2026-08-26 19:25:33 -07:00
Teknium e4451ec6e5 feat(browser): close-with-approval flow for Windows real-profile (toggle arms, agent asks, blocked if still locked) [proof do-not-merge]
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
   browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
   always returns the [profile-locked] signal and the copy is refused. A later
   attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
    (new CLI subcommand) runs
   close_browser_holding_profile only when the agent has the user's OK. The
   locked error tells the agent to ask first, then run it, then retry.

- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
  message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py:  subcommand (identity+binding-verified
  tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.

Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
2026-08-26 19:25:33 -07:00
Teknium 00d5632249 chore: remove Windows real-profile PROOF workflow + live test
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 75f402c302 test(ci): assert cookie DB at either location + dump Default contents on miss [do-not-merge]
Auto-close live test asserted the legacy Default/Cookies path, but modern Chrome
writes Default/Network/Cookies. Accept either; on miss, print the copy's Default
listing so a real copy gap (vs a path-assertion bug) is visible.
2026-08-26 19:25:33 -07:00
Teknium 9e9e1b2245 feat(browser): consented auto-close of a running browser for Windows real-profile [proof workflow do-not-merge]
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.

- close_browser_holding_profile: graceful terminate → kill → poll until the
  cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
  unrelated same-name process on a different dir).
- Config key + docs admonition.

Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.

Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
2026-08-26 19:25:33 -07:00
Teknium b73f78714a chore: remove Windows real-profile PROOF workflow + live tests
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 931bf613b1 fix(browser): Windows real-profile fails fast when the browser is running
Live windows-latest proof settled it: a running Chrome opens its cookie DB
deny-all (even CreateFile with FILE_SHARE_READ|WRITE|DELETE fails; sqlite
mode=ro/immutable/nolock all 'unable to open'), so copy-while-running is
impossible on Windows without VSS/admin — and the prior code HUNG ~24min on the
locked file.

Fix: a fast up-front lock probe (_profile_is_locked: one open() of the active
profile's cookie DB; PermissionError = locked) runs BEFORE any copy in
snapshot_real_profile. If locked, bail immediately with 'fully quit the browser
(incl. background/tray) and retry, or turn browser.use_real_profile off'. Never
hangs, never a silent signed-out copy. POSIX has no mandatory locking so the
probe never trips there — copy-while-running still works on macOS/Linux.

Docs: admonition stating Windows needs the browser fully closed (background
apps included); the live-drive-while-running path is #95669.

Tests: lock-probe unit coverage (readable/no-db/PermissionError), snapshot
fails-fast-no-copytree when locked. Windows live E2E asserts the fast-fail
contract (returns <30s with the quit message) + the read-strategy diagnostic.
2026-08-26 19:25:33 -07:00
Teknium bd504bee6d test(browser): PROOF diag — probe which read strategy beats Chrome's Windows lock [do-not-merge]
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
2026-08-26 19:25:33 -07:00
Teknium 1740abbe64 test(browser): PROOF round 2 — Windows locked-profile fails CLOSED (live-corrected)
The first live Windows run DISPROVED the sqlite-online-backup claim: Chrome's
share lock on Windows is strong enough that even a read-only SQLite open is
refused by the OS (raw-copy precondition fired, _copy_auth_file still returned
False). Copy-while-Chrome-runs is impossible on Windows — the earlier fix was
theatre that only passed on Linux (no mandatory locking).

Corrected contract, now asserted live: with a running Chrome holding the cookie
DB, snapshot_real_profile FAILS CLOSED with 'could not read ... login data
(N locked). Close <browser> and retry' — never a silent signed-out/torn copy.
Second test proves the supported path (Chrome closed) copies cleanly. So
real-profile browsing on Windows requires the browser closed; Linux/macOS
unaffected; live-drive-the-real-profile is tracked in #95669.

PROOF branch evidence only — workflow + test reverted before merge.
2026-08-26 19:25:33 -07:00
Teknium 2ecb18c4a7 test(browser): PROOF — Windows live E2E for locked-DB real-profile copy [do-not-merge]
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.

This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
2026-08-26 19:25:33 -07:00
Jan-Stefan Janetzky 7e2c2b1b08 fix(browser): resolve snap and Flatpak Chromium profiles on Linux
real_profile_data_dir hard-wired Linux to $XDG_CONFIG_HOME/<name>, and the
xdg fragment map only knew the native package names. Ubuntu's default snap
Chromium (xdg reports chromium_chromium.desktop, profile under
~/snap/chromium/common/chromium) and Flatpak builds (~/.var/app/<id>/config/…)
therefore ended in 'profile directory was not found' for a browser the user
runs every day, and Flatpak Chrome (com.google.Chrome.desktop) was reported as
'not a supported Chromium browser'.

Try the native, snap and Flatpak locations and return the first that exists;
fall back to the native path so the error message still names a concrete
directory. Map the Flatpak application ids in the xdg lookup.

Tests cover the xdg names for all four browsers in native and Flatpak form,
and the directory preference order with a temp HOME.
2026-08-26 19:25:33 -07:00
Jan-Stefan Janetzky f5e6028417 fix(browser): read the macOS https handler per entry and drop the installed-browser fallback
_detect_default_darwin matched a Chromium bundle id and the literal 'https'
anywhere in the whole LSHandlers dump, so a browser registered for ftp or a
content type was reported as the https default, and map order decided ties.
When nothing matched it fell back to the first installed Chromium app — with
Safari or Firefox as the actual default that drove a browser the user never
consented to, contradicting the docstring, the config comment and the desktop
copy ('a non-Chromium default fails with a clear message').

Parse the dump entry by entry, take the LSHandlerRoleAll/Viewer of the entry
whose LSHandlerURLScheme is https, and fail closed on anything else — an empty
handler list is what macOS stores while Safari is still the implicit default.

Tests feed real 'defaults read' output shapes instead of patching the detector
(reviewer fixture from the PR discussion: Safari on https, Chrome on ftp).
2026-08-26 19:25:33 -07:00
Teknium adb29d8527 test: invert the resume-token fleet-probe pin to the new contract
Main's pin (added after this branch was cut) froze the #93406 bug as the
contract: resume-token services demanded probe rows that SCM-paused
services can never produce, stalling every healthy Windows desktop update
(#95589). The pin now asserts the exclusion; restart-phase and
pre-restart-pid signals keep failing closed.
2026-08-26 18:23:33 -07:00