Commit Graph

197 Commits

Author SHA1 Message Date
kshitijk4poor 86221181da fix(browser): inline windows-footgun annotation for os.kill(pid, 0) (#85125 CI) 2026-08-27 16:11:16 +05:30
kshitijk4poor 011b2cf4d0 fix(browser): psutil.pid_exists liveness probe — Windows footgun (#85125 CI) 2026-08-27 16:11:16 +05:30
kshitijk4poor 3a94101524 fix(browser): suspect-session recycle + wedged-daemon tree-kill on timeout (#85125 3b-browser/4c)
Builds on e11187208f (salvage of #72206 by @luyifan, authorship
preserved) which added the post-timeout session reset. This commit
completes the Phase 3b contract:

- suspect flag: a command timeout marks the session key suspect;
  the NEXT _get_session_info for that key health-checks and recycles
  the session instead of handing back the poisoned handle (#72205)
- flag cleared on fresh-session store so a healthy new session is
  never spuriously recycled; cross-key leakage fixed
- wedged-vs-alive rule (#68139): after a timeout, if the daemon's
  socket still answers, recycle the session only; if it is
  unresponsive (or no socket exists to probe — conservative), tree-
  kill via agent.deadline.kill_process_tree and evict
- negative probe: successful commands never mark or recycle

Tests: tests/tools/test_browser_suspect_recycle.py (20 tests:
mark-once, recycle-then-succeed, success-never-recycles, tree-kill
invoked on the wedged path with pid assertion, flag lifecycle).

Co-authored-by: luyifan <al3060388206@gmail.com>
2026-08-27 16:11:16 +05:30
luyifan 0be9b7cd57 fix(browser): reset sessions after command timeout 2026-08-27 16:11:16 +05:30
Teknium e4451ec6e5 feat(browser): close-with-approval flow for Windows real-profile (toggle arms, agent asks, blocked if still locked) [proof do-not-merge]
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
   browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
   always returns the [profile-locked] signal and the copy is refused. A later
   attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
    (new CLI subcommand) runs
   close_browser_holding_profile only when the agent has the user's OK. The
   locked error tells the agent to ask first, then run it, then retry.

- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
  message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py:  subcommand (identity+binding-verified
  tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.

Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
2026-08-26 19:25:33 -07:00
Teknium 6e854595e2 fix(browser): real-profile review round 3 — overlay-ordering race, torn-copy marker, consent cleanup, active-only copy
Addresses the round-3 findings from @Adolanium + @kshitijk4poor on #95620:

1. Overlay-before-reuse race (blocker): _real_profile_cdp ran snapshot_real_profile
   BEFORE the session-reuse check, so a cold resolve that ends in reuse rewrote
   Cookies/Login Data under a live Chromium holding the user-data-dir open (torn
   DBs, locked txns, phantom logouts). Now: resolve copy dir as a PATH, probe
   reuse first, return early on a hit; snapshot/overlay only on the relaunch
   path when no live browser owns the dir.
2. Torn first copy poisoned freshness forever: freshness keyed on isdir(Default),
   so a half-written copy (disk full / Ctrl+C) was treated as populated and only
   ever got auth overlays. Now gated on a .hermes-snapshot-complete marker
   written only after a full copy succeeds; a torn copy is rebuilt from scratch.
3. Consent revocation left copied credentials on disk: turning use_real_profile
   off now deletes ~/.hermes/browser-profile/ on next browser use
   (cleanup_real_profile_snapshots), so cookies/logins don't outlive consent.
4. Stale non-active profile copies: only the ACTIVE profile (last_used) is copied
   into the copy's Default now — other Chrome profiles are never snapshotted
   (smaller copy, no stale credential dirs lingering).
5. Docs/config/desktop wording aligned to actual behavior (active-profile only,
   refresh on fresh session, consent-off cleanup).

Tests: overlay-skipped-on-reuse + overlay-runs-on-relaunch, done-marker gating +
torn-copy rebuild, active-only copy, consent-off cleanup (removes store +
idempotent + triggered from _real_profile_cdp). 206 browser + 222
backup/file_safety pass. Live: reuse skips re-snapshot; direct launch on the
active-only copy loads the real signed-in Gmail inbox.
2026-08-26 19:25:33 -07:00
Teknium 7fb52c14ba fix(browser): real-profile review round 2 — last_used profile, sidecar isolation, macOS26 parser, perms, lightpanda
Addresses the five findings from @kshitijk4poor + @GottZ on #95620:

1. macOS 26 LSHandlers parser returned a version number ('7559.97') from the
   nested LSHandlerPreferredVersions block instead of the bundle id — detection
   returned None on a machine whose default IS Chrome. Strip the nested block
   before the role regex.
2. Wrong profile launched (the LinkedIn/Gmail 'logged out' bug): Chrome opens
   Default, but the session lives in Local State profile.last_used (e.g.
   'Profile 6'). Resolve last_used and mirror its auth files into the copy's
   Default on both fresh and refresh paths, so the launched browser is signed in.
3. Private-URL sidecar carried the real cookie jar to arbitrary LAN hosts:
   _create_local_session gains allow_real_profile (default True); the
   force_local sidecar passes False → always a throwaway profile, and a
   real-profile resolve failure no longer breaks private-URL routing.
4. Snapshot permissions were set once (fresh only): now secure the snapshot dir
   AND its browser-profile parent on every consented launch.
5. browser.engine=lightpanda + consent gave an unactionable error: guard with
   _using_lightpanda_engine() before detection, naming the setting and the fix.

Tests: last_used mirroring (fresh+refresh+fallback), sidecar throwaway + error
isolation, macOS26 parser + detect, perms-on-refresh, lightpanda guard. 198
browser tests pass. Live: Profile-6 cookie DB lands in copy Default (file-level);
real Gmail (Default profile) still signed in.
2026-08-26 19:25:33 -07:00
Teknium f1d05ce7d8 fix(browser): real-profile snapshot is a first-class secret store + preserve channel identity
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:

Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
  (_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
  same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
  (honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod

Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
  Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
  stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots

Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.
2026-08-26 19:25:33 -07:00
Teknium 1f4d095fd8 feat(browser): real-profile browsing via agent-browser copy + browser-use CDP
Copy the user's default-Chromium profile (auth state only) into a managed
snapshot, launch Hermes' packaged Chromium on it via agent-browser, and hand
the CDP endpoint to the Browser Use CLI (and built-in tools) to drive. The
snapshot is a non-default dir, so it sidesteps Chrome 136+'s default-profile
remote-debugging block and never contends with the user's running browser;
launched without mock-keychain switches so keyring-encrypted cookies decrypt.

- consent-gated browser_exec 'local' arg (schema only appears with consent)
- fail-closed on non-Chromium default / snapshot failure
- stale-session guard: reuse only when the live session is on our copy dir
- snapshot excludes extensions/service-workers (renderer wedge) + caches
2026-08-26 19:25:33 -07:00
Jan-Stefan Janetzky 0f56937cd7 fix(browser): keep local_browser inside the private-URL policy and the existing local session
_navigation_session_key returned the ::local sidecar key for local_browser
before the cloud-provider and auto_local_for_private_urls checks. Two
consequences:

- Every private-URL gate in browser_navigate (credential-bearing query,
  _is_safe_url pre-nav, post-redirect) is keyed off the sidecar key, so a
  model-supplied browser_navigate(url, local_browser=True) opened LAN
  addresses in a host-side Chromium — with the real profile's cookies — even
  when the user had set browser.auto_local_for_private_urls: false. Consent to
  the profile is not consent to override the LAN routing opt-out; a private
  URL now follows auto_local_for_private_urls exactly as without the flag.

- Without a cloud provider the bare session already is the local Chromium
  (and already carries the real profile when consented), so the flag created a
  second session for the same task on the same user-data-dir, which Chromium's
  process singleton refuses. The flag is now a no-op there.

The existing consent/CDP-precedence tests keep their assertions; the sidecar
test now states the cloud-provider precondition it silently relied on.
2026-08-26 19:25:33 -07:00
Teknium 830e4a29be feat(browser): consent-gated real default-Chromium profile for local browsing + local_browser arg 2026-08-26 19:25:33 -07:00
Leegenux 6ce7ab8bfb feat(browser): make snapshot threshold configurable 2026-08-24 21:51:44 -07:00
pierrenode bf8b28f27a fix(tools): route browser snapshot storage through the symlink-safe writer
Today's spill/cache-writer hardening (tools/spill_safety.py,
write_text_exclusive/ensure_spill_dir with O_CREAT|O_EXCL|O_NOFOLLOW)
migrated tools/web_tools.py::_store_full_text() — which writes to the same
cache/web directory with the same content-hash filename scheme — but left
its near-identical sibling, tools/browser_tool.py::_store_full_snapshot(),
on the pre-fix plain open()/write_text() pattern. A pre-planted symlink at
the content-hash path redirected the write onto an arbitrary user-owned
file, same as the sites that commit fixed.

Reproduced live: with a symlink planted at the exact
browser-snapshot-<digest>.txt path (predictable from the snapshot content
hash), the pre-fix write followed the link and overwrote the link's
target with the (secret-redacted but otherwise user/page-controlled)
snapshot content.

Fix mirrors _store_full_text's exact usage: ensure_spill_dir(private=False)
+ write_text_exclusive(private=False, overwrite=True) — not private since
cache/web is bind-mounted into remote backends whose container UID must
read it; overwrite=True because re-snapshotting the same page state
legitimately reuses the same content-hash name (the overwrite path
lstat-unlinks the link itself, never following it to write through).

Added a regression test planting a symlink at the exact digest path and
asserting the link's target is untouched (only the link itself gets
safely replaced by a real file). Mutation-verified: with the fix stashed,
the pre-fix code wrote the snapshot content into the symlink's target
file, reproducing the vulnerability exactly.
2026-08-24 21:45:56 -07:00
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
kshitijk4poor 547f985286 refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
Per-site decisions:

1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
   MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
   via a function-local import; keeps the swallow-everything fail-open
   contract and the (proc) -> None signature (agent/shell_hooks.py imports
   it by name; _kill_git_process_tree alias preserved). The old body is kept
   verbatim as _legacy_kill_process_tree and used as fallback when the
   delegation import/call fails. A final proc.kill() is retained on the
   happy path so Popen bookkeeping sees the exit (matches old behavior).

2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
   (delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
   body sent SIGTERM then SIGKILL with zero grace between them; the shared
   primitive sends SIGKILL only. With no grace period the observable effect
   is identical, and the psutil descendant sweep now also reaches
   agent-browser's setsid'd daemon grandchild, which killpg alone missed.
   tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
   the legacy fallback (its assertions describe the fallback's internals).

3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
   MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
   kill when escalate=True) — expressed as two delegated calls:
   kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
   kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
   proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
   body terminated children before the parent; the shared primitive
   signals the group atomically (child is a session leader via
   start_new_session=True) plus an identity-aware descendant sweep —
   strictly wider coverage, same signals.

4. gateway/status.py: KEPT BOTH SITES.
   - terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
     incompatible with the shared primitive — it must RAISE OSError with
     taskkill's stderr on non-zero exit (callers branch on that), falls back
     to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
     single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
     fail-soft primitive would invert the error contract.
   - reap_gateway_children (~l2029): NOT migrated. It operates on a
     pre-snapshotted child list from a parent that is already dead
     (psutil.Process(pid) on the parent would fail), and every signal is
     wrapped in identity/ownership checks the primitive lacks: is_running()
     identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
     plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
     The coupling is the feature; migrating would delete the safety logic.

5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
   Dev tooling that intentionally kills by CAPTURED pgid because the direct
   child is usually already reaped (psutil/pid-based primitive cannot find
   it), and it avoids the psutil import on the test-runner hot path. Its
   docstring already documents why psutil is the wrong tool there.

New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
2026-08-25 01:34:56 +05:30
kshitijk4poor c02fe6501b fix(vision): shrink oversized history embeds via JPEG quality, not halving
PNG has no quality ladder, so a text-dense screenshot over the 256KB
history-embed cap (#92699 / #92783) could only shrink by halving
dimensions — 1568px dropped to ~784px and on-screen text became
unreadable, the exact fidelity screenshot QA depends on.

Add force_jpeg to _resize_image_for_vision: the two history-embed call
sites (vision_analyze native, browser_vision native) re-encode
resize-needing screenshots as JPEG so the quality ladder (85/70/50)
absorbs the byte pressure and the readable resolution survives.
Under-cap images are untouched and stay PNG; one-shot/reactive paths
keep their existing format behavior.

Flagged during the #92783 salvage review.
2026-08-23 15:06:03 +05:30
kshitijk4poor dff84f1890 fix(browser): cap browser_vision native embeds for history reuse
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.

Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.

Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 13:27:04 +05:30
abundantbeing d524cc9a16 fix(browser): harden extension controller routing
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.

Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
2026-08-21 22:33:45 +05:30
abundantbeing 5df1d0e113 feat(browser): add authenticated control broker 2026-08-21 22:33:45 +05:30
Teknium 7f83d3808c fix(browser): strict cloud-provider selection; camofox becomes a selection
An explicitly stored browser.cloud_provider that names no registered
plugin now raises the honest selection-naming error instead of warning
and silently auto-detecting; the auto-detect walk (including the managed
gateway entitlement probe) runs only when no cloud_provider key was ever
written. The 'nous' selection routes to the Browser Use provider, whose
config resolver is now a strict switch: 'nous' => managed only, stored
vendor => direct BROWSER_USE_API_KEY only with a selection-naming error
when missing. Camofox is selected via browser.cloud_provider: camofox;
CAMOFOX_URL stays the server ADDRESS only and can no longer override an
explicit different selection (never-configured installs keep the legacy
env-var activation).
2026-08-19 16:10:01 -07:00
andy c99a45b28e fix(browser): reap leaked agent-browser daemons whose owner is still alive
The orphan reaper had two gaps that let agent-browser daemons accumulate
indefinitely inside a single long-lived hermes process:

1. `_reap_orphaned_browser_sessions()` ran exactly once, before the cleanup
   loop started, so a leak appearing after boot could never be recovered.

2. `owner_alive is True` skipped unconditionally. In-memory session tracking
   is lost on any exception path between spawn and registration, but the
   owner PID stays up — so such a daemon was skipped forever.

The daemon-side `AGENT_BROWSER_IDLE_TIMEOUT_MS` is not a backstop for (2):
it does not fire when the daemon itself is wedged, e.g. after Chrome's
framework was replaced underneath it by an auto-update.

Observed on macOS: five agent-browser daemons (96 Chrome processes) built up
over 10 days inside an 18-day-uptime hermes process, holding roughly 5 CPU
cores busy and driving the load average past 100. Four of those processes
were still running a Chrome framework version that had since been replaced
on disk, spinning at ~85% CPU each.

Changes:

- Re-run the reaper every `BROWSER_ORPHAN_REAP_INTERVAL` (300s) from inside
  the cleanup loop. Cycle 0 preserves the existing startup reap.

- When the owner is alive but the session is untracked, fall back to idle
  age: reap past `BROWSER_ORPHAN_GRACE_SECONDS`, defined as
  `max(1h, 20 x inactivity_timeout)`. Unknown age fails safe.

- Add `_socket_dir_idle_seconds()` — the newest mtime under a session's
  socket dir. Every browser command writes `_stdout_<cmd>` / `_stderr_<cmd>`
  there, making it a last-activity marker that survives hermes restarts and
  does not depend on in-memory bookkeeping surviving an exception path. It
  scans directory entries rather than reading the directory mtime alone:
  command names repeat, and rewriting an existing `_stdout_click` updates
  that file's mtime but not the directory's, so a dir-mtime-only check would
  report a busy session as idle and reap it.

Sessions still present in `_active_sessions` are never touched at any age,
and the new path still goes through `_verify_reapable_browser_daemon`, so
the anti-spoof / anti-PID-recycle guarantees from #14073 are unchanged.

Adds 9 tests: idle-age unit tests (including the dir-mtime regression),
spared/reaped/fail-safe cases for a live owner, the identity-guard gate on
the new path, and a periodic-reap test asserting more than one reap per
cleanup-thread lifetime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 00:33:01 -07:00
adikpb dcc2f3de1d fix(vision): stop capping aux vision output with hardcoded max_tokens
The vision tools' call_kwargs hardcode max_tokens caps (2000 for
vision_analyze/browser_vision, 4000 for video analysis), truncating
descriptions of complex images at the cap. The centralized aux client
already omits max_tokens by default (#34845) so providers use their
model max output; these three call sites were the leftovers that
bypassed that policy.

Remove the hardcoded caps entirely — the aux client handles the
mandatory-max_tokens Anthropic wire via _resolve_anthropic_messages_max_tokens
(model output ceiling) and Gemini native omits maxOutputTokens (65K ceiling),
so no wire needs an explicit cap.
2026-08-15 12:53:37 +05:30
Zak B. Elep 03cdc3b20c fix(browser): harden npx agent-browser resolution
- --ignore-scripts on every real npx agent-browser invocation.
  AGENT_BROWSER_NPX_SPEC is a floating ^0.26.0 range, not an exact
  pin, and none of these sites passed it (unlike install.sh/
  install.ps1's own npm install of the same package). Verified against
  the real CLI: `npx --ignore-scripts --prefer-offline -y
  "agent-browser@^0.26.0" --version` resolves cleanly on npm
  11.19.0/node 26.
- _resolve_npx_bin() now checks the Hermes-managed/extended search
  before a bare ambient PATH lookup, validating each candidate with
  node_tool_runnable before trusting it — a bare PATH-first lookup let
  a broken system npx shadow a healthy managed one with no recovery.
- warm_agent_browser_npx_cache() now runs a credential-scrubbed,
  PATH-propagated environment (matching every other agent-browser
  subprocess spawn) instead of inheriting the full parent environment
  including every provider/gateway credential Hermes holds, and kills
  the whole process tree (not just the top-level npx PID) on timeout
  via the new _kill_process_tree helper, since a surviving descendant
  can otherwise hold a capture pipe open past the nominal deadline.
2026-08-13 02:38:28 -07:00
Zak B. Elep 675d41fb25 fix(browser): pin npx agent-browser resolution and share a sentinel constant
Git-clone installs resolving agent-browser via bare npx floated latest
with no integrity check, while install.sh/install.ps1 installs stayed
pinned to ^0.26.0. Pin the npx spec to match. Also extract the
"npx agent-browser" sentinel comparison (6 call sites across two
packages) into a named constant/predicate, fix a PATH-priority
inversion where a broken system npx could shadow a healthy
Hermes-managed one at the two real npx launch sites, and stop
`hermes doctor --fix` from counting a bonus npx cache warm as a fixed
issue on an otherwise-healthy run.
2026-08-13 02:38:28 -07:00
Zak B. Elep c196e0f08f fix(browser): hide console window for npx cache warm-up on Windows
warm_agent_browser_npx_cache() spawns a resolved npx.cmd via
subprocess.run with a list arg and no shell=True, which Windows still
routes through cmd.exe. Without creationflags=windows_hide_flags(),
that can flash a console window during hermes update/doctor --fix,
same as the existing agent-browser subprocess spawn elsewhere in this
file already guards against.

Adds a regression test to the cross-cutting Windows no-window-flags
audit suite so a future refactor can't silently drop the flag again.
2026-08-13 02:38:28 -07:00
Zak B. Elep 5f5f8d5b62 fix(cli): drop agent-browser/@streamdown-math from root npm deps
`hermes update` was pruning root-level Node dependencies (agent-browser)
because npm ci always wipes and reifies node_modules according to its
active filter -- no root-first/workspace-first ordering or flag
combination (--workspaces=false, --include-workspace-root, etc.) can
reliably keep a root-only package.json dependency from being pruned by
a subsequent workspace-scoped npm ci. Confirmed empirically and via
npm/cli source (isArboristCmd hardcodes includeWorkspaceRoot=false for
ci/install), so no amount of install-order juggling fixes this for good.

Instead of chasing install order, remove the root-only dependencies
that made the npm step fragile in the first place:

- agent-browser is no longer a root package.json dependency. It
  resolves lazily via `npx agent-browser` (tools/browser_tool.py
  already had this as a fallback; it's now the primary path).
  warm_agent_browser_npx_cache() is called fire-and-forget from both
  `hermes update` and `hermes doctor --fix` to keep npx's cache warm,
  preserving the "available before any session starts" property
  agent-browser had as an eager dependency without re-entangling it
  with the npm workspace graph.
- @streamdown/math moves to apps/desktop/package.json, where it's
  actually imported (markdown-text.tsx, katex-memo.ts) -- it was
  never used anywhere else and was subject to the same pruning risk.
- _update_node_dependencies() collapses to a single
  `npm ci --workspace ui-tui --workspace web` call now that root has
  no dependencies of its own to protect, and keeps its original spot
  ahead of `_build_web_ui()` at both call sites in update_cmd.py --
  with no root-only dependencies left to protect, there's no reason
  for the Node refresh and the web build to run in any particular
  order relative to each other.
- hermes_cli/tools_config.py's post-setup Chromium-install path and
  hermes_cli/doctor.py's agent-browser check both now resolve through
  the same PATH -> Homebrew/Hermes-managed-node -> npx cascade
  (_find_agent_browser / _resolve_npx_bin) instead of hand-rolling
  their own node_modules/.bin lookups, so they can't diverge from what
  browser tools actually invoke at runtime.
- tests-js/package-json-lazy-deps.test.ts gets a lockfile-level check
  mirroring the existing camofox one, so a future regression that
  reintroduces agent-browser into package-lock.json fails this test
  directly instead of relying on manual review to catch it.

Fixes #43564.
2026-08-13 02:38:28 -07:00
doncazper 85020f2238 fix(plugins): isolate ownership by profile 2026-08-12 19:13:32 -07:00
Laith Weinberger 9e5e1740ed fix(browser): apply safety checks to browser_exec URLs 2026-08-10 10:45:44 -07:00
Laith Weinberger a1835c8c17 feat(browser): integrate Browser Use CLI 3.0 2026-08-10 10:45:44 -07:00
kshitij 0044876a11 fix(browser): guard against concurrent expired-session replacement race
After cleaning up an expired cloud browser session, re-check under
_cleanup_lock whether another thread has already created a replacement.
If so, return the live replacement instead of falling through to create
yet another session (which would be orphaned, or worse, a second
concurrent cleanup could destroy the replacement).

Follow-up to helix4u's expired-session renewal fix in PR #76356.
2026-08-02 11:18:41 +05:30
Gille 58e85f4314 fix(browser): replace expired cloud sessions 2026-08-02 11:18:41 +05:30
teknium1 490056d6d3 perf: lazy mcp SDK import + tool-discovery mtime cache + browser_tool import diet
Three cold-start cuts, measured with .venv python, PYTHONPATH=worktree,
median of 3-5 fresh subprocesses:

1. tools/mcp_oauth.py — availability now via importlib.util.find_spec("mcp");
   SDK classes (OAuthClientProvider et al.) import lazily on first use via
   _ensure_sdk_loaded(). Module-level names kept as None placeholders so the
   test-patch surface (patch.object(mcp_oauth, "OAuthClientProvider", ...))
   still works. import tools.mcp_oauth: 242ms -> 56ms (mcp SDK no longer
   loaded at import time).

2. tools/registry.py — discover_builtin_tools AST scan memoized in an
   mtime_ns+size-keyed disk cache at ~/.hermes/cache/tool_discovery_cache.json
   (atomic write via utils.atomic_json_write, best-effort/never raises;
   corrupt or missing cache -> full rescan + rewrite; per-file stat mismatch
   -> rescan just that file). Scan of 100 files: 158ms cold -> 3ms warm.

3. tools/browser_tool.py — top-level `import requests` and
   agent.auxiliary_client.call_llm moved to lazy first-use (PEP 562
   __getattr__ preserves patch("tools.browser_tool.requests.get") and
   patch("tools.browser_tool.call_llm") surfaces; internal call sites go
   through _lazy_call_llm which reads module globals so patches are honored).

Entry-point imports (median ms, before -> after):
  import model_tools   392 -> 245   (warm discovery cache)
  import cli           152 -> 151   (unchanged; cli doesn't hit these paths)
  import gateway.run   234 -> 230

Functional verification:
- get_tool_definitions(quiet_mode=True) under temp HERMES_HOME: identical
  sorted 30-tool name set before vs after (empty diff).
- Discovery cache: delete cache -> 158ms, second run -> 3ms; corrupt cache
  -> clean full rescan; touching one file -> single-file rescan (12ms).
- Tests green: tests/tools/test_registry.py, all test_mcp_oauth*,
  test_mcp_dashboard_oauth, test_mcp_tool_401_handling, all
  tests/tools/test_browser*.py, tests/hermes_cli/test_mcp_{config,startup,
  dashboard_oauth}.py, test_skills_tool_discovery_cache.py.
2026-07-29 10:02:03 -07:00
atakan g 7b18de5f40 fix: scope private URL policy per profile 2026-07-28 14:17:45 -07:00
Teknium 731aa0ccc9 fix(browser): stop stale cdp_url from stalling every startup by 10+ seconds
Tool-schema assembly at CLI/Desktop startup runs the browser-family
check_fns (browser, browser_cdp, browser_dialog, browser_vision). Each
of those gates called _get_cdp_override(), which resolves the configured
endpoint over HTTP (GET /json/version, timeout=10) — so a *stale*
browser.cdp_url pointing at a dead debug browser cost ~7 serial blocking
socket connects before the banner rendered. Measured on a real Windows
install with a dead http://[::1]:9222 config: 15.1s of an 18s launch,
with no warning or error — just mystery slowness. The value is easy to
leave behind: /browser connect writes a session-scoped env override, but
'hermes config set browser.cdp_url' persists forever while the debug
Chrome it pointed at dies on the next browser restart.

Split the helper:

- _get_cdp_override_raw() — returns the configured value (env var or
  config.yaml) with zero network I/O. Used by every is-it-configured
  gate: check_browser_requirements, _browser_cdp_check, _is_local_mode,
  _is_local_backend, _navigation_session_key, _should_inject_engine
  (via _is_local_mode), and the hermes doctor chromium-skip check.
- _get_cdp_override() — unchanged contract (raw + /json/version
  resolution), now only called on paths that are about to connect:
  session creation and the dialog-supervisor attach.

This follows the existing rule in check_browser_requirements ('do not
execute agent-browser --version here') and the browser.manage status
path, which already banned _get_cdp_override for exactly this reason
(test_browser_manage_status_does_not_call_get_cdp_override): schema
assembly must not perform blocking I/O.

A/B on the same machine, same dead endpoint: get_tool_definitions()
15.08s unpatched -> 1.89s patched, with browser_cdp/browser_dialog
still advertised (gate now keys off configuration, not reachability —
matching the documented lazy-supervisor contract in
_browser_dialog_check).

Adds a regression test asserting the browser_cdp check_fn never touches
the network.
2026-07-27 14:32:05 -07:00
Stoltemberg c89481db5e fix: add explicit UTF-8 encoding to all subprocess text=True calls (#53428)
On Windows with Chinese locale (GBK), subprocess.run(text=True) without
explicit encoding causes UnicodeDecodeError crashes. This fix adds
encoding='utf-8', errors='replace' to all subprocess.run() and
subprocess.Popen() calls that use text=True across 76 non-test Python files.

Fixes #53428 (master tracker for Windows GBK locale crash).

Note: credential_pool.py and electron changes excluded per reviewer request —
those will be submitted as separate focused PRs.
2026-07-24 11:45:57 -07:00
Vishnu 29899c2aa9 Fix headed browser sessions being killed after every turn
The per-turn `_cleanup_task_resources` unconditionally calls
`cleanup_browser`, closing the browser window immediately after each
bot reply. This makes headed mode (`AGENT_BROWSER_HEADED=1`) unusable
— the window flashes up and disappears on every response.

This mirrors the existing VM persistence pattern: skip per-turn
cleanup when headed mode is active and let the inactivity reaper
handle idle sessions instead. Full-session teardown on gateway
shutdown remains unconditional.

Also adds `browser.headed` config.yaml support and passes `--headed`
to agent-browser in local mode when configured, so users don't need
to rely solely on the `AGENT_BROWSER_HEADED` env var.

Closes #11020 (lead bug)
2026-07-18 10:07:46 -07:00
Teknium 0f102fa4dc feat(browser): store full snapshots on truncation; make eval denylist opt-in (#65923)
* feat(browser): store full snapshots on truncation; make eval denylist opt-in

Two harness fixes motivated by BU_Bench results where fixed-verb + lossy
observation cost Hermes heavily vs code-driven browser agents:

1. Snapshot truncation no longer loses content. When a snapshot exceeds
   the 8000-char threshold, the complete accessibility tree is saved to
   cache/web (same truncate-and-store pattern as web_extract) and the
   truncated view / LLM summary includes the file path plus a ready-made
   read_file call. Element refs beyond the cut are recoverable without
   re-snapshotting. Stored copies are force-redacted and capped at 2MB;
   content-hash filenames dedupe repeated snapshots of the same page.

2. The browser_console(expression=...) sensitive-primitive denylist is
   now opt-in via browser.restrict_evaluate (default false). The
   names-based denylist blocked legitimate DOM extraction — any selector
   or expression containing 'fetch', 'cookie', 'input', etc. — which
   crippled the agent's only programmatic page-inspection path. The
   SSRF/private-URL egress guards in _browser_eval are independent of
   this policy and remain always-on. browser.allow_unsafe_evaluate keeps
   its meaning (bypass the denylist) for configs that already set it.

* test: update None-guard test for stored-snapshot pointer in _extract_relevant_content

test_normal_content_returned pinned the exact return value; the summary
now carries a pointer to the stored full snapshot. Assert the summary
passes through and the pointer is present instead.

* feat(browser): align snapshot threshold with web_extract's 15k char budget

SNAPSHOT_SUMMARIZE_THRESHOLD 8000 -> 15000, matching
web_tools.DEFAULT_EXTRACT_CHAR_LIMIT so the snapshot and web_extract
truncate-and-store paths give the model the same per-page budget.
_truncate_snapshot's default max_chars now follows the constant.
Invariant test added; docs (en+zh) and CLI tip updated.
2026-07-16 23:41:26 -07:00
Ray 6a58badfdc fix(browser): guard Camofox eval private pages
Extends the browser private-network eval guard to the Camofox backend.
On main, _browser_eval() returned early in Camofox mode before running the
shared private-URL literal pre-scan and before re-checking the page URL
after eval, leaving Camofox as a sibling backend that could execute
browser_console(expression=...) against private/internal targets.

- move the eval private-URL literal pre-scan before the Camofox early return
- add a Camofox current-page private-URL probe via the evaluate endpoint
- withhold Camofox eval results when the page is now private/internal

Follow-up to browser private-network hardening in #56173, #56526, #56664.

Salvage of #56764 by @rayjun (rayoo), cherry-picked to preserve authorship.
2026-07-02 13:10:30 +05:30
srojk34 4612ee9464 security(browser): re-check private-network guard after browser_back navigation
Every other content-returning browser tool entry point
(browser_snapshot/vision/console/eval, and click/type/press via
_blocked_private_page_action) re-checks window.location.href against the
private/internal/cloud-metadata floor after the page could have changed --
because a redirect chain or client-side navigation can land on an address
the initial browser_navigate preflight never saw. browser_back was the one
navigation-triggering entry point missing this: it called
_run_browser_command(..., "back", []) and returned the resulting URL
straight to the model with no re-check.

On a cloud/CDP (non-local) backend, if browser history contains a
private/internal address (e.g. a prior redirect touched an internal host),
browser_back would navigate the live browser there and hand the URL back
to the model with no guard -- the exact class of gap the private-page
guard exists to close, just on the one entry point it hadn't reached yet.

Re-check happens after the navigation succeeds (not before, unlike
click/type/press) since it's the resulting page -- not the one being left
-- whose safety matters. A failed back navigation (no history) skips the
check entirely since nothing changed. Verified live: the new regression
test fails (returns the private URL instead of a blocked payload) on the
pre-fix code and passes after.
2026-07-01 20:01:55 +03:00
dsad a4af257a6d fix(browser): extend private-network guard to browser_console 2026-07-01 05:23:17 -07:00
claudlos 0a75616514 security(browser): enforce cloud-metadata floor on all backends; CDP is non-local
browser_navigate's always-blocked cloud-metadata floor (169.254.169.254,
metadata.google.internal, ECS/Azure/GCP IMDS) was gated on
`not _is_local_backend()`, contradicting both the adjacent comment and the
is_always_blocked_url docstring ("denied regardless of backend"). A default
local headless Chromium on a cloud VM — or an off-host CDP browser — could
navigate to IMDS and read instance credentials into the model context. Make the
floor unconditional on the initial-nav and post-redirect paths.

Also: _is_local_backend() ignored a CDP override while _is_local_mode() honors
it, so an off-host CDP browser was treated as "local" and skipped the broader
private/internal SSRF check too. Treat a CDP override as non-local.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 05:09:35 -07:00
yongyong 937e56be92 fix(browser): block bracketed sensitive eval primitives 2026-07-01 05:04:41 -07:00
yongjin a0beb52a50 fix(browser): harden browser tool safety boundaries
Add policy gates and output redaction for browser/CDP surfaces, strengthen session ownership tracking, and block credential-like query parameters before third-party browser/web backends receive URLs.

Inspired by the agbrowse review: keep local browser magic-link flows possible while preventing cloud reader/browser escalation from receiving opaque token, code, signature, or key query parameters.
2026-07-01 05:04:41 -07:00
kshitijk4poor 83ae65487e test(browser): cover guard-inactive + camofox short-circuit paths; fix blank lines
Review follow-up on the private-page action guard:
- Add test_guard_inactive_does_not_block_or_probe: when the SSRF guard is
  inactive (local backend / allow_private_urls), click/type/press must proceed
  WITHOUT probing the page URL. This is the branch most likely to silently
  regress if the guard condition is inverted; a mutation check (flipping the
  condition) confirms the test fails as designed.
- Add test_camofox_short_circuits_before_guard: camofox mode returns from the
  dedicated camofox_* path before the guard runs; guards never consulted.
- Fix PEP8: 3 -> 2 blank lines before _blocked_private_page_action.
2026-07-01 13:56:49 +05:30
dsad 3e4c138251 fix(browser): block private-page interactions after eval navigation 2026-07-01 13:56:49 +05:30
kshitijk4poor c626dded13 refactor(redact): consolidate CDP-URL log redaction into one chokepoint
The session-log fix (browser_tool._sanitize_url_for_logs) and the
supervisor attach-timeout fix (CDPSupervisor.start) both composed the
same three redactors (redact_sensitive_text -> _redact_url_query_params
-> _redact_url_userinfo) to mask CDP endpoint credentials. Two copies of
one policy drift: tune one site (e.g. add fragment masking) and the other
silently re-leaks.

Promote that composition to a single public helper redact_cdp_url() in
agent/redact.py -- the one place the CDP-URL redaction policy lives -- and
route both call sites through it (_sanitize_url_for_logs becomes a thin
wrapper; the supervisor imports the helper instead of re-composing the
private redactors). Add direct unit tests for the seam covering query
tokens, multiple credentials, userinfo passwords, plain-URL passthrough,
non-string/exception coercion, and None.

No behavior change at the call sites; both leak paths remain closed.
2026-07-01 13:43:58 +05:30
srojk34 265da9cadb fix(browser): redact CDP URL token in _create_cdp_session log and supervisor timeout
PR #54851 added _sanitize_url_for_logs() and wired it into the three log
sites inside _resolve_cdp_override(). A fourth site was missed:
_create_cdp_session() logs the already-resolved cdp_url unconditionally,
and CDPSupervisor.start() interpolates the raw cdp_url[:80] into the
attach-timeout TimeoutError (which _ensure_cdp_supervisor() logs with %s).
Both leak query-string credentials (e.g. ?token=secret from hosted CDP
providers) into Hermes logs.

Sanitize the URL at both remaining sites. The raw URL is preserved
unmodified in the returned session dict and used for the real connection;
only the logged/error representation is redacted.

Salvaged from #55883.

Co-authored-by: srojk34 <286497132+srojk34@users.noreply.github.com>
2026-07-01 13:43:58 +05:30
Ruzzgar 576424cc1c fix(security): redact browser CDP endpoint logs 2026-06-29 04:25:26 -07:00
teknium1 9f97915163 fix(browser): route open-timeout base through _safe_command_timeout
Wire the salvaged _safe_command_timeout() guard into the surviving
open-timeout call site. _get_open_command_timeout() feeds the
browser_navigate 'open' path; this closes the last call site that
could observe a None timeout from a torn cache (#14331), since the
original PR's max(_get_command_timeout(), 60) site no longer exists
on main (now routed through _get_open_command_timeout).
2026-06-29 02:24:57 -07:00
Sanjay Santhanam c79e6bceae fix(browser_tool): resolve race in _get_command_timeout cache returning None (#14331)
# Conflicts:
#	tools/browser_tool.py
2026-06-29 02:24:57 -07:00