Python Fire serializes the return value of a wrapped function but does not
use that value as the process exit code. Error paths in main() that used
or therefore caused the process to exit 0, swallowing
fatal errors and argument-validation failures.
Raise SystemExit(1) on every error path so batch_runner returns a non-zero
exit code when it cannot run. Success paths (e.g. --list_distributions) are
left unchanged.
ClosesNousResearch/hermes-agent#86524.
Docker Desktop writes fpath=(~/.docker/completions ...) into .zshrc.
The referenced-script walk then opened that directory, saw a non-regular
file, and fail-closed — blocking source ~/.zshrc on every terminal
command. Directories are not scripts; devices stay fail-closed.
Fixes#86864.
Legacy custom_providers configs commonly used short/placeholder
api_keys ('123', 'm') for local no-auth services like Ollama --
harmless for the endpoint itself, since Ollama accepts any key or no
key. A stricter has_usable_secret(value, min_length=4) gate added
later now rejects these, but only the credential-POOL resolution path
lacked the same "no-key-required" exemption every OTHER resolution
path in this file already has for exactly this scenario:
- The config-based custom_providers fallback (non-pool path) already
ends with `api_key or "no-key-required"`.
- The "actual" provider's local-offline path already injects
ACTUAL_LOCAL_NOAUTH_PLACEHOLDER before the usable-secret gate for a
loopback base_url.
- _try_resolve_from_custom_pool() was the one gap: it returned the raw
short pool credential unchanged, which then failed the downstream
has_usable_secret() gate with a generic "No usable credentials found
for custom" error that contradicts setup.status ("configured
credentials" vs "runtime failed"), sending users hunting in the
wrong direction.
Fixed by substituting the same "no-key-required" placeholder when the
pool's stored credential fails has_usable_secret() AND the base_url
resolves to a loopback hostname (using the existing _loopback_hostname
helper, matching the exemption scope the issue itself requested:
localhost/127.0.0.1/::1 only, not arbitrary remote endpoints with a
genuinely-too-short key).
Added 4 regression tests extending the existing
test_runtime_provider_resolution.py file, following its established
credential-pool mocking pattern: the exact reported 3-char repro
('123'), a 1-char case, a non-loopback sanity check confirming the
exemption stays scoped (a short key for a remote endpoint is NOT
silently exempted), and a sanity check that a genuinely usable
loopback key passes through unmodified. Verified as a genuine
regression by reverting the fix and confirming 2 tests fail with the
exact raw short key leaking through unchanged.
59/59 pass in the extended test file; 14/14 across two more related
custom-provider test files (no regression).
Apply findings from /simplify-code 3-agent review:
1. Extract _install_paired() inner helper — the Ctrl, Alt, and Shift
sections all repeated the same mok+csiu sequence generation pattern
(~30 lines of duplication). Now each section builds a dict and
delegates to _install_paired(modifier, mapping).
2. Replace 10 hardcoded Ctrl+digit lines with a loop matching the
Ctrl+letter pattern above it.
3. Fix misleading comment: claimed 'Ctrl+0 doesn't produce a control
byte' but chr(ord('0') & 0x1F) = 0x10 = ControlP. The code was
correct (maps directly to Keys.Control0..9); only the comment was
wrong.
4. Add comment explaining why Shift+letter maps both lowercase and
uppercase codepoints (some terminals send the already-shifted
codepoint with modifier=2).
5. Test fixture: snapshot/restore ANSI_SEQUENCES in teardown so 294
mappings don't leak into sibling test files (global mutable state).
Commit 4c34eeb416 stopped pushing the Kitty keyboard protocol (CSI >1u)
because Ctrl+C arrived as ESC[99;5u instead of \x03, breaking SIGINT.
But modifyOtherKeys level 2 (CSI >4;2m) was kept so Shift+Enter stays
distinguishable from Enter.
Under modifyOtherKeys=2, terminals re-encode EVERY Ctrl+key combo as
ESC[27;5;<codepoint>~ instead of the raw control byte. prompt_toolkit
3.x only maps ESC[27;5;13~ (Ctrl+Enter = Ctrl+M); all other Ctrl+letter
combos are unmapped and leak as literal text or get swallowed — breaking
Ctrl+A, Ctrl+C, Ctrl+D, Ctrl+E, Ctrl+K, Ctrl+R, Ctrl+U, Ctrl+W, Ctrl+Z,
etc. Shift+letter combos (ESC[27;2;<codepoint>~) have the same problem,
causing the 'caps locked sessions' symptom where typed text appears
corrupted or stuck.
Fix: add install_modify_other_keys_aliases() to pt_input_extras.py that
populates prompt_toolkit's ANSI_SEQUENCES dict with 294 mappings covering:
- Ctrl+letter (a-z): ESC[27;5;<code>~ and ESC[<code>;5u -> Keys.ControlA..Z
- Ctrl+digit (0-9): same formats -> Keys.Control0..9
- Ctrl+symbol ([ \ ] ^ _ @ Space): same formats -> matching Keys.Control*
- Alt+letter (a-z, A-Z): both formats -> (Escape, <letter>) tuple
- Shift+letter (a-z, A-Z): both formats -> uppercase character
Uses setdefault semantics — never clobbers existing mappings from
install_shift_enter_alias or install_ctrl_enter_alias. The Ink TUI
(Node.js) already handles this via a regex parser; prompt_toolkit 3.x
uses dict lookup only, so we populate the dict.
Refs #56684, #87711.
`cargo build 2>&1 | tail -20` exits with tail's 0 even when the build
failed — bash without pipefail reports the last pipeline command's
status, and `cmd || echo failed` swallows the status the same way. The
model reads exit_code: 0 as a strong success signal and can conclude a
build passed while the visible output says it failed (community report,
Windows Rust builds; not platform-specific).
Two-part fix, mirroring OpenCode's prompt-side approach plus a
result-side backstop they don't have:
- Tool description now forbids piping builds/tests through
tail/head/cat (output is already auto-truncated + spilled to a file)
and warns that pipes/|| fallbacks mask exit codes.
- New annotate_masked_success() in tools/terminal_hints.py: when
exit_code == 0, the command shape can mask an upstream status
(top-level pipe into a passthrough consumer, or || echo/true), AND
the output carries strong tool-specific failure shapes (rustc,
cargo, pytest, gcc, npm, make, ninja), attach an advisory 'hint'
telling the model to treat the run as failed and re-run bare.
exit_code itself is never modified. Search/content heads
(grep/rg/echo/printf/...) are excluded to avoid false positives on
pipelines whose output legitimately contains error text.
E2E-verified through the real terminal tool path: hint fires on masked
cargo-style failures, silent on bare commands, clean pipes, and
grep/printf pipelines. 42 targeted tests pass.
CI caught two rotation-path regressions from the unbounded clone: the #47202
pre-publish flush writes the rotator's OWN input transcript to the parent
(above the start-watermark), and the clone was duplicating it into the child
alongside the handoff. publish_compression_child gains watermark_ceiling —
the MAX(id) captured immediately BEFORE that flush — so only rows in
(watermark, ceiling] (genuinely foreign concurrent appends) clone across.
Ceiling capture failure falls back to no tail preservation (historical
behavior) rather than risking duplication. Ceiling-exclusion test added.
CI caught the sibling site the in-place fix missed: legacy (non-in-place)
compression rotates via publish_compression_child, where a mid-summary
append previously stranded in the closed parent. Same watermark + pure-SQL
column clone as archive_and_compact, with session_id rewritten to the child.
Lineage-guard test flipped to pin the appends-flow-freely contract; rotation
watermark tests added (tail follows the child; None = historical behavior).
Redesign of the #75316 class (supersedes the approach in PR #87307).
Root cause family: the compression lock fenced ORDINARY transcript appends
for the whole slow provider-summary call. Turns died as
session_persistence_failed whenever a message overlapped a compression
(#74568, #77386, #75083), stale dead-PID locks blocked writes for the full
TTL, and the busy-wait mitigation (#75264) was an order of magnitude shorter
than real summaries. Separately, the commit archived from a pre-call
snapshot, so rows appended mid-compression were swept into the archive.
Design: the commit transaction is already exclusive — no lock phases needed.
1. Appends never check compression_locks. The lock's only job is stopping
two compressions colliding; it keeps that job. The whole stale-lock /
busy-wait symptom family dies as a class.
2. Watermark captured in the DB at compression start
(get_active_message_watermark = MAX(id) of active rows) — not from
in-memory message dicts, which carry no row ids in production.
3. archive_and_compact(watermark=, lock_holder=): one transaction verifies
the holder still owns an unexpired lease (a reclaimed lease cannot
publish a stale compaction), archives the snapshot, inserts the compacted
set, and re-sequences the concurrent tail (id > watermark) via a
pure-SQL column clone — every column except id survives byte-exact
(api_content, platform_message_id, reasoning sidecars, token counts),
FTS triggers index the clones naturally, originals stay archived and
recoverable. watermark=None preserves the historical behavior.
Removed: the append-side compression fence in _check_transcript_write_guards
(with rationale note), making the _COMPRESSION_BUSY_WAIT_S retry lane
unreachable from append paths (kept for other callers).
Tests: 12 new (watermark contract, column-exact clone, commit fence incl.
lease-lost/expired/rollback failure injection, append-vs-commit race);
busy-retry suite flipped to pin the new contract; sabotage-verified (5 fail
with the watermark disabled, 12 pass restored); E2E through the real
compress_context seam with a mid-summary append landing and surviving.
Some MCP OAuth providers (notably Supabase) return a client_secret from
dynamic client registration but omit token_endpoint_auth_method. The MCP
SDK defaults the missing method to "none", so the token exchange omits
client_secret and the server rejects it (HTTP 422 "Required parameter:
client_secret"), looping the browser consent page.
This resolves the whole class, not just one provider:
- Storage layer (HermesTokenStorage): coerce secret-bearing client info
with missing/none auth method to client_secret_post on both read and
write, persisting the corrected shape.
- Both live provider paths (tools/mcp_oauth.py HermesOAuthClientProvider
and tools/mcp_oauth_manager.py HermesMCPOAuthProvider): coerce
in-memory client info immediately before token exchange and refresh.
- Accept the full 2xx range on token and refresh responses (Supabase
returns 201 Created), instead of the SDK's exact-200 check.
- Redact token response bodies from error messages and logs on
malformed responses.
The Figma-specific request-time default (apply_oauth_provider_defaults)
remains; this generalizes the same bug class for every DCR provider.
Fixes#29680. Supersedes #34274 and #35700 (201-only variants).
Follow-up to the #28953 salvage:
- Extract _resolve_block_from_details() so resolve_pre_tool_block and
_dispatch_pre_tool_call_hooks share ONE fail-closed approval-gate
implementation. This also gives the new dispatcher the observability
context wrapping around request_tool_approval that the original PR's
inlined copy lacked.
- Update sibling tests that patched resolve_pre_tool_block at the three
migrated dispatch sites to patch _dispatch_pre_tool_call_hooks with the
(block_message, modified_args) tuple contract.
Verified: 448 targeted tests green; E2E with a real shell hook in an
isolated HERMES_HOME rewrote a live write_file call (path + content)
through handle_function_call, with block and negative paths intact.
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.
- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
and returns (block_message, modified_args); modify directives
shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
{"action": "modify", "args": {...}} and Claude Code-compatible
{"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).
Salvaged from PR #28953. Best fix for #18988.
KDE/Qt apps report [0,0,0,0] bounds for elements that are perfectly
clickable by index (live QA: all 50 of kcalc's zero-rect elements,
including every radio button). Serializing that as a plausible rect
invites a model to derive coordinate=[0,0] and click the screen corner.
- _element_to_dict: zero rect -> bounds: null
- _format_elements: '@ bounds-unknown (click by element index)' instead
of the fake rect in the summary line
- malformed bounds fail open (unchanged serialization)
Live-proven on real kcalc (cua-driver 0.20.0): 50 elements now null, 0
zero-rect leftovers, summary annotated, real rects preserved, and a
null-bounds radio button still clicks fine by index.
Live complex-action QA on a real KDE desktop (kcalc + kate multi-app
flows) found two dispatch gaps:
1. Wrong-window input reported as success. Input actions deliver to the
backend's sticky target (last capture/focus_app); the app= argument
models routinely pass on the input call itself was silently dropped.
Proven live: with kcalc sticky, type(text='777', app='kate') returned
ok:true and typed 777 INTO KCALC. New guard: provable mismatch
(both names known, neither substring of the other — list_windows
names are localized/variant) refuses with input_target_mismatch and
a one-call fix instruction. Unknown current target fails open so
legacy no-app flows are untouched.
2. Near-miss unknown actions were dead ends. A model emitting 'hotkey'
got a bare unknown-action error. Suggestion map now names the real
action ('did you mean key?') without aliasing — we never repair bad
model output, we just point at the schema.
Also documents the verified-lost-keystroke rung in the computer-use
skill: KTextEditor (Kate/KWrite) discards synthetic X keystrokes at the
toolkit level — foreground type reports ok but AX shows nothing arrived,
and a raw XTest control fails identically outside our stack. Guidance:
after one verified-lost round trip, switch to file/DBus I/O instead of
looping the ladder.
Live proof on the fixed build: mismatch refused, kcalc display clean,
same call after capture(app=kate) succeeds, 'hotkey' suggests 'key'.
11 new tests; 158 sibling tests green.
A headless Mac or asleep built-in panel leaves ScreenCaptureKit with 0
shareable displays while TCC grants pass — health_report stays ok and
every capture silently returns 0x0 (#67165). Guard at the report seam
(_apply_display_count_guard, both real and fallback paths): flips the
screen_capture_capability check to fail with recovery actions (wake
display / HDMI dummy / virtual display) and downgrades ok -> degraded.
The empty-discovery reason ladder gains the matching darwin rung.
Composed from #52949 (sujeet111) and #67259 (webtecnica); both PRs
predate the doctor rewrite and the envelope normalization on main, so
this reimplements their shared intent at the current seams.
Co-authored-by: Sujeet <64351924+sujeet111@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
capture() read any non-None pid/window_id as a request for exact-window
targeting. Several providers emit every declared schema property on every
tool call, zero-filling unused optional integers, so those calls arrive as
pid=0, window_id=0. The exact-target branch was then entered, the caller's
app= was discarded, _positive_int(0) returned None for both ids, and the
capture failed with a message pointing at pid/window_id. For that class of
model capture(app=...) and frontmost capture never worked at all.
Normalize non-positive ids to None before the branch decision so dispatch
falls through to app/frontmost discovery. Malformed non-numeric ids are
deliberately not treated as placeholders: they still reach the existing
validation error instead of being silently ignored.
Fixes#81333
The first hardening (deterministic worker start) still went red on its
own PR's CI: the 0.1s deadline can ALSO expire between worker start and
handle_function_call — argument parsing, hooks, and cold imports run in
that window, and the timeout interrupt then returns the middleware
without dispatching (first_started unset; observed 0.15s tool error vs
0.1s budget in the slice-4 log).
Deadline raised to a 1.0s floor: still fast, ~7x the worst observed
preamble, and the submit() sync from the first commit keeps the
countdown anchored to worker start. The clarify human-wait test's sleep
rises to 1.3s so it still outlasts the deadline (the relationship the
test exists to pin). Sabotage v2 injects BOTH races (0.3s thread start
+ 0.5s preamble): hardened passes, un-hardened reproduces the exact CI
failure.
test_sequential_tool_timeout_emits_result_and_continues failed twice in
two days on unrelated computer_use PRs (slice 4, xdist): assert
first_started.is_set() -> False. Root cause: the sequential timeout path
computes deadline = now + timeout_s right after executor.submit(); with
the test's 0.1s deadline, a loaded CI worker can take longer than the
whole deadline just to START the pool thread, so the future is cancelled
before the tool ever dispatches.
Fix: an autouse fixture subclasses DaemonThreadPoolExecutor so submit()
blocks (bounded 10s) until the worker callable has begun — the deadline
now races the tool, not the thread scheduler, which is what these tests
mean to pin. Timeouts stay tight (0.05s), so the suite stays fast.
Sabotage-proven: a 0.3s injected thread-start delay reproduces the exact
CI failure without the fixture and passes with it.
Two live-QA findings from a locked KDE desktop (real cua-driver 0.20.0):
1. capture() with zero discovered windows returned a bare
'capture mode=ax 0x0' — no hint that the desktop session was LOCKED,
which freezes renderers and hides windows. New
_empty_discovery_reason() names the dominant causes in order: locked
session (loginctl LockedHint probe, fail-safe), missing DISPLAY,
else a pointer at hermes computer-use doctor. Surfaced through the
existing window_title -> summary path, so the model and the user see
it inline.
2. _call_tool_via_cli retried 'daemon is not running' 4x with ~3.5s of
backoff sleeps — a permanent condition for that invocation (the CLI
transport needs the machine-wide daemon; Hermes' MCP runtime does
not). Now fails fast on the first attempt with a message naming the
split. Transient empty output (EAGAIN congestion) keeps the retry
loop — pinned by test.
Live-verified on the locked desktop: capture now reports the lock and
the unlock action; CLI fallback errors immediately with the transport
explanation. 7 new tests; 191 sibling tests green.
Regression test for the _ref_map merge (salvaged from #79515): the live
0.19.3 driver splits action refs into refs[] while content_refs re-lists
every node with empty actions; the empty entries must not clobber the
action-bearing ones. Caught live: every typed click refused with
browser_ref_stale until the merge fix.
The Desktop's content-based truncation-target resolution (and reactions)
address persisted turns by row_id, but session.history loaded the
transcript without include_row_ids=True, so _history_to_messages had no
stamp to forward and the projection silently stripped the one durable
address clients can use. Discovered live-testing the #87294 client flow:
resolveDurableRowId saw 0 stamped rows and degraded every edit to a
plain resubmit.
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):
- bounded serve flags corrected: the daemon accepts
--session-policy/--approve-session-policy, not the docs'
--capability-manifest names (which it rejects). Verified end-to-end:
a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
TTY) and its token is a legacy compatibility path disabled by default
on current drivers (per the live browser_prepare schema). Kept as a
passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
with cua-driver's trusted-launcher grant. config opt-in
computer_use.grant_existing_profile: true appends
--grant existing-profile to the standard-mode MCP spawn (MCP
initialize verified accepting the flag). Default false = attachment
keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
ladder: config grant > bounded manifest > YOLO; token = legacy.
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:
- hermes computer-use browser-approve: CLI passthrough that mints
cua-driver's five-minute single-use attachment token for one exact
(pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
browser_route), forwarded only for existing_profile and only as a
non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
private per-session embedded daemon launched with
--capability-manifest/--approve-capability-manifest; missing manifest
fails loudly. 'unrestricted' is deliberately NOT a config value —
it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
rungs and the isolated-profile-first default.
E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
The plugin's _Runtime.run_in_session wrapper serves every mark/event it
emits (turn start/end, approvals, subagent marks) and runs synchronously
on the agent's conversation thread. It passed no timeout, so the host's
run_in_session default (timeout=None) made each mark an UNBOUNDED native
call. With a wedged native Relay pipeline the agent blocked between API
calls with zero activity ticks — observed live 2026-08-15: two cron jobs
died at the 600s inactivity kill and a gateway chat session at 1800s,
all with last_activity="API call #N completed".
The core's scope push/pop/flush/close sites were bounded with
_SCOPE_OP_TIMEOUT after the 2026-08-10 delegation stall; the plugin's
event marks were the missed sibling class.
Changes:
- plugins/observability/nemo_relay: the wrapper always passes
timeout=relay_runtime._SCOPE_OP_TIMEOUT (10s) to the host. A breach
costs one telemetry span, never the agent; it also sets scope_errored
(so close_session skips the ATIF export for the wedged session) and
warns once so the sick pipeline is visible.
- tests/plugins/test_nemo_relay_bounded_marks.py: proves the budget
reaches the host (fails on the pre-fix code — sabotage-verified),
a TimeoutError flags the session and disables its export, and the
generic error path keeps its scope_errored contract.
Both quarantine wrappers (_run_quarantined_install in main.py and
_run_install_cmd in _install_repair.py) renamed live hermes*.exe shims
aside before invoking the installer, but only renamed them back on
FAILURE. A SUCCESSFUL install that never rewrites entry points — uv
audits an already-satisfied editable install as a no-op — left the
shims quarantined as hermes.exe.old.<ms> and `hermes` disappeared from
PATH after a green install (#75584; reproduced live on a Windows
install recovering from the #86735 self-lock deferral).
Switch both sites from except/re-raise to try/finally so restore runs
on every path. _restore_quarantined_exes already skips shims the
installer actually replaced, so fresh output is never clobbered and
failure behavior is unchanged.
Regression tests cover both wrappers x {no-op success, rewriting
success, failure}; the no-op cases fail on the previous code.
Session-finalize hooks ran synchronously on the gateway event loop from
three call sites (shutdown drain, session-expiry watcher, /new reset).
A plugin hook doing heavy blocking work froze the whole loop: adapter
heartbeats stopped, the drain machinery could not run, and systemd
eventually SIGKILLed the process mid-export. Observed live on a
multi-day 4.7G session where the nemo_relay observability plugin
serialized a full-session ATIF trace inside on_session_finalize.
Changes:
- gateway/run.py: new GatewayRunner._finalize_session_off_loop()
dispatches hermes_cli.lifecycle.finalize_session via the gateway
executor under asyncio.wait_for (10s budget), mirroring
_cleanup_agent_resources_off_loop (#53175). Shutdown finalize and
the session-expiry watcher now use it.
- gateway/slash_commands.py: /new reset path uses the same helper.
- plugins/observability/nemo_relay: ATIF export is now bounded
(HERMES_NEMO_RELAY_ATIF_EXPORT_TIMEOUT_S, default 30s) and skipped
entirely for sessions whose Relay scope operations already errored
(their exporter state is unreliable and the export can be
pathologically slow).
- tests/gateway/test_finalize_session_off_loop.py: regression tests
proving the loop stays live under a wedged hook and the budget is
enforced.
Two related failure modes after a crashed/interrupted fetch on a shallow
clone (git clone --depth 1 installs):
1. STALE LOCK WEDGES EVERY FETCH. A killed fetch can leave .git/shallow.lock
behind; every later 'git fetch' then fails with 'Unable to create
.../shallow.lock: File exists'. 'hermes update --check' reported a hard
fetch failure, and the passive banner check swallowed the exception and
compared stale refs. Add hermes_cli.gitlock.clear_stale_git_locks(), a
guarded sweep (age + git-process check so a live fetch is never yanked)
wired into the check path, the apply path, and the banner's passive check.
2. SHALLOW TIP-SHA COMPARE FALSE-POSITIVES. On a shallow clone the check
cannot count commits, so it compares tip SHAs. Local cherry-picks on top
of the remote tip (e.g. re-applied local patches) make HEAD differ from
origin/main even though HEAD already contains it — a false 'update
available' banner. Add hermes_cli.gitlock.is_ancestor_of_head() and use
'git merge-base --is-ancestor' in the CLI check and banner paths before
reporting an update. Mirror in the desktop (update-count.ts gains an
isAncestor input; main.ts probes merge-base --is-ancestor).
Tests: tests/test_gitlock.py (9) covering stale/young/no-lock/no-repo sweeps
and ancestry true/false; update-count.test.ts +3 for the isAncestor path.
Follow-up to #86916. That fix gave named sessions their own daemon
(socket/log/pid) and their own provider browser — but on a SHARED local
Chrome / CDP browser, a fresh named daemon still attaches to the first
existing page, the same page a sibling daemon may hold. A named session
that never calls new_tab() could still stomp another's tab.
browser_exec now prepends a small preamble to the model's code for named
sessions on shared browsers: once per daemon process (marker keyed by
uid + BU_NAME + daemon pid), it creates a fresh tab via
Target.createTarget and switch_tab()s onto it before any model code
runs. Private per-name browsers (provider-keyed bu-named-<name>, or
direct-API Browser Use cloud) skip the preamble via an internal env
sentinel popped before launch — there's nobody to collide with, and the
extra tab would leak.
Best-effort by design: if the preamble's CDP calls fail, behavior
degrades to pre-fix, never blocks the exec.
E2E against a shared headless Chrome with the STOCK harness: two named
sessions issuing bare js() writes (no new_tab) kept distinct state
(EDGE-A/EDGE-B read back intact); the sabotage run without the preamble
reproduced the clobber (both read EDGE-B). Removes the dependency on the
upstream browser-harness tab-isolation PR for correctness.
Desktop's send path pre-analyzed every attached image with the auxiliary
vision model serially, BEFORE dispatching the turn (_enrich_with_attached_
images). Users saw the progress box sit idle 25s-4min for messages that
take ~4s in the CLI; failures were silently swallowed, and touching
another session during the window killed the turn with zero API calls
(#83291). The prepended description also poisoned session auto-titles
(#82339).
Replace pre-analysis with _build_image_ref_message: reference the image
paths in the message and let the agent analyze them in-loop with
vision_analyze — its own retries, visible tool progress, and the turn
starts immediately. This is exactly how the @folder: reference path
already behaves, which responds in seconds for the same images.
Native-vision routing is unchanged; only the "text" mode (non-vision
main model / codex_app_server) loses the blocking submit-path calls.
Tests: tests/tui_gateway/test_image_ref_message.py (6 cases) including
a guard asserting the submit path never invokes the vision tool;
sabotage-verified (restoring the old blocking body fails 5/6).
_select_plugin_image_gen_provider hardcoded image_gen.use_gateway = False.
The managed (Nous-subscription) flow writes use_gateway = True via
_write_provider_config, then this selector runs AFTER it — so picking FAL
through Nous Portal silently persisted provider: fal, use_gateway: false
and every generation billed the user's personal FAL_KEY instead of the
subscription (real incident: key drained to zero-balance lock while the
managed route sat unused).
Fix the class, not the site:
- _select_plugin_image_gen_provider gains the same use_gateway kwarg its
video twin (_select_plugin_video_gen_provider) already had; all four
call sites pass use_gateway=bool(managed_feature), matching the video
call sites, TTS, STT, browser, and web.
- Active-provider detection (the checkmark in `hermes tools`): the
image_gen_plugin_name branch now defers managed entries to the
managed_feature branch and requires use_gateway OFF for direct-key
entries — mirroring the video branch's existing guard, so a managed
FAL pick and a direct-key FAL pick no longer both report active.
Runtime side (prefers_gateway("image_gen")) was already correct; the bug
was purely the setup-time writer.
Tests: new tests/hermes_cli/test_imagegen_managed_gateway.py (3 cases:
managed flag survives, direct pick still clears, image/video selector
contract parity). Sabotage-verified: restoring the hardcoded False fails
2/3. Neighboring hermes_cli provider/managed suites: 180 passed.