Commit Graph

13123 Commits

Author SHA1 Message Date
Brooklyn Nicholson 6559448306 Merge origin/main into bb/bot-mode-design-system
Keeps plugin.js and its new .mjs test deleted. main's closed-chat fix
(7c91079) landed in both; it is a real behaviour change, so the commit that
follows ports it onto the split modules rather than dropping it with the
files. Its core half — focusWorkspaceOwnerSessionTile and the
host.focusOpenWorkspaceSession verb — merged cleanly and is used as-is.
2026-08-27 22:51:30 -05:00
kshitijk4poor 0976ceaa98 fix(macos): harden anchor alias failures — warn, unique staging, marker-last
Follow-up to the #95605 salvage, closing the review findings:

- _copy_alias no longer swallows OSError silently: it warns (a leftover
  alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
  (update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
  it now asserts the whole layout (anchor + aliases) is complete, so a
  partially-materialized alias set can never read 'active' in doctor —
  the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
  managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.

5 new regression tests.
2026-08-28 09:05:19 +05:30
Zheqing Zeng 37bccf343e fix(macos): scrub gate env, refuse EACCES, normalize marker paths
Review fixes from kokhlo's live-hardware review:

- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
  PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
  PYTHONHOME=<venv> boots a staged copy that would otherwise die with
  "No module named 'encodings'", papering over the exact prefix
  failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
  images) still skip; EACCES after our own chmod now refuses the
  install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
  so the managed-runtime layout (cpython-3.11-macos-* symlinked to
  cpython-3.11.15-macos-*) no longer reports stale on a fresh install.

Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
2026-08-28 09:05:19 +05:30
Zheqing Zeng 1c92b95a12 test(macos): cover every boot-gate refusal branch directly
The gate is the never-brick guarantee, so each direction gets its own
test: nonzero exit (dyld/encodings crash), build-time prefix leak,
timeout, and the deliberate OSError skip (a binary that cannot execute
here means the symlinked venv was equally dead — installing cannot make
things worse).
2026-08-28 09:05:19 +05:30
Zheqing Zeng aa72df4b42 fix(macos): re-land dylib-complete TCC interpreter anchor
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).

Re-land:

- Keep the signed real-file copy of bin/python (identifier-pinned
  via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
  symlinks. Copies boot on every build we could reproduce and keep
  the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
  devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
  plus the venv prefix, abort and leave the live venv untouched
  on failure.

Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.

Closes #95596.
2026-08-28 09:05:19 +05:30
Brooklyn Nicholson 595ee92289 fix(commands): put desktop slash metadata on the CommandDef
Inference is now any args_hint without subcommands → text. Mixed is the
only remaining hint-token path. desktop= and the few argument_mode
overrides live on the registry entry; the side tables are gone. Catalog
aliases get their own dict copy. Composer tests seed the catalog so
/goal stays mixed without an overlay row.
2026-08-27 22:05:40 -05:00
Brooklyn Nicholson b61408e95e feat(commands): attach desktop slash metadata to the registry
New commands and plugins declare argument_mode and desktop availability
on CommandDef / register_command. commands.catalog ships that map so
desktop does not need a second command list.
2026-08-27 22:05:40 -05:00
Brooklyn Nicholson bbcf5691a6 Merge origin/main into bb/bot-mode-design-system
main's Bots-home flash fix (6f8be61) landed in plugin.js, which this branch
deletes. Both files stay deleted: the guard it adds (botOpenInFlight gating
botsHomeMayOpen) protects a surface this branch removes, and the two sites
where it renamed the generation bump to cancelBotOpen already bump here via
bumpBotOpenGeneration in shared.ts.
2026-08-27 21:52:40 -05:00
Victor Kyriazakos 5cc47c994b fix(cron): resolve api_server host for the manual-run forward (bind parity)
The forward dialed a hardcoded 127.0.0.1. The api_server adapter binds
extra.host -> API_SERVER_HOST -> 127.0.0.1, so mirror that chain when
dialing. Wildcard binds (0.0.0.0/::) listen on loopback, so keep dialing
loopback for those; bracket bare IPv6 literals.
2026-08-27 19:52:17 -07:00
Victor Kyriazakos ad0e522362 fix(cron): forward manual run to the gateway for relay-fronted delivery (NS-773)
A manual 'hermes cron run' on a relay-fronted target has no live relay adapter
and no standalone sender, so it now forwards to the running gateway's
POST /api/jobs/{id}/run (marks due for the gateway ticker, which delivers via
the live relay adapter). Gateway unreachable -> the accurate 'start the gateway
or use cron trigger' error. Native topologies are untouched.
2026-08-27 19:52:17 -07:00
Victor Kyriazakos 2d1d65de46 fix(cron): accurate error for relay-fronted delivery with no live gateway (NS-773)
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
2026-08-27 19:52:17 -07:00
Brooklyn Nicholson 0357982696 feat(desktop): make the idle tip rotation opt-in, ungate the tool
The two halves of tips were behind one switch, which meant the app
volunteering commentary at idle shipped on by default. Split them along
the line that matters: the rotation talks unprompted, so it now waits to
be asked for, while an agent tip stays ungated like the tour it mirrors
— Hermes raises one mid-conversation, in answer to something the user
said.

Drops the tool's config gate along with the config key it read. The
renderer mirrored that key with config.set, which has no branch for it
and answered "unknown config key" into a swallowed catch, so the opt-out
never reached the backend in the first place.
2026-08-27 21:50:18 -05:00
Brooklyn Nicholson 911c6c50d3 feat(tools): let Hermes point at one thing with the tip tool
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.

Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
2026-08-27 21:50:18 -05:00
686f6c61 e6b4f3750b fix(agent): count native Responses preflight against pruned wire
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.
2026-08-27 18:56:21 -07:00
Brooklyn Nicholson 01a3e9a44c fix(cli): give a profile created without a clone a usable model block
Creating a bot from the desktop dialog builds the profile tree but no
config.yaml, so the profile resolves no provider and its first turn dies
with "No LLM provider configured" — created, but unable to run. Every bot
made that way was dead on arrival.

Seed the active profile's model block at creation. It is a copy, not a
link: profiles stay independent islands and editing either afterwards never
touches the other. "Fresh" means fresh skills and SOUL, not unreachable.
2026-08-27 19:08:28 -05:00
Ben Barclay 6dcebea7fc Merge pull request #96341 from fangliquanflq/fix/computer-use-notarised-cua-paths
fix(computer-use): launch notarised CUA Driver from standard macOS installs
2026-08-28 09:04:46 +10:00
Gille 4956ff0cb9 fix(cli): keep journey labels readable 2026-08-27 17:45:43 -05:00
kshitijk4poor 80ab7d2b1c fix(compression): dedupe current-turn rows when rotation splits the session mid-turn
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.

Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
  row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
  scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback

Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.

Credit: @fedosis original PR #94996.
2026-08-28 02:47:34 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
fangliquanflq dd401e0f15 fix(bot-mode): keep delivery runner on host backend 2026-08-27 13:49:04 -07:00
Teknium 4882184e95 fix(state): guest durability barriers also apply configured database.synchronous
apply_durability_barriers() is the guest-connection entry point for
secondary state.db users (async delegation ledger) that must not run
journal-mode setup. database.synchronous (#90892) rides on the
journal-mode/pragma paths guests now skip, so apply it here directly —
otherwise guest connections silently run at the compile-time default.

Follow-up to the #93012 salvage.
2026-08-27 13:48:51 -07:00
fangliquanflq 7e6eda7bbb fix(delegation): expose state durability barriers 2026-08-27 13:48:51 -07:00
fangliquanflq 6548177eb1 fix(delegation): restore ledger durability barriers 2026-08-27 13:48:51 -07:00
fangliquanflq 4a8b4d43a3 fix(delegation): preserve state database journal mode 2026-08-27 13:48:51 -07:00
liuhao1024 e40f1be759 fix(state): route the post-repair journal-mode restore through the canonical path
Review feedback on #89681: the restore issued the switch pragma
directly via _set_journal_mode_no_wait, which bypassed the
vulnerable-SQLite WAL-reset gate — on the reporter's own runtime
(SQLite 3.50.4) it re-enabled WAL on the one file that had just
demonstrated it can corrupt, exactly what the gate exists to prevent
(a rebuilt file is a new database). It also skipped the WAL companions
(size limit, checkpoint barrier, synchronous=FULL) and used the
leave-WAL helper for an enter-WAL switch.

The restore now calls apply_wal_with_fallback, inheriting the gate,
the macOS-NFS silent-refusal handling and the companions from the
front door, and the pre-surgery mode is probed (best-effort; a
malformed file may refuse) and recorded in the report so the
before/after WARNING only fires when the comparison is honest.
2026-08-27 13:48:51 -07:00
liuhao1024 786e65bf4f fix(state): re-apply the configured journal mode after corruption repair
Corruption can drop the WAL bit from the database header, and every
repair strategy rebuilds or rewrites the file in place — so a repaired
store comes back in the default journal mode (delete). The WAL-reset
gate at open time never sees the flip because it happens inside the
repair path, not at open (the open-time flip #89393 warns about is a
different door), leaving the operator no signal that a WAL store moved
to DELETE (#89674).

After a successful repair, re-apply the canonical database.journal_mode
through _set_journal_mode_no_wait (concurrent openers abort the flip
instead of sneaking it between transactions) and log a WARNING naming
the old and restored modes. Best-effort: a refused restore is logged,
never raised — the repair itself already succeeded. Probing the damaged
file for a pre-repair mode is deliberately not trusted: the corruption
itself is what drops the WAL bit, so the configured setting is the only
reliable target.
2026-08-27 13:48:51 -07:00
pierrenode ba4c2d5253 fix(cron): make a recurring job stuck in state=error recoverable again
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:

* state=completed: a one-shot that genuinely has no more occurrences,
  ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
  fails to produce a next occurrence (e.g. the croniter package is
  missing at runtime). _mark_job_run_locked's own comment is explicit:
  "Recurring jobs must NEVER be silently disabled" (issue #16265) — the
  job is left enabled=True specifically so it keeps being a live,
  recoverable job once the underlying issue resolves.

Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:

* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
  its own is_terminal_job() check) never runs, because the check itself
  skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
  ... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
  itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
  and claim steps the scheduler's own dispatch loop calls immediately
  after get_due_jobs() for anything that DOES make it into the due list —
  both also refuse the job, so even a manually-recovered next_run_at
  would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
  broken recurring job through the normal path.

The only way out was deleting the job and recreating it.

Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.

Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.

New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.

Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
2026-08-28 02:18:25 +05:30
Teknium d91e4376f1 Merge pull request #65108 from NousResearch/hermes/hermes-793f4fd9
feat(skills): rewrite AgentMail optional skill CLI-first (salvages #60811)
2026-08-27 12:45:01 -07:00
hope b39d76d902 feat(tools): session-persistent kernels for execute_code (kernel_mode: session) (#94647)
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)

execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.

Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.

Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.

Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.

Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).

* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority

Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.

1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.

2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.

Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.

Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).

* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)

---------

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-08-27 12:26:31 -07:00
Gille 0dfba37b11 fix(dashboard): trust configured reverse proxies (#94126)
* fix(dashboard): trust configured reverse proxies

* fix(dashboard): trust IPv6 loopback proxies
2026-08-27 10:35:22 -07:00
Brooklyn Nicholson 39f1e1881a fix(tui-gateway): spare durable rows while a sibling backend holds them
Automatic Desktop ends (ws_orphan_reap, disconnect, idle, LRU, shutdown)
now drop the local runtime but keep the state.db row and durable-key
delegations when another live lease still owns the session.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
Brooklyn Nicholson 51e67babca fix(cli): keep Desktop liveness leases when the session cap is off
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
kshitijk4poor e941be7a81 test(gateway): adapt witness-composition harness to the SIGUSR1 in-place drain path
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
2026-08-27 22:06:17 +05:30
kshitijk4poor 5abe2e1880 docs+test(gateway): pin Windows witness-absent behavior and the probe-budget math
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
  there, so the witness is permanently absent, the payload records
  loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
  WEDGED — deliberate fail-safe (graceful drain remains the backstop).
  WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
  stays inside the documented probe budget so retuning can't silently
  blow past the 10s subprocess query tier.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor b48540701c refactor(gateway): simplify witness-probe ambiguity arms; sweep stale tick-socket nodes; clean test tempdirs
/simplify-code follow-ups on the 90502 salvage:

- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
  byte-identical — saw_node was effectively write-only. Collapsed to one
  arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
  from dead PIDs at arm time (POSIX-only liveness probe; Windows never
  creates AF_UNIX nodes) so state/ does not accumulate nodes across
  os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
  was DISPROVED for this call site — asyncio's create_unix_server
  os.remove()s an existing node before binding — but the contract is now
  pinned by test_producer_rebinds_over_stale_socket_node (a live
  producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
  longer leaks a directory per test run.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor db154edbaf test: short tmp_path for loop-tick witness sockets (macOS AF_UNIX limit)
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
2026-08-27 22:06:17 +05:30
rodrigo ca4a9ec686 fix(gateway): a single tick-socket miss must not authorize the wedge kill
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.

WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.

New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
  required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
  the loop is frozen for longer than one tick timeout but shorter than
  the wedge window -> probe is ALIVE and launchd_restart drains, never
  escalates; loop frozen for longer than the window -> WEDGED.

The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
2026-08-27 22:06:17 +05:30
rodrigo a1c83ef901 fix(gateway): interlock the stale-heartbeat wedge verdict with a loop-scheduling witness
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.

The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).

The classifier is now two-witness:
- socket answers            -> ALIVE (file age irrelevant: a stalled write
  or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
  longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
  agree the loop stopped scheduling)
- legacy payload (no flag)  -> unchanged single-witness contract: the
  legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity    -> UNKNOWN, never escalate

Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
2026-08-27 22:06:17 +05:30
rodrigo f39931afd9 fix(gateway): the loop watchdog's own heartbeat can freeze the loop it watches
`loop_heartbeat_forever` wrote the heartbeat inline on the gateway loop. That
write ends in `atomic_json_write` -> `os.fsync`, and on a stalling filesystem the
fsync blocks whichever thread runs it — which was the loop the liveness watchdog
exists to monitor. So the watchdog timed out its probe
(DEFAULT_LOOP_WATCHDOG_TIMEOUT_S = 10s, MAX_STRIKES = 3, a ~90-120s budget) and
took the hard exit, for a loop that was unresponsive because it was blocked
inside the watchdog's own liveness write.

Measured on the install that prompted this: a WSL2 VHDX under io pressure
(/proc/pressure/io full avg300 = 7.60) stalled a trivial stat-and-fsync probe at
p99 31s and max 112s — longer than the entire watchdog budget. Two of that day's
three gateway restarts carry byte-identical stack dumps parked in the heartbeat
writer.

The write now goes to a thread. Awaited, not fire-and-forget, and that distinction
is the whole design: the docstring promises that a frozen loop lets the file age,
because that staleness is how an external supervisor notices. The loop still
initiates and awaits the write, so a wedged loop still stops refreshing the file —
while a blocked fsync no longer stops the loop from answering the probe. One write
in flight at a time, so a 112s stall cannot queue a thread per interval behind it.

Scope: the timeout and strike defaults are untouched. Raising them would only
delay the same kill, and picking a budget above this box's p99 is a deployment
decision, not a fix.

2 tests. The behavioural one patches a slow write and asserts the loop still
completes ~15 ticks while it blocks; it is bounded by a fixed sleep rather than an
Event handshake so a regression fails on the tick count instead of hanging. The
second pins that the write stays awaited, since fire-and-forget would pass the
first test while destroying the staleness signal.

Verified by mutation: reverting only gateway/shutdown_watchdog.py fails both new
tests and leaves the 5 existing ones green. 15 passed across
test_loop_liveness_watchdog / test_shutdown_watchdog /
test_systemd_watchdog_lifecycle / test_watchdog_review_76354.
2026-08-27 22:06:17 +05:30
Finn763 cae58be1f5 fix(state.db): cross-backend heartbeat gates orphan sweep
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.

Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.

Credit: @Finn763
2026-08-27 11:31:59 -05:00
Brooklyn Nicholson 2119ed7b4a test(model): cover numeric YAML provider keys in picker and CRUD
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
kshitij 8966b0a700 review: tighten pre-call gate comment; drop redundant _ReadyAdapter test stub
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
2026-08-27 21:36:35 +05:30
kshitij 8aae2ea539 fix(mcp): signal reconnect from the mid-call fast-fail site too + regression tests
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.

Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.

Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
2026-08-27 21:36:35 +05:30
kshitijk4poor f3cbb262c1 fix(update): valid --ignored=matching mode; rename-only path split; shared preserve constant
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):

- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
  ignored mode'); with it, every ZIP update was refused as 'could not
  check the working tree'. The mocked tests could not see this — a new
  real-git test creates an actual repo + .gitignore and asserts the guard
  runs clean, blocks on an ignored user file, and exempts ignored
  preserved entries. --ignored=matching also reports an ignored dir as
  one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
  rename/copy status codes. Porcelain v1 does not quote plain filenames
  with spaces, so an ignored file literally named 'venv -> node_modules'
  parsed as two preserved tops and slipped past the guard into the
  destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
  instead of a comment-synced duplicate set (change-detector test added).
2026-08-27 21:07:34 +05:30
joaomarcos e64db76982 fix(update): gitignored user files also block the ZIP overlay
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.

Credit: @JoaoMarcos44, whose #87392 included this hardening.
2026-08-27 21:07:34 +05:30
hbentel 2673d5f5bd fix(gemini): embed images in Gemini 3.x functionResponse.parts for multimodal tool results
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.

Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.

Original PR #32352 by @hbentel, salvaged onto current main.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
2026-08-27 21:07:21 +05:30
fangliquanflq 164db901e8 test(computer-use): cover CUA signature security paths
Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
2026-08-27 23:26:40 +08:00
liuhao1024 23f597a8f5 fix(cron): verify a persisted final assistant message before booking complete
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).

Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.

Fixes #93820
2026-08-27 20:39:30 +05:30
liuhao1024 ad4e61f4c9 test(cron): stale hand-edit without manual marker still re-anchors
Carried from #94034 (closed as duplicate of #94033): explicit regression
that a jobs.json-edited stale next_run_at with NO manual_run_at marker
still re-anchors without firing — the direct #93049 protection case.
2026-08-27 20:39:20 +05:30
fangliquanflq 7e64a48303 fix(cron): preserve recurring manual run intent 2026-08-27 20:39:20 +05:30