Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium 9522c4e8e2 chore: map contributor email for salvage of #95956 2026-08-27 13:48:45 -07:00
Thomas Bekkers dbca7a4f02 fix(hermes-bots): keep the Cronjobs tile registered while it holds focus in Bot Mode
Clicking the Cronjobs tile shifts focus onto the tile itself, momentarily
dropping bot-chat workspace ownership — syncRoutinesPane then unregistered
the pane out from under the user's own click, with no way back. Keep the
tile while Bot Mode is on screen and the tile is the focused surface;
leaving Bot Mode still unregisters as designed. Live-verified.
2026-08-27 13:48:45 -07:00
Thomas Bekkers 584f3a748b fix(desktop): keep a restore tab when a pane or strip collapses (#91223)
Hiding the Sessions/Bots strip, or tapping the header of a lone docked
tile (Cronjobs and any plugin pane beside the workspace), left no mouse
path back: the restore menu lived on chrome the gesture just unmounted,
and a row-collapsed rail could size to 0px.

Treat hide-only chrome as stranded so `never` cannot hide those chips.
Stop collapsing on header tap (chevron only). Size a minimized zone to
MINIMIZED_TRACK and keep the horizontal strip when two or more tabs
remain.
2026-08-27 13:48:45 -07:00
kshitijk4poor 939dec1348 fix: harden _is_recoverable_error_job against schedule=None
job.get("schedule", {}).get("kind") crashes with AttributeError when
schedule is present but explicitly None (disk corruption edge case).
Use (job.get("schedule") or {}).get("kind") instead, which safely
returns False for None. This pattern is already used at other sites
in the file (e.g. cron/scheduler.py line 170).
2026-08-28 02:18:25 +05:30
pierrenode ba4c2d5253 fix(cron): make a recurring job stuck in state=error recoverable again
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:

* state=completed: a one-shot that genuinely has no more occurrences,
  ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
  fails to produce a next occurrence (e.g. the croniter package is
  missing at runtime). _mark_job_run_locked's own comment is explicit:
  "Recurring jobs must NEVER be silently disabled" (issue #16265) — the
  job is left enabled=True specifically so it keeps being a live,
  recoverable job once the underlying issue resolves.

Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:

* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
  its own is_terminal_job() check) never runs, because the check itself
  skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
  ... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
  itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
  and claim steps the scheduler's own dispatch loop calls immediately
  after get_due_jobs() for anything that DOES make it into the due list —
  both also refuse the job, so even a manually-recovered next_run_at
  would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
  broken recurring job through the normal path.

The only way out was deleting the job and recreating it.

Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.

Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.

New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.

Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
2026-08-28 02:18:25 +05:30
kshitijk4poor 29d1c02699 chore: map contributor emails — shauneccles (#95433) + fedosis (#94996)
Adds contributor email→username mappings via the new
contributors/emails/ system (one file per email) for two
compression contributors whose salvage PRs need attribution CI.
2026-08-28 02:10:33 +05:30
Brooklyn Nicholson a24c12d14f fix(desktop): gate transcript budget cap so Show earlier works
The render-phase cap snapped a visible pane's Show-earlier growth back
on the next render, so the button did nothing. Clamp only hot-hidden
panes, and grow the DOM budget when expanding the store window too.

Supersedes #87686.

Co-authored-by: Kirk <317508070+chukirk-svg@users.noreply.github.com>
Co-authored-by: Per0 <175494353+Per0-1@users.noreply.github.com>
2026-08-27 14:51:13 -05:00
Teknium d91e4376f1 Merge pull request #65108 from NousResearch/hermes/hermes-793f4fd9
feat(skills): rewrite AgentMail optional skill CLI-first (salvages #60811)
2026-08-27 12:45:01 -07:00
hope b39d76d902 feat(tools): session-persistent kernels for execute_code (kernel_mode: session) (#94647)
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)

execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.

Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.

Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.

Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.

Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).

* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority

Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.

1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.

2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.

Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.

Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).

* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)

---------

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-08-27 12:26:31 -07:00
Gille 0dfba37b11 fix(dashboard): trust configured reverse proxies (#94126)
* fix(dashboard): trust configured reverse proxies

* fix(dashboard): trust IPv6 loopback proxies
2026-08-27 10:35:22 -07:00
Brooklyn Nicholson 39f1e1881a fix(tui-gateway): spare durable rows while a sibling backend holds them
Automatic Desktop ends (ws_orphan_reap, disconnect, idle, LRU, shutdown)
now drop the local runtime but keep the state.db row and durable-key
delegations when another live lease still owns the session.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
Brooklyn Nicholson 51e67babca fix(cli): keep Desktop liveness leases when the session cap is off
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
kshitijk4poor f6f707b783 chore: suppress posix-gated os.kill probe in footgun scan; map rodrigogs in AUTHOR_MAP
- The stale tick-socket sweep's os.kill(pid, 0) liveness probe sits inside
  an explicit os.name == 'posix' gate (AF_UNIX nodes never exist on
  Windows) — suppress with the standard inline marker.
- contributors/emails/: rodrigo.smscom@gmail.com -> rodrigogs (author of
  the salvaged #92315 commits), unblocking check-attribution.
2026-08-27 22:06:17 +05:30
kshitijk4poor e941be7a81 test(gateway): adapt witness-composition harness to the SIGUSR1 in-place drain path
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
2026-08-27 22:06:17 +05:30
kshitijk4poor 5abe2e1880 docs+test(gateway): pin Windows witness-absent behavior and the probe-budget math
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
  there, so the witness is permanently absent, the payload records
  loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
  WEDGED — deliberate fail-safe (graceful drain remains the backstop).
  WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
  stays inside the documented probe budget so retuning can't silently
  blow past the 10s subprocess query tier.
2026-08-27 22:06:17 +05:30
kshitijk4poor 8a32baafc5 fix(gateway): disarm the heartbeat writer in _stop_loop_liveness_guards
Ported from #95808 (@rtcopenclawgh, closed as duplicate of this PR):
once shutdown starts loading the loop, a heartbeat task that keeps
refreshing state/gateway.heartbeat can make a draining gateway look
healthy to external probes. Cancel it alongside the watchdog and floor
timer.

The other additive piece in #95808 — passing a confirm_s confirmation
window at the update_cmd.py probe call sites — is not needed here: this
PR's probe_gateway_loop_liveness already builds the sustained window in
(tick_strikes=3 consecutive socket misses), so every caller gets it
without a new parameter.

Co-authored-by: rtcopenclawgh <rafael@amora.com.br>
2026-08-27 22:06:17 +05:30
Kshitij Kapoor b48540701c refactor(gateway): simplify witness-probe ambiguity arms; sweep stale tick-socket nodes; clean test tempdirs
/simplify-code follow-ups on the 90502 salvage:

- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
  byte-identical — saw_node was effectively write-only. Collapsed to one
  arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
  from dead PIDs at arm time (POSIX-only liveness probe; Windows never
  creates AF_UNIX nodes) so state/ does not accumulate nodes across
  os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
  was DISPROVED for this call site — asyncio's create_unix_server
  os.remove()s an existing node before binding — but the contract is now
  pinned by test_producer_rebinds_over_stale_socket_node (a live
  producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
  longer leaks a directory per test run.
2026-08-27 22:06:17 +05:30
Kshitij Kapoor db154edbaf test: short tmp_path for loop-tick witness sockets (macOS AF_UNIX limit)
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
2026-08-27 22:06:17 +05:30
rodrigo ca4a9ec686 fix(gateway): a single tick-socket miss must not authorize the wedge kill
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.

WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.

New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
  required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
  the loop is frozen for longer than one tick timeout but shorter than
  the wedge window -> probe is ALIVE and launchd_restart drains, never
  escalates; loop frozen for longer than the window -> WEDGED.

The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
2026-08-27 22:06:17 +05:30
rodrigo a1c83ef901 fix(gateway): interlock the stale-heartbeat wedge verdict with a loop-scheduling witness
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.

The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).

The classifier is now two-witness:
- socket answers            -> ALIVE (file age irrelevant: a stalled write
  or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
  longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
  agree the loop stopped scheduling)
- legacy payload (no flag)  -> unchanged single-witness contract: the
  legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity    -> UNKNOWN, never escalate

Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
2026-08-27 22:06:17 +05:30
rodrigo f39931afd9 fix(gateway): the loop watchdog's own heartbeat can freeze the loop it watches
`loop_heartbeat_forever` wrote the heartbeat inline on the gateway loop. That
write ends in `atomic_json_write` -> `os.fsync`, and on a stalling filesystem the
fsync blocks whichever thread runs it — which was the loop the liveness watchdog
exists to monitor. So the watchdog timed out its probe
(DEFAULT_LOOP_WATCHDOG_TIMEOUT_S = 10s, MAX_STRIKES = 3, a ~90-120s budget) and
took the hard exit, for a loop that was unresponsive because it was blocked
inside the watchdog's own liveness write.

Measured on the install that prompted this: a WSL2 VHDX under io pressure
(/proc/pressure/io full avg300 = 7.60) stalled a trivial stat-and-fsync probe at
p99 31s and max 112s — longer than the entire watchdog budget. Two of that day's
three gateway restarts carry byte-identical stack dumps parked in the heartbeat
writer.

The write now goes to a thread. Awaited, not fire-and-forget, and that distinction
is the whole design: the docstring promises that a frozen loop lets the file age,
because that staleness is how an external supervisor notices. The loop still
initiates and awaits the write, so a wedged loop still stops refreshing the file —
while a blocked fsync no longer stops the loop from answering the probe. One write
in flight at a time, so a 112s stall cannot queue a thread per interval behind it.

Scope: the timeout and strike defaults are untouched. Raising them would only
delay the same kill, and picking a budget above this box's p99 is a deployment
decision, not a fix.

2 tests. The behavioural one patches a slow write and asserts the loop still
completes ~15 ticks while it blocks; it is bounded by a fixed sleep rather than an
Event handshake so a regression fails on the tick count instead of hanging. The
second pins that the write stays awaited, since fire-and-forget would pass the
first test while destroying the staleness signal.

Verified by mutation: reverting only gateway/shutdown_watchdog.py fails both new
tests and leaves the 5 existing ones green. 15 passed across
test_loop_liveness_watchdog / test_shutdown_watchdog /
test_systemd_watchdog_lifecycle / test_watchdog_review_76354.
2026-08-27 22:06:17 +05:30
Finn763 cae58be1f5 fix(state.db): cross-backend heartbeat gates orphan sweep
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.

Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.

Credit: @Finn763
2026-08-27 11:31:59 -05:00
Brooklyn Nicholson 2119ed7b4a test(model): cover numeric YAML provider keys in picker and CRUD
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
Brooklyn Nicholson d83dcb4c3a fix(model): coerce YAML integer provider names before picker/CRUD
PyYAML loads unquoted names like provider: 2070 as int. GET /api/model/options
then called .lower() on the providers dict key and 500ed, and activate/delete
looked up "2070" and missed. Stringify identity fields so Desktop can assign
and remove those endpoints.

Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
2026-08-27 11:31:21 -05:00
hermes-seaeye[bot] ca2a0d4d6f fmt(js): npm run fix on merge (#96506)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-27 16:29:27 +00:00
HexLab98 c086bb6f71 test(desktop): cover Voxtral JSON unwrap on client-direct STT
Pin the Groq plain-text path and the Mistral envelope so dictation
keeps spoken words, not the raw transcription object, in the composer.
2026-08-27 11:23:40 -05:00
HexLab98 bc737576ba fix(desktop): unwrap Mistral Voxtral JSON in client-direct STT
Mistral ignores response_format=text and returns the full transcription
object. Desktop was dumping that JSON into the composer; pull .text out
so dictation shows the spoken words instead.
2026-08-27 11:23:40 -05:00
fangliquanflq f54d015470 fix(desktop): retain remote owner after session resume 2026-08-27 11:22:35 -05:00
Gille 46f091b93e fix(desktop): recover cloud auth through portal (#96170) 2026-08-27 11:22:09 -05:00
kshitij 8966b0a700 review: tighten pre-call gate comment; drop redundant _ReadyAdapter test stub
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
2026-08-27 21:36:35 +05:30
kshitij 8aae2ea539 fix(mcp): signal reconnect from the mid-call fast-fail site too + regression tests
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.

Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.

Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
2026-08-27 21:36:35 +05:30
deadczarvc 2663117f72 fix(mcp): correct inverted liveness check in _stdio_children_dead
_stdio_children_dead() returned True ('all children dead') when a child
process was ALIVE — the liveness predicate was inverted:

    for pid in pids:
        if not psutil.pid_exists(pid):
            continue   # dead — skip
        return True     # BUG: an ALIVE child reported as 'all dead'

Consequence: every tools/call on a stdio MCP server with a healthy
subprocess failed instantly (~0.01-0.3s) with
'MCP stdio subprocess has exited; failing the call fast' (#81995
fast-fail path), while hermes mcp test kept passing (it never reaches
tools/call). Servers appeared dead regardless of restarts.

Fix: return False as soon as one tracked child is alive; True only when
every child has exited:

    for pid in pids:
        if psutil.pid_exists(pid):
            return False  # at least one child alive
    return True

Also: on a genuinely-dead stdio (session object still present), signal
reconnect instead of a bare fast-fail TimeoutError, so the manager
respawns the subprocess instead of stranding the call slot.

Verified: alive child -> False, all exited -> True, no tracked pids ->
False (unknown, don't fail fast).
2026-08-27 21:36:35 +05:30
kshitij 59805b1dd2 chore: map contributor email for deadczarvc 2026-08-27 21:36:35 +05:30
kshitijk4poor f3cbb262c1 fix(update): valid --ignored=matching mode; rename-only path split; shared preserve constant
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):

- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
  ignored mode'); with it, every ZIP update was refused as 'could not
  check the working tree'. The mocked tests could not see this — a new
  real-git test creates an actual repo + .gitignore and asserts the guard
  runs clean, blocks on an ignored user file, and exempts ignored
  preserved entries. --ignored=matching also reports an ignored dir as
  one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
  rename/copy status codes. Porcelain v1 does not quote plain filenames
  with spaces, so an ignored file literally named 'venv -> node_modules'
  parsed as two preserved tops and slipped past the guard into the
  destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
  instead of a comment-synced duplicate set (change-detector test added).
2026-08-27 21:07:34 +05:30
joaomarcos e64db76982 fix(update): gitignored user files also block the ZIP overlay
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.

Credit: @JoaoMarcos44, whose #87392 included this hardening.
2026-08-27 21:07:34 +05:30
hbentel 2673d5f5bd fix(gemini): embed images in Gemini 3.x functionResponse.parts for multimodal tool results
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.

Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.

Original PR #32352 by @hbentel, salvaged onto current main.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
2026-08-27 21:07:21 +05:30
fangliquanflq 164db901e8 test(computer-use): cover CUA signature security paths
Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
2026-08-27 23:26:40 +08:00
kshitijk4poor 42ac29eacc docs(cron): comment accuracy — booking is fail-open on probe errors; cross-ref status vocabulary
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
2026-08-27 20:39:30 +05:30
liuhao1024 23f597a8f5 fix(cron): verify a persisted final assistant message before booking complete
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).

Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.

Fixes #93820
2026-08-27 20:39:30 +05:30
kshitijk4poor 046a47a301 docs(cron): pin the manual_run_at comparison as intentionally string-exact
Review follow-up on the #94033 salvage: guard the equality gate against a
future 'helpful' datetime normalization — any rewrite of next_run_at must
invalidate the marker, and normalizing would weaken that.
2026-08-27 20:39:20 +05:30
liuhao1024 ad4e61f4c9 test(cron): stale hand-edit without manual marker still re-anchors
Carried from #94034 (closed as duplicate of #94033): explicit regression
that a jobs.json-edited stale next_run_at with NO manual_run_at marker
still re-anchors without firing — the direct #93049 protection case.
2026-08-27 20:39:20 +05:30
fangliquanflq 7e64a48303 fix(cron): preserve recurring manual run intent 2026-08-27 20:39:20 +05:30
liuhao1024 b0a8d16c60 test(gateway): end-to-end rendered-unit coverage for the cron drain floor
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
2026-08-27 20:39:11 +05:30
HexLab98 1282362803 test(gateway): cover TimeoutStopSec including the cron drain floor 2026-08-27 20:39:11 +05:30
HexLab98 a3f1cc002d fix(gateway): size systemd TimeoutStopSec from the full stop budget
The generated unit only counted restart_drain_timeout, so a default
cron drain (30s + 10s cleanup) could still be inside budget when
systemd SIGKILLed the cgroup. Size TimeoutStopSec from max(drain,
cron floor + reserve) plus headroom so an in-budget stop is not killed.
2026-08-27 20:39:11 +05:30
Teknium 82e1856720 docs: document the database section (journal_mode, synchronous, WAL sizing)
Covers the new operator signals from #85609/#89393 and the
database.synchronous key from #90892.
2026-08-27 07:52:26 -07:00
Jack Lau 1944570708 docs(state): call the macOS synchronous rule a floor, not a pin
cli-config.yaml.example said macOS is "always held at FULL regardless",
which reads as "your setting is ignored on this platform" and would talk
an operator out of choosing EXTRA. _apply_synchronous_pragma only refuses
values BELOW FULL on Darwin; EXTRA is applied normally.

The existing doc guard only asserted the key's presence, so it could not
have caught this. Pin the distinction instead.

Reported by @Enough1122 in review.
2026-08-27 07:52:26 -07:00
Jack Lau ef29fc63d7 fix(state): make state.db synchronous configurable on every platform
`apply_database_pragmas()` reads five sizing pragmas from `database:` and no
durability one, and `_enforce_macos_synchronous_full()` returns early when
`sys.platform != "darwin"`. Between them, nothing in the process ever executes
`PRAGMA synchronous` against state.db on Linux or Windows.

The effective level there is therefore `SQLITE_DEFAULT_WAL_SYNCHRONOUS`, a
compile-time constant of whichever SQLite the interpreter links. Debian and
Ubuntu builds commonly ship it as NORMAL; the bundled build and a plain
source build use FULL. So the durability of state.db is decided by which
python3 the installer found, is invisible from config, and cannot be pinned.

#90837 is three weeks of corruption forensics conducted on Ubuntu under the
stated premise `synchronous=FULL`, with every other cause eliminated live.
That premise is not something the reporter could have verified from config,
because there was no config key to set and no log line to read back.

- `resolve_synchronous_level()` maps the spellings operators actually write
  (OFF/NORMAL/FULL/EXTRA, any case, or 0-3) to the PRAGMA integer, and
  returns None for anything else. Kept out of the sizing loop on purpose: an
  unrecognised `cache_size` harmlessly falls back to a default, an
  unrecognised durability level must not.
- `_apply_synchronous_pragma()` applies it, and on Darwin refuses to lower
  below FULL. `_enforce_macos_synchronous_full()` runs during
  `apply_wal_with_fallback()`, which is earlier than `apply_database_pragmas()`,
  so without an explicit floor a configured NORMAL would silently undo #64355
  by the accident of running last. Raising to EXTRA on macOS is allowed.
- Unset changes nothing, so no existing install moves.

Tests: 34, covering the parser, application, the unset path, the typo path,
a guardrail that #77630's five keys still apply, and the Darwin floor in both
directions. Removing the wiring fails 6; removing the floor alone fails 1.

Related to #90837
2026-08-27 07:52:26 -07:00
Jack Lau d5d42b96ef fix(state): warn when an existing database's journal_mode is flipped to WAL
apply_wal_with_fallback treats an on-disk WAL database as authoritative and
says so twice: it never live-downgrades one. The mirror case had no
protection at all. When the on-disk mode is DELETE and the configured mode
is wal, the function flips the database and logs nothing.

journal_mode is a property of the FILE, so that rewrites the header and
persists after the process exits. Setting the mode directly on the file is
something operators do; it was the documented mitigation for the SQLite
3.50.4 WAL-reset bug. A config key that makes the choice durable already
exists (database.journal_mode, #68545), but nothing named it at the moment
the PRAGMA was being undone.

#89293 reports the cost: after upgrading past the vulnerable SQLite,
is_sqlite_wal_reset_vulnerable() stopped short-circuiting into
_apply_delete_for_wal_reset_bug, the flip path went live, and 4 of 5
databases silently returned to WAL with no log line anywhere.

Add a deduped WARNING at both points where the switch succeeds, decided
before the pragma runs since both inputs are only readable while the file is
still in its original state. Log-only: the flip still happens, the return
value is unchanged, and the never-live-downgrade rule is untouched.

WARNING rather than ERROR is deliberate. The reverse direction is ERROR
because dropping to DELETE costs concurrency; this direction is normally the
desirable one (managed_uv treats a database stuck on DELETE as a bug worth
repairing on update). The problem was never the change, it was that the
change was invisible.

The page_count guard is the load-bearing half. A brand-new database also
reports journal_mode=delete and is also about to be switched to WAL, and
every opener applies WAL before creating schema, so without it the warning
would fire on the first run of every install.

Refs #89293
2026-08-27 07:52:26 -07:00
Paolo Antinori 9aeb582e38 fix(state): warn when configured journal_mode=delete is overridden by on-disk WAL
When database.journal_mode=delete is configured but the on-disk DB is already
WAL, apply_wal_with_fallback honors the never-live-downgrade rule and keeps WAL.
That is correct (a live downgrade under open connections causes mixed-mode
corruption), but the operator's configured mode silently has no effect, and on
a WAL-incompatible filesystem (virtiofs/NFS/SMB) the DB then corrupts on the
next crash/sleep exactly what they configured to prevent.

Two code paths return WAL in this situation; both now emit a once-per-process-
per-db_label ERROR telling the operator the config did not apply and they must
convert the DB header offline (stop connections, PRAGMA journal_mode=DELETE):

1. The WAL-reset-vulnerable path (_apply_delete_for_wal_reset_bug): previously
   warned only about the vulnerability with an "upgrade SQLite" remedy, which
   does not help when the real cause is the filesystem. Emitted after that
   warning so the actionable message is last.
2. The read-only probe path (non-vulnerable runtime): previously returned WAL
   with no signal at all.

The never-live-downgrade behavior is unchanged (existing test now also asserts
the warning). New tests cover both paths, the per-db_label dedup, and the
require_wal=True edge case.

Real-world impact: a Hermes deployment with state.db on a Podman virtiofs
bind-mount (or any NFS/SMB home) that upgrades across a version where WAL was
the default, then sets journal_mode=delete, sees no corruption protection
until the DB header is converted. This makes the gap visible. See #68545.
2026-08-27 07:52:26 -07:00