A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
Two review nits on the clarify ambiguity fix:
- _approval_send_outcome swallowed the failure detail the old inline
callers logged (scheduling exception text / SendResult error). Both
the approval and clarify lanes now share one warning with the detail,
logged in the classifier itself.
- The clarify tests pinned the disposition helper but nothing proved
the ambiguous branch actually reaches the bounded wait. The
send-then-wait sequence is extracted to _clarify_send_then_wait (the
callback closure now just binds context onto it) and the suite gains
caller-path tests: ambiguous/sent -> wait_for_response with the
generated clarify_id and configured timeout; definitive failure ->
sentinel without waiting; no-response timeout sentinel preserved;
plus caplog assertions that failed sends log their detail.
Relay + gateway sweep 289/289.
The clarify caller treated a send-scheduling timeout as a definitive
failure: clear_session() + '[clarify prompt could not be delivered]'.
Same physics as the approval card fixed earlier in this PR — the card
may well have posted with a late connector ack — so the teardown ran
out from under a rendered clarify card and the user's answer resolved
nothing.
New _clarify_send_disposition() routes the outcome through
_approval_send_outcome: only a DEFINITIVE failure (error result,
non-timeout exception, no future) clears the registration and aborts;
ambiguous logs a warning and falls through to wait_for_response, whose
existing bounded wait already handles the truly-lost-card case. This
makes the boundary rule stated in the ambiguity test docstring hold
for the clarify lane, not just approvals.
Tests: 5 disposition tests mirroring the approval suite, including
clear_session-not-called on timeout. Mutation-verified: folding
ambiguous into the failed branch sends
test_timeout_keeps_registration_armed_and_proceeds_to_wait red.
Relay + gateway sweep 283/283.
test_expired_own_prompt_notifies_instead_of_unknown_command asserted the
expiry notice synchronously after _consume_prompt_response returned. The
notice now rides a background task (read-loop self-deadlock fix: awaiting
a send from the prompt_response handler blocks the very read loop that
resolves the send's result future), so the test yields one tick before
asserting egress. Behavior contract unchanged: exactly one notice, no
chat dispatch.
Round 2 of the approval-turn stuck-stream hunt. Round 1 (interim-marked
acks) fixed the draft-hijack-by-matching path — live logs confirm the
absorption fallback no longer fires — but the freeze persisted because
of a second, deeper defect on the same codepath:
_consume_prompt_response executes ON the transport read loop (inbound
frame -> _handle_frame -> _inbound handler). The handler awaited
self.send() for its '✅ Approved once' ack — but send() blocks on an
outbound_result future that ONLY the read loop can resolve, and the
read loop is blocked inside this very handler. Guaranteed self-deadlock
for the full outbound timeout (30s) on EVERY button tap. While wedged,
everything on the transport starved: draft appends (the frozen stream
right after approving), sibling approval-card sends (timed out into
'possibly-delivered' — the observed double-approval ambiguity), and the
turn's seal (timed out ambiguous -> plain-send fallback -> duplicate
final). Log signature was the tell: card-send timeout at tap time, no
absorption INFO, no seal-failed WARNING, no suppression line.
Fix: _send_lifecycle_ack() — acks ride a background task with strong
ref retention; the handler returns immediately and the read loop keeps
consuming, so the ack's own result frame resolves normally. Applied to
all six lifecycle sends (approval ack, slash-confirm ack + result text,
clarify acks, expiry notice). Acks are cosmetic by contract; failure
logs at debug and never breaks the reader.
Tests: new deadlock-shape test (gated transport send; handler must
return within 1s and the ack must still egress afterwards — RED on the
awaited version via TimeoutError at the exact deadlock), prior 3 tests
green with a yield for the background task. Targeted sweep 203/203.
Live finding (rc.4 staging, 100% reproducible on approval turns): after
resolving an exec-approval prompt_response, the adapter sends a short
ack ('Approved once'). _prompt_reply_metadata carried only placement
metadata (thread_id) — no per-turn identity, no interim marker — so
send()'s single-open-stream fallback (review B2) matched the approval
turn's OWN live draft and sealed it with the ack text. From there,
silently: every later append died on the post-seal tombstone (built for
millisecond stragglers, deliberately quiet), freezing the visible draft
mid-word; the turn-final found no open draft and fell through to a
plain send — the duplicate 'fallback' message. No suppression line, no
seal-failed warning: the log signature was pure absence.
Fix: _prompt_reply_metadata stamps _interim_send=True, which send()
already honors by bypassing draft matching. One source covers the whole
lifecycle class (approval ack, slash-confirm ack, prompt-expired
notice — all six call sites route through it).
Observability (the quiet parts, out loud):
- single-open-stream absorption now logs at INFO with the absorbed key;
- the FIRST post-seal tombstone swallow per draft key logs at WARNING
(bounded FIFO dedup) — one swallow is the normal straggler race, a
burst means a live stream was sealed mid-flight by someone else.
Tests (RED-first: both ack tests failed on the unfixed adapter at the
'draft still armed' assertion): approval ack leaves the open draft
armed and egresses as a plain send op; expiry notice same; regression
control pins the B2 contract — a real identity-less turn-final still
absorbs into its single open stream.
Three live findings from rc.4 staging, all on the relay-fronted Slack
path, all with the failure observed in live logs before the fix:
1. Approval-send timeout is AMBIGUOUS, not failed (no re-ask).
send_exec_approval through the connector can time out with the card
already rendered — the connector may ack after the deadline (slow
platform API call, transient backpressure, event-loop stall) — and
the timeout-as-failure path re-sent and produced duplicate cards.
The outcome is now tri-state: sent / failed / ambiguous. Ambiguous =
no re-send, no text fallback; the prompt registration stays armed so
a late tap still resolves. Only a definite send error falls back to
text.
2. pending_approval tool results forbid re-issuing the command.
With one card correctly armed, the agent could still mint a SECOND
card by re-running a rephrased variant of the gated command after
reading the pending_approval tool result (observed live: same
command re-issued in a different form, two cards). The tool message
now instructs: do not re-run/rephrase; wait or report pending.
Applied to both the terminal and execute_code arms.
3. Draft interim AND seal frames carry format_hints.
format_hints are stamped on send, edit, and send_for_platform, but
both draft-frame builders (send_draft interim + _seal_open_draft
seal) shipped bare metadata. A streamed final therefore arrived at
the connector hintless and sealed as a plain code block while
non-streamed sends rendered native markdown blocks (observed live:
language-tagged block on send/edit, downgrade on streamed seal).
Both sites now stamp _with_format_hints_for_chat
(destination-resolved, same pattern as the existing lanes).
Verified live after the fix against the platform's stored message
payload: rich_text_preformatted with language field on a streamed
seal.
Tests: tri-state outcome unit tests (5), draft/seal hint stamping + knobs-
off regression control (2, RED-first), existing format-hints suite intact
(14/14). Mutation-verified: reverting the adapter hunk sends
test_draft_interim_and_seal_frames_carry_hints red; restore -> green.
Boundary sweep (text egress lanes crossing the frame contract): send ✓
(pre-existing) edit ✓ (pre-existing) send_for_platform ✓ (pre-existing)
draft-interim ✓ (this PR) draft-seal ✓ (this PR); task_card lane carries
no text content — exempt.
sol-reviewer findings: multiplex secret-scope regression test (global
stamp must not leak into a scoped profile; scoped stamp must trigger
the sweep), WARNING-vs-INFO log-level assertions, and a check that
_enabled_explicit never survives config load.
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.
Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.
- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
Deployments that intentionally mix connector-fronted and direct ingress
can set GATEWAY_RELAY_ALLOW_DIRECT_PLATFORMS=true to keep directly-
connected messaging adapters enabled beside the relay. Unset, the
GATEWAY_RELAY_URL env stamp keeps its exclusive behavior. Like the
trigger, the opt-out is a deploy-stamp env var, not config.yaml.
A GATEWAY_RELAY_URL set in the process environment marks a
connector-fronted deployment where the connector owns every platform
connection. A directly-connected messaging adapter in the same process
is a second, unmanaged ingress path: it causes duplicate deliveries and
split sessions, and its live socket disarms scale-to-zero.
At the end of _apply_env_overrides, after all enablement passes, the
env stamp now disables every other enabled messaging platform:
- Explicitly-enabled platforms (config.yaml enabled: true) are disabled
with a WARNING that names the platform.
- Credential-auto-enabled platforms are disabled with an INFO line.
- Non-messaging surfaces (local, api_server, webhook) are untouched --
the same exclusion set as the scale-to-zero arm gate.
- gateway.relay_url in config.yaml alone (no env stamp) keeps the old
additive behavior: relay runs beside direct adapters.
With the gateway owning the suspend, the idle predicate covers every
work source (agent turns, cron jobs, API-server runs, background work,
fail-awake on unreadable sources) and the relay drains + flips before
the freeze, so a long timeout no longer buys safety — real work always
blocks the suspend and resume is sub-second. 5 idle minutes just bills
idle RAM. Per-instance override stays config.yaml
gateway.scale_to_zero.idle_timeout_minutes (D2).
New behavior-contract test: invalid config values degrade to the module
default (whatever it is), never zero/negative — asserts the RELATION,
not the literal, per the no-change-detector-tests rule.
Address sol-reviewer findings:
- The shared shutdown-drain counters swallow exceptions to 0 — fine for
a drain, unsafe for a suspend predicate (a transient read failure made
live work look idle, reopening the mid-job freeze). The suspend path
now reads both sources itself and treats an unreadable source as work
(sentinel 1, fail-awake) with a debug log. A MISSING api_server
adapter remains a normal not-work state.
- is_idle()'s parameter renamed running_agent_count -> active_work_count:
it receives the broad aggregate, and the old name invited future
callers to pass only agents again.
- New failure-path tests: unreadable cron source and unreadable API
source each hold the machine awake (both fail against the fail-open
shape); missing adapter stays idle-capable.
The config key is declared in DEFAULT_CONFIG and shown in the auxiliary
config UI, but the judge path never read it — a user raising the timeout
for a slow-but-healthy endpoint got the same 30s cap, and the loop
auto-paused on transport failures advising a provider/key check. Mirror
the _goal_judge_max_tokens reader and resolve the timeout at call time;
explicit timeout= arguments still win.
Fixes#91022
`uv pip install -e .` never audits an editable target. It reinstalls on every
invocation and rewrites the console-script shims each time, which is the only
reason `hermes update` has to quarantine the running `hermes.exe` on Windows —
and a quarantine that loses its race is the whole `os error 32` family.
Gate the reinstall on whether the pull actually touched a file that defines the
install. It's safe to skip because the editable finder is pinned to a static
module list (`py-modules` + `packages.find.include`), so the one source-only
change that could stale it — a new top-level module or package — cannot land
without a `pyproject.toml` diff. Dependencies and `[project.scripts]` live
there too, and new submodules inside an already-mapped package resolve through
the real directory.
The predicate fails closed: no pre-pull SHA, an unresolvable one, or a failed
`git diff` all reinstall as before. On the skip path the two verifiers that
normally run inside the install run directly, so a wrong skip self-heals into a
real install rather than leaving an unchecked venv.
This is the pattern the file already uses everywhere else — `_tui_need_npm_install`
diffs node_modules against package-lock.json, and the desktop build is gated on a
content hash so `hermes update` "will skip if nothing actually changed". The
Python editable install was the one path with no such gate.
* fix(update): bound the Windows update hand-off's step pipe drain
Invoke-HermesStep collected each step's output with ReadToEndAsync().Result.
That task does not complete when the step exits; it completes when the pipe
reaches EOF. On Windows the write end of a redirected pipe goes to the child as
an inheritable handle, so every descendant spawned without its own redirection
holds a duplicate and EOF waits for the last of them to close it. hermes update
deliberately runs its build steps with stdout inherited, so the tree under a
step is arbitrarily deep and not something this script can enumerate. When one
of those descendants is a resident gateway, the pipe stays open for the life of
the gateway and the hand-off blocks forever.
Everything the hand-off owes the Desktop is downstream of that call:
.hermes-update-result.json is never written, .hermes-update-in-progress is never
cleared, and the Desktop is never relaunched. The app sits on "Updating Hermes"
until the user kills the gateway by hand, and the stale marker then refuses the
next update too.
Read both pipes in chunks into a StringBuilder and bound the drain once the step
process itself has exited. The bound cannot truncate a slow step: the clock only
starts after the process is gone, at which point everything it wrote is already
in the pipe buffer waiting to be read, so the grace only has to cover the final
drain. Chunked reads are what make abandoning safe at all, since .Result cannot
hand back a partial read.
Also switch to the bounded WaitForExit overload. The argument-less one waits on
redirected streams as well, which is the same unbounded wait by another name.
An abandoned drain logs one line to logs/desktop-update-handoff.log naming the
cause, so a truncated step log is never mistaken for a step that printed
nothing.
Measured on Windows 11 / PowerShell 5.1 against a step whose grandchild
inherits its stdout and outlives it by 45s: 47.4s before, 4.3s after, with the
step's exit code and output preserved in both.
Fixes#90455
* test(update): prove the hand-off survives a step that leaks its pipe
Four source-level guards on Invoke-HermesStep, scoped to that function so the
legitimate WaitForExit and .Result uses elsewhere in the script cannot mask a
regression: no ReadToEndAsync, a drain bound keyed on the step having exited,
no argument-less WaitForExit, and a log line when a drain is abandoned. All
four fail against the previous drain. They are source-level for the same
reason the sibling python-handoff guard is: Linux CI cannot execute the
PowerShell hand-off.
Source-level is not enough for a deadlock, though, so the script also grows a
-SelfTestPipeDrain fixture alongside the existing -SelfTestUi one. It needs no
checkout, no install and no update: it starts a step that spawns a grandchild
with UseShellExecute = $false and no redirection, which is exactly the shape
that makes the grandchild inherit the step's stdout and stderr, then exits 7
while the grandchild sleeps on. The fixture asserts the grandchild was still
alive when Invoke-HermesStep returned, so a pass cannot be a timing
coincidence, and that the exit code and the step's output both survived the
abandonment. A windows_only test drives it, so the OS lane runs the real
drain rather than a text match.
Measured on Windows 11 / PowerShell 5.1: 4.3s with the fix, 47.4s (the
grandchild's full lifetime) with the previous drain restored.
The python-handoff guard now reads the script with its -SelfTest* blocks
removed. Those blocks exercise the machinery deliberately and exit before any
marker, venv or desktop work, so the "every step drives python.exe, never the
hermes.exe shim" rule does not apply to them. Scoping the source that way
rather than allow-listing a target keeps that rule absolute for every real
step.
Refs #90455
* fix(update): don't meter the step drain that #90455's bound introduced
Chunked reads make the bounded drain possible, but the loop idled 150ms
after every chunk it consumed, so a step's output moved at one 16 KiB
buffer per tick (~107 KB/s). The pipe then backs up, which is
backpressure on the *running* step rather than a slow read: a chatty
step blocks on write() waiting for the reader.
`hermes update` is exactly that shape -- the Electron/vite build alone
is megabytes -- so the layer that fixed "the hand-off waits forever"
would have shipped "the hand-off is slow" in its place.
Idle only when both pipes came up empty, and idle on the reads
themselves (WaitAny with the same 150ms cap) rather than on the clock:
a freshly issued ReadAsync is rarely complete by the very next pass, so
a bare `if (-not $moved)` still sleeps between chunks. WaitAny expires
on its own, so a silent step keeps the marquee animating and keeps the
abandon deadline advancing.
Measured against the drain as submitted, same harness, one variable:
4 MiB of step stdout 38.99s -> 0.07s
1 MiB stdout + 1 MiB err 18.22s -> 0.27s
leaked grandchild (20s) 3.24s -> 3.20s, exit code + output kept
quiet step, exits at 4s 4.29s -> 4.04s, 29 passes (not spinning)
* test(update): make the pipe-drain fixture cover metering, not just deadlock
The fixture proved the drain returns while a descendant holds the pipe.
It could not have caught the opposite failure -- a drain slow enough to
backpressure the step it is reading -- and that is the regression the
first version of this fix shipped.
Add a flood arm: a step that writes megabytes and holds nothing, with a
wall-clock budget far under what a sleep-per-chunk drain needs. The two
arms bracket the contract from both sides: bounded when a descendant
holds the pipe open, never slower than the step can write.
Few large lines rather than many small ones, deliberately --
Write-HandoffLog is one Add-Content per line and runs inside the
measured window, so line-heavy output would time the logger.
Also drops the four source-grep guards. Reading windows.ps1's text to
assert it contains `$abandonAt` tests the shape of the source, not its
behavior: it passes on a drain that is wired wrong but spelled right,
fails on a correct refactor, and blocks the extraction it should
survive. AGENTS.md bans the pattern outright, and all four pass on the
metered drain. The executable arms cover the same contract and actually
run the code -- the Windows lane is where this is verified either way.
---------
Co-authored-by: Jack Lau <72348727+jackulau@users.noreply.github.com>
test_seed_supervise_skeleton_creates_expected_layout has been failing on every
macOS checkout. The helper is correct — it chmods explicitly, so this isn't a
umask problem. BSD drops S_ISGID from a directory chmod unless the caller is
root or in the directory's group, so the same call that yields 03730 on Linux
yields 01730 on macOS.
s6 only ever runs on Linux, inside s6-overlay's stage2 as root with umask 0, so
Linux is the host whose answer matters. Split the mode assertion into its own
linux_only test rather than marking the whole case: the layout the test also
covers (dirs present, supervise/ 0755, control is a 0660 FIFO) is host-
independent and worth keeping on the machines developers actually run.
Adaptive Claude models think by default, so omitting the `thinking`
parameter left thinking ON for users who had turned it off. Send
`thinking: {"type": "disabled"}` instead, and keep the omission for
reasoning-mandatory families that answer a disable with HTTP 400.
The lane shipped dead. `classify_changes.py` emitted `rust`, the composite
action re-exported it, and ci.yaml's `rust-tests` job gated on
`needs.detect.outputs.rust` — but the `detect` job never declared that
output, so the expression was the empty string and the job reported
"skipping" on the very PR that added it. GitHub does not error on a
reference to an output a job never declared, so nothing went red.
Adds the missing line plus the invariant that catches the whole class:
every `needs.detect.outputs.X` referenced by a job's `if` must be
declared by `detect`. Verified it fails with the line removed.
The related check — every lane reaching the composite action — is
separate on purpose: nix.yml and docker.yml own their triggers and
re-export different subsets, so `docker` and `nix` are legitimately not
ci.yaml detect outputs.
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.
Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
`hermes update` runs in the pre-pull interpreter. The auto-restart phase
imports freshly-pulled gateway source, which resolves sibling imports
against the OLD sys.modules cache — so any update where an already-cached
module gained a new export ImportErrored the whole phase and left the
gateway serving pre-update code (2026-08-20 field failure: new gateway.py
needs cli_output.line_input, cached cli_output predates it).
Class fix replacing the per-symptom _UPDATE_RUNTIME_RELOAD_MODULES
approach: _purge_stale_hermes_modules() evicts every cached module under
the Hermes package prefixes (hermes_cli/gateway/tools/tui_gateway/agent)
right before the restart phase, so later lazy imports rebuild a
self-consistent module graph from the updated checkout. The updater's own
executing modules are exempt (purging them buys nothing; reload-in-place
is the unsafe op, and we never reload). Root-segment check spares
prefix-lookalike packages. Best-effort, never raises.
5 new tests incl. an end-to-end repro of the field failure shape
(stale module missing symbol -> ImportError -> purge -> import resolves).
tests/tools/test_website_policy.py and tests/cli/test_surrogate_sanitization.py
repeatedly failed under the parallel CI runner (process-teardown timeouts /
async-timeout flakes) across multiple unrelated PRs this cycle, while passing
locally. Removed per maintainer direction to stop the flake taxing every PR.
posix.sh now probes `update --help` before the real update call; the fake
counted the probe as call #1, shifting the exits.N mapping so the retry
gate never fired. Answer the probe out-of-band so counted calls remain
actual update attempts.
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').
New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.
Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
Addresses teknium1 sweeper review on PR #67696:
1. Gateway bridge: Skip str(None) bridging when YAML value is Python None
(from or bare ). Previously str(None) → None → unlimited
instead of default 90. Now clears stale env var so resolver applies default.
2. TUI: Route _cfg_max_turns through resolve_turn_limit instead of bare
int(). Old code crashed on none/unlimited and swallowed 0 via
. HERMES_TUI_MAX_TURNS env var also routed through
resolver.
3. Docs: Document unlimited spellings (none/unlimited/infinite/0/-1) in
configuration.md.
4. Tests: Add TestGatewayBridgeNullHandling (4 tests) and TestTUIResolver
(8 tests) covering null handling, string spellings, env var override,
and legacy root-level config.
All 50 tests pass.
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.
This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:
- int/float → int(raw) (floats truncated)
- numeric string ('120') → int(raw)
- 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
whitespace-tolerant) → sys.maxsize sentinel
- YAML None/null → default (90)
- bool/list/dict/garbage → default (with debug log)
All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.
The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.
Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:
- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
spawn-ledger.json self-registration keyed on (pid, create_time) — PID
reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
themselves at startup and attach to the job; Desktop legacy
HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
_ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
backends (purpose reapable + recorded spawner provably dead) in ANY update
context. Ledger-unknown holders fall through to the existing rungs.
22 new tests, sabotage-verified.
The one-shot keyless extract rescue (d1eefe6ac) treated ANY whole-batch
failure as a backend outage. A website-policy refusal also arrives as a
failed batch, so blocked URLs were routed through the free-tier ring:
in CI the ring's live fetch attempt returned a result for the wrong URL
or a bare None error, turning test_website_policy reds on main (slices
8/12 and 12/12) — and in production it would fetch content the user
explicitly blocked.
_rescue_extract now partitions policy blocks (blocked_by_policy flag or
policy error text) out of the rescue set: they are preserved verbatim,
only genuine failures ride the ring, and order/merge parity is kept.
Two sabotage-verified regression tests pin the class.
Field incident (2026-08-20): a Windows Desktop update hand-off
(update --yes --gateway --force) left a swarm of per-profile serve
backends (mr-tester, probe-inherit, turqoise, clippy, maroon, …) holding
cryptography/_rust.pyd. Some still had a live parent (the tearing-down
Electron process, or the venv launcher->worker two-hop chain mid-exit),
so the strict orphan-only reap (_orphaned_desktop_backend_pids, which
bails the instant ANY holder has a live parent) disqualified the whole
set and the venv-holder guard dead-ended. The user saw a ~12-minute hang,
force-closed, and the half-done state stranded bot sessions.
New rung: _handoff_reapable_backend_pids reaps surviving Hermes
serve/dashboard backends from this venv — live parent or not — but ONLY
in the hand-off context the caller gates on: args.gateway AND the
update-incomplete marker present AND no live hermes.exe shim. In that
window nothing legitimate supervises or respawns a serve backend (the
Desktop tree-kills its backends and parks any relaunch behind the marker,
#50238), so a surviving backend is a leak, not a race. A non-backend
holder (operator REPL, stray script) still disqualifies the whole set;
psutil-unavailable returns None (keep refusing). Wired as the final rung
before the existing dead-end, after the orphan-only reap.
One conflict, gateway/relay/adapter.py send_for_platform: main added the
turn-final draft-seal interception (_sfp_metadata with the _interim_send
marker stripped, seal-or-fall-through); this branch added format-hint
stamping on the same frame. COMPOSED: the plain-send frame now stamps
_with_format_hints_for_platform over _sfp_metadata (the stripped copy),
so both the seal fall-through contract and the cron-lane block hints
hold. Note: the seal frame itself (op:draft final) does not stamp hints
— cron sends are never open drafts, so the flagship path is unaffected;
noted as a connector-PR follow-up for streamed interactive finals.
_scale_to_zero_is_idle() consumed _running_agent_count(), but cron jobs
run through a standalone AIAgent on the scheduler's own thread pool and
API-server runs live on the adapter — both outside _running_agents (the
same blind spot the #60432 shutdown-drain fix addressed with
_active_work_count()). The idle predicate therefore read True DURING a
running cron job; a suspend at that moment freezes the job mid-flight.
Observed live on staging 2026-08-20: is_idle held True throughout the
10:45:04-22 cron run — only watcher-tick timing (next tick 9s after
completion) avoided a mid-job freeze.
Use _active_work_count() (agents + cron + API runs). New tests cover a
running cron job and an active API run each blocking idle, plus the
all-quiet True case; both blocking tests fail without the fix.
seed_mock.assert_not_called() alone could pass for the wrong reason — a
harness failure before delivery also leaves the seed uncalled. Assert
the real adapter recorded exactly one live send to the origin chat, so
the test pins the D6 thread-fallback decision, not an accidental
no-delivery.
Unspecced MagicMock/AsyncMock adapters fabricate
supports_inchannel_continuable_for_platform as a truthy callable, so the
scheduler's duck-typed D6 gate silently took the relay accessor branch in
every in-channel test — the native scalar fallback the fixtures describe
was never exercised, and setting supports_inchannel_continuable=False on
a mock could not force thread mode. Pin the accessor to None on both
mock fixtures (matching a real native adapter, which never defines the
method), and add a fallback-boundary test with a real minimal adapter
class: scalar False -> in_channel fails safe to thread, flat seed never
fires.
`hermes update` on Windows detached on every run, including the
`Already up to date!` no-op that never touches the venv. emozilla hit
the visible half: the shim exits, PowerShell takes the console back,
and a child prints the result under a fresh prompt — it reads as a
frozen update. The invisible half is worse: the hand-off sat ahead of
the fetch, so it also carried off the stash and branch-switch
questions, which #90205 then had to answer by closing stdin. Nobody
who mods Hermes got asked about their local changes again.
The shim lock is real and the child is still required — a launcher
holds venv\Scripts\hermes.exe open without FILE_SHARE_DELETE for the
whole command, so the quarantine rename is refused and uv fails with
os error 32. A parent that waits deadlocks against the handle it is
itself holding, and Windows has no exec to escape with.
But that lock only binds one step. Move the hand-off to the dependency
sync boundary, beside the native-module deferral that solves the same
"this process holds a file the sync must replace" problem — and for
the reason that placement already exists (#86735: a preflight ahead of
the fetch re-bricked the flow it was meant to protect). Everything
before the sync now runs foreground in the user's console: the
preflight, the stash question, the branch switch, git pull. An
up-to-date run never hands off at all.
Deferring to the next launch cannot substitute here the way it does
for a mapped .pyd: every future `hermes` launch is also the shim, so
the marker would defer forever. The child re-runs the update to keep
the node/web/lazy-refresh tail, and takes the sync it was spawned for
rather than the up-to-date early return.
The seed-key fix made the SESSION scoped, but the delivery leg still
dropped the scope: cron route_metadata carried only job_id (+thread),
DeliveryRouter stamps scope_id only for the configured HOME channel, and
the RelayAdapter's per-chat scope cache is cold after a gateway restart
(learned from inbound only). A scoped Slack origin that is not the home
chat therefore egressed with NO tenant discriminator, and the connector's
fail-closed guard could reject the brief before delivery — the
delivery-leg sibling of the seed-key scope gap.
Copy origin.scope_id into the live text and media routing metadata for
ORIGIN-MATCHING targets only (setdefault — never overrides router/home
stamping). Fan-out/broadcast targets are excluded by the origin gate: a
fan-out target's tenant is not the origin's, and a wrong scope is worse
than none.
Tests: restart-shaped positive (scoped non-home origin -> scope_id on
routed metadata, RED before this fix) and legacy negative (scope-less
origin stamps nothing).
Every drive_preview action answered with the entire inventory — around 120
elements of ref, role, label, and an up-to-eight-rung `:nth-child` selector
chain. On a real app shell that was ~24.5k characters, re-sent after every
click, so a ten-step task paid for ten copies of a page that had barely moved.
Handles are now durable and legible. An element is named after what it is and
what it says — `btn-sign-in`, `inp-email`, `srch-search-projects` — minted once
per page and never reused, with duplicates disambiguated as `btn-edit`,
`btn-edit-1`. Each one remembers a stable attribute, its role, its accessible
name, and the nearest landmark it sits in, so when a framework destroys the
node and builds a new one the handle moves across and the agent is told
`rebound` rather than being handed a removal it has to react to and an addition
it has to re-read. The re-bind ladder is anchortree's (Apache-2.0), minus its
geometry rung, which can never clear the threshold on its own.
Because the handles hold, the first look at a page returns the inventory and
every look after it returns only what moved. `changed` carries the ref and
whichever of label/value/disabled actually shifted — role and selector are
absent by construction, since a change in either would mean the re-bind ladder
was looking at a different element. A delta gives way to a full re-read when
half the page is new, where there is nothing left to reuse.
The selector column is gone with it. It was 74% of the inventory on an
85-element page, nothing downstream ever read it, and a positional chain is
wrong the moment a sibling appears. An `#id` or `[data-testid]` survives when
the page offers one; everything else is addressed by handle.
Legibility is what makes the delta work rather than a nicety. `+ btn-sign-in`
on turn nine reads on its own, where `+ @e42` sends the model back to an
inventory twenty thousand tokens ago.
Measured on an 85-element app shell: 18,693 -> 4,930 characters for a baseline,
and a steady turn that moved two things costs ~200.
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.