OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.
- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
The config key is declared in DEFAULT_CONFIG and shown in the auxiliary
config UI, but the judge path never read it — a user raising the timeout
for a slow-but-healthy endpoint got the same 30s cap, and the loop
auto-paused on transport failures advising a provider/key check. Mirror
the _goal_judge_max_tokens reader and resolve the timeout at call time;
explicit timeout= arguments still win.
Fixes#91022
`uv pip install -e .` never audits an editable target. It reinstalls on every
invocation and rewrites the console-script shims each time, which is the only
reason `hermes update` has to quarantine the running `hermes.exe` on Windows —
and a quarantine that loses its race is the whole `os error 32` family.
Gate the reinstall on whether the pull actually touched a file that defines the
install. It's safe to skip because the editable finder is pinned to a static
module list (`py-modules` + `packages.find.include`), so the one source-only
change that could stale it — a new top-level module or package — cannot land
without a `pyproject.toml` diff. Dependencies and `[project.scripts]` live
there too, and new submodules inside an already-mapped package resolve through
the real directory.
The predicate fails closed: no pre-pull SHA, an unresolvable one, or a failed
`git diff` all reinstall as before. On the skip path the two verifiers that
normally run inside the install run directly, so a wrong skip self-heals into a
real install rather than leaving an unchecked venv.
This is the pattern the file already uses everywhere else — `_tui_need_npm_install`
diffs node_modules against package-lock.json, and the desktop build is gated on a
content hash so `hermes update` "will skip if nothing actually changed". The
Python editable install was the one path with no such gate.
test_seed_supervise_skeleton_creates_expected_layout has been failing on every
macOS checkout. The helper is correct — it chmods explicitly, so this isn't a
umask problem. BSD drops S_ISGID from a directory chmod unless the caller is
root or in the directory's group, so the same call that yields 03730 on Linux
yields 01730 on macOS.
s6 only ever runs on Linux, inside s6-overlay's stage2 as root with umask 0, so
Linux is the host whose answer matters. Split the mode assertion into its own
linux_only test rather than marking the whole case: the layout the test also
covers (dirs present, supervise/ 0755, control is a 0660 FIFO) is host-
independent and worth keeping on the machines developers actually run.
`hermes update` runs in the pre-pull interpreter. The auto-restart phase
imports freshly-pulled gateway source, which resolves sibling imports
against the OLD sys.modules cache — so any update where an already-cached
module gained a new export ImportErrored the whole phase and left the
gateway serving pre-update code (2026-08-20 field failure: new gateway.py
needs cli_output.line_input, cached cli_output predates it).
Class fix replacing the per-symptom _UPDATE_RUNTIME_RELOAD_MODULES
approach: _purge_stale_hermes_modules() evicts every cached module under
the Hermes package prefixes (hermes_cli/gateway/tools/tui_gateway/agent)
right before the restart phase, so later lazy imports rebuild a
self-consistent module graph from the updated checkout. The updater's own
executing modules are exempt (purging them buys nothing; reload-in-place
is the unsafe op, and we never reload). Root-segment check spares
prefix-lookalike packages. Best-effort, never raises.
5 new tests incl. an end-to-end repro of the field failure shape
(stale module missing symbol -> ImportError -> purge -> import resolves).
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').
New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.
Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
Addresses teknium1 sweeper review on PR #67696:
1. Gateway bridge: Skip str(None) bridging when YAML value is Python None
(from or bare ). Previously str(None) → None → unlimited
instead of default 90. Now clears stale env var so resolver applies default.
2. TUI: Route _cfg_max_turns through resolve_turn_limit instead of bare
int(). Old code crashed on none/unlimited and swallowed 0 via
. HERMES_TUI_MAX_TURNS env var also routed through
resolver.
3. Docs: Document unlimited spellings (none/unlimited/infinite/0/-1) in
configuration.md.
4. Tests: Add TestGatewayBridgeNullHandling (4 tests) and TestTUIResolver
(8 tests) covering null handling, string spellings, env var override,
and legacy root-level config.
All 50 tests pass.
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.
This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:
- int/float → int(raw) (floats truncated)
- numeric string ('120') → int(raw)
- 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
whitespace-tolerant) → sys.maxsize sentinel
- YAML None/null → default (90)
- bool/list/dict/garbage → default (with debug log)
All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.
The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.
Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:
- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
spawn-ledger.json self-registration keyed on (pid, create_time) — PID
reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
themselves at startup and attach to the job; Desktop legacy
HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
_ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
backends (purpose reapable + recorded spawner provably dead) in ANY update
context. Ledger-unknown holders fall through to the existing rungs.
22 new tests, sabotage-verified.
Field incident (2026-08-20): a Windows Desktop update hand-off
(update --yes --gateway --force) left a swarm of per-profile serve
backends (mr-tester, probe-inherit, turqoise, clippy, maroon, …) holding
cryptography/_rust.pyd. Some still had a live parent (the tearing-down
Electron process, or the venv launcher->worker two-hop chain mid-exit),
so the strict orphan-only reap (_orphaned_desktop_backend_pids, which
bails the instant ANY holder has a live parent) disqualified the whole
set and the venv-holder guard dead-ended. The user saw a ~12-minute hang,
force-closed, and the half-done state stranded bot sessions.
New rung: _handoff_reapable_backend_pids reaps surviving Hermes
serve/dashboard backends from this venv — live parent or not — but ONLY
in the hand-off context the caller gates on: args.gateway AND the
update-incomplete marker present AND no live hermes.exe shim. In that
window nothing legitimate supervises or respawns a serve backend (the
Desktop tree-kills its backends and parks any relaunch behind the marker,
#50238), so a surviving backend is a leak, not a race. A non-backend
holder (operator REPL, stray script) still disqualifies the whole set;
psutil-unavailable returns None (keep refusing). Wired as the final rung
before the existing dead-end, after the orphan-only reap.
`hermes update` on Windows detached on every run, including the
`Already up to date!` no-op that never touches the venv. emozilla hit
the visible half: the shim exits, PowerShell takes the console back,
and a child prints the result under a fresh prompt — it reads as a
frozen update. The invisible half is worse: the hand-off sat ahead of
the fetch, so it also carried off the stash and branch-switch
questions, which #90205 then had to answer by closing stdin. Nobody
who mods Hermes got asked about their local changes again.
The shim lock is real and the child is still required — a launcher
holds venv\Scripts\hermes.exe open without FILE_SHARE_DELETE for the
whole command, so the quarantine rename is refused and uv fails with
os error 32. A parent that waits deadlocks against the handle it is
itself holding, and Windows has no exec to escape with.
But that lock only binds one step. Move the hand-off to the dependency
sync boundary, beside the native-module deferral that solves the same
"this process holds a file the sync must replace" problem — and for
the reason that placement already exists (#86735: a preflight ahead of
the fetch re-bricked the flow it was meant to protect). Everything
before the sync now runs foreground in the user's console: the
preflight, the stash question, the branch switch, git pull. An
up-to-date run never hands off at all.
Deferring to the next launch cannot substitute here the way it does
for a mapped .pyd: every future `hermes` launch is also the shim, so
the marker would defer forever. The child re-runs the update to keep
the node/web/lazy-refresh tail, and takes the sync it was spawned for
rather than the up-to-date early return.
'hermes --version' (and -V) now prints the full version report — banner
version line with upstream SHA, install directory, authoritative install
method, Python and OpenAI SDK versions, and update status — making the
separate 'hermes version' subcommand redundant. The subcommand is removed.
- _startup_fast.print_fast_version_info() is now THE canonical version
printer: static lines print instantly from stdlib probes, then the
banner label, install-method resolver, and update check lazy-import
after the first line is on screen (each degrades gracefully).
- main.py _print_version_info() delegates to it (used by /version in the
CLI chat surface and the --version flag path); the old duplicate
implementation is deleted.
- hermes_cli/subcommands/version.py removed; parser wiring, subcommand
sets, console-engine extraction entry, and tests updated. Hermes
Console keeps a 'version' command wired to the shared printer.
- Termux fast paths now include update status too (previously
check_updates=False).
- Docs/i18n, CONTRIBUTING, SECURITY, and nix checks updated to
'hermes --version'.
A GitHub-side HTTP 429 during 'hermes update' printed only
'Failed to fetch updates from origin.' — and the curl
'unable to access ... returned error: 429' shape even matched the
network-error branch, blaming the user's connection for a GitHub
outage.
- new _classify_fetch_failure(): 429/rate-limit -> 'GitHub is rate
limiting requests or having an outage — try again in 5 minutes';
5xx -> outage message with githubstatus.com; ordered BEFORE the
generic 'unable to access' network check
- both fetch-failure sites (update apply + --check) now share the
classifier via _print_fetch_failure(), and both always print the
first raw stderr line so the wire error stays diagnosable
- tests: classifier matrix + E2E against a live local HTTP server
returning 429 through real git
Fixes#89287
The startup pruner is deliberately conservative (unattended, pre-banner),
so real installs accumulate what it can never touch: trees preserved for
untracked-only scratch, and orphaned local branches beyond the two
auto-generated prefixes it deletes. A measured multi-agent box: 35 trees /
15GB / 244 local branches, 120 of them fully merged.
New attended surface (hermes_cli/worktree_gc.py + worktree_cmd.py):
- hermes worktree list — audit every tree: age, size, verdict, reason,
plus deletable-branch count
- hermes worktree prune [--dry-run|--trees-only|--branches-only]
- /worktree prune [--dry-run] — same engine in-session; never touches the
session's own active tree
- startup escalation: one WARNING when .worktrees/ exceeds 10 trees or
5GB, naming the reclaim commands (silence is how boxes hit 15GB)
Safety invariants (shared with the startup pruner via cli.py primitives):
tracked modifications and unique unpushed commits never deleted at any
age; live-locked trees untouched; branch deletion gated on worktree
removal success; untracked-only scratch ARCHIVED to
~/.hermes/archive/worktree-prune/ before its tree is reaped.
Branch GC is content-gated, not name-gated: any local branch fully merged
or git-cherry patch-equivalent upstream is safe to delete (rebase merges
rewrite SHAs, so --merged alone misses the dominant leak); unique-commit,
checked-out, protected, and stale-base (>50 ahead) branches are kept.
Classification is parallel (8 workers) — 244 branches audit in ~64s live.
git timeouts degrade to keep (returncode 124) instead of crashing the
audit — live-verified failure on a 746MB .git repo.
16 behavior-contract tests against real git fixtures; live dry-run on the
production repo: 12 trees reclaimable, 120 branches deletable, 0 false
positives among kept trees.
Salvaged from #78434 by @Slobaka (also the issue reporter; earlier than
the competing #78436). hermes doctor no longer paints a green web check
when the explicitly selected provider cannot initialize — web splits
into per-capability rows (web search / web extract) resolved through
the same registry resolvers the dispatchers use, with readiness from a
true availability probe (_provider_is_ready).
Keyless-tier integration on top of the salvage:
- _provider_is_ready counts is_keyless_available() as ready — keyless
mode is a working state, not a misconfiguration (zero-config installs
and selected-keyless Tavily/Firecrawl show ok, not warn)
- Tavily/Firecrawl gain is_keyless_available() (True only when
explicitly selected — they stay out of the zero-config fallback)
- doctor triggers plugin discovery before reading the registry (fresh
doctor processes saw an empty registry and warned on everything)
E2E: searxng-selected-without-URL warns (the #78412 repro);
zero-config, tavily-keyless, firecrawl-keyless all read ok;
parallel pinned paid without a key warns.
Both the wire path and the picker only consulted the catalog's
`mandatory` flag, so a route the Portal lists as accepting no reasoning
parameter at all still got sent a disable, and still offered a Thinking
toggle in the model picker.
For a route it serves, the aggregator's own catalog outranks the
models.dev inference: `supports_reasoning: false` now suppresses the
disable on the wire and drops reasoning controls from the picker
entirely, so there is no disable left to describe.
Portal reasoning capabilities were held only in memory, so a process that
had not yet fetched them answered "unknown" — and on that answer the Nous
profile drops the disable rather than risk a 400. A short-lived process
(`hermes -p`, a cron job, a freshly booted gateway) is always in that
state, so every one of those runs silently ignored "thinking off" and
billed the user for reasoning they had turned off.
The parsed catalog is now mirrored to `cache/reasoning_caps.json`, keyed
by the URL it came from, and hydrated on a cold lookup without touching
the network. Every picker and pricing fetch already pulls that same
document, so they seed the mirror for free.
The catalog URL itself now resolves through the same ladder as the rest
of the Nous catalog reads (`NOUS_INFERENCE_BASE_URL` → credential base →
production) instead of being pinned to production, which had a staging
profile deciding the reasoning-mandatory question from prod's answers.
Keying the mirror by URL keeps those deployments apart.
The picker offered an off switch for every reasoning model, including routes
whose upstream answers a disable with HTTP 400 — so "thinking off" was a
control that could not work. Carry the catalog's mandatory verdict through
model.options as can_disable_reasoning and hide the toggle when it is false.
Effort levels are left alone. The catalog's supported_efforts under-reports
what the Portal serves (z-ai/glm-5.3 publishes max, high, low yet honors
minimal at its lowest thinking), so filtering the scale by it would hide
levels that work.
The Portal serves OpenRouter's catalog schema, so the existing parser and
cache-only tri-state contract carry over unchanged. Only the HTTP fetch is
generalized across the two catalogs; each keeps its own cache because they
list different models.
The Portal 403s a catalog read with no User-Agent, so the shared fetch now
sends one.
/api/model/options probed Copilot auth via `gh auth token` four separate
times per payload build. When gh has no credential store for the backend's
HOME (fresh profile, desktop-spawned backend, CI), each probe blocks its
full 5s subprocess timeout on keyring/D-Bus, so every open of the Desktop
Models or Providers settings page took 20s — past the renderer's 15s IPC
budget, painting 'Error invoking remote method hermes:api: Timed out'.
Fix: cache the gh-CLI probe result (hit or miss) for 5 minutes with an
invalidation hook, feed gh stdin=DEVNULL, and disable gh interactive
prompts/update notifier in the probe env.
Measured on the failing profile: 20.5s -> 5.3s cold (one bounded probe),
0.03s warm.
Update the sibling tests that pinned the old use_gateway-writing
contract: image/video selector and reconfigure rows now assert the
single provider string ('nous' managed / 'fal' BYOK) plus legacy-key
popping, the stt/video picker writes drop the use_gateway expectation,
the web_server managed-browser select asserts the persisted 'nous'
cloud_provider, and explicit-local STT pins no-cloud-fallback against
a stored raw-config selection.
New tests/tools/test_strict_provider_selection.py covers read_selection
semantics (legacy use_gateway interpretation, seeded stt local, empty
strings, browser.backend vs cloud_provider) and the three strict
behaviors per category: managed 'nous' selection wins over present
direct keys, a vendor selection with missing credentials raises the
selection-naming error with NO managed call, and never-configured
installs keep today's autodetect. Updated the tests that pinned the old
credential-first precedence (TTS resolver gateway override, STT silent
managed fallback, web invalid-backend reroute, video_gen picker writes).
Sabotage-verified: reverting the image FAL strict switch makes the new
managed-selection tests fail.
`_spawn_gateway_restart` already reuses an in-flight `hermes gateway
restart` child so a double-clicked button cannot start two racing
restarts. That guard evaporates exactly when it is needed most: the
child exits as soon as it has handed the restart to the supervisor (or
to the running gateway), long before the gateway is actually back, so a
stale cached dashboard frontend re-firing its own restart every few
seconds cleared the guard on every attempt and started a fresh restart
each time.
#89034 measured the result on an s6-supervised container: 77
`gateway-restart started` entries, 17 of them inside one minute. Each
one SIGHUPs a gateway that is still coming up, and killing it
mid-FTS5-write corrupted `state.db` ("database disk image is
malformed", 203x in agent.log) until the operator recreated the file by
hand.
Requests for the same profile within GATEWAY_RESTART_COOLDOWN_SECONDS of
the last spawn are now coalesced onto that spawn and logged, so a storm
produces one restart instead of one per request. The window is fixed
rather than health-gated on purpose: a gateway that never comes back
would leave a health-gated restart action permanently inert, which is a
worse failure than the flood it prevents. The cooldown state is kept
outside `_ACTION_PROCS` because completed action children are reaped out
of that table, and a guard that disappears when the child exits is the
bug being fixed.
Only the *frontend-flood* half of #89034 is addressed here. The s6
`finish` death-cap the report also asks for is a separate change to
`hermes_cli/service_manager.py` with a much larger blast radius, and is
left for a maintainer decision.
The re-exec'd child inherits the console, so sys.stdin.isatty() still reported
a terminal and the update asked its local-changes question. By then the parent
shim had exited and the shell had taken the console back, so the prompt could
not be answered and the update sat there forever — worse than the lock it
replaced, because nothing recovers without closing the window.
Spawn the child with stdin closed. It then takes the same path the gateway and
Desktop updates take: honour updates.non_interactive_local_changes, which
stashes by default so nothing is lost, and keep going without asking.
Detection across every launch variant (argv[0], the zipapp __main__.py, the
main-module spec origin, the ancestor chain) plus the venv scoping that keeps
an unrelated hermes.exe from triggering a hand-off; the re-exec's argv, env
marker, loop guard and both fall-through paths; the pending-rename filter;
and the venv/.venv layout split.
Retires the reboot-deferred quarantine assertion along with the fallback.
Launch-variant cases from #89970 by @Akloenx123, pending-rename cases from
#88121 by @fangliquanflq.
Covers the new --yes/-y flag on , asserting the
parsed value reaches do_uninstall(skip_confirm=True) via the real
main() -> cmd_skills -> skills_command dispatch path. Mirrors the
install-flag test pattern in test_skills_install_flags.py.