Follow-up structural pass on the review fix:
- Runtime provider, auxiliary resolution, model validation
(hermes_cli/models.py), live discovery (bedrock_model_ids_or_none),
and the Mantle URL/SigV4 fallbacks all resolve their region through
resolve_bedrock_runtime_region() — one canonical implementation of the
config-first priority instead of three hand-rolled copies.
- agent_init: drop the 'if "client_kwargs" in locals()' guard by
initializing client_kwargs unconditionally at the top of the else
branch; the Mantle kwargs hook is a documented no-op for non-Mantle
base URLs.
Address review feedback on #65076:
- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
config-first region resolution (bedrock.region in config.yaml, then
AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
branch) to the new helper. Previously it derived its region with bare
resolve_bedrock_region() (env-first), so when config.yaml pinned
bedrock.region to a different region than the ambient AWS env, auxiliary
calls (compression, memory, vision) left the primary runtime's region.
Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
uses the OpenAI-compatible endpoint, which the Mantle route made stale.
Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
Converse), the Mantle auth model (bearer token or SigV4), and add the
GPT-5.5/5.6 model IDs to the models table.
GPT-5.6 Sol, Terra, and Luna went GA on Amazon Bedrock on 2026-07-13.
Like GPT-5.5, they are served exclusively from the Bedrock Mantle
OpenAI-compatible Responses endpoint (the model cards list
bedrock-runtime/Converse as unsupported), so they ride the allowlist
routing introduced for GPT-5.5:
- Add openai.gpt-5.6-{sol,terra,luna} to BEDROCK_OPENAI_RESPONSES_MODEL_IDS
so runtime resolution, auxiliary calls, and MoA slots all take the
SigV4/bearer Mantle Responses path.
- Surface the family in the curated Bedrock picker list.
- Record the 272K context window from the AWS model cards for all four
Mantle OpenAI models (previously fell back to the 128K default).
- Generalize picker tests from the hardcoded single-model checks to the
BEDROCK_OPENAI_RESPONSES_MODEL_IDS allowlist so future Mantle model
additions do not require test surgery; add routing, picker, and
context-length coverage for the 5.6 family.
Docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-openai.html
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
`_guard_named_profile_under_multiplexer` correctly refuses a named-profile
gateway while the default gateway is multiplexing — starting a second one would
double-bind that profile's platforms. The refusal is right; its exit code was
not.
The refusal is decided entirely by configuration (`multiplex_profiles` plus the
allowlist), so it is permanent: no number of retries can change the answer.
Exiting 1 made it look transient to a service manager.
That matters because this module generates the systemd unit, and the template
pairs `Restart=always` / `RestartSec=5` with `StartLimitIntervalSec=0` — it
deliberately trades systemd's generic start-rate limiter for the specific
`RestartPreventExitStatus=GATEWAY_FATAL_CONFIG_EXIT_CODE` backstop declared
three lines below it. Returning 1 left that backstop unarmed with the limiter
already disabled, so a correct, permanent refusal became an unbounded restart
loop. Observed on a host running `multiplex_profiles: true` with a leftover
per-profile unit: 136 refusals in ~13 minutes, stopped only by hand.
`GATEWAY_FATAL_CONFIG_EXIT_CODE` (78, EX_CONFIG) is this codebase's existing
answer for exactly this case — `gateway/restart.py` documents it as the fatal
configuration error that the s6 finish script translates into 125 "permanent
failure" (#51228). This adopts that contract rather than inventing one, so the
fix also works on s6 hosts, not just systemd.
After: one refusal, `status=78/CONFIG`, `NRestarts=0`, unit settles in `failed`.
Also strengthens the two guard tests. They asserted
`pytest.raises(SystemExit, match="1")`, but `match=` is a regex search over
`str(exc)`, so it passed for 1, 21, 100 and 111 alike — it read like an exit-code
assertion while pinning nothing. They now assert
`excinfo.value.code == GATEWAY_FATAL_CONFIG_EXIT_CODE`. The exit code is the
contract here: it is the only thing that tells a supervisor the failure is
permanent.
hermes backup freezes mid-archive when a .db file under HERMES_HOME is
locked by another process — e.g. a live Chromium profile database held
with an exclusive lock by a running browser. sqlite3.Connection.backup()
retries SQLITE_BUSY indefinitely and never honors the connection's busy
timeout, while a plain statement on the same source fails cleanly after
~5s with "database is locked".
Fixes:
- Probe the source with a cheap read before snapshotting, so a locked
database fails fast instead of hanging the whole backup.
- Add a watchdog that interrupts the source connection after 15 minutes
as a last resort for pathological cases.
- Exclude browser-profiles/ from full backups: the CDP browser profile is
live, regenerable (cache + re-login), and unsafe to snapshot while
running. On a real install this cut the backup from 28,396 files /
1.1 GB to ~4,000 files / 548 MB, completing in ~33s instead of hanging.
The pre-update automatic backup shares this code path and was equally
at risk.
Adds a regression test that holds an EXCLUSIVE transaction in a separate
process and asserts _safe_copy_db returns False in bounded time.
`hermes backup` already skips `backups/` so a full zip never re-ships
earlier pre-update zips. `state-snapshots/` (written by `hermes backup
--quick`, `/snapshot create`, and the pre-update safety net) has the same
shape — every retained snapshot holds its own copy of state.db — but was
not in `_EXCLUDED_DIRS`, so a full backup shipped the DB once per
retained snapshot on top of the live one.
Two places hit this in practice:
- `hermes update` in `full` mode takes the quick snapshot *before* the
full zip, so the pre-update zip always nests the snapshot it just made
(state.db twice in every pre-update-*.zip).
- Any recurring `hermes backup --quick` (default keep=20) makes a daily
`hermes backup` grow by roughly one compressed state.db per retained
snapshot; a 750 MB state.db with two snapshots on disk pushed a daily
zip from 1.8 GB to 2.3 GB.
Add `_QUICK_SNAPSHOTS_DIR` to `_EXCLUDED_DIRS` (moving the constant up
next to the exclusion rules so there is one source of truth). Both walk
sites and `_should_exclude` share the set, so `hermes backup`, the
pre-update zip and the auto-backup path all pick it up. Restoring
snapshots after a machine move was never the point of the full backup —
`profiles.py` already excludes `state-snapshots/` from `--clone-all` for
the same reason.
Tests: unit case next to the `backups/` one, plus two end-to-end cases
that use the real `create_quick_snapshot` producer and assert the zip
carries exactly one state.db (full backup and pre-update-order).
Free ($0/$0) Nous Portal models sat with a blank discount column and no
sale star (stealth/ox-alpha, upstage/solar-pro4:free), reading as missing
data next to the -20% sale rows. compute_sale_discount now returns a flat
100% for free models; was_* raws pass through only when the gateway served
a pricing.original, so natively-free models render bare '-100%' with no
fabricated 'was ?/?'. CLI picker star follows on_sale automatically;
inventory feed carries discount_percent=100 to Desktop, whose FREE badge
row now renders the amber -100% pill beside it.
Follow-up to the salvaged GLM-5.3 support commit: drop z-ai/glm-5.1 from
both curated lists per Teknium's direction (glm-5.2 keeps the 'default'
tag), and regenerate the docs manifest. glm-5.1 remains available via
live discovery and on out-of-scope surfaces (zai plugin, setup defaults,
opencode-go) — named leftovers, not silently swept.
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.
- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
by the endpoint, HTTP 200)
The OpenCode Zen wire slug for the Ox Alpha stealth model is opaque
(x-preview-f-free); users searching the picker for 'ox' or 'ox-alpha'
found nothing. Adds the search alias across all four synced alias
tables (CLI, desktop, web, TUI) plus tests. Wire id is unchanged and
still what renders and gets sent to the provider, matching the k3 →
kimi-k3 precedent. No canonical-dedup collision with opencode-go's
keyed ox-alpha-free slug.
resolve_exec_command wrote the repo hermes script (env-python shebang)
straight into Exec=; spawned by the DE that shebang escapes the venv and
dies on the first import, invisibly (Terminal=false, entry rewritten
every launch). A python-script launcher whose shebang points outside the
running interpreter's env now gets Exec={sys.executable} {script} desktop;
native binaries, bash wrappers, and venv-shebang scripts are untouched.
Third surface for the Ox Alpha stealth reasoning model (after the
OpenCode Zen rollout in #91250 and the OpenRouter listing in #91284).
Adds stealth/ox-alpha to the curated Nous list and regenerates the docs
manifest. Free on the portal ($0/$0), 1M context, 131K max output —
verified against the live inference-api.nousresearch.com/v1/models.
Provider-agnostic metadata already resolves via the bare ox-alpha slug
(DEFAULT_CONTEXT_LENGTHS 1,048,576; reasoning_timeouts 300s floor), and
the nous route bills via official_models_api, so no pricing snapshot is
needed.
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.
- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
snapshot set, per-file 1GiB cap, and keep policy as the invoking
profile (no partial tier, no new restore-coherence class), each into
the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
runs the #34600 cron-loss safety net per profile against its OWN
snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
profile's (best-effort, receipt-recorded); post-update cron restore
extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
The feature flag was only documented in cli-config.yaml.example; every other
browser.* key is declared in DEFAULT_CONFIG so config tooling (dashboard
editor, hermes config get) can see it. Defaults unchanged: enabled=False,
developer_mode=False. Surfaced during review of PR #85351.
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:
Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
(absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free
Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
+ laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
keyless Ox Alpha; keyed — Go relay 401s anonymous requests)
Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
glm-5.2:free 256K (the free variant is capped below the 1M paid entry)
Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
MEMBERSHIP in the verified opencode-free catalog, not the -free
suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
suffix-based healing would have routed it to a Zen relay that
doesn't serve it (verified: zen 401s 'not supported', go 401s
'Missing API key'). New regression test pins this.
Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
A keyless provider has no credential to lack, but every auth-gated
surface treated 'no key' as 'not authenticated', so opencode-free was
invisible in /model, provider:model listing, and the desktop model
pickers unless a user had unrelated OpenCode env vars set.
One policy, three gates, all derived from the HermesOverlay keyless
flag (#91358):
- auth.py get_api_key_provider_status: keyless providers report
configured/logged_in=True with key_source 'keyless' — flows through
get_auth_status to every status consumer (hermes status, dashboards,
list_available_providers).
- model_switch.py list_authenticated_providers: keyless overlay rows
get has_creds=True before any env/pool/auth-store checks — this is
the source for /model, the TUI picker, and the desktop
/api/model/options payload.
- inventory.py explicit-only filter (desktop chat pickers): keyless
providers are kept — there is nothing to 'explicitly configure', and
hiding a zero-setup provider defeats its purpose.
E2E (temp HERMES_HOME, all keys stripped): get_auth_status logged_in,
list_available_providers authenticated, picker row with 6 models,
desktop payload default AND explicit_only both include the provider,
and the full switch pipeline (parse free:x-preview-f-free →
switch_model) resolves to the keyless runtime. 4 new tests.
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.
- hermes_cli/update_inventory.py (new): side-effect-free runtime
inventory — install kind via detect_install_method (git / docker / nix
/ apt, updatable-in-place or not, with the correct external update
command for image/package-managed installs), all profiles, every live
gateway with its supervisor (systemd / launchd / manual via the
fleet-wide _get_service_pids), running code_sha/code_version from the
#91283 gateway_state.json stamps, and the restart mechanism each
runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
docker/nix refusal gates so image-managed installs get a useful
'not updatable in place + right command' report instead of a bare
refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
('plan' key) and prints a one-line fleet summary, so post-mortems can
compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
never-raises, JSON round-trip for the receipt, print output shapes,
receipt integration.
Live verification (2026-08-21): big-pickle and mimo-v2.5-free return 429
FreeUsageLimitError for ANY User-Agent except the opencode CLI's own
'opencode/latest' — same IP, no cooldown effect, while the other six free
models serve our honest HermesAgent UA freely. Hermes sends deliberate
attribution headers and does not impersonate other clients, so these two
models are broken for our users by policy on OpenCode's side; delist them
rather than ship dead picker entries.
- opencode-free catalog: 8 -> 6 models (both curated lists)
- plugin default_aux_model: big-pickle -> laguna-s-2.1-free (fastest
non-gated free model)
- keyless predicate keeps big-pickle (it IS free-tier; correct routing if
a user enters it manually — they get the relay's own 429, not our 401)
E2E: picker shows 6, laguna aux default completes a keyless agent turn.
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:
- _restart_macos_launchd_gateways(): the invoking profile keeps the
existing launchd_restart() path; every other gateway of this install
is drained via SIGUSR1 (same as systemd siblings), then hard-
kickstarted unless KeepAlive already respawned it, then verified on a
fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
counts toward failed_or_stale_units — including timeouts during
liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
derives labels from THIS install's profiles (get_default_hermes_root),
not by globbing the shared per-user ~/Library/LaunchAgents — a
sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
installs) must never enumerate, let alone restart, another install's
fleet. This also keeps the hermetic test suite blind to a dev
machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
liveness, kickstart, and fresh-PID verification all use the domain the
service was actually located in (gui/<uid> vs user/<uid> probed per
label via `launchctl print`). This addresses the #41403 review defect:
the process-wide _launchd_domain() cache resolves the current profile's
domain and must never be reused for a sibling. _launchd_domain() itself
becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
sweep excludes every gateway service PID (mirror of the systemd
hermes-gateway* pattern) so it cannot mistake a freshly respawned
sibling service for a stale manual gateway. Default-scope callers
(gateway status, cron checks, stop_profile_gateway's orphan reaper —
which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
hints for launchd labels alongside the systemctl ones.
Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).
Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review on #91283: begin_update_receipt() fires early in _cmd_update_impl,
but finalization only existed on the success/ZIP/CalledProcessError
paths. Early sys.exit paths (Windows concurrent-instance preflight,
venv-holder refusal, head-pinned no-op, fetch failure) terminated with
the receipt started but never written — losing exactly the refused/
failed runs the receipt matters most for.
- update_receipt.py: finalize_pending_update_receipt(exit_code,
stop_reason) — boundary safety net; maps exit 2 → 'refused', other
non-zero → 'failed'; records exit_code + stop_reason. Exactly-once by
construction (singleton popped in finalize_update_receipt), so runs
the inner paths already finalized are untouched.
- main.py cmd_update: SystemExit/BaseException/else arms around
_cmd_update_impl persist any still-open receipt with the real exit
code, then re-raise unchanged. Future early exits are covered without
per-site finalize patches.
- 5 regression tests incl. end-to-end through the real cmd_update
wrapper (exit 2 preserved, outcome 'refused', stop reason recorded,
singleton cleared, exactly one receipt file).
opencode-free broke two provider-surface contract tests: 'api_key
providers must expose a credential env var' and 'GUI ⊇ hermes model
universe'. Both premises assume a credential exists. Add a keyless flag
to HermesOverlay + ProviderDescriptor (same derived-exemption pattern as
virtual providers) so any future anonymous provider is covered without
hardcoded slugs. Nothing to configure = no Providers-tab card, by design;
the model picker remains the selection surface.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
Follow-up to the #74075 salvage: _reap-path _get_service_pids() call now
passes all_profiles=True. With the ps scan fixed, the reaper's process
scan surfaces sibling-profile launchd gateways on macOS; excluding only
the current profile's label would misclassify them as unsupervised
orphans and reap them (same class as the update-sweep sites the
contributor fixed). Also refresh the stale 'ps -A eww' comment.
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
on macOS/BSD ps, making the fallback silently return [] on every macOS
machine. The matcher only needs argv (not env vars), so e is
unnecessary. -ww keeps unlimited-width output on both BSD and
procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
on macOS, enumerate every ai.hermes.gateway* launchd agent across
profiles via bare launchctl list instead of only the current
profile's label. This prevents the update sweep from misclassifying
sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
_get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
pass all_profiles=True so the update fleet sweep excludes every
service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
launchctl list <label>, all_profiles uses bare launchctl list
with prefix filtering, handles empty/broken output gracefully, and
preserves systemd behavior.
Tranquil-Flow
Adds OpenRouter's free "Ox Alpha" stealth reasoning model
(stealth/ox-alpha) to the OpenRouter fallback snapshot, plus the
provider-agnostic metadata it needs:
- OPENROUTER_MODELS: free-tier entry (1M ctx)
- DEFAULT_CONTEXT_LENGTHS: ox-alpha -> 1,048,576 (verified against
OpenRouter live /api/v1/models; without this the slug fell through
to no match)
- reasoning_timeouts.py: 300s stale floor for ox-alpha and the
OpenCode Zen twin slug x-preview-f-free (reasoning model,
long-horizon agentic work per its model card)
- model-catalog.json regenerated
Pricing snapshot skipped: openrouter bills via official_models_api
(live pricing; model is free anyway).
The Zen relay serves *-free models (x-preview-f-free / Ox Alpha) ONLY
anonymously: any Authorization bearer it doesn't recognize is a 401
'Invalid API key' — including our no-key-required placeholder and valid
OpenCode GO subscription keys. The Go relay doesn't serve the free tier
at all ('Model x is not supported'). So the free model failed for every
Hermes user: keyless setups got the placeholder bearer, and OpenCode
subscribers sent a Go key to a relay that rejects it.
Fix (class-wide for all 8 current *-free Zen slugs, not just Ox Alpha):
- hermes_cli/models.py: is_opencode_zen_free_model / opencode_zen_free_runtime
/ opencode_zen_free_headers — one shared policy: free slugs pin to the
Zen relay with a keyless placeholder and an empty Authorization header
that overrides the OpenAI SDK's 'Bearer <key>'.
- runtime_provider.py: free slugs route through the keyless runtime before
the credential-pool/explicit/api_key paths (no key required; Go
selections heal to Zen). Paid models still fail closed without a key.
- agent_init.py + auxiliary_client.py: the placeholder key swaps in the
empty-Authorization headers at both client-build chokepoints.
Verified live (2026-08-21): anonymous chat/completions 200 incl. tools,
streaming, parallel; bad bearer 401; full E2E AIAgent turn with a real
terminal tool round-trip completes keyless under both opencode-zen and
opencode-go providers. Sabotage run: routing tests fail without the fix.
get_code_identity() shelled 'git rev-parse HEAD', which broke two tightly
mocked test suites (sequenced subprocess.run side effects in the
head-moved gate, call-count asserts in the Windows taskkill test) and
added process-spawn cost to gateway runtime-status writes.
_resolve_git_head_sha() now reads HEAD/refs/packed-refs directly,
handling regular checkouts and worktree/submodule pointer files.
Also: skip the 2s fleet settle wait when the restart phase touched no
gateways, and hoist killed_pids init outside the restart try-block.
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
sol-reviewer round-4 IMPORTANT (reproduced by execution): the enroll
warning read multiplex_profiles via load_gateway_config() under the
SECONDARY profile's HERMES_HOME, but the flag normally lives in the
DEFAULT root's config.yaml — so the warning never fired in the real
topology, preserving the round-3 defect it claimed to fix.
The topology decision now mirrors the multiplexer-conflict guard in
hermes_cli/gateway.py: secondary detection is the resolved-path
relationship to <default_root>/profiles/ (not a directory-name
heuristic — also fixes the round-4 MINOR false positive on unrelated
dirs named 'profiles'), and the multiplex flag comes from the
GATEWAY_MULTIPLEX_PROFILES env override or a raw read of the default
root's config.yaml. The raw read also avoids running the full
enablement pass (round-4 MINOR: load_gateway_config() emitted the
relay-exclusive sweep's own warnings into enroll output).
The warning now replaces the generic 'restart to pick up the new env'
line instead of following it (round-4 NIT: the two messages were
contradictory), and the helper returns whether it fired.
New test file pins all six topology cases, including the exact
false-negative reproduction (flag in default root only) and env
override in both directions.
sol-reviewer round-3 IMPORTANT: gateway enroll --connector-url /
--wake-url persist GATEWAY_RELAY_URL / GATEWAY_RELAY_WAKE_URL into the
active profile's .env and tell the user a restart activates them. For
a SECONDARY profile of a multiplexed gateway that is silently untrue:
the routing stamps are process-global, and a secondary profile's .env
is loaded into an isolated secret scope, never exported to os.environ,
so the gateway can never read them from there. Enrollment (the
credential exchange) still succeeds and the creds are still valid, so
warn rather than refuse, pointing at the process environment or the
default profile. Single-profile gateways and the launch profile are
unaffected (load_hermes_dotenv exports their .env at startup).
The named-custom-provider runtime path returned a static api_mode, so a
providers: entry like opencode-go-bridge -> https://opencode.ai/zen/go/v1
sent responses-only models (grok-4.5, gpt-5.6-luna) to /chat/completions
and got HTTP 503 (#85589 repro). Now: when the provider name is in the
OpenCode family or the base_url is hosted on opencode.ai, derive api_mode
from the effective model and run the symmetric /v1 normalization — unless
the user declared an explicit transport, which stays authoritative.
5 new regression tests against a real temp HERMES_HOME config.
Builds on @Lesnak1's #85619 (issue #85589):
- New opencode_provider_family() single-owner predicate in
hermes_cli/models.py — resolves built-in AND custom family providers
(opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
x4 from the salvaged commits) plus 4 sibling sites the PR missed:
cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
model_normalize.py flat-namespace strip, model_switch.py base_url
normalization.
- Responses transport: alias OpenCode-reserved function names
(web_search, search_files -> hermes_*) on the wire and map them back on
dispatch — same pattern as the xAI web_search collision fix. Matches
family providers and any base_url on opencode.ai. Fixes the HTTP 400
'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.
Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.
- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
With the gateway owning the suspend, the idle predicate covers every
work source (agent turns, cron jobs, API-server runs, background work,
fail-awake on unreadable sources) and the relay drains + flips before
the freeze, so a long timeout no longer buys safety — real work always
blocks the suspend and resume is sub-second. 5 idle minutes just bills
idle RAM. Per-instance override stays config.yaml
gateway.scale_to_zero.idle_timeout_minutes (D2).
New behavior-contract test: invalid config values degrade to the module
default (whatever it is), never zero/negative — asserts the RELATION,
not the literal, per the no-change-detector-tests rule.
The config key is declared in DEFAULT_CONFIG and shown in the auxiliary
config UI, but the judge path never read it — a user raising the timeout
for a slow-but-healthy endpoint got the same 30s cap, and the loop
auto-paused on transport failures advising a provider/key check. Mirror
the _goal_judge_max_tokens reader and resolve the timeout at call time;
explicit timeout= arguments still win.
Fixes#91022
`uv pip install -e .` never audits an editable target. It reinstalls on every
invocation and rewrites the console-script shims each time, which is the only
reason `hermes update` has to quarantine the running `hermes.exe` on Windows —
and a quarantine that loses its race is the whole `os error 32` family.
Gate the reinstall on whether the pull actually touched a file that defines the
install. It's safe to skip because the editable finder is pinned to a static
module list (`py-modules` + `packages.find.include`), so the one source-only
change that could stale it — a new top-level module or package — cannot land
without a `pyproject.toml` diff. Dependencies and `[project.scripts]` live
there too, and new submodules inside an already-mapped package resolve through
the real directory.
The predicate fails closed: no pre-pull SHA, an unresolvable one, or a failed
`git diff` all reinstall as before. On the skip path the two verifiers that
normally run inside the install run directly, so a wrong skip self-heals into a
real install rather than leaving an unchecked venv.
This is the pattern the file already uses everywhere else — `_tui_need_npm_install`
diffs node_modules against package-lock.json, and the desktop build is gated on a
content hash so `hermes update` "will skip if nothing actually changed". The
Python editable install was the one path with no such gate.