Following the grant's source on every save made a fresh device-code
login (or `hermes auth import`) under a profile that had been borrowing
root's Codex grant overwrite root's account instead of creating the
profile's own. Redirecting a save into another file is the exception, so
it is opt-in: the refresh path passes write_through=True; login, import
and recovery keep saving locally. The two save branches collapse into one
(store, path, set_active) triple.
Test: root discovery on Windows comes from LOCALAPPDATA — set it so the
fixture's root is the resolved root on every host.
Codex refresh tokens are single-use with rotation-family reuse
detection. _save_codex_tokens resolved the state via the profile's
root fallback but always persisted into the ACTIVE (profile) store, so
a profile-scoped refresh left the global store holding the consumed
refresh token — the next process to read it replayed it and OpenAI
revoked the whole rotation family, forcing a manual device-code
re-auth (#87503; observed four times on one multi-profile deployment).
Mirror the xAI source-aware save (#43589/#74339): resolve the state
with _load_provider_state_with_source; when the grant came from the
global root, write the rotated chain back to root only — singleton AND
credential_pool entries, under the root store's own lock, without
creating a shadowing profile key. Best-effort, with the same pytest
seat belt as the xAI path.
Fixes#87503
The shebang read-and-classify wrapper was a byte-for-byte twin of
_needs_interpreter, which already delegates to _shebang_escapes_running_env
and carries its own edge-case tests; import the whole predicate instead of
half of it. Lazy import: linux_desktop_entry lazily imports resolve_hermes_bin.
The launcher guard rejected every python shebang, so a pip/uv console
script pinned to the running venv (#!<venv>/bin/python) was also
discarded in favour of `python -m hermes_cli.main`. Reuse
linux_desktop_entry._shebang_escapes_running_env, which already knows
that `env` shebangs escape and a shebang inside the running
interpreter's directory does not; only the escaping launcher loses the
venv. Also drops the second shebang classifier the fix had introduced.
attested_gateway_died() re-ran find_gateway_pids() (current profile only)
although both callers had just proven the process table empty with
all_profiles=True, and it re-implemented check_start_attestation's
liveness rule. Callers now pass the liveness they hold (current_pids=[])
and both probes share _attested_dead(), so the consuming and read-only
twins cannot drift.
attested_gateway_died() is deliberately read-only so the CLI-start warning
still fires, but that left the dead marker in place after the update path
acted on it. If the restored gateway never became ready (or died again
before the next CLI start consumed the marker), the same stale crash marker
would re-authorize another cold start against Desktop ownership on the
next update. Clear it via the existing _clear_start_attestation() path
right after _spawn_detached() succeeds - the marker has done its job at
that point; a new one is written once the spawn is confirmed ready.
_attested_pids_from now returns [] when "pids" is missing, null, or any
non-list value. The previous `data.get("pids", [])` only guarded a missing
key: `{"pids": null}` raised TypeError on iteration and a scalar would too.
attested_gateway_died() drives a Desktop-ownership override for the update
cold-start, so a malformed marker must never be able to authorize a spawn.
A Desktop self-update hand-off exits the app before the updater runs and can
kill the messaging gateway in those same seconds (#109538), so the updater's
discovery finds no live PID while the one-shot start attestation still
vouches for the dead one. Both Desktop-ownership checks then read "nothing
running" as "nothing to restore" and the bot stayed down until a manual
start.
Consult the attestation non-destructively before Desktop-owned lifecycle
suppresses a cold-start: a vouched-for PID gone without a clean ledger exit
keeps the plan and is restored; no attested death preserves the #76129 skip
unchanged.
Two remaining halves of #89184 (Desktop Settings saves rewriting unrelated
config):
- The `fallback_providers` structured editor normalized every entry down to
`{provider, model}`, so any edit (remove a row, pick a model) re-emitted a
hand-written local-gateway chain without its `base_url` / `api_key` /
`key_env` / `api_mode` — the next autosave persisted bare pairs and the
fallbacks silently routed to the public provider. Entries now carry every
key through; the editor only owns the two selects.
- `PUT /api/model/moa` did `cfg = load_config(); cfg["moa"].update(...);
save_config(cfg)`: the whole default-expanded snapshot went back to disk,
so a Desktop MoA autosave re-persisted every other section too (the
2026-09-10 repro: `fallback_providers: []` written alongside the MoA block
the user had just edited). It now saves `{"moa": ...}` with
`merge_existing=True`, the same section-scoped write every other sparse
writer uses since #110535. Hand-edited moa keys (#58819) still survive.
The `model.default not persisted / base_url cleared` symptom from the 0.20.4
report no longer reproduces on main through the real REST path (Config page
diffs against a baseline since 5361867c6d32; `_denormalize_config_from_web`
keeps the on-disk `model:` block).
SessionDB(read_only=True) cannot create a missing store, so a fresh install
running hermes insights / /insights errored instead of reporting no data
(reported by @ehz0ah on #110718; guard shape from @kshitijk4poor's #110026).
Co-authored-by: kshitijk4poor <kshitijk4poor@users.noreply.github.com>
hermes insights and /insights only read; taking the default writer runs
schema init and contends the write lock on a live store.
Salvaged from #109737.
auxiliary.background_review.reasoning_effort was a silent no-op on the
same-model path: the fork inherits the parent's reasoning_config verbatim to
keep prompt-cache parity (#30532), and nothing told the user. Emit a one-time
user-visible warning (parent-scoped, so a nudge-per-turn session warns once,
not per fork) when the key is actually set, document the no-op in the config
reference, and leave the fork-birth request bytes unchanged.
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.
- gateway/platforms/webhook.py: cron_job route mode — after the same
HMAC auth / rate limit / filters / script / idempotency as agent
routes, the rendered prompt becomes transient per-run context and the
job fires through execute_job_for_event on a worker thread (202
Accepted immediately). Startup validation rejects cron_job +
deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
the shared claimed-run body (_execute_job_now), so event fires share
at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
subscriptions; job ref validated (and canonicalized to the job ID) at
create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
CLI tests in test_webhook_cli.py.
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.
- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
and it now honors the base_url argument instead of hardcoding
api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
incl. single-page and stuck-cursor termination; updated the two URL-pinning
pool-discovery tests for the ?limit=1000 contract
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).
- hermes_cli/models_catalog_static.py: remove both slugs from the
opencode-free offline floor and the opencode-zen discovery floor;
document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
invariant to cover both.
The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
Passive update checks no longer run git fetch (338bf9ea9a); the thread
still shells out to rev-parse/remote get-url, which is what the pytest
no-op guards against. Also explain why the predicate checks sys.modules.
The prefetch_update_check daemon thread (started at tui_gateway.server
import time) shells out to git via the shared subprocess singleton at an
arbitrary point after import. Tests that patch subprocess.run/Popen
process-wide can capture that stray spawn in call_args, flaking their
assertions: on 2026-08-28 CI, test_slash_worker_popen_uses_utf8_replace
saw encoding=None from the thread's un-encoded 'git fetch' (red on main,
run 33175879563) and test_deliver_validates_profile_and_runs_transport
captured argv ['rev-parse', 'FETCH_HEAD'] from the shallow-checkout
banner path (FLAKY frame, run 33183215857).
Fix the class at the source: _skip_background_prefetch() makes both
prefetch_update_check and prefetch_banner_data no-ops under pytest
(nothing in tests needs a live update check; the done event is set so
get_update_result callers don't burn their timeout). Tests exercising
the prefetch itself monkeypatch the predicate. Sabotage-verified
regression tests pin both no-ops.
gh CLI 2.98+ removed the 'authenticated' field from 'gh auth status --json'
(only 'hosts' remains), causing the command to exit 1 even when the user
is authenticated. The doctor then falsely reports 'No GITHUB_TOKEN' despite
the user being logged in via 'gh auth login'.
Since the code only checks the return code (it never parses stdout), the
'--json' flag is unnecessary. 'gh auth status' without it works across all
gh versions and returns exit 0 when authenticated.
Fixes: gh auth status --json authenticated → gh auth status
omo's omo-agent-toolkit worktree-sweep (their PR #7151) added three
capabilities our hermes worktree command lacked:
- --json on list and prune: machine-readable audit/result payloads so
scripts and agents can consume verdicts without scraping table output.
- --older-than DAYS: an age floor that only ever RESTRICTS reaping
(young-but-reapable trees are kept); it never widens eligibility, so
the existing safety invariants are untouched.
- External-tree visibility: linked worktrees registered outside
.worktrees/ are now reported read-only in the audit (branch, locked,
missing) instead of being invisible, and registrations whose
directory has vanished are dropped via git worktree prune (metadata
only, no files touched) during prune.
Not ported: omo's ancestor-of-default-branch merge test (our git cherry
patch-equivalence is strictly stronger under rebase/squash merges), and
their hardcoded external-root exclusion list (we exclude by location:
everything outside .worktrees/ is hands-off).
Tests: 9 new contracts in tests/hermes_cli/test_worktree_gc.py (age gate
restrict-only, external trees never reaped, stale-registration prune
dry-run/real, JSON shapes, negative --older-than rejected). Live E2E on
a scratch repo verified all three flags end to end.
Tests collapse 13 change-detectors into 7 invariants (silent below threshold /
without context, fires above, same-model re-select silent, config override and
0-disables, registry threading, agent context derivation). The legacy 5-arg
guard test goes with the TypeError fallback it covered: that fallback was
defence for a case nobody has (every in-tree guard and test double is
*args-tolerant) and would re-run a guard whose real TypeError it masked, so the
rebased port passes the context positionally like every other argument.
Docs no longer claim the confirm fires on the Telegram/Discord pickers or the
dashboard: those surfaces call combined_selection_warning() without a live
agent, so the context-cache guard is (correctly) silent there.
Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).
- hermes_cli/model_selection_guards.py: new context_cache guard +
SelectionContext carrier + selection_context_for_agent() helper;
registry threads live-session facts to guards (6-arg signature with a
TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
resolve_runtime_provider returns the bare billing class 'custom' for every
named providers:/custom_providers: entry; the configured id only survives in
requested_provider. All three fallback resolvers (gateway, TUI/desktop, cron)
persisted runtime['provider'] as the agent identity, so an automatic fallback
labeled the session 'custom' in the UI and billing rows, while a manual
/model switch to the same provider showed the configured name.
New shared helper hermes_cli.fallback_config.effective_runtime_provider()
upgrades the bare class back to the entry's configured identity (ad-hoc
provider: custom entries stay unchanged), applied at all three sites —
same class as the delegate_tool fix.
The ChatGPT Codex models endpoint interprets client_version as a Codex
CLI compatibility version and filters out any model whose
minimal_client_version is newer than the value sent. Hermes hardcoded
client_version=1.0.0 at both catalog request sites, so model visibility
was accidentally coupled to a version scheme Hermes doesn't follow —
future models gated behind a higher minimal version would silently
vanish from the account catalog.
The backend accepts the exact sentinel 0.0.0 as an ungated request
returning the complete account catalog (verified live: 0.0.0 and
current versions return identical model sets today, while omitting the
parameter is HTTP 400 and out-of-sequence values like 0.0.1 return no
models). Both request sites (hermes_cli/codex_models.py and the
context-length probe in agent/model_metadata.py) now share one
CODEX_UNGATED_CLIENT_VERSION constant.
Clean-room port of the observed behavior in zed-industries/zed#62729;
no GPL code translated.
Snapshot-then-replace let a voice transcript or interrupt re-queue that landed
between the two steps be dropped by edit/rm/move/clear. Every mutation is now one
critical section under queue.Queue's own mutex.
A leading management word alone no longer hijacks prompts: "/queue clear the logs"
and "/queue edit the config" enqueue as before; only "/queue clear", "/queue edit N ...",
"/queue rm N", "/queue move A B" (numbers present) manage the queue. Handlers live on
CLILoopsMixin next to the existing _cmd_queue (cli.py is a facade), dispatched via a
verb table, and the CommandDef lists the subcommands for tab completion.
Droid v0.203 (Aug 25 2026) added 'edit queued messages' — a queued steering
message can be pulled back and changed before it is sent. Hermes /queue could
only append blindly: no way to see, fix, drop, or reorder queued prompts.
/queue now supports management subcommands in the CLI:
- /queue — list pending prompts (bare prompt still enqueues)
- /queue list — same
- /queue edit N <p> — replace item N (keeps voice sentinel, #65827)
- /queue rm N — remove item N
- /queue move A B — reorder
- /queue clear — drop everything
- /queue add <p> — force-enqueue prompts starting with a management word
Queue mutations hold queue.Queue's mutex and rebuild unfinished_tasks so
join()/task_done bookkeeping stays consistent. Paste references expand on
enqueue and edit, matching the old inline path.
Reimplementation of PR #18833 by @abhinav11082001-stack (commit was authored
under a fabricated 'Hermes Agent' noreply identity that cannot be carried
into history; engineering credit is theirs), hardened for current main:
voice-sentinel-aware previews/edit, paste-reference expansion, queue
bookkeeping asserts, out-of-range no-op tests, and docs.
_is_sensitive_path documents itself as the read-side guard for list/read/
download (#57505), but only fs_read_data_url and fs_download called it.
fs_read_text returned .env / auth.json / mcp-tokens/* contents to an
authenticated dashboard session and fs_list enumerated them.
Move the check into _fs_regular_file, the resolver every fs reader goes
through, and drop the two per-handler copies. fs_list filters on the same
predicate alongside _FS_READDIR_HIDDEN.
Reported-by: Brian Grablin <bgrablin@gmail.com>
Backend half of the per-task effort control, on today's layout: POST /api/model/set
distinguishes omitted (leave the task's override alone) from explicit null (clear →
inherit) via model_fields_set, canonicalises a level through parse_reasoning_effort
(400 on an unknown one), and "Reset all to main" also drops every override. GET
/api/model/auxiliary returns reasoning_effort per task and the row summary shows it.
The inherit row reads "inherit · main model effort" (own i18n key in all six locales)
rather than reusing the provider's "auto · use main model" copy — the two mean
different things and the reused string read as "use the main model" for the effort.
Runtime already consumes auxiliary.<task>.reasoning_effort (agent/auxiliary_client.py)
and hermes model writes the same key (#110346), so Desktop and CLI now edit one value.
Closes#89259. Salvages #90649 by @higgs1729.
Settings → Model → Auxiliary gets a reasoning-effort selector per task next to the
provider/model pick (inherit / Off / level), sent as reasoning_effort on
POST /api/model/set and read back from GET /api/model/auxiliary.
(cherry picked from commit f09d008f10a81f57ed2426f835898c8e8ae595d7, resolved onto main; the backend half lives in
hermes_cli/web_server_config.py since the routers split and lands in the next commit)
The empty-value row means 'a child inherits the parent agent's effort' for the
delegation task, not a provider default. Wording adopted from #105431 by
@fangliquanflq, which added the same step for delegation alone.
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.
One request now carries a model pick AND its effort on every surface:
- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
<level>` (validated against `parse_reasoning_effort`; unknown level ->
`MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
`ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
high` applies the effort AFTER the agent swap (`switch_model` re-resolves
`reasoning_config` from config.yaml, so an earlier write is clobbered) with the
pick's scope (session; config on `--global`; `--once` snapshots and restores it).
The `/model` picker gains a third stage, "Reasoning effort for <model>", built
from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
`config.set model "X --reasoning high"` applies after the swap; session pin
(`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
one-turn restore carries `reasoning_config`; re-emits `session_info` so the
status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
`<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
`_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
the Copilot-only inline prompt; Copilot keeps its per-model level set via
`github_model_reasoning_efforts`, other routes get the ladder, catalog
`supports_reasoning=False` skips it) plus a "Reasoning effort for the current
model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
with the same step (+ "Provider default"), stored as
`auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
the task list ("openrouter · model · high"), cleared by "Reset all to auto";
tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
it.
Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
"Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
errors "Model names cannot contain spaces"; after switches and `config.get
reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
"reasoning: high", status bar "fable 5.1 high".
Gate the tombstone+notify on "a live default gateway has recorded a served set"
(recorded_served_profiles() is not None) rather than on the per-profile
_served_by_running_multiplexer probe: a multiplexer serves every dir under
profiles/, the signal is cheap, and the narrower probe falls back to config
derivation the CLI process cannot see. Trim the salvaged tests to two
invariants — ordering (unroute while the old home still exists and a stale
mkdir_under_hermes_home of it is refused; hot-serve after the move; no
tombstone left) and no-signal-without-multiplexer. Rollback on a failed move is
kept and covered by the same code path.
Builds on #109269 (xielevi). Fixes#109267.
The salvaged flag was spelled --profile, which collides with the global
-p/--profile that hermes_cli.main scans BEFORE argparse: `hermes webhook
subscribe x --profile compta` would switch this CLI process to compta's
HERMES_HOME and write the subscription into compta's webhook_subscriptions.json
— a file the default gateway's webhook adapter never reads — while the route
still lacked the profile key. #109020 special-cased the scanner for the webhook
subcommand; naming the flag --route-profile removes the ambiguity without
touching _scan_profile_flag: -p picks the gateway whose subscriptions file is
written, --route-profile picks which /p/<profile>/ prefix may hit the route.
Docs: cli-commands reference row, multi-profile-gateways webhook section, the
route `profile` field. Builds on #109020 (fangliquanflq). Fixes#109016.
A WS 'profile' param like '../../foo' normalized to a path component that
escaped the profiles root, letting a connected client bind an arbitrary
existing directory as a profile home (state.db opened there, and session
delete chains into per-id file cleanup under <dir>/sessions/).
get_profile_dir now validates the canonical name against the profile id
regex before joining it under profiles/. The regex only, not the reserved
list, so pre-reserved-list dirs like profiles/hermes keep resolving.
Callers that probe existence (profile_exists, _profile_home, the 4064
resolvers) treat ValueError as 'not found'.
Follow-up to the #109930 salvage (#109901). The probe endpoint was the reported
site, but the same class covers every router path that expands a secondary
profile's `${VAR}` refs while only a home override is installed:
`GET /api/mcp/servers` (a `${VAR}` in `url` expanded from this process's env)
and the `/auth` config read, whose expanded entry is handed to the OAuth worker.
Hoist the PR's inline wrapper into one `_profile_secret_scope` context manager
(mirrors `_run_dashboard_mcp_oauth`'s wrapping) and use it at all three sites.
Policy unchanged: scope miss still falls through to os.environ outside
multiplexing; under multiplexing a miss is a miss, never another profile's value.
Tests: the salvaged probe test now uses monkeypatch.setenv (no raw os.environ
mutation); one invariant test for the list endpoint, red on origin/main.
The /api/mcp/servers/{name}/test endpoint reads config and probes with no
profile secret scope installed, so config.yaml's ${VAR} expansion
(_env_ref_lookup) and the probe's interpolation resolve against the
dashboard process's own os.environ — the default profile's values (or
nothing) on a shared remote dashboard. A secondary profile whose
credential comes only from an external secret source (Bitwarden/
1Password) never resolves and the probe sends the literal placeholder,
so the server answers 400 while a fresh profile-scoped CLI process
works (#109901).
Wrap both the config read and the probe in _config_profile_scope +
hydrate_profile_secret_sources + set_secret_scope so refs resolve
against the requested profile's .env plus its per-home hydrated secret
sources, matching the multiplexed turn path (#84079 semantics).
Both default-profile process matchers (`gateway.status._command_line_belongs_to_profile`
and `hermes_cli.gateway._scan_gateway_pids._matches_current_profile`) rejected a named
gateway with a substring test for `--profile ` / ` -p `, which the equals spelling the
CLI pre-parser accepts (`--profile=ops`) slipped past. The default home's identity check
then adopted that gateway's PID, and a default-profile `gateway stop` with no pid file
scanned the process table and could SIGTERM the named gateway (review of #108352,
finding E). Both sites now ask `profile_flag_value()`, the same tokenizer the named
branch already uses.
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.
One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
`live_default_gateway_pid()` (hermes_cli/gateway_multiplex_served.py) read only the
pid record, so it returned None for a gateway that is alive but has no gateway.pid.
Consumers of the helper then reported the gateway as down:
- `hermes -p <profile> cron list` printed "Gateway is not running" with "jobs won't
fire automatically" while the multiplexer was firing that profile's jobs
- `hermes -p <profile> status` dropped its "running (via the default-profile
multiplexer)" line
- `named_profile_served_by_running_multiplexer()` returned False for a profile the
live gateway serves
The rest of the liveness surface already handles a missing pid file: the
`runtime_pid_probe` seam of `resolve_gateway_liveness()` exists for "launch-service
gateways with no live PID file" (hermes_cli/profiles.py, hermes_cli/web_routers/),
and `hermes_cli/gateway_migrate._live_gateway_pid()` reads "pid file, then runtime
status". This probe was the one call site that never got either.
Read the pid record first, then the PID in `gateway_state.json` validated against the
process table, matching `_live_gateway_pid()`. A record naming a dead pid still
resolves to None, so a stopped gateway keeps reporting stopped and cron keeps warning.
Related to #99631.
hermes profile delete calls hermes_state_registry.close_all_under(profile_dir)
before rmtree, which force-closes the shared handle goals.py cached for that
home. A same-name recreate in the long-lived dashboard process then reused the
closed object: save_goal swallowed the closed-db error and the replacement
state.db was never created. Drop the cache entry once the registry has torn
the handle down (it clears _shared_registry_owned at teardown) so the next
call acquires a live generation for the recreated profile.
Reshape the two salvaged commits onto current main (#109954):
- Move the boundary guard out of the gateway_migrate facade into a new sibling
hermes_cli/gateway_migrate_guards.py as a table of guard functions
(_AUTO_MIGRATION_GUARDS: service domain, UNIX user, HERMES_HOME tree) plus the
identity resolver. The facade grows by ~20 lines only (uid/runtime_home on
ProfileGateway, one seam, the hook wiring).
- Compare uids, not strings: live pid owner via /proc (ps fallback only on
macOS, where /proc does not exist), else the system unit's User= via
_read_systemd_user_from_unit (root when absent), else the home directory's
owner. None means unknown and never blocks.
- The home-tree guard reads the HERMES_HOME the installed unit pins, not the
directory the plan enumerated: that is where the gateway really runs and is
exactly the "stale copies under profiles/" shape from the report.
- When the default is detached, a service-managed secondary is a different
domain for the AUTO path (it must not elect the secondary's manager); the
explicit command keeps electing it as before.
- The explicit command surfaces the same findings as notices (dry run shows
them) and is never blocked by them; only the update hook refuses.
- Rename the opt-out key to gateway.auto_multiplex_migration (nested only, no
top-level alias) and read it before a plan is built, so false prints nothing
and touches nothing. The explicit command ignores it.
- Tests trimmed to the invariants: one parametrized boundary test that exercises
the real hook end to end (refuses, touches nothing, dry run shows the notice),
one "same user / same scope still migrates" control, one opt-out test.
- Docs: boundary table + renamed opt-out section in multi-profile-gateways.md;
one line in hermes_cli/AGENTS.md.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Athena <athena@olympus.local>
`hermes update` folds an eligible multi-profile install onto one multiplexed
gateway on its own, and there is currently no way to say no. The only lever,
`gateway.multiplex_profiles: false`, is also the default: `_read_multiplex_flag`
returns `False` for "absent" and for an explicit `false` alike, so an operator
who has already decided to stay on per-profile gateways has no way to record
that decision. The migration runs again on the next update.
Add `gateway.auto_migrate` (bool, default `true`). Read from the default
profile's config, it gates the automatic path only:
- absent or `true` -> today's behaviour exactly, no change
- `false` -> `maybe_auto_migrate_after_update()` returns before
building a plan; no output, no changes
`hermes gateway migrate --multiplex` is an explicit request and still migrates
regardless of the flag, so it stays the supported way to opt back in.
One early return, one schema entry with the reasoning inline, one invariant
test (opt-out blocks the hook, absent/true do not, explicit command still
applies), one section in the multi-profile gateways guide.
Both stale-pin detections exempt only '' and 'auto':
- desktop persistentStaleAux banner (model-settings.tsx)
- switch-time stale_aux response (hermes_cli/web_server.py)
'main' is a backend-supported alias (auxiliary_client._normalize_aux_provider)
meaning "follow the active main provider", so aux slots pinned to it can
never be stale. The false positive fires for users following Moonshot's
official Hermes integration guide, which prescribes
auxiliary.vision.provider: main.
Exempt the alias in both places and add a regression test.
`hermes sessions export --session-id X <dir>/` crashed with IsADirectoryError
because jsonl/html/trace opened the positional as a file while --help called it
an "output path" and md/qmd really do take a directory. An existing directory
(or one spelled with a trailing separator) now receives a default-named file
(`hermes_session_<id>.<fmt>`), and the help text spells out per-format what
OUTPUT means.
`/model` fed `validate_requested_model()` a key resolved through
`agent.secret_scope.get_secret`, which (multiplexing off) reads only
`os.environ`. Hermes does not export `$HERMES_HOME/.env` into the process
environment, so a `custom_providers` entry whose `key_env` lives only in `.env`
probed `/v1/models` unauthenticated, got 401 and printed a spurious "could not
reach this custom endpoint's model listing" note while chat worked fine.
Resolve through `get_env_prefer_dotenv` — the chain `client_lifecycle` uses for
the real request — when no profile scope is installed. With a scope installed or
multiplexing active the scope stays authoritative: a scoped miss still returns
"" and never borrows another profile's `.env`/process value.
Slimmed from the contributor's two commits (same mechanism, fewer branches,
tests trimmed to two invariants).
Fixes#109315