get_mcp_status reports status='disabled' (never 'lazy') for a disabled
server and derives the 'disabled' flag from that same status, so the
extra 'and not srv.get("disabled")' check could never change the branch.
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:
- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
lazily registered server. It now reports `lazy` with the cached tool
count (`connected: False`); an in-flight or failed first-use connect
still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
`_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
failed)` right after registering every cached tool, and re-announced
the same "failure" on every repeat discovery. Lazy servers are now
reported as `(N lazy, not spawned yet)` and an already-lazy server is
not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
two sites, so every startup logged `Background MCP discovery completed
with zero connected servers` and every later call re-spawned the
discovery thread as a retry. One predicate,
`_discovery_registered_servers`, treats a lazy registration as a
usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
red "could not connect" line; it now shows the cached tool count with
`(lazy, starts on first use)`.
Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).
Fixes#111717
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.
Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.
Refs #111707
Follow-up to the cherry-picked #111584 (@chelsealong):
- website/docs/user-guide/docker.md: new warning block next to the existing
"do not override the entrypoint" note explaining WHY (with `/init` gone the
hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
children), the Compose `init: true` / `docker run --init` remedy, and that
supervision is still lost on that path; plus a Troubleshooting entry for
`<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
check and drops the `platform.system()` gate and the blanket
`try/except Exception: pass` — a user process is never PID 1 on any host OS
(PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).
Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
A deployment that overrides the image's `entrypoint:` to invoke hermes
directly skips docker/entrypoint-dispatch.sh entirely, so hermes itself
becomes PID 1 with no s6-overlay /init (or any other init) above it.
Nothing then reaps orphaned grandchildren (browser tooling, MCP
subprocesses, shell-tool children) reparented to PID 1, and they
accumulate as zombies without bound.
entrypoint-dispatch.sh already warns on its own non-PID-1 fallback
path, but that script never runs in the entrypoint-override case, so
there was no signal at all. Add the same style of warning inside
hermes_cli.main, gated on being PID 1 on Linux, pointing users at the
image's default ENTRYPOINT or `docker run --init` / `init: true`.
Fixes#111577
Review finding: hermes_cli/kanban_db_dispatch.py::_resolve_hermes_argv still resolved which('hermes') before sys.executable -m hermes_cli.main while claiming to mirror gateway.run._resolve_hermes_bin, which this PR made module-first (#111569). Keep the explicit HERMES_BIN override first, then the module argv whenever hermes_cli is importable, PATH only as fallback; docstring updated.
Two diagnostics diverged after a session-only `/model` switch (#111436).
/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).
Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).
Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.
Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.
Fixes#111695
With the shared single-flight probe (previous commit) a `refresh=true`
request that landed while a probe was already running simply joined it. That
probe may have started before `gh auth login` completed, so the refresh
returned "not authenticated" and cached it for the full 5-minute TTL — the
composer pill kept offering /github-auth right after a successful login.
A refresh now accepts only a probe that started at or after the refresh was
requested: it awaits the in-flight one, then starts (or joins) the next.
Still only one `gh` runs at a time. A finished task whose done-callback has
not run yet is treated as absent so the loop cannot spin on it.
Live probe (real route, fake `gh` reading login state at start, 1 s answer):
PR head: refresh=true right after login -> authenticated False, cached False
fixed: refresh=true right after login -> authenticated True, cached True;
5 concurrent requests -> peak concurrent probes 1
Co-authored-by: aron-intframe <aron-intframe@users.noreply.github.com>
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.
Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.
Part of #111804
Two false positives in `hermes plugins validate` that block catalog admission
for plugins that are correct at runtime:
- `kind: model-provider` plugins register at import via
providers.register_provider(ProviderProfile); the PluginManager never calls a
register(ctx) on them (plugins_discovery skips the kind). The probe demanded
register() anyway, so every provider plugin -- including the in-tree
plugins/model-providers/* -- failed with "no register() function". The probe
now records register_provider calls for that kind and fails only when the
import registers nothing.
- RecordingContext returned a no-op callable for ANY attribute, so
`getattr(ctx, "profile_path", None)` was truthy under validation alone and
register() crashed with an error the real PluginContext never produces. The
parent now passes the real PluginContext method names into the probe; other
names raise AttributeError exactly like the real object.
Surfaced by the 2026-09-15 catalog sweep (Gondola provider, hermes-persona).
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.
Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).
Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.
Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.
Fixes#111791
Only Task.from_row coerced BLOB-typed cells; a BLOB task_comments.body (or
event payload / run summary) still came back as bytes and crashed
`hermes kanban show <id> --json` with "Object of type bytes is not JSON
serializable". Apply _lossy_text in the other from_row constructors.
Test: BLOB comment body and event payload -> str with U+FFFD and
JSON-serialisable (red before).
A tasks row whose TEXT body holds invalid UTF-8 made sqlite3 raise
"Could not decode to UTF-8 column 'body'" inside fetchall, so `hermes kanban
list` (and `show`, and every other reader of that row) failed board-wide
until the row was deleted by hand; a BLOB-typed body came back as bytes and
crashed `--json` (#111743).
Fix it once at the connection: every board connection (`_open_configured`
and the read-only descendant path in `connect`) installs a lossy
text_factory that substitutes U+FFFD, and `Task.from_row` runs BLOB cells
through the same helper so a corrupt row renders with replacement
characters instead of taking its neighbours down.
Fixes#111743
The PR docstring said complete_task applies "the same fence request_review
applies", but the two disagreed: complete_task keyed on a live worker
process while request_review still refused any running task with a
claim_lock, so a claim whose worker is gone (or a CLI/library claim that
never spawned one) could be completed but not sent to review without
force. Factor the liveness test into _claim_is_live and use it in both.
TTL expiry is deliberately not part of "live": reclaim_stale_tasks extends
(not reclaims) the claim of a live worker, so the process stays the
liveness authority.
Test: claim -> request_review without a worker PID now succeeds (red
before); a live worker's claim is still refused without expected_run_id.
The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).
Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.
Fixes#111764
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.
- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
task plus rate_limited / skipped_locked / memory_pressure for one or more
DispatchResults, so the CLI daemon and gateway warnings share one wording:
`Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.
Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>
Part of #111910
`hermes kanban dispatch --json` only emitted `spawned` and the skip buckets it
already knew about, so a ready card held by the respawn guard (`active_pr`,
`recent_success`, ...), a quota-released worker, a lost dispatch lock or a
memory-pressure hold all looked like `spawned: []` with no reason. Emit
`respawn_guarded`, `rate_limited`, `skipped_locked` and `memory_pressure`
from the DispatchResult the tick already returns.
Salvaged from #111917 by @KoNit-K. Dropped hunk: the `_ACTIVE_PR_RECOVERY_LANES
= frozenset({"review"})` rename in kanban_db_dispatch.py, which is behaviour-
identical to the existing `lane == "review"` check and does not implement the
role-aware exemption the issue asks for.
Part of #111910
`hermes kanban specify|decompose`, the dashboard specify/decompose routes and
the gateway auto-decomposer all reach the LLM through
hermes_cli/kanban_specify.py::_call_aux outside any agent turn. No
conversation affinity scope is bound there, so agent/opencode_affinity.py
resolved an empty key and sent no `x-opencode-session`; the OpenCode Go relay
rejects such requests with 400 MissingSessionID and the user sees
"Specify failed: LLM error: BadRequestError".
Declare a per-task affinity scope (`kanban:<task_id>`) around the call —
the same host-declared scope the main turn, compression and the
OpenRouter/Portal sticky keys already resolve first — but only when no scope
is bound, so an in-turn caller keeps its conversation's key. Reset in a
finally so nothing leaks past the call.
Live: real httpx transport capture against auxiliary.triage_specifier
provider=opencode-go — before: no x-opencode-session header; after:
`kanban:t_45567533` on specify and decompose, stable per task, distinct per
task; a pre-declared scope is preserved; an openai route gets no header.
Fixes#112043
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.
- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
be untrue — the process provably runs old code and the Desktop app does not respawn it after
a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
`unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
origin/main.
Fixes#111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Reshape of the cherry-picked fix from #104194:
- The SOUL.md gate + `register_profile_gateway(start_now=False)` now live in
`hermes_cli/service_manager.py::register_unregistered_profile_gateway`, next to the s6
manager and `_profile_dir_for_gateway_service` it needs, instead of a private reach-in
from the 6.5k-line `hermes_cli/gateway.py` facade. The facade only decides "start
repairs, stop/restart re-raise" and keeps ONE error handler (S6Error is a RuntimeError;
register's ValueError/RuntimeError/OSError surface as the same `✗` + exit 1).
- Tests trimmed from four to two invariants: start on a real profile registers `down`
and then starts; stop on an unregistered profile / start on a directory without
SOUL.md keep the original error and mint nothing (parametrized). Dropped: the
registration-failure traceback test (covered by the single except clause) and the
duplicate no-marker/stop split. The test now resolves the profile dir through the real
HERMES_HOME mapping instead of monkeypatching `_profile_dir_for_gateway_service`.
- Docs: docker.md multi-profile section says `gateway start` inside the container
registers a slot for a profile created from the host.
`profile create` registers an s6 gateway slot only when the creating process is
itself inside the container: `detect_service_manager()` reads `/proc/1/comm` in
the CALLER's PID namespace, so on a host whose `~/.hermes` is bind-mounted into
the container the hook is a silent no-op. The profile directory lands exactly where
the container reads it, but no slot exists, and `hermes -p <name> gateway start`
inside the container fails with "not registered" until the operator restarts the
container so the boot reconciler notices. Same symptom as #54174 reaches `profile
install` by the same route.
Register the slot on demand instead. When `start` hits `GatewayNotRegisteredError`
and the profile directory carries `SOUL.md` — the boot reconciler's own "real
profile" marker — create the slot and start it. That makes
`_maybe_register_gateway_service`'s promise true without a restart.
Deliberately narrow:
- Only `start` self-heals. `stop` and `restart` on an unregistered profile keep the
original error; registering a slot in order to stop it would be absurd.
- `SOUL.md` gates it, so a mistyped `-p` name or a stray directory cannot mint a
phantom slot for a profile that does not exist.
- Registers with `start_now=False` and then goes through the ordinary `start` path,
so the `desired_state` write that lets boot reconciliation restore want-up after
a container restart keeps a single owner.
- `ValueError` (slot appeared underneath us) and `RuntimeError` (s6-svscanctl
failed) surface as the existing actionable error, never a traceback.
Refs #54174. PR #54182 fixes the narrower `profile install` case by adding the same
registration call; this closes the symptom for any existing profile directory and
without a container restart.
gnome-shell moves a launched ShellApp from STARTING to STOPPED when the
startup-notification sequence completes or times out (mutter, ~15 s), not
when the process exits. finish() healing right after an exit-without-reveal
(boot crash, early quit) therefore wrote the entry during STARTING — the
exact #111906 arming condition. Only the reveal byte from Electron heals
now; the wake byte finish() writes just unblocks the reader. A skipped heal
is picked up by the next terminal/updater launch or revealed grid launch.
`hermes desktop` used to create/refresh `~/.local/share/applications/hermes.desktop`
synchronously before spawning Electron. When the entry is ABSENT (first run after an
update, deleted by a cleaner, tombstoned by AV) that write lands while gnome-shell still
has the grid-launched ShellApp in STARTING; unpatched shells (before GNOME MR !4428)
drop the app's last strong reference on `installed_changed` and the next idle GC kills
the whole Wayland session, minutes to an hour later (#111906, residual after #111396).
Now a launch that carries `DESKTOP_STARTUP_ID` (app grid / menu) defers the write:
the launcher opens a pipe, hands Electron its write end as `HERMES_DESKTOP_READY_FD`
(pass_fds), and a worker thread installs the entry once Electron reports the main
window revealed (`onRevealed` in createWindow, via the new
`apps/desktop/electron/linux-launcher-ready.ts`), plus a 2 s settle so the compositor
has mapped the surface. If Electron exits without ever revealing a window, `finish()`
installs the entry after the exit (STOPPED app, no STARTING object) — self-heal
semantics survive. Terminal launches, the updater's detached relaunch and
`--build-only` have no DESKTOP_STARTUP_ID / spawn no app and keep writing immediately.
Why not simply write after Electron exits (#111915's approach): a daemon thread
started right before `sys.exit` is killed with the interpreter, so the heal is lost,
and a heal that waits for the user to quit the app leaves the menu entry missing for
the whole session.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Turn the SIGTERM→SIGKILL deadline in `_kill_pids_posix` into `_POSIX_TERM_GRACE_SECONDS`
(10.0s) documented against what it has to outlast: `web_server.py::_lifespan` blocks on
`stop_hosted_room_service(timeout=5.0)` + `join(1.0)` before `PTY_REGISTRY.close_all()`
(≤1.5s per attached Chat PTY). The 3.0s deadline predates the hosted-room stop (added
2026-08-30) and SIGKILLed the backend mid-teardown, so its ui-tui / tui_gateway.entry
children were never closed and kept the deleted state.db-wal inode open — the next
`hermes` start refused with FATAL DeletedWalGenerationError (#111912).
Two invariant tests against real child processes: a teardown as long as the lifespan
budget finishes gracefully (red on the old 3.0s deadline); a SIGTERM-ignoring process is
still SIGKILLed at the deadline. The orphan reaper's 1.5s grace is left alone on purpose
— it runs on the Desktop boot path under a 10s ready-probe — and says so in the comment.
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
compression.model_thresholds.<model>, terminal.docker_env.<VAR>, lsp.servers.<lang>.*,
auxiliary.<task>.extra_body.<k> and similar free-form mappings are declared as {} in
DEFAULT_CONFIG; the fail-closed gate walked into the empty dict, found the user-chosen key
missing and refused the write. An empty dict now accepts the rest of the path, like a scalar
leaf or a platforms container does. Populated sections keep the did-you-mean refusal.
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.
Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.
`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.
Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
`hermes config set FEISHU_HOME_CHANNEL oc_x` wrote the top level of config.yaml
while the platform setup flows and /sethome write the same name to .env via
save_env_value, so two writers fed two readers: the gateway bridges the yaml copy
into the environment only when .env lacks the name, one-shot CLI readers never
bridge, and the two copies diverged silently (#111848). Only credential-shaped
names were routed to .env because `_is_env_config_key` is the provider-credential
predicate.
Follow-up to KoNit-K's cherry-picked fix (#111850), which routed the
`setup_hidden_env` suffix family: the predicate now lives in the topical sibling
`hermes_cli/config_env_routing.py` and covers every bare name Hermes itself
registers as an environment variable (OPTIONAL_ENV_VARS, _EXTRA_ENV_KEYS — "env
var names written to .env" — plus the setup-hidden suffixes for plugin adapters
nobody enumerated), so `*_ALLOWED_USERS`, `WHATSAPP_MODE`, `MATRIX_PASSWORD` and
the rest of the adapter-saved family take the same file. `set` and `unset` also
drop a stale same-named top-level config.yaml copy so the reporter's drift cannot
come back, and `get` resolves .env first then that copy — the gateway's own read
order. Provider credentials keep the credential_lifecycle rotation path.
Docs: environment-variables.md tip, hermes_cli/AGENTS.md config rule.
`launchctl kickstart system/<label>` needs root. When a LaunchDaemon-supervised
backend does not come back and `hermes update` runs as a regular user, the
manual hint now reads `sudo launchctl kickstart -k system/<label>`, so the
operator gets a command that works from the shell they are in. LaunchAgent
(gui/user domain) hints are unchanged.
`hermes dashboard --stop` decided its exit code by re-scanning the process
table after the kill. On macOS a launchd KeepAlive job brings the backend
back on a fresh PID within the grace window, so the re-scan found it and
--stop exited 1 right after printing that the stop succeeded and that the job
restarts itself. Judge the exit code from the kill result instead: exit 1
only when a matched pid could not be stopped.
_launchd_job_owning_backend matched a process to a job only on exact argv
equality with the plist ProgramArguments, but _respawn_dashboard_processes
appends --no-open to dashboard commands. A detached copy created by an
earlier update therefore never matched, was treated as a manual backend and
was respawned detached again on every update. Ignore --no-open on both sides
of the comparison.
`hermes doctor` printed the same informational "WAL journal mode" line for a database
that is still WAL although the operator configured `database.journal_mode: delete`. The
runtime never live-downgrades an existing WAL database (a downgrade under open
connections corrupts it) and says so only once per process in the gateway log, so the
one surface operators check told them they were protected when every process was still
writing WAL on the filesystem they configured `delete` for (#111729).
Doctor now compares the configured mode (`resolve_journal_mode`) with the on-disk header
and warns `<db> is in WAL mode despite database.journal_mode=delete`, naming the
never-live-downgraded rule and the one-time offline `PRAGMA journal_mode=DELETE`
remedy; a vulnerable SQLite still counts the database as WAL-reset exposed. This check
outranks the cross-VM hint (whose remedy, "set journal_mode: delete", is already
applied). A configured `wal` is unchanged.
Core hunk ported from #104714 (@jonpol01); its holder enumeration
(`foreign_state_db_holders(include_scan_gaps=True)` + per-PID report) is not included
here — that half is a 600-line change under separate review.
Adopted from the sibling PRs in this cluster, with thanks:
- #99807 (@Chevron7Locked) and #101541 (@ivcdigital26): a kickstart that returns 0 is only
"restart requested" — success now requires launchd to report a live PID other than the one
that was stopped (the gateway's _wait_for_launchd_service_pid, 15 s), and a job that never
comes back is reported with the manual `launchctl kickstart -k` command instead of a tick.
- #89793 (@mettamyron) and #101541: a plist may wrap the backend in `/bin/sh -c …` without
exec, so launchd's live PID is the shell and the stale backend is its child — ownership now
follows the parent chain (bounded, cycle-safe).
- #89793: `hermes dashboard --stop` on a launchd-owned backend now says a KeepAlive job
restarts itself and gives the `launchctl bootout` command, instead of the misleading
"restart it yourself" hint. The job scan runs for --stop too, on macOS only.
`hermes update` snapshots each stale dashboard/serve backend's systemd unit before the
kill and respawns everything else from its captured argv (#40449, #68934, #78821). On
macOS that "everything else" includes a backend supervised by a launchd job: the
respawn comes back detached and sits on the job's port, launchd's KeepAlive then fails
every restart of the job with "port already in use" (exit 75, one attempt per ~10 s,
forever), and the backend that IS serving is no longer supervised. Every later update
kills the copy and respawns it again, so the loop outlives the update that started it.
The update now snapshots the loaded launchd backend jobs before the kill (LaunchAgents
in the gui/user domains, LaunchDaemons in system; only plists whose ProgramArguments
are a dashboard/serve backend are probed), attributes a PID to a job when launchd
reports it as the job's live process or the process runs the job's exact
ProgramArguments (the detached copy an earlier respawn left behind), and brings such a
PID back with `launchctl kickstart <domain>/<label>` instead of an argv respawn. A
failed kickstart is reported with the manual command and counted as unrecovered, like
a failed systemctl restart. Linux and Windows paths are unchanged.
Follow-up to the two salvaged commits (#111419 @Ckarey007, #111521 @wangtaotaotao95),
which both edit `_stage_candidate_venv` and were merged keep-both:
- `_record_runtime_repair` (new topical helper next to `_run_runtime_repair`): a
`failed`/`safe`/`repaired` repair is one `sqlite_runtime_repair` step; a
`skipped`/`not-applicable` repair is one skip WITH its reason instead of a red
step plus a skip. Every pip/non-venv install runs the hook and would otherwise
carry a failed-looking step in every receipt.
- The sync rejection's reason now travels into `RuntimeRepairResult.detail`, which
`_report_runtime_repair_failure` prints and the receipt records, so the extra
console print in `_stage_candidate_venv` is dropped (it duplicated the ℹ line).
- Tests trimmed to invariants: the child-reason test asserts the reason on the
rejection the repair returns (the exception carries it now); the receipt test
covers failed step / repaired step / skipped skip and the no-receipt no-op.
Dropped: the 4-case stage-rejection walk (the same chain is covered by the
existing `_repair_under_lock` tests whose rejection now carries the reason).
The SQLite runtime repair builds its replacement venv with
`uv sync --extra all --locked` and, when that fails, reported only
"replacement environment did not pass dependency and import smoke tests".
The child's diagnosis went to inherited stdout, so it survived in console
scrollback and nowhere else: the rejection line the logger records carried the
bare exit code, the failure detail the repair returns is generic, and update
receipts are built from explicit record_step calls — none of which covers this
repair. On an install where the repair cannot succeed, `hermes update`
therefore ended at "partially complete" with no reason — the reason being a
one-line uv message the user never saw.
Forward that sync's output live (unchanged contract: stderr merged into
stdout and drained while the child runs, so pre-desktop-update hand-offs
cannot deadlock uv on a full stderr pipe) and keep its tail: reject with the
child's `error:`/`hint:` lines and announce them, so the console and the log
say what to fix. Carrying the reason into RuntimeRepairResult.detail, and
therefore into a receipt step, is deliberately out of scope here — separate
change, separate PR.
Test: the reported reason keeps uv's error + hint and drops the progress
noise that precedes them (red before this change, where the reason was the
bare exit code).
`hermes -p <profile> setup gateway` (and `hermes setup` / `hermes import`) reach the
service step through `ensure_gateway_service`, which only knew "is THIS profile's
unit running" — a satellite served by the default multiplexer has no unit of its
own, so the step printed "Installing the gateway background service ..." and
registered a launchd plist / systemd unit that the #97120 start guard then refused,
leaving a stray dead service the user had to find and remove by hand (#111958).
Route both setup surfaces through one shared predicate: `_served_profile_needs_no_service`
wraps `named_profile_served_by_running_multiplexer` (the same probe `profile create`,
cron liveness and the run/start/install guards use), prints the "already served"
note and returns True so `ensure_gateway_service` and the `hermes gateway setup`
wizard (#111962's hunk) skip the install. Default profile and non-multiplex hosts
are unchanged.
Adds the invariant for the `ensure_gateway_service` path (served → no install,
unserved → still installs). Docs: multi-profile-gateways.md names the skipped step.
Co-authored-by: kvnloo <7121943+kvnloo@users.noreply.github.com>