Commit Graph

7749 Commits

Author SHA1 Message Date
John Paul Soliva 9c36cafcec fix(gateway): in-container gateway start registers a missing s6 slot
`profile create` registers an s6 gateway slot only when the creating process is
itself inside the container: `detect_service_manager()` reads `/proc/1/comm` in
the CALLER's PID namespace, so on a host whose `~/.hermes` is bind-mounted into
the container the hook is a silent no-op. The profile directory lands exactly where
the container reads it, but no slot exists, and `hermes -p <name> gateway start`
inside the container fails with "not registered" until the operator restarts the
container so the boot reconciler notices. Same symptom as #54174 reaches `profile
install` by the same route.

Register the slot on demand instead. When `start` hits `GatewayNotRegisteredError`
and the profile directory carries `SOUL.md` — the boot reconciler's own "real
profile" marker — create the slot and start it. That makes
`_maybe_register_gateway_service`'s promise true without a restart.

Deliberately narrow:

- Only `start` self-heals. `stop` and `restart` on an unregistered profile keep the
  original error; registering a slot in order to stop it would be absurd.
- `SOUL.md` gates it, so a mistyped `-p` name or a stray directory cannot mint a
  phantom slot for a profile that does not exist.
- Registers with `start_now=False` and then goes through the ordinary `start` path,
  so the `desired_state` write that lets boot reconciliation restore want-up after
  a container restart keeps a single owner.
- `ValueError` (slot appeared underneath us) and `RuntimeError` (s6-svscanctl
  failed) surface as the existing actionable error, never a traceback.

Refs #54174. PR #54182 fixes the narrower `profile install` case by adding the same
registration call; this closes the symptom for any existing profile directory and
without a container restart.
2026-09-15 18:30:05 -07:00
teknium1 49b9bbb6fc fix(desktop): an exit without a window reveal no longer writes hermes.desktop
gnome-shell moves a launched ShellApp from STARTING to STOPPED when the
startup-notification sequence completes or times out (mutter, ~15 s), not
when the process exits. finish() healing right after an exit-without-reveal
(boot crash, early quit) therefore wrote the entry during STARTING — the
exact #111906 arming condition. Only the reveal byte from Electron heals
now; the wake byte finish() writes just unblocks the reader. A skipped heal
is picked up by the next terminal/updater launch or revealed grid launch.
2026-09-15 18:29:37 -07:00
teknium1 169c48fa13 fix(desktop): app-grid launches write hermes.desktop only after the window is on screen
`hermes desktop` used to create/refresh `~/.local/share/applications/hermes.desktop`
synchronously before spawning Electron. When the entry is ABSENT (first run after an
update, deleted by a cleaner, tombstoned by AV) that write lands while gnome-shell still
has the grid-launched ShellApp in STARTING; unpatched shells (before GNOME MR !4428)
drop the app's last strong reference on `installed_changed` and the next idle GC kills
the whole Wayland session, minutes to an hour later (#111906, residual after #111396).

Now a launch that carries `DESKTOP_STARTUP_ID` (app grid / menu) defers the write:
the launcher opens a pipe, hands Electron its write end as `HERMES_DESKTOP_READY_FD`
(pass_fds), and a worker thread installs the entry once Electron reports the main
window revealed (`onRevealed` in createWindow, via the new
`apps/desktop/electron/linux-launcher-ready.ts`), plus a 2 s settle so the compositor
has mapped the surface. If Electron exits without ever revealing a window, `finish()`
installs the entry after the exit (STOPPED app, no STARTING object) — self-heal
semantics survive. Terminal launches, the updater's detached relaunch and
`--build-only` have no DESKTOP_STARTUP_ID / spawn no app and keep writing immediately.

Why not simply write after Electron exits (#111915's approach): a daemon thread
started right before `sys.exit` is killed with the interpreter, so the heal is lost,
and a heal that waits for the user to quit the app leaves the menu entry missing for
the whole session.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:29:37 -07:00
teknium1 f93d33f93f fix(cli): name the dashboard kill grace and pin it against the lifespan teardown
Turn the SIGTERM→SIGKILL deadline in `_kill_pids_posix` into `_POSIX_TERM_GRACE_SECONDS`
(10.0s) documented against what it has to outlast: `web_server.py::_lifespan` blocks on
`stop_hosted_room_service(timeout=5.0)` + `join(1.0)` before `PTY_REGISTRY.close_all()`
(≤1.5s per attached Chat PTY). The 3.0s deadline predates the hosted-room stop (added
2026-08-30) and SIGKILLed the backend mid-teardown, so its ui-tui / tui_gateway.entry
children were never closed and kept the deleted state.db-wal inode open — the next
`hermes` start refused with FATAL DeletedWalGenerationError (#111912).

Two invariant tests against real child processes: a teardown as long as the lifespan
budget finishes gracefully (red on the old 3.0s deadline); a SIGTERM-ignoring process is
still SIGKILLed at the deadline. The orphan reaper's 1.5s grace is left alone on purpose
— it runs on the Desktop boot path under a 10s ready-probe — and says so in the comment.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-15 18:29:16 -07:00
AzurePii 83f1b88305 fix(cli): increase dashboard kill timeout to prevent PTY child leaks and WAL corruption 2026-09-15 18:29:16 -07:00
teknium1 c1115a7166 fix(config): treat empty-dict DEFAULT_CONFIG sections as open containers in the typo gate
compression.model_thresholds.<model>, terminal.docker_env.<VAR>, lsp.servers.<lang>.*,
auxiliary.<task>.extra_body.<k> and similar free-form mappings are declared as {} in
DEFAULT_CONFIG; the fail-closed gate walked into the empty dict, found the user-chosen key
missing and refused the write. An empty dict now accepts the rest of the path, like a scalar
leaf or a platforms container does. Populated sections keep the did-you-mean refusal.
2026-09-15 18:28:49 -07:00
teknium1 0e63a1bc5c fix(config): refuse an unknown path under a known section before writing
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.

Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.

`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.

Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:28:49 -07:00
teknium1 70e4938c07 fix(config): route every registered env setting through .env from config set/get/unset
`hermes config set FEISHU_HOME_CHANNEL oc_x` wrote the top level of config.yaml
while the platform setup flows and /sethome write the same name to .env via
save_env_value, so two writers fed two readers: the gateway bridges the yaml copy
into the environment only when .env lacks the name, one-shot CLI readers never
bridge, and the two copies diverged silently (#111848). Only credential-shaped
names were routed to .env because `_is_env_config_key` is the provider-credential
predicate.

Follow-up to KoNit-K's cherry-picked fix (#111850), which routed the
`setup_hidden_env` suffix family: the predicate now lives in the topical sibling
`hermes_cli/config_env_routing.py` and covers every bare name Hermes itself
registers as an environment variable (OPTIONAL_ENV_VARS, _EXTRA_ENV_KEYS — "env
var names written to .env" — plus the setup-hidden suffixes for plugin adapters
nobody enumerated), so `*_ALLOWED_USERS`, `WHATSAPP_MODE`, `MATRIX_PASSWORD` and
the rest of the adapter-saved family take the same file. `set` and `unset` also
drop a stale same-named top-level config.yaml copy so the reporter's drift cannot
come back, and `get` resolves .env first then that copy — the gateway's own read
order. Provider credentials keep the credential_lifecycle rotation path.

Docs: environment-variables.md tip, hermes_cli/AGENTS.md config rule.
2026-09-15 18:28:49 -07:00
KoNit-K 49bbc738d4 fix(config): route platform setup values through .env 2026-09-15 18:28:49 -07:00
teknium1 de93cb05fd fix(update): sudo-prefix the manual kickstart hint for a LaunchDaemon run as a user
`launchctl kickstart system/<label>` needs root. When a LaunchDaemon-supervised
backend does not come back and `hermes update` runs as a regular user, the
manual hint now reads `sudo launchctl kickstart -k system/<label>`, so the
operator gets a command that works from the shell they are in. LaunchAgent
(gui/user domain) hints are unchanged.
2026-09-15 18:28:26 -07:00
teknium1 469e5b6a2e fix(dashboard): --stop exits 0 when a launchd KeepAlive job respawns the backend
`hermes dashboard --stop` decided its exit code by re-scanning the process
table after the kill. On macOS a launchd KeepAlive job brings the backend
back on a fresh PID within the grace window, so the re-scan found it and
--stop exited 1 right after printing that the stop succeeded and that the job
restarts itself. Judge the exit code from the kill result instead: exit 1
only when a matched pid could not be stopped.
2026-09-15 18:28:26 -07:00
teknium1 29979b8c25 fix(update): attribute an earlier --no-open respawn to its launchd job
_launchd_job_owning_backend matched a process to a job only on exact argv
equality with the plist ProgramArguments, but _respawn_dashboard_processes
appends --no-open to dashboard commands. A detached copy created by an
earlier update therefore never matched, was treated as a manual backend and
was respawned detached again on every update. Ignore --no-open on both sides
of the comparison.
2026-09-15 18:28:26 -07:00
John Paul Soliva b4a6958129 fix(doctor): warn when database.journal_mode=delete never applied to a WAL database
`hermes doctor` printed the same informational "WAL journal mode" line for a database
that is still WAL although the operator configured `database.journal_mode: delete`. The
runtime never live-downgrades an existing WAL database (a downgrade under open
connections corrupts it) and says so only once per process in the gateway log, so the
one surface operators check told them they were protected when every process was still
writing WAL on the filesystem they configured `delete` for (#111729).

Doctor now compares the configured mode (`resolve_journal_mode`) with the on-disk header
and warns `<db> is in WAL mode despite database.journal_mode=delete`, naming the
never-live-downgraded rule and the one-time offline `PRAGMA journal_mode=DELETE`
remedy; a vulnerable SQLite still counts the database as WAL-reset exposed. This check
outranks the cross-VM hint (whose remedy, "set journal_mode: delete", is already
applied). A configured `wal` is unchanged.

Core hunk ported from #104714 (@jonpol01); its holder enumeration
(`foreign_state_db_holders(include_scan_gaps=True)` + per-PID report) is not included
here — that half is a 600-line change under separate review.
2026-09-15 18:28:26 -07:00
John Paul Soliva efc2b8e28d fix(update): verify the launchd job comes back on a fresh PID, follow wrapper plists, warn on --stop
Adopted from the sibling PRs in this cluster, with thanks:

- #99807 (@Chevron7Locked) and #101541 (@ivcdigital26): a kickstart that returns 0 is only
  "restart requested" — success now requires launchd to report a live PID other than the one
  that was stopped (the gateway's _wait_for_launchd_service_pid, 15 s), and a job that never
  comes back is reported with the manual `launchctl kickstart -k` command instead of a tick.
- #89793 (@mettamyron) and #101541: a plist may wrap the backend in `/bin/sh -c …` without
  exec, so launchd's live PID is the shell and the stale backend is its child — ownership now
  follows the parent chain (bounded, cycle-safe).
- #89793: `hermes dashboard --stop` on a launchd-owned backend now says a KeepAlive job
  restarts itself and gives the `launchctl bootout` command, instead of the misleading
  "restart it yourself" hint. The job scan runs for --stop too, on macOS only.
2026-09-15 18:28:26 -07:00
John Paul Soliva 99b727bd96 fix(update): a launchd-supervised backend restarts through launchd, never as a detached respawn
`hermes update` snapshots each stale dashboard/serve backend's systemd unit before the
kill and respawns everything else from its captured argv (#40449, #68934, #78821). On
macOS that "everything else" includes a backend supervised by a launchd job: the
respawn comes back detached and sits on the job's port, launchd's KeepAlive then fails
every restart of the job with "port already in use" (exit 75, one attempt per ~10 s,
forever), and the backend that IS serving is no longer supervised. Every later update
kills the copy and respawns it again, so the loop outlives the update that started it.

The update now snapshots the loaded launchd backend jobs before the kill (LaunchAgents
in the gui/user domains, LaunchDaemons in system; only plists whose ProgramArguments
are a dashboard/serve backend are probed), attributes a PID to a job when launchd
reports it as the job's live process or the process runs the job's exact
ProgramArguments (the detached copy an earlier respawn left behind), and brings such a
PID back with `launchctl kickstart <domain>/<label>` instead of an argv respawn. A
failed kickstart is reported with the manual command and counted as unrecovered, like
a failed systemctl restart. Linux and Windows paths are unchanged.
2026-09-15 18:28:26 -07:00
teknium1 a1b997beb9 fix(update): one receipt entry per repair outcome; reason no longer printed twice
Follow-up to the two salvaged commits (#111419 @Ckarey007, #111521 @wangtaotaotao95),
which both edit `_stage_candidate_venv` and were merged keep-both:

- `_record_runtime_repair` (new topical helper next to `_run_runtime_repair`): a
  `failed`/`safe`/`repaired` repair is one `sqlite_runtime_repair` step; a
  `skipped`/`not-applicable` repair is one skip WITH its reason instead of a red
  step plus a skip. Every pip/non-venv install runs the hook and would otherwise
  carry a failed-looking step in every receipt.
- The sync rejection's reason now travels into `RuntimeRepairResult.detail`, which
  `_report_runtime_repair_failure` prints and the receipt records, so the extra
  console print in `_stage_candidate_venv` is dropped (it duplicated the ℹ line).
- Tests trimmed to invariants: the child-reason test asserts the reason on the
  rejection the repair returns (the exception carries it now); the receipt test
  covers failed step / repaired step / skipped skip and the no-receipt no-op.
  Dropped: the 4-case stage-rejection walk (the same chain is covered by the
  existing `_repair_under_lock` tests whose rejection now carries the reason).
2026-09-15 18:28:03 -07:00
wangtao 2f3d0b5cc2 fix(update): record SQLite runtime repair outcomes in receipts 2026-09-15 18:28:03 -07:00
Ckarey 882c59ee93 fix(update): report why the candidate runtime sync failed
The SQLite runtime repair builds its replacement venv with
`uv sync --extra all --locked` and, when that fails, reported only
"replacement environment did not pass dependency and import smoke tests".

The child's diagnosis went to inherited stdout, so it survived in console
scrollback and nowhere else: the rejection line the logger records carried the
bare exit code, the failure detail the repair returns is generic, and update
receipts are built from explicit record_step calls — none of which covers this
repair. On an install where the repair cannot succeed, `hermes update`
therefore ended at "partially complete" with no reason — the reason being a
one-line uv message the user never saw.

Forward that sync's output live (unchanged contract: stderr merged into
stdout and drained while the child runs, so pre-desktop-update hand-offs
cannot deadlock uv on a full stderr pipe) and keep its tail: reject with the
child's `error:`/`hint:` lines and announce them, so the console and the log
say what to fix. Carrying the reason into RuntimeRepairResult.detail, and
therefore into a receipt step, is deliberately out of scope here — separate
change, separate PR.

Test: the reported reason keeps uv's error + hint and drops the progress
noise that precedes them (red before this change, where the reason was the
bare exit code).
2026-09-15 18:28:03 -07:00
teknium1 71aa0d635e fix: setup gateway skips the standalone service for a multiplex-served profile
`hermes -p <profile> setup gateway` (and `hermes setup` / `hermes import`) reach the
service step through `ensure_gateway_service`, which only knew "is THIS profile's
unit running" — a satellite served by the default multiplexer has no unit of its
own, so the step printed "Installing the gateway background service ..." and
registered a launchd plist / systemd unit that the #97120 start guard then refused,
leaving a stray dead service the user had to find and remove by hand (#111958).

Route both setup surfaces through one shared predicate: `_served_profile_needs_no_service`
wraps `named_profile_served_by_running_multiplexer` (the same probe `profile create`,
cron liveness and the run/start/install guards use), prints the "already served"
note and returns True so `ensure_gateway_service` and the `hermes gateway setup`
wizard (#111962's hunk) skip the install. Default profile and non-multiplex hosts
are unchanged.

Adds the invariant for the `ensure_gateway_service` path (served → no install,
unserved → still installs). Docs: multi-profile-gateways.md names the skipped step.

Co-authored-by: kvnloo <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:27:35 -07:00
KoNit-K edca3c2653 fix(gateway): skip served profile setup install 2026-09-15 18:27:35 -07:00
teknium1 0a3e792942 fix(worktree): judge no-remote repos against the local trunk instead of reaping everything
`hermes worktree prune` (and the startup/cron pruner) classified every clean tree in a
repository with no remote as "clean and fully merged/pushed" and force-deleted its branch,
even when the branch carried commits that exist nowhere else. `audit_branches` returned []
in the same repos, so a unique local-only branch was invisible to the audit as well.

Root cause: `_worktree_has_unpushed_commits` answered False when `refs/remotes` was empty
("nothing to be unpushed against") and both consumers — `worktree_gc._classify_tree` and
`worktree_ops._classify_prune_candidates` — read False as "safe to reap".

The preceding commit (#111897) flips that branch to True, which is safe but also means a
no-remote repo can never reclaim anything (a tree sitting at trunk, or squash-merged into
it, stays "unpushed" forever because `_worktree_commits_all_merged_upstream` finds no
origin/* base). This commit replaces the unconditional True with a real baseline:

- `_worktree_local_trunk`: `main`/`master`, else the branch checked out in the main
  worktree; None when no trunk exists.
- `_worktree_merge_base_ref`: origin/HEAD|origin/main|origin/master, falling back to the
  local trunk ONLY when the repo has no remote-tracking refs at all. Single resolver used
  by `_worktree_commits_all_merged_upstream` and `worktree_gc.audit_branches`.
- `_worktree_has_unpushed_commits`: with no remote refs, `git log HEAD --not <trunk>`;
  no trunk -> True (preserve), matching the docstring's fail-safe promise.

Net effect in a no-remote repo: unique work is kept ("unpushed commits not found
upstream"), trees at/merged into the local trunk still reap, branch audit reports unique
branches as keep and merged ones as delete, and `git branch -D` can only run on a branch
whose every commit is reachable from or patch-equivalent to the trunk.

Tests: the salvaged no-remote keep test now uses a shared `local_repo` fixture; a control
test proves trunk-merged trees still reclaim and the branch audit reports in the same
repo; `test_merged_predicate_fails_safe_without_upstream` now pins both halves of the
contract (trunk resolves -> merged; no trunk at all -> False/preserve).
2026-09-15 18:27:06 -07:00
KoNit-K 1039c15a03 fix(worktree): preserve commits without remote refs 2026-09-15 18:27:06 -07:00
teknium1 cf94a29319 fix(profiles): share the runtime-tree trio with backup via hermes_constants
backup.py imports hermes_cli.profiles only lazily and profiles.py never imports backup, so
there was no cycle to justify two literals. LOCAL_RUNTIME_ROOT_DIRS now feeds both
backup._EXCLUDED_ROOT_DIRS and the clone-all root gate; one invariant test pins the identity.
2026-09-15 18:26:38 -07:00
John Paul Soliva f11be8979c fix(profiles): --clone-all skips the local-models runtime trees (models/, runtimes/, node/)
43e67d872 (local models) put three machine-scoped trees under the default
HERMES_HOME — models/ (GGUF weights, tens of GB), runtimes/ (managed
llama.cpp binaries) and node/ (managed Node) — and taught backup.py to
exclude them (_EXCLUDED_ROOT_DIRS). `hermes profile create X --clone-all`
from the default profile did not follow: it copied all three into the new
profile, which never reads them (models_dir() resolves from the default
root only) and cannot use them (the binaries are re-downloaded on demand),
turning a clone into a tens-of-GB copy.

Add the three names to _CLONE_ALL_DEFAULT_EXCLUDE_ROOT. The set is gated
on the source being the default profile, so a named profile that really
carries a models/ directory of its own keeps it; and the exclusion applies
at the source root only, so a skill's nested models/ directory is copied
as user data. Kept as a separate literal from backup.py's set on purpose
(backup also matches profiles/<name>/ and importing it here would be a
circular import) with a comment tying the two together.
2026-09-15 18:26:38 -07:00
John Paul Soliva 7af3515c57 fix(config): validate the timezone key so a bad IANA name is reported, not silently ignored
hermes_time reads `timezone` from config.yaml (or HERMES_TIMEZONE) and, when
ZoneInfo cannot load the name, logs one warning and falls back to
server-local time. Nothing else ever looks at the value: not doctor, not
validate_config_structure. So a typo ("Asia/Tokio", "EST5", "Tokyo") puts
the agent clock and every cron schedule on the server's zone, with the
only evidence a single line in the gateway log.

validate_config_structure now checks the key: a non-string is an error, a
non-empty string that ZoneInfo cannot load is an error, and blank/missing
stays silent (server-local is the documented default). Because the check
runs at startup too, it is skipped when the interpreter has no tz database
at all (bare Windows without tzdata) — there is nothing to judge against,
and flagging every valid name there would be worse than the bug. The hint
names the IANA form and points out that HERMES_TIMEZONE overrides the key.
2026-09-15 18:26:15 -07:00
teknium1 536a07bec2 fix: sessions list prints a footer when --limit cuts the listing
`hermes sessions list` applied --limit inside the SQL query and rendered
whatever came back, so a user with 35 sessions saw 20 rows and a prompt
and had no way to tell the page was cut. The lister now asks for one row
past the cap, drops that probe row, and ends a truncated page with
"… more not shown (use --limit N to see more)". `--limit 0` is LIMIT 0
(no rows), so the copy never suggests it.

The footer goes through one shared helper, `cli_output.print_truncated`,
and the three "... N more" / "… N more" / "… and N more" variants already
in sessions_cmd.py (export dry-run preview, never-active cleanup, prune/
archive preview) now use it too, so every capped listing in the command
reads the same. Migrating checkpoints / curator / skills search / console
listers is follow-up work.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:25:46 -07:00
teknium1 d1bd778a5e fix(checkpoints): clear-legacy trims to two real-path tests, aligns failure line with the log
Replace the three monkeypatch-heavy tests from the salvaged commit (which
faked clear_legacy's return dict, so they could not catch the manager hunk
regressing) with two invariant tests that drive the real cmd_clear_legacy
against a temp checkpoint base: an undeletable legacy-* dir yields exit 2
plus the "Could not delete" line while the archive stays on disk; a clean
sweep keeps exit 0 and the unchanged success line. The green-path guard is
harvested from #111789.

Reword the CLI failure line to "Could not delete N archive(s) (see logs)."
so it matches the manager's WARNING wording and the text proposed in
#111776, and document the exit code in the CLI reference.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-15 18:24:22 -07:00
KoNit-K 065bc91846 fix(checkpoints): report legacy archive deletion failures 2026-09-15 18:24:22 -07:00
teknium1 9aaa71367b fix(catalog): apply the opencode free exclusion on the live-first keyed Zen/Go picker 2026-09-15 18:23:29 -07:00
teknium1 bb5d745ab1 fix(catalog): keep the delisted deepseek-v4-flash-free out of the live OpenCode picker too
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).

Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.
2026-09-15 18:23:29 -07:00
KoNit-K 4766cb2c0c fix(catalog): remove delisted opencode free model 2026-09-15 18:23:29 -07:00
teknium1 a9fabe43c4 fix(runtime): keep the bare-custom fail-fast off local aliases
Key the post-ladder AuthError on the literal `custom` request again, in
addition to the dead runtime shape (provider=custom, empty api_key). The
previous commit widened it to every alias that resolves to custom (ollama,
vllm), which broke `/model <direct-alias>` switching: _creds_for_switched_provider
resolves the alias provider tolerantly and _apply_direct_alias_endpoint
supplies the alias endpoint AFTER that call, so the early raise turned a
working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery, red in CI).

The issue's own scope (#111741) is the bare non-routable placeholder; a
credential-less local alias is a legitimate intermediate state there.
2026-09-15 18:19:50 -07:00
teknium1 e9a54c48f2 fix(runtime): bare-custom fail-fast keys on the dead runtime shape, covers custom aliases
Tighten the post-ladder guard from #111744 to the exact failing shape: a
resolved ``custom`` runtime with an EMPTY api_key. Every other custom rung
(named entry, direct alias, local bypass, pool, key_cmd) yields a key, a
callable or the ``no-key-required`` placeholder, so the loopback heuristic and
the has_usable_secret() re-check were dead branches — and keying on the
requested name alone missed aliases that resolve to custom (``ollama``,
``vllm`` with nothing configured), which still returned the credential-less
OpenRouter fallback and died at agent construction as "No LLM provider
configured". The AuthError now names the requested provider, so cron's
runner/preflight, the CLI, the gateway and the TUI gateway (all of which
format AuthError or walk their fallback chain on it) show the culprit.

Tests folded to two invariants: raise-and-name (bare + alias, plus the
OpenRouter-key control from #111750) and the loopback no-auth control.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-15 18:19:50 -07:00
KoNit-K 879d65ec78 fix(runtime): fail fast for bare custom credentials 2026-09-15 18:19:50 -07:00
teknium1 416a8177c2 perf(plugins): scan each plugin's source for removed imports once per process, not once per profile
A multiplex gateway runs plugin discovery for every served profile, and each
discovery ran plugin_compat.scan_plugin (an ast.parse + two ast.walk passes per
file) over every external plugin: on an 11-profile host that was ~0.4s per
profile of identical work on the gateway boot path, ahead of adapter connect.
The scan result is now cached process-wide on the plugin dir's (relpath,
mtime_ns, size) signature whenever the loaded manifest is used; a changed file
rescans, a caller-supplied manifest bypasses the cache.

Measured over the 11-profile MCP discovery loop on the reporting host:
3.78s -> 3.01s; per empty profile 0.25-0.30s -> 0.11s.
2026-09-15 15:11:33 -07:00
teknium1 f81c33cb6a fix(update): stop the fleet settle poll early when the restarted unit is dead
Salvage of #111385 (@JoaoMarcos44): the 30s -> 120s settle window is kept so a
slow host's gateway can publish its state stamp. This commit bounds the other
side of that trade: when every restarted systemd unit reports neither active
nor activating, the successor has died and nothing will ever publish, so the
poll fails closed at once instead of spending the full 120s. Unknown states
(no units, systemctl missing or slow) keep waiting. Also drops the test
assertion pinning the 2s poll cadence.
2026-09-15 15:10:41 -07:00
joaomarcos c99fee4e0a fix(update): wait for fleet state publication after supervised restart
A systemd unit can be active before the replacement gateway finishes
bootstrap and publishes gateway_state.json. The 30s poll in
_collect_fleet_snapshot() then returned [] with rows_expected=true,
causing _verify_fleet_after_update() to mark the restart incomplete and
exit(1) before clearing fleet_restart_pending. Every later CLI
startup/doctor therefore printed a false "did not restart" warning even
though the live gateway already served expected_sha.

Allow the default systemd startup budget plus publication slack by
extending the bounded settle window to 120s. Keeps fail-closed on
stale/down/empty after the deadline, only widens the window for slow
bootstraps (e.g. Raspberry Pi).

Fixes #111272
2026-09-15 15:10:41 -07:00
teknium1 f9c3a8a186 feat: hermes update and doctor tell you when /rollback checkpoints are on and large
Checkpoints were enabled by default from 9e845a6e (2026-03-16) until #20709
(2026-05-06) flipped the default back to off; the migration of that window
wrote `checkpoints.enabled: true` into user configs, where a later default
flip cannot reach it. Users who never type /rollback have carried a GB-scale
`~/.hermes/checkpoints/store` since (one live install: 1.2 GB across 250
projects, mostly disposable worktrees and /tmp dirs), and the cap cannot
bring it down because every project keeps at least one snapshot.

Silently flipping the key back is indistinguishable from overriding a real
opt-in, so this surfaces it instead: `checkpoint_footprint_notice()` returns
one line when checkpoints are enabled AND the store is at or above
`max_total_size_mb`, naming the opt-out (`hermes config set
checkpoints.enabled false` + `hermes checkpoints clear`) and the
retention knob. `hermes update` prints it with the post-update notices;
`hermes doctor` reports it as a warning after the state.db check.
2026-09-15 12:06:08 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1 804707bea6 fix: checkpoint store gc never runs inside a tool call or gateway startup
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.

The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.

- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
  drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
  pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
  so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
  calls it from the housekeeping tick (last chore), the CLI from a daemon
  thread. Nothing on either startup path waits for git.

Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
2026-09-15 10:57:16 -07:00
kshitijk4poor 5341f135a5 fix(update): the purge keeps hermes_constants; it is refreshed in place instead
hermes_constants owns the _HERMES_HOME_OVERRIDE ContextVar. Evicting it hands later
imports a fresh var: a reset token taken through the old module raises, and an
override set before the purge silently disappears. _reload_updated_runtime_modules
already re-executes it in place before every purge, so new symbols still arrive.
Caught by tests/hermes_cli/test_update_config_reload_tools_config.py on the stack.
2026-09-15 21:57:59 +05:30
kshitijk4poor d6ea7aa002 refactor(update): inline the tests exclusion; fix comments the wider purge made stale
`_STALE_PURGE_EXCLUDED_TOP_LEVEL` had one reader and a re-export nothing patched;
the config-check comment claimed root modules stay cached across the purge (no
longer true); the purge test module docstring still described package prefixes.
2026-09-15 21:57:59 +05:30
kshitijk4poor 9395f2f0d4 refactor(update): the purge scan has no fallback; trim to two invariants
The checkout root is where the update just pulled into, so "root unreadable" cannot
happen after a successful pull — drop the OSError fallback and the static five-name
tuple it fell back to (the tuple was the drift that caused the bug). Drop the phantom
`hermes_cli.hermes_logging` protection entry: no such module exists; the real
`hermes_logging` is root-level and now protected by name. Keep two tests: the stale
root `utils` scenario (red on base) and the hermes_logging protection.
2026-09-15 21:57:59 +05:30
Hubert Nimitanakit 29758a4eb2 fix(cli): purge every top-level checkout module, not just five packages
`_STALE_PURGE_PREFIXES` listed five package names, so every top-level module in
the checkout root survived the post-pull purge. `hermes update` then imported
new source against a cached pre-pull `utils`, `hermes_constants`, `plugins` or
`providers`.

Field failure (2026-09-12, macOS): updating 0.20.6 -> 0.21.2 crossed 3145986c20,
which added `base_url_origin` to `utils.py`. The restart phase's up-front
`from hermes_cli.gateway import ...` pulled `agent.auxiliary_client`, whose
`from utils import base_url_origin` hit the cached 0.20.6 `utils`:

    Update incomplete - gateway auto-restart failed: cannot import name
    'base_url_origin' from 'utils' (.../hermes-agent/utils.py)

The purge docstring already claims it evicts EVERY cached Hermes module; the
hardcoded tuple was the same "re-fixed per symptom" shape it replaced. Scan
`PROJECT_ROOT` instead: top-level `.py` files plus directories with an
`__init__.py`. 50 names here, and a newly added module can no longer drift out.

Two exclusions, both deliberate:

- `hermes_logging` joins `_STALE_PURGE_PROTECTED`. Its queue listener, handler
  list and `_logging_initialized` flag are module globals, so a fresh copy
  starts a second QueueListener over the same log files while the first runs.
- `tests` is never purged. pytest resolves fixtures through the identity of its
  already-imported test modules.

Falls back to the old tuple when the root is unreadable.
2026-09-15 21:57:59 +05:30
mr-r0b0t 3d2842d84f fix(cli): import file_signature in TUI run-state init
#111408 widened the MCP config watcher seed from mtime to
utils.file_signature but omitted the import in cli_tui_mixin.
The NameError only fires when config.yaml exists, so isolated-home
tests short-circuited past it and every real CLI launch crashed.

Signed-off-by: mr-r0b0t <adam.manning@gmail.com>
2026-09-15 20:47:06 +05:30
kshitijk4poor 95dba8d9a5 fix(free-tier): classify_mint_exception keeps an uncoded AuthError's wait hint
The pure rewrite re-raises an uncoded `AuthError` as a server-error twin but
dropped its `retry_after`, so the mint memo and a sign-in `Failed` fell back
to the ladder / default wait instead of the wait the raiser named.
2026-09-15 20:44:42 +05:30
kshitijk4poor 658f319147 fix(free-tier): setup.ready carries the failure block flat, the shape setup.status already spreads
`SetupRecord.as_payload()` serialised the record verbatim, so the broadcast
nested `failure: {...}` while `setup.status` spread the same four keys flat.
A client keyed on `error_code` saw it on one surface and not the other. Flatten
it in `as_payload`, declare the three optional keys on `SetupReadyPayload`, and
regenerate the TS/OpenRPC contract.
2026-09-15 20:44:42 +05:30
Robin Fernandes 1034215ae8 docs(free-tier): drop the rehearsal page and its references
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes a241f42fc8 copy(free-tier): "it's free" without "keeps the free model" — signed-in free models are not the same model
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30