With the shared single-flight probe (previous commit) a `refresh=true`
request that landed while a probe was already running simply joined it. That
probe may have started before `gh auth login` completed, so the refresh
returned "not authenticated" and cached it for the full 5-minute TTL — the
composer pill kept offering /github-auth right after a successful login.
A refresh now accepts only a probe that started at or after the refresh was
requested: it awaits the in-flight one, then starts (or joins) the next.
Still only one `gh` runs at a time. A finished task whose done-callback has
not run yet is treated as absent so the loop cannot spin on it.
Live probe (real route, fake `gh` reading login state at start, 1 s answer):
PR head: refresh=true right after login -> authenticated False, cached False
fixed: refresh=true right after login -> authenticated True, cached True;
5 concurrent requests -> peak concurrent probes 1
Co-authored-by: aron-intframe <aron-intframe@users.noreply.github.com>
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.
Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.
Part of #111804
Two false positives in `hermes plugins validate` that block catalog admission
for plugins that are correct at runtime:
- `kind: model-provider` plugins register at import via
providers.register_provider(ProviderProfile); the PluginManager never calls a
register(ctx) on them (plugins_discovery skips the kind). The probe demanded
register() anyway, so every provider plugin -- including the in-tree
plugins/model-providers/* -- failed with "no register() function". The probe
now records register_provider calls for that kind and fails only when the
import registers nothing.
- RecordingContext returned a no-op callable for ANY attribute, so
`getattr(ctx, "profile_path", None)` was truthy under validation alone and
register() crashed with an error the real PluginContext never produces. The
parent now passes the real PluginContext method names into the probe; other
names raise AttributeError exactly like the real object.
Surfaced by the 2026-09-15 catalog sweep (Gondola provider, hermes-persona).
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.
Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).
Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.
Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.
Fixes#111791
Only Task.from_row coerced BLOB-typed cells; a BLOB task_comments.body (or
event payload / run summary) still came back as bytes and crashed
`hermes kanban show <id> --json` with "Object of type bytes is not JSON
serializable". Apply _lossy_text in the other from_row constructors.
Test: BLOB comment body and event payload -> str with U+FFFD and
JSON-serialisable (red before).
A tasks row whose TEXT body holds invalid UTF-8 made sqlite3 raise
"Could not decode to UTF-8 column 'body'" inside fetchall, so `hermes kanban
list` (and `show`, and every other reader of that row) failed board-wide
until the row was deleted by hand; a BLOB-typed body came back as bytes and
crashed `--json` (#111743).
Fix it once at the connection: every board connection (`_open_configured`
and the read-only descendant path in `connect`) installs a lossy
text_factory that substitutes U+FFFD, and `Task.from_row` runs BLOB cells
through the same helper so a corrupt row renders with replacement
characters instead of taking its neighbours down.
Fixes#111743
The PR docstring said complete_task applies "the same fence request_review
applies", but the two disagreed: complete_task keyed on a live worker
process while request_review still refused any running task with a
claim_lock, so a claim whose worker is gone (or a CLI/library claim that
never spawned one) could be completed but not sent to review without
force. Factor the liveness test into _claim_is_live and use it in both.
TTL expiry is deliberately not part of "live": reclaim_stale_tasks extends
(not reclaims) the claim of a live worker, so the process stays the
liveness authority.
Test: claim -> request_review without a worker PID now succeeds (red
before); a live worker's claim is still refused without expected_run_id.
The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).
Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.
Fixes#111764
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.
- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
task plus rate_limited / skipped_locked / memory_pressure for one or more
DispatchResults, so the CLI daemon and gateway warnings share one wording:
`Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.
Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>
Part of #111910
`hermes kanban dispatch --json` only emitted `spawned` and the skip buckets it
already knew about, so a ready card held by the respawn guard (`active_pr`,
`recent_success`, ...), a quota-released worker, a lost dispatch lock or a
memory-pressure hold all looked like `spawned: []` with no reason. Emit
`respawn_guarded`, `rate_limited`, `skipped_locked` and `memory_pressure`
from the DispatchResult the tick already returns.
Salvaged from #111917 by @KoNit-K. Dropped hunk: the `_ACTIVE_PR_RECOVERY_LANES
= frozenset({"review"})` rename in kanban_db_dispatch.py, which is behaviour-
identical to the existing `lane == "review"` check and does not implement the
role-aware exemption the issue asks for.
Part of #111910
`hermes kanban specify|decompose`, the dashboard specify/decompose routes and
the gateway auto-decomposer all reach the LLM through
hermes_cli/kanban_specify.py::_call_aux outside any agent turn. No
conversation affinity scope is bound there, so agent/opencode_affinity.py
resolved an empty key and sent no `x-opencode-session`; the OpenCode Go relay
rejects such requests with 400 MissingSessionID and the user sees
"Specify failed: LLM error: BadRequestError".
Declare a per-task affinity scope (`kanban:<task_id>`) around the call —
the same host-declared scope the main turn, compression and the
OpenRouter/Portal sticky keys already resolve first — but only when no scope
is bound, so an in-turn caller keeps its conversation's key. Reset in a
finally so nothing leaks past the call.
Live: real httpx transport capture against auxiliary.triage_specifier
provider=opencode-go — before: no x-opencode-session header; after:
`kanban:t_45567533` on specify and decompose, stable per task, distinct per
task; a pre-declared scope is preserved; an openai route gets no header.
Fixes#112043
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.
- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
be untrue — the process provably runs old code and the Desktop app does not respawn it after
a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
`unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
origin/main.
Fixes#111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Reshape of the cherry-picked fix from #104194:
- The SOUL.md gate + `register_profile_gateway(start_now=False)` now live in
`hermes_cli/service_manager.py::register_unregistered_profile_gateway`, next to the s6
manager and `_profile_dir_for_gateway_service` it needs, instead of a private reach-in
from the 6.5k-line `hermes_cli/gateway.py` facade. The facade only decides "start
repairs, stop/restart re-raise" and keeps ONE error handler (S6Error is a RuntimeError;
register's ValueError/RuntimeError/OSError surface as the same `✗` + exit 1).
- Tests trimmed from four to two invariants: start on a real profile registers `down`
and then starts; stop on an unregistered profile / start on a directory without
SOUL.md keep the original error and mint nothing (parametrized). Dropped: the
registration-failure traceback test (covered by the single except clause) and the
duplicate no-marker/stop split. The test now resolves the profile dir through the real
HERMES_HOME mapping instead of monkeypatching `_profile_dir_for_gateway_service`.
- Docs: docker.md multi-profile section says `gateway start` inside the container
registers a slot for a profile created from the host.
`profile create` registers an s6 gateway slot only when the creating process is
itself inside the container: `detect_service_manager()` reads `/proc/1/comm` in
the CALLER's PID namespace, so on a host whose `~/.hermes` is bind-mounted into
the container the hook is a silent no-op. The profile directory lands exactly where
the container reads it, but no slot exists, and `hermes -p <name> gateway start`
inside the container fails with "not registered" until the operator restarts the
container so the boot reconciler notices. Same symptom as #54174 reaches `profile
install` by the same route.
Register the slot on demand instead. When `start` hits `GatewayNotRegisteredError`
and the profile directory carries `SOUL.md` — the boot reconciler's own "real
profile" marker — create the slot and start it. That makes
`_maybe_register_gateway_service`'s promise true without a restart.
Deliberately narrow:
- Only `start` self-heals. `stop` and `restart` on an unregistered profile keep the
original error; registering a slot in order to stop it would be absurd.
- `SOUL.md` gates it, so a mistyped `-p` name or a stray directory cannot mint a
phantom slot for a profile that does not exist.
- Registers with `start_now=False` and then goes through the ordinary `start` path,
so the `desired_state` write that lets boot reconciliation restore want-up after
a container restart keeps a single owner.
- `ValueError` (slot appeared underneath us) and `RuntimeError` (s6-svscanctl
failed) surface as the existing actionable error, never a traceback.
Refs #54174. PR #54182 fixes the narrower `profile install` case by adding the same
registration call; this closes the symptom for any existing profile directory and
without a container restart.
gnome-shell moves a launched ShellApp from STARTING to STOPPED when the
startup-notification sequence completes or times out (mutter, ~15 s), not
when the process exits. finish() healing right after an exit-without-reveal
(boot crash, early quit) therefore wrote the entry during STARTING — the
exact #111906 arming condition. Only the reveal byte from Electron heals
now; the wake byte finish() writes just unblocks the reader. A skipped heal
is picked up by the next terminal/updater launch or revealed grid launch.
`hermes desktop` used to create/refresh `~/.local/share/applications/hermes.desktop`
synchronously before spawning Electron. When the entry is ABSENT (first run after an
update, deleted by a cleaner, tombstoned by AV) that write lands while gnome-shell still
has the grid-launched ShellApp in STARTING; unpatched shells (before GNOME MR !4428)
drop the app's last strong reference on `installed_changed` and the next idle GC kills
the whole Wayland session, minutes to an hour later (#111906, residual after #111396).
Now a launch that carries `DESKTOP_STARTUP_ID` (app grid / menu) defers the write:
the launcher opens a pipe, hands Electron its write end as `HERMES_DESKTOP_READY_FD`
(pass_fds), and a worker thread installs the entry once Electron reports the main
window revealed (`onRevealed` in createWindow, via the new
`apps/desktop/electron/linux-launcher-ready.ts`), plus a 2 s settle so the compositor
has mapped the surface. If Electron exits without ever revealing a window, `finish()`
installs the entry after the exit (STOPPED app, no STARTING object) — self-heal
semantics survive. Terminal launches, the updater's detached relaunch and
`--build-only` have no DESKTOP_STARTUP_ID / spawn no app and keep writing immediately.
Why not simply write after Electron exits (#111915's approach): a daemon thread
started right before `sys.exit` is killed with the interpreter, so the heal is lost,
and a heal that waits for the user to quit the app leaves the menu entry missing for
the whole session.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Turn the SIGTERM→SIGKILL deadline in `_kill_pids_posix` into `_POSIX_TERM_GRACE_SECONDS`
(10.0s) documented against what it has to outlast: `web_server.py::_lifespan` blocks on
`stop_hosted_room_service(timeout=5.0)` + `join(1.0)` before `PTY_REGISTRY.close_all()`
(≤1.5s per attached Chat PTY). The 3.0s deadline predates the hosted-room stop (added
2026-08-30) and SIGKILLed the backend mid-teardown, so its ui-tui / tui_gateway.entry
children were never closed and kept the deleted state.db-wal inode open — the next
`hermes` start refused with FATAL DeletedWalGenerationError (#111912).
Two invariant tests against real child processes: a teardown as long as the lifespan
budget finishes gracefully (red on the old 3.0s deadline); a SIGTERM-ignoring process is
still SIGKILLed at the deadline. The orphan reaper's 1.5s grace is left alone on purpose
— it runs on the Desktop boot path under a 10s ready-probe — and says so in the comment.
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
compression.model_thresholds.<model>, terminal.docker_env.<VAR>, lsp.servers.<lang>.*,
auxiliary.<task>.extra_body.<k> and similar free-form mappings are declared as {} in
DEFAULT_CONFIG; the fail-closed gate walked into the empty dict, found the user-chosen key
missing and refused the write. An empty dict now accepts the rest of the path, like a scalar
leaf or a platforms container does. Populated sections keep the did-you-mean refusal.
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.
Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.
`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.
Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
`hermes config set FEISHU_HOME_CHANNEL oc_x` wrote the top level of config.yaml
while the platform setup flows and /sethome write the same name to .env via
save_env_value, so two writers fed two readers: the gateway bridges the yaml copy
into the environment only when .env lacks the name, one-shot CLI readers never
bridge, and the two copies diverged silently (#111848). Only credential-shaped
names were routed to .env because `_is_env_config_key` is the provider-credential
predicate.
Follow-up to KoNit-K's cherry-picked fix (#111850), which routed the
`setup_hidden_env` suffix family: the predicate now lives in the topical sibling
`hermes_cli/config_env_routing.py` and covers every bare name Hermes itself
registers as an environment variable (OPTIONAL_ENV_VARS, _EXTRA_ENV_KEYS — "env
var names written to .env" — plus the setup-hidden suffixes for plugin adapters
nobody enumerated), so `*_ALLOWED_USERS`, `WHATSAPP_MODE`, `MATRIX_PASSWORD` and
the rest of the adapter-saved family take the same file. `set` and `unset` also
drop a stale same-named top-level config.yaml copy so the reporter's drift cannot
come back, and `get` resolves .env first then that copy — the gateway's own read
order. Provider credentials keep the credential_lifecycle rotation path.
Docs: environment-variables.md tip, hermes_cli/AGENTS.md config rule.
`launchctl kickstart system/<label>` needs root. When a LaunchDaemon-supervised
backend does not come back and `hermes update` runs as a regular user, the
manual hint now reads `sudo launchctl kickstart -k system/<label>`, so the
operator gets a command that works from the shell they are in. LaunchAgent
(gui/user domain) hints are unchanged.
`hermes dashboard --stop` decided its exit code by re-scanning the process
table after the kill. On macOS a launchd KeepAlive job brings the backend
back on a fresh PID within the grace window, so the re-scan found it and
--stop exited 1 right after printing that the stop succeeded and that the job
restarts itself. Judge the exit code from the kill result instead: exit 1
only when a matched pid could not be stopped.
_launchd_job_owning_backend matched a process to a job only on exact argv
equality with the plist ProgramArguments, but _respawn_dashboard_processes
appends --no-open to dashboard commands. A detached copy created by an
earlier update therefore never matched, was treated as a manual backend and
was respawned detached again on every update. Ignore --no-open on both sides
of the comparison.
`hermes doctor` printed the same informational "WAL journal mode" line for a database
that is still WAL although the operator configured `database.journal_mode: delete`. The
runtime never live-downgrades an existing WAL database (a downgrade under open
connections corrupts it) and says so only once per process in the gateway log, so the
one surface operators check told them they were protected when every process was still
writing WAL on the filesystem they configured `delete` for (#111729).
Doctor now compares the configured mode (`resolve_journal_mode`) with the on-disk header
and warns `<db> is in WAL mode despite database.journal_mode=delete`, naming the
never-live-downgraded rule and the one-time offline `PRAGMA journal_mode=DELETE`
remedy; a vulnerable SQLite still counts the database as WAL-reset exposed. This check
outranks the cross-VM hint (whose remedy, "set journal_mode: delete", is already
applied). A configured `wal` is unchanged.
Core hunk ported from #104714 (@jonpol01); its holder enumeration
(`foreign_state_db_holders(include_scan_gaps=True)` + per-PID report) is not included
here — that half is a 600-line change under separate review.
Adopted from the sibling PRs in this cluster, with thanks:
- #99807 (@Chevron7Locked) and #101541 (@ivcdigital26): a kickstart that returns 0 is only
"restart requested" — success now requires launchd to report a live PID other than the one
that was stopped (the gateway's _wait_for_launchd_service_pid, 15 s), and a job that never
comes back is reported with the manual `launchctl kickstart -k` command instead of a tick.
- #89793 (@mettamyron) and #101541: a plist may wrap the backend in `/bin/sh -c …` without
exec, so launchd's live PID is the shell and the stale backend is its child — ownership now
follows the parent chain (bounded, cycle-safe).
- #89793: `hermes dashboard --stop` on a launchd-owned backend now says a KeepAlive job
restarts itself and gives the `launchctl bootout` command, instead of the misleading
"restart it yourself" hint. The job scan runs for --stop too, on macOS only.
`hermes update` snapshots each stale dashboard/serve backend's systemd unit before the
kill and respawns everything else from its captured argv (#40449, #68934, #78821). On
macOS that "everything else" includes a backend supervised by a launchd job: the
respawn comes back detached and sits on the job's port, launchd's KeepAlive then fails
every restart of the job with "port already in use" (exit 75, one attempt per ~10 s,
forever), and the backend that IS serving is no longer supervised. Every later update
kills the copy and respawns it again, so the loop outlives the update that started it.
The update now snapshots the loaded launchd backend jobs before the kill (LaunchAgents
in the gui/user domains, LaunchDaemons in system; only plists whose ProgramArguments
are a dashboard/serve backend are probed), attributes a PID to a job when launchd
reports it as the job's live process or the process runs the job's exact
ProgramArguments (the detached copy an earlier respawn left behind), and brings such a
PID back with `launchctl kickstart <domain>/<label>` instead of an argv respawn. A
failed kickstart is reported with the manual command and counted as unrecovered, like
a failed systemctl restart. Linux and Windows paths are unchanged.
Follow-up to the two salvaged commits (#111419 @Ckarey007, #111521 @wangtaotaotao95),
which both edit `_stage_candidate_venv` and were merged keep-both:
- `_record_runtime_repair` (new topical helper next to `_run_runtime_repair`): a
`failed`/`safe`/`repaired` repair is one `sqlite_runtime_repair` step; a
`skipped`/`not-applicable` repair is one skip WITH its reason instead of a red
step plus a skip. Every pip/non-venv install runs the hook and would otherwise
carry a failed-looking step in every receipt.
- The sync rejection's reason now travels into `RuntimeRepairResult.detail`, which
`_report_runtime_repair_failure` prints and the receipt records, so the extra
console print in `_stage_candidate_venv` is dropped (it duplicated the ℹ line).
- Tests trimmed to invariants: the child-reason test asserts the reason on the
rejection the repair returns (the exception carries it now); the receipt test
covers failed step / repaired step / skipped skip and the no-receipt no-op.
Dropped: the 4-case stage-rejection walk (the same chain is covered by the
existing `_repair_under_lock` tests whose rejection now carries the reason).
The SQLite runtime repair builds its replacement venv with
`uv sync --extra all --locked` and, when that fails, reported only
"replacement environment did not pass dependency and import smoke tests".
The child's diagnosis went to inherited stdout, so it survived in console
scrollback and nowhere else: the rejection line the logger records carried the
bare exit code, the failure detail the repair returns is generic, and update
receipts are built from explicit record_step calls — none of which covers this
repair. On an install where the repair cannot succeed, `hermes update`
therefore ended at "partially complete" with no reason — the reason being a
one-line uv message the user never saw.
Forward that sync's output live (unchanged contract: stderr merged into
stdout and drained while the child runs, so pre-desktop-update hand-offs
cannot deadlock uv on a full stderr pipe) and keep its tail: reject with the
child's `error:`/`hint:` lines and announce them, so the console and the log
say what to fix. Carrying the reason into RuntimeRepairResult.detail, and
therefore into a receipt step, is deliberately out of scope here — separate
change, separate PR.
Test: the reported reason keeps uv's error + hint and drops the progress
noise that precedes them (red before this change, where the reason was the
bare exit code).
`hermes -p <profile> setup gateway` (and `hermes setup` / `hermes import`) reach the
service step through `ensure_gateway_service`, which only knew "is THIS profile's
unit running" — a satellite served by the default multiplexer has no unit of its
own, so the step printed "Installing the gateway background service ..." and
registered a launchd plist / systemd unit that the #97120 start guard then refused,
leaving a stray dead service the user had to find and remove by hand (#111958).
Route both setup surfaces through one shared predicate: `_served_profile_needs_no_service`
wraps `named_profile_served_by_running_multiplexer` (the same probe `profile create`,
cron liveness and the run/start/install guards use), prints the "already served"
note and returns True so `ensure_gateway_service` and the `hermes gateway setup`
wizard (#111962's hunk) skip the install. Default profile and non-multiplex hosts
are unchanged.
Adds the invariant for the `ensure_gateway_service` path (served → no install,
unserved → still installs). Docs: multi-profile-gateways.md names the skipped step.
Co-authored-by: kvnloo <7121943+kvnloo@users.noreply.github.com>
`hermes worktree prune` (and the startup/cron pruner) classified every clean tree in a
repository with no remote as "clean and fully merged/pushed" and force-deleted its branch,
even when the branch carried commits that exist nowhere else. `audit_branches` returned []
in the same repos, so a unique local-only branch was invisible to the audit as well.
Root cause: `_worktree_has_unpushed_commits` answered False when `refs/remotes` was empty
("nothing to be unpushed against") and both consumers — `worktree_gc._classify_tree` and
`worktree_ops._classify_prune_candidates` — read False as "safe to reap".
The preceding commit (#111897) flips that branch to True, which is safe but also means a
no-remote repo can never reclaim anything (a tree sitting at trunk, or squash-merged into
it, stays "unpushed" forever because `_worktree_commits_all_merged_upstream` finds no
origin/* base). This commit replaces the unconditional True with a real baseline:
- `_worktree_local_trunk`: `main`/`master`, else the branch checked out in the main
worktree; None when no trunk exists.
- `_worktree_merge_base_ref`: origin/HEAD|origin/main|origin/master, falling back to the
local trunk ONLY when the repo has no remote-tracking refs at all. Single resolver used
by `_worktree_commits_all_merged_upstream` and `worktree_gc.audit_branches`.
- `_worktree_has_unpushed_commits`: with no remote refs, `git log HEAD --not <trunk>`;
no trunk -> True (preserve), matching the docstring's fail-safe promise.
Net effect in a no-remote repo: unique work is kept ("unpushed commits not found
upstream"), trees at/merged into the local trunk still reap, branch audit reports unique
branches as keep and merged ones as delete, and `git branch -D` can only run on a branch
whose every commit is reachable from or patch-equivalent to the trunk.
Tests: the salvaged no-remote keep test now uses a shared `local_repo` fixture; a control
test proves trunk-merged trees still reclaim and the branch audit reports in the same
repo; `test_merged_predicate_fails_safe_without_upstream` now pins both halves of the
contract (trunk resolves -> merged; no trunk at all -> False/preserve).
backup.py imports hermes_cli.profiles only lazily and profiles.py never imports backup, so
there was no cycle to justify two literals. LOCAL_RUNTIME_ROOT_DIRS now feeds both
backup._EXCLUDED_ROOT_DIRS and the clone-all root gate; one invariant test pins the identity.
43e67d872 (local models) put three machine-scoped trees under the default
HERMES_HOME — models/ (GGUF weights, tens of GB), runtimes/ (managed
llama.cpp binaries) and node/ (managed Node) — and taught backup.py to
exclude them (_EXCLUDED_ROOT_DIRS). `hermes profile create X --clone-all`
from the default profile did not follow: it copied all three into the new
profile, which never reads them (models_dir() resolves from the default
root only) and cannot use them (the binaries are re-downloaded on demand),
turning a clone into a tens-of-GB copy.
Add the three names to _CLONE_ALL_DEFAULT_EXCLUDE_ROOT. The set is gated
on the source being the default profile, so a named profile that really
carries a models/ directory of its own keeps it; and the exclusion applies
at the source root only, so a skill's nested models/ directory is copied
as user data. Kept as a separate literal from backup.py's set on purpose
(backup also matches profiles/<name>/ and importing it here would be a
circular import) with a comment tying the two together.
hermes_time reads `timezone` from config.yaml (or HERMES_TIMEZONE) and, when
ZoneInfo cannot load the name, logs one warning and falls back to
server-local time. Nothing else ever looks at the value: not doctor, not
validate_config_structure. So a typo ("Asia/Tokio", "EST5", "Tokyo") puts
the agent clock and every cron schedule on the server's zone, with the
only evidence a single line in the gateway log.
validate_config_structure now checks the key: a non-string is an error, a
non-empty string that ZoneInfo cannot load is an error, and blank/missing
stays silent (server-local is the documented default). Because the check
runs at startup too, it is skipped when the interpreter has no tz database
at all (bare Windows without tzdata) — there is nothing to judge against,
and flagging every valid name there would be worse than the bug. The hint
names the IANA form and points out that HERMES_TIMEZONE overrides the key.
`hermes sessions list` applied --limit inside the SQL query and rendered
whatever came back, so a user with 35 sessions saw 20 rows and a prompt
and had no way to tell the page was cut. The lister now asks for one row
past the cap, drops that probe row, and ends a truncated page with
"… more not shown (use --limit N to see more)". `--limit 0` is LIMIT 0
(no rows), so the copy never suggests it.
The footer goes through one shared helper, `cli_output.print_truncated`,
and the three "... N more" / "… N more" / "… and N more" variants already
in sessions_cmd.py (export dry-run preview, never-active cleanup, prune/
archive preview) now use it too, so every capped listing in the command
reads the same. Migrating checkpoints / curator / skills search / console
listers is follow-up work.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Replace the three monkeypatch-heavy tests from the salvaged commit (which
faked clear_legacy's return dict, so they could not catch the manager hunk
regressing) with two invariant tests that drive the real cmd_clear_legacy
against a temp checkpoint base: an undeletable legacy-* dir yields exit 2
plus the "Could not delete" line while the archive stays on disk; a clean
sweep keeps exit 0 and the unchanged success line. The green-path guard is
harvested from #111789.
Reword the CLI failure line to "Could not delete N archive(s) (see logs)."
so it matches the manager's WARNING wording and the text proposed in
#111776, and document the exit code in the CLI reference.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).
Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.