Commit Graph

7790 Commits

Author SHA1 Message Date
xielevi 81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00
xielevi 4ba717df12 fix(profiles): migrate session/routing identity on profile rename
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.

The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:

- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
  tables, matching the agent:<name>: namespace by exact prefix (substr, not
  LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
  the profile inside routing/origin JSON, and REFUSING on a target collision
  (routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
  (keys + origin.profile) then persist — the half a DB write cannot reach.
  Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
  params only to handlers that declare them, bare handlers unchanged) so a
  live gateway rekeys its in-memory copy AND both durable stores (routing home
  + the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
  does NOT fall back to a racing CLI-side write: it prints a warning telling
  the operator to restart the gateway and retry. With no live gateway it
  performs the durable rewrite itself (safe: nothing else holds the store
  open).

Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.

Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
2026-09-16 00:32:15 -07:00
chelsealong 3b9c1118cd fix(kanban): surface the worker's own last output on a dead-worker reap
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.

Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
2026-09-15 22:47:03 -07:00
teknium1 08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1 fd74bf6f58 fix(config): drop the runtime read of the docs env registry
Production code must not depend on website/docs being present; the
"Hermes does not read this name" note now keys only on the code
registry (OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS / setup-hidden suffixes).
2026-09-15 21:47:48 -07:00
teknium1 f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1 db25a7852e fix(plugins): install refuses to ship an unreadable plugin tree (#111804)
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.

After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.

Fixes #111804 (its discovery half landed in #112293).
2026-09-15 21:47:18 -07:00
teknium1 2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1 2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural 7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
dmelkk-secondbrain be9d4369a7 fix(cli): one-shot -q runs report their outcome in the exit code
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.

`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.

`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.

This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.

Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.

Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
2026-09-15 19:28:32 -07:00
teknium1 fff10484d6 fix: drop the dead disabled guard on the lazy MCP banner line
get_mcp_status reports status='disabled' (never 'lazy') for a disabled
server and derives the 'disabled' flag from that same status, so the
extra 'and not srv.get("disabled")' check could never change the branch.
2026-09-15 19:06:54 -07:00
John Paul Soliva a17d0409be fix(mcp): report lazily registered servers as lazy, not configured or failed
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:

- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
  lazily registered server. It now reports `lazy` with the cached tool
  count (`connected: False`); an in-flight or failed first-use connect
  still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
  `_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
  failed)` right after registering every cached tool, and re-announced
  the same "failure" on every repeat discovery. Lazy servers are now
  reported as `(N lazy, not spawned yet)` and an already-lazy server is
  not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
  two sites, so every startup logged `Background MCP discovery completed
  with zero connected servers` and every later call re-spawned the
  discovery thread as a retry. One predicate,
  `_discovery_registered_servers`, treats a lazy registration as a
  usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
  red "could not connect" line; it now shows the cached tool count with
  `(lazy, starts on first use)`.

Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).

Fixes #111717
2026-09-15 19:06:54 -07:00
teknium1 204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
teknium1 a2837ec088 docs(docker): overriding entrypoint: drops the zombie reaper — document init: true; trim the PID-1 warning
Follow-up to the cherry-picked #111584 (@chelsealong):

- website/docs/user-guide/docker.md: new warning block next to the existing
  "do not override the entrypoint" note explaining WHY (with `/init` gone the
  hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
  children), the Compose `init: true` / `docker run --init` remedy, and that
  supervision is still lost on that path; plus a Troubleshooting entry for
  `<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
  check and drops the `platform.system()` gate and the blanket
  `try/except Exception: pass` — a user process is never PID 1 on any host OS
  (PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
  check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).

Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
2026-09-15 19:04:00 -07:00
chelsealong 8d5cce4d94 fix(cli): warn when hermes runs unsupervised as PID 1
A deployment that overrides the image's `entrypoint:` to invoke hermes
directly skips docker/entrypoint-dispatch.sh entirely, so hermes itself
becomes PID 1 with no s6-overlay /init (or any other init) above it.
Nothing then reaps orphaned grandchildren (browser tooling, MCP
subprocesses, shell-tool children) reparented to PID 1, and they
accumulate as zombies without bound.

entrypoint-dispatch.sh already warns on its own non-PID-1 fallback
path, but that script never runs in the entrypoint-override case, so
there was no signal at all. Add the same style of warning inside
hermes_cli.main, gated on being PID 1 on Linux, pointing users at the
image's default ENTRYPOINT or `docker run --init` / `init: true`.

Fixes #111577
2026-09-15 19:04:00 -07:00
teknium1 da18c20226 fix: kanban dispatcher prefers module argv over PATH hermes
Review finding: hermes_cli/kanban_db_dispatch.py::_resolve_hermes_argv still resolved which('hermes') before sys.executable -m hermes_cli.main while claiming to mirror gateway.run._resolve_hermes_bin, which this PR made module-first (#111569). Keep the explicit HERMES_BIN override first, then the module argv whenever hermes_cli is importable, PATH only as fallback; docstring updated.
2026-09-15 19:03:08 -07:00
teknium1 de132792f6 fix(gateway): /status resolves the switched-to model's context window; validator names the missing slug
Two diagnostics diverged after a session-only `/model` switch (#111436).

/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).

Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).

Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:58:02 -07:00
teknium1 6062ad5aee fix(platforms): QR fallback tip names Hermes' own uv when it is not on PATH
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
2026-09-15 18:55:47 -07:00
teknium1 c2b93ae5ac fix(platforms): QR fallback tip uses uv against the running interpreter
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.

Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:55:47 -07:00
Kevin Rajan 505b36b7af fix(platforms): make QR fallback install tip target the active interpreter
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.

Fixes #111695
2026-09-15 18:55:47 -07:00
teknium1 8017dfa4a8 fix(web): drop dead subprocess import from git router 2026-09-15 18:54:51 -07:00
teknium1 ac829e8dee fix(web): gh auth refresh waits out a probe that started before it was asked for
With the shared single-flight probe (previous commit) a `refresh=true`
request that landed while a probe was already running simply joined it. That
probe may have started before `gh auth login` completed, so the refresh
returned "not authenticated" and cached it for the full 5-minute TTL — the
composer pill kept offering /github-auth right after a successful login.

A refresh now accepts only a probe that started at or after the refresh was
requested: it awaits the in-flight one, then starts (or joins) the next.
Still only one `gh` runs at a time. A finished task whose done-callback has
not run yet is treated as absent so the loop cannot spin on it.

Live probe (real route, fake `gh` reading login state at start, 1 s answer):
PR head: refresh=true right after login -> authenticated False, cached False
fixed:   refresh=true right after login -> authenticated True,  cached True;
         5 concurrent requests -> peak concurrent probes 1

Co-authored-by: aron-intframe <aron-intframe@users.noreply.github.com>
2026-09-15 18:54:51 -07:00
KoNit-K 460768056d fix(web): bound and deduplicate gh auth probes 2026-09-15 18:54:51 -07:00
KoNit-K 68e4833134 fix(desktop): bound background profile hydration retries 2026-09-15 18:54:28 -07:00
teknium1 22716fab2a fix(plugins): one unreadable plugin dir no longer aborts loader or list discovery
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.

Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.

Part of #111804
2026-09-15 18:48:59 -07:00
KoNit-K 9af62e7ef0 fix(dashboard): skip unreadable plugin manifests 2026-09-15 18:48:59 -07:00
teknium1 5f9042bec6 fix(plugins): validate accepts model-provider plugins and mirrors the real PluginContext surface
Two false positives in `hermes plugins validate` that block catalog admission
for plugins that are correct at runtime:

- `kind: model-provider` plugins register at import via
  providers.register_provider(ProviderProfile); the PluginManager never calls a
  register(ctx) on them (plugins_discovery skips the kind). The probe demanded
  register() anyway, so every provider plugin -- including the in-tree
  plugins/model-providers/* -- failed with "no register() function". The probe
  now records register_provider calls for that kind and fails only when the
  import registers nothing.
- RecordingContext returned a no-op callable for ANY attribute, so
  `getattr(ctx, "profile_path", None)` was truthy under validation alone and
  register() crashed with an error the real PluginContext never produces. The
  parent now passes the real PluginContext method names into the probe; other
  names raise AttributeError exactly like the real object.

Surfaced by the 2026-09-15 catalog sweep (Gondola provider, hermes-persona).
2026-09-15 18:39:54 -07:00
KoNit-K addcf7d898 fix(cli): accept native npm paths under mnt 2026-09-15 18:35:59 -07:00
teknium1 bd63866253 fix(kanban): give a finished worker a grace window before the terminal reaper signals it
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.

Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
2026-09-15 18:35:32 -07:00
teknium1 aa5817d9be fix(kanban): reap workers that outlive their finished run
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).

Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.

Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.

Fixes #111791
2026-09-15 18:35:32 -07:00
teknium1 4465b8d7d5 fix(kanban): BLOB cells in comment/event/run rows degrade like a BLOB task body
Only Task.from_row coerced BLOB-typed cells; a BLOB task_comments.body (or
event payload / run summary) still came back as bytes and crashed
`hermes kanban show <id> --json` with "Object of type bytes is not JSON
serializable". Apply _lossy_text in the other from_row constructors.

Test: BLOB comment body and event payload -> str with U+FFFD and
JSON-serialisable (red before).
2026-09-15 18:35:09 -07:00
teknium1 a647c6cb2d fix(kanban): one undecodable task field no longer breaks the whole board listing
A tasks row whose TEXT body holds invalid UTF-8 made sqlite3 raise
"Could not decode to UTF-8 column 'body'" inside fetchall, so `hermes kanban
list` (and `show`, and every other reader of that row) failed board-wide
until the row was deleted by hand; a BLOB-typed body came back as bytes and
crashed `--json` (#111743).

Fix it once at the connection: every board connection (`_open_configured`
and the read-only descendant path in `connect`) installs a lossy
text_factory that substitutes U+FFFD, and `Task.from_row` runs BLOB cells
through the same helper so a corrupt row renders with replacement
characters instead of taking its neighbours down.

Fixes #111743
2026-09-15 18:35:09 -07:00
teknium1 85d4415bed fix(kanban): request_review shares complete_task's live-worker fence
The PR docstring said complete_task applies "the same fence request_review
applies", but the two disagreed: complete_task keyed on a live worker
process while request_review still refused any running task with a
claim_lock, so a claim whose worker is gone (or a CLI/library claim that
never spawned one) could be completed but not sent to review without
force. Factor the liveness test into _claim_is_live and use it in both.

TTL expiry is deliberately not part of "live": reclaim_stale_tasks extends
(not reclaims) the claim of a live worker, so the process stays the
liveness authority.

Test: claim -> request_review without a worker PID now succeeds (red
before); a live worker's claim is still refused without expected_run_id.
2026-09-15 18:34:40 -07:00
teknium1 72916de360 fix(kanban): live-claim guard keys on a live worker process, not on any claim
The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
2026-09-15 18:34:40 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1 c7f4bc5bd7 fix(kanban): text dispatch output and both "dispatcher stuck" warnings name the hold reason
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.

- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
  task plus rate_limited / skipped_locked / memory_pressure for one or more
  DispatchResults, so the CLI daemon and gateway warnings share one wording:
  `Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
  tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.

Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>

Part of #111910
2026-09-15 18:34:11 -07:00
KoNit-K ec64ec0d24 fix(kanban): dispatch --json reports respawn_guarded and other suppression reasons
`hermes kanban dispatch --json` only emitted `spawned` and the skip buckets it
already knew about, so a ready card held by the respawn guard (`active_pr`,
`recent_success`, ...), a quota-released worker, a lost dispatch lock or a
memory-pressure hold all looked like `spawned: []` with no reason. Emit
`respawn_guarded`, `rate_limited`, `skipped_locked` and `memory_pressure`
from the DispatchResult the tick already returns.

Salvaged from #111917 by @KoNit-K. Dropped hunk: the `_ACTIVE_PR_RECOVERY_LANES
= frozenset({"review"})` rename in kanban_db_dispatch.py, which is behaviour-
identical to the existing `lane == "review"` check and does not implement the
role-aware exemption the issue asks for.

Part of #111910
2026-09-15 18:34:11 -07:00
teknium1 058e620bef fix(kanban): headless specify/decompose aux calls carry a relay session key
`hermes kanban specify|decompose`, the dashboard specify/decompose routes and
the gateway auto-decomposer all reach the LLM through
hermes_cli/kanban_specify.py::_call_aux outside any agent turn. No
conversation affinity scope is bound there, so agent/opencode_affinity.py
resolved an empty key and sent no `x-opencode-session`; the OpenCode Go relay
rejects such requests with 400 MissingSessionID and the user sees
"Specify failed: LLM error: BadRequestError".

Declare a per-task affinity scope (`kanban:<task_id>`) around the call —
the same host-declared scope the main turn, compression and the
OpenRouter/Portal sticky keys already resolve first — but only when no scope
is bound, so an in-turn caller keeps its conversation's key. Reset in a
finally so nothing leaks past the call.

Live: real httpx transport capture against auxiliary.triage_specifier
provider=opencode-go — before: no x-opencode-session header; after:
`kanban:t_45567533` on specify and decompose, stable per task, distinct per
task; a pre-declared scope is preserved; an openai route gets no header.

Fixes #112043
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:33:43 -07:00
teknium1 7c6b8c7a8d fix(update): Desktop-owned serve no longer fails the update or holds fleet_restart_pending
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.

- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
  reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
  be untrue — the process provably runs old code and the Desktop app does not respawn it after
  a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
  `unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
  that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
  hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
  excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
  gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
  in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
  serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
  origin/main.

Fixes #111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:30:32 -07:00
teknium1 336227bf00 refactor(gateway): move the on-demand s6 slot helper into service_manager, trim tests
Reshape of the cherry-picked fix from #104194:

- The SOUL.md gate + `register_profile_gateway(start_now=False)` now live in
  `hermes_cli/service_manager.py::register_unregistered_profile_gateway`, next to the s6
  manager and `_profile_dir_for_gateway_service` it needs, instead of a private reach-in
  from the 6.5k-line `hermes_cli/gateway.py` facade. The facade only decides "start
  repairs, stop/restart re-raise" and keeps ONE error handler (S6Error is a RuntimeError;
  register's ValueError/RuntimeError/OSError surface as the same `✗` + exit 1).
- Tests trimmed from four to two invariants: start on a real profile registers `down`
  and then starts; stop on an unregistered profile / start on a directory without
  SOUL.md keep the original error and mint nothing (parametrized). Dropped: the
  registration-failure traceback test (covered by the single except clause) and the
  duplicate no-marker/stop split. The test now resolves the profile dir through the real
  HERMES_HOME mapping instead of monkeypatching `_profile_dir_for_gateway_service`.
- Docs: docker.md multi-profile section says `gateway start` inside the container
  registers a slot for a profile created from the host.
2026-09-15 18:30:05 -07:00
John Paul Soliva 9c36cafcec fix(gateway): in-container gateway start registers a missing s6 slot
`profile create` registers an s6 gateway slot only when the creating process is
itself inside the container: `detect_service_manager()` reads `/proc/1/comm` in
the CALLER's PID namespace, so on a host whose `~/.hermes` is bind-mounted into
the container the hook is a silent no-op. The profile directory lands exactly where
the container reads it, but no slot exists, and `hermes -p <name> gateway start`
inside the container fails with "not registered" until the operator restarts the
container so the boot reconciler notices. Same symptom as #54174 reaches `profile
install` by the same route.

Register the slot on demand instead. When `start` hits `GatewayNotRegisteredError`
and the profile directory carries `SOUL.md` — the boot reconciler's own "real
profile" marker — create the slot and start it. That makes
`_maybe_register_gateway_service`'s promise true without a restart.

Deliberately narrow:

- Only `start` self-heals. `stop` and `restart` on an unregistered profile keep the
  original error; registering a slot in order to stop it would be absurd.
- `SOUL.md` gates it, so a mistyped `-p` name or a stray directory cannot mint a
  phantom slot for a profile that does not exist.
- Registers with `start_now=False` and then goes through the ordinary `start` path,
  so the `desired_state` write that lets boot reconciliation restore want-up after
  a container restart keeps a single owner.
- `ValueError` (slot appeared underneath us) and `RuntimeError` (s6-svscanctl
  failed) surface as the existing actionable error, never a traceback.

Refs #54174. PR #54182 fixes the narrower `profile install` case by adding the same
registration call; this closes the symptom for any existing profile directory and
without a container restart.
2026-09-15 18:30:05 -07:00
teknium1 49b9bbb6fc fix(desktop): an exit without a window reveal no longer writes hermes.desktop
gnome-shell moves a launched ShellApp from STARTING to STOPPED when the
startup-notification sequence completes or times out (mutter, ~15 s), not
when the process exits. finish() healing right after an exit-without-reveal
(boot crash, early quit) therefore wrote the entry during STARTING — the
exact #111906 arming condition. Only the reveal byte from Electron heals
now; the wake byte finish() writes just unblocks the reader. A skipped heal
is picked up by the next terminal/updater launch or revealed grid launch.
2026-09-15 18:29:37 -07:00
teknium1 169c48fa13 fix(desktop): app-grid launches write hermes.desktop only after the window is on screen
`hermes desktop` used to create/refresh `~/.local/share/applications/hermes.desktop`
synchronously before spawning Electron. When the entry is ABSENT (first run after an
update, deleted by a cleaner, tombstoned by AV) that write lands while gnome-shell still
has the grid-launched ShellApp in STARTING; unpatched shells (before GNOME MR !4428)
drop the app's last strong reference on `installed_changed` and the next idle GC kills
the whole Wayland session, minutes to an hour later (#111906, residual after #111396).

Now a launch that carries `DESKTOP_STARTUP_ID` (app grid / menu) defers the write:
the launcher opens a pipe, hands Electron its write end as `HERMES_DESKTOP_READY_FD`
(pass_fds), and a worker thread installs the entry once Electron reports the main
window revealed (`onRevealed` in createWindow, via the new
`apps/desktop/electron/linux-launcher-ready.ts`), plus a 2 s settle so the compositor
has mapped the surface. If Electron exits without ever revealing a window, `finish()`
installs the entry after the exit (STOPPED app, no STARTING object) — self-heal
semantics survive. Terminal launches, the updater's detached relaunch and
`--build-only` have no DESKTOP_STARTUP_ID / spawn no app and keep writing immediately.

Why not simply write after Electron exits (#111915's approach): a daemon thread
started right before `sys.exit` is killed with the interpreter, so the heal is lost,
and a heal that waits for the user to quit the app leaves the menu entry missing for
the whole session.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:29:37 -07:00
teknium1 f93d33f93f fix(cli): name the dashboard kill grace and pin it against the lifespan teardown
Turn the SIGTERM→SIGKILL deadline in `_kill_pids_posix` into `_POSIX_TERM_GRACE_SECONDS`
(10.0s) documented against what it has to outlast: `web_server.py::_lifespan` blocks on
`stop_hosted_room_service(timeout=5.0)` + `join(1.0)` before `PTY_REGISTRY.close_all()`
(≤1.5s per attached Chat PTY). The 3.0s deadline predates the hosted-room stop (added
2026-08-30) and SIGKILLed the backend mid-teardown, so its ui-tui / tui_gateway.entry
children were never closed and kept the deleted state.db-wal inode open — the next
`hermes` start refused with FATAL DeletedWalGenerationError (#111912).

Two invariant tests against real child processes: a teardown as long as the lifespan
budget finishes gracefully (red on the old 3.0s deadline); a SIGTERM-ignoring process is
still SIGKILLed at the deadline. The orphan reaper's 1.5s grace is left alone on purpose
— it runs on the Desktop boot path under a 10s ready-probe — and says so in the comment.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-15 18:29:16 -07:00
AzurePii 83f1b88305 fix(cli): increase dashboard kill timeout to prevent PTY child leaks and WAL corruption 2026-09-15 18:29:16 -07:00
teknium1 c1115a7166 fix(config): treat empty-dict DEFAULT_CONFIG sections as open containers in the typo gate
compression.model_thresholds.<model>, terminal.docker_env.<VAR>, lsp.servers.<lang>.*,
auxiliary.<task>.extra_body.<k> and similar free-form mappings are declared as {} in
DEFAULT_CONFIG; the fail-closed gate walked into the empty dict, found the user-chosen key
missing and refused the write. An empty dict now accepts the rest of the path, like a scalar
leaf or a platforms container does. Populated sections keep the did-you-mean refusal.
2026-09-15 18:28:49 -07:00
teknium1 0e63a1bc5c fix(config): refuse an unknown path under a known section before writing
`hermes config set gateway.discord.gateway_restart_notification true` wrote the
typo into config.yaml and only then printed the "not a recognized config key — it
was saved anyway" notice (#112003). Under a KNOWN section an unknown sub-key can
only be a typo, so `set_config_value` now exits non-zero via `_exit_invalid`
before reading or writing config.yaml, with the did-you-mean hint.

Scope preserved from ed3a0b3 (warn-after-write): unknown TOP-LEVEL keys are
still written with the post-write notice, because top-level scalars are bridged
into os.environ for skills/external apps and that namespace is open by design;
the `_OPEN_SUBKEY_TOP_LEVEL_KEYS` / platform-container exemptions in
`_validate_config_key` are untouched, and `--force` keeps writing anything. This
is the fail-fast piece the maintainer scoped in the close comment on #111133.

`_validate_config_key` also suggests the path minus its wrong prefix
(`gateway.discord.x` -> `discord.x`) when no same-level sibling is close; the
headline typo previously produced no hint at all.

Docs: cli-commands.md `set`/`unset` rows, configuration.md tip, `--force` help.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:28:49 -07:00
teknium1 70e4938c07 fix(config): route every registered env setting through .env from config set/get/unset
`hermes config set FEISHU_HOME_CHANNEL oc_x` wrote the top level of config.yaml
while the platform setup flows and /sethome write the same name to .env via
save_env_value, so two writers fed two readers: the gateway bridges the yaml copy
into the environment only when .env lacks the name, one-shot CLI readers never
bridge, and the two copies diverged silently (#111848). Only credential-shaped
names were routed to .env because `_is_env_config_key` is the provider-credential
predicate.

Follow-up to KoNit-K's cherry-picked fix (#111850), which routed the
`setup_hidden_env` suffix family: the predicate now lives in the topical sibling
`hermes_cli/config_env_routing.py` and covers every bare name Hermes itself
registers as an environment variable (OPTIONAL_ENV_VARS, _EXTRA_ENV_KEYS — "env
var names written to .env" — plus the setup-hidden suffixes for plugin adapters
nobody enumerated), so `*_ALLOWED_USERS`, `WHATSAPP_MODE`, `MATRIX_PASSWORD` and
the rest of the adapter-saved family take the same file. `set` and `unset` also
drop a stale same-named top-level config.yaml copy so the reporter's drift cannot
come back, and `get` resolves .env first then that copy — the gateway's own read
order. Provider credentials keep the credential_lifecycle rotation path.

Docs: environment-variables.md tip, hermes_cli/AGENTS.md config rule.
2026-09-15 18:28:49 -07:00
KoNit-K 49bbc738d4 fix(config): route platform setup values through .env 2026-09-15 18:28:49 -07:00