Commit Graph

7798 Commits

Author SHA1 Message Date
m4 3ca91c55f3 feat(desktop): openExternalFileForIpc opens files via OS handler
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-09-17 22:28:07 +08:00
teknium1 db54f5448d fix(kanban): worker fingerprint carries a boot witness; an uncaptured fingerprint never authorizes a signal
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):

- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
  is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
  a reboot, and that counter does not, so an unrelated process on a later boot with the
  same PID and the same tick value passed _start_times_agree(). The fingerprint is now
  "<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
  the witness the drain marker already uses); both halves must match. Integer values on
  rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
  row and falls back to bare PID existence - a new spawn silently recreated the #89614/
  #99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
  claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
  by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
  once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
  the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
  and must not outlive its pid.

Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.

Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
2026-09-16 00:35:00 -07:00
teknium1 8609743389 fix(serve): launch-profile scope decided at entry; send keeps scope authority; per-reset release
Three edges of the fail-closed multi-profile host (#111620 review, andrexibiza P1 + P2,
kvnloo finding 1):

- `send` under a routed profile's scope `update()`d the installed scope from raw `.env`,
  reversing build_profile_secret_scope's precedence (user .env, then external secret
  sources) for the rest of the request; a stale user value beat the secret-manager one.
  The installed scope is authoritative as-is; only the config.yaml setdefault bridge runs.
- The launch profile's body was scoped only when `is_multiplex_active()` was already true
  at entry, while get_secret consults that global on every read. A launch RPC / dashboard
  request entering single-profile and resuming after a concurrent first `?profile=B`
  activation raised UnscopedSecretError mid-request. The launch profile's secret scope
  (its .env + external sources over the launch env: live while single-profile, the frozen
  snapshot once multiplexing is active) is now bound for every launch-profile body, so the
  credential source is fixed at entry. The terminal policy overlay stays multiplex-only
  (standalone terminal execution keeps its os.environ bridge). _publish_env_value mirrors a
  same-request .env write into that scope AND os.environ for the launch profile, only into
  the scope for a routed one (serves_routed_profile).
- _release_profile_runtime_scope_tokens reset terminal → secret → home in sequence under
  one outer suppress; a failing terminal reset left the previous profile's secrets and
  HERMES_HOME installed for the next body in that context. Each reset is now independent;
  the first failure is re-raised after every scope is released.

tests/tui_gateway/test_multi_profile_hosting_transitions.py: manager-vs-dotenv precedence
through _load_hermes_env, TUI-RPC and dashboard barrier tests (launch enters single-profile,
B activates on another thread, launch resumes and still resolves its injected credential,
never B's), forced terminal-reset failure still releases secret + home. 4/4 red on base.
2026-09-16 00:35:00 -07:00
teknium1 7dde7a2424 fix(gateway): migrate --multiplex resumes from live state, compensates the whole destructive phase, and refuses an unknown default principal
Three P1 findings from the review of #111062 (fixes #110850 remainder):

- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
  installed-but-dead unit; `systemd_install` writes the unit before the start that
  can still be killed, so an apply interrupted at default/start (or mid-restart of
  a stopped default) was reported as "already multiplexed" and never recovered.
  `interrupted` is now derived from the manifest vs the LIVE default's recorded
  served_profiles: flag on + manifest + not every migrated profile served = resume.

- the compensation `try` began after `_remove_secondary_gateways` and the flag
  write, so a later secondary's stop, a unit's daemon-reload or the config write
  failing left the first secondary removed with no rollback. The flag write,
  removals and default bring-up now all sit inside one boundary that rolls back
  through the manifest; `_preflight_apply` refuses failures knowable from the plan
  (root system unit without a recorded User=, unresolvable recorded user, an
  unwritable config.yaml) before any working gateway is stopped.

- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
  `default_uid is None`; a default system unit whose User= this host cannot
  resolve is the same boundary from the other side and now blocks the update
  hook (same known uid still folds, different known uid still refuses).
2026-09-16 00:33:03 -07:00
Sora-bluesky 9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
teknium1 5d97d5ed6d refactor(profiles): move rename identity migration off the facades; trim tests to invariants
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
  migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
  now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
  profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
  gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
  hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
  tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.

Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
2026-09-16 00:32:15 -07:00
KoNit-K 91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
xielevi 81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00
xielevi 4ba717df12 fix(profiles): migrate session/routing identity on profile rename
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.

The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:

- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
  tables, matching the agent:<name>: namespace by exact prefix (substr, not
  LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
  the profile inside routing/origin JSON, and REFUSING on a target collision
  (routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
  (keys + origin.profile) then persist — the half a DB write cannot reach.
  Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
  params only to handlers that declare them, bare handlers unchanged) so a
  live gateway rekeys its in-memory copy AND both durable stores (routing home
  + the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
  does NOT fall back to a racing CLI-side write: it prints a warning telling
  the operator to restart the gateway and retry. With no live gateway it
  performs the durable rewrite itself (safe: nothing else holds the store
  open).

Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.

Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
2026-09-16 00:32:15 -07:00
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
chelsealong 3b9c1118cd fix(kanban): surface the worker's own last output on a dead-worker reap
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.

Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
2026-09-15 22:47:03 -07:00
teknium1 08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1 fd74bf6f58 fix(config): drop the runtime read of the docs env registry
Production code must not depend on website/docs being present; the
"Hermes does not read this name" note now keys only on the code
registry (OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS / setup-hidden suffixes).
2026-09-15 21:47:48 -07:00
teknium1 f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1 db25a7852e fix(plugins): install refuses to ship an unreadable plugin tree (#111804)
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.

After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.

Fixes #111804 (its discovery half landed in #112293).
2026-09-15 21:47:18 -07:00
teknium1 2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1 2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural 7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
dmelkk-secondbrain be9d4369a7 fix(cli): one-shot -q runs report their outcome in the exit code
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.

`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.

`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.

This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.

Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.

Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
2026-09-15 19:28:32 -07:00
teknium1 fff10484d6 fix: drop the dead disabled guard on the lazy MCP banner line
get_mcp_status reports status='disabled' (never 'lazy') for a disabled
server and derives the 'disabled' flag from that same status, so the
extra 'and not srv.get("disabled")' check could never change the branch.
2026-09-15 19:06:54 -07:00
John Paul Soliva a17d0409be fix(mcp): report lazily registered servers as lazy, not configured or failed
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:

- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
  lazily registered server. It now reports `lazy` with the cached tool
  count (`connected: False`); an in-flight or failed first-use connect
  still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
  `_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
  failed)` right after registering every cached tool, and re-announced
  the same "failure" on every repeat discovery. Lazy servers are now
  reported as `(N lazy, not spawned yet)` and an already-lazy server is
  not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
  two sites, so every startup logged `Background MCP discovery completed
  with zero connected servers` and every later call re-spawned the
  discovery thread as a retry. One predicate,
  `_discovery_registered_servers`, treats a lazy registration as a
  usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
  red "could not connect" line; it now shows the cached tool count with
  `(lazy, starts on first use)`.

Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).

Fixes #111717
2026-09-15 19:06:54 -07:00
teknium1 204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
teknium1 a2837ec088 docs(docker): overriding entrypoint: drops the zombie reaper — document init: true; trim the PID-1 warning
Follow-up to the cherry-picked #111584 (@chelsealong):

- website/docs/user-guide/docker.md: new warning block next to the existing
  "do not override the entrypoint" note explaining WHY (with `/init` gone the
  hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
  children), the Compose `init: true` / `docker run --init` remedy, and that
  supervision is still lost on that path; plus a Troubleshooting entry for
  `<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
  check and drops the `platform.system()` gate and the blanket
  `try/except Exception: pass` — a user process is never PID 1 on any host OS
  (PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
  check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).

Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
2026-09-15 19:04:00 -07:00
chelsealong 8d5cce4d94 fix(cli): warn when hermes runs unsupervised as PID 1
A deployment that overrides the image's `entrypoint:` to invoke hermes
directly skips docker/entrypoint-dispatch.sh entirely, so hermes itself
becomes PID 1 with no s6-overlay /init (or any other init) above it.
Nothing then reaps orphaned grandchildren (browser tooling, MCP
subprocesses, shell-tool children) reparented to PID 1, and they
accumulate as zombies without bound.

entrypoint-dispatch.sh already warns on its own non-PID-1 fallback
path, but that script never runs in the entrypoint-override case, so
there was no signal at all. Add the same style of warning inside
hermes_cli.main, gated on being PID 1 on Linux, pointing users at the
image's default ENTRYPOINT or `docker run --init` / `init: true`.

Fixes #111577
2026-09-15 19:04:00 -07:00
teknium1 da18c20226 fix: kanban dispatcher prefers module argv over PATH hermes
Review finding: hermes_cli/kanban_db_dispatch.py::_resolve_hermes_argv still resolved which('hermes') before sys.executable -m hermes_cli.main while claiming to mirror gateway.run._resolve_hermes_bin, which this PR made module-first (#111569). Keep the explicit HERMES_BIN override first, then the module argv whenever hermes_cli is importable, PATH only as fallback; docstring updated.
2026-09-15 19:03:08 -07:00
teknium1 de132792f6 fix(gateway): /status resolves the switched-to model's context window; validator names the missing slug
Two diagnostics diverged after a session-only `/model` switch (#111436).

/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).

Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).

Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:58:02 -07:00
teknium1 6062ad5aee fix(platforms): QR fallback tip names Hermes' own uv when it is not on PATH
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
2026-09-15 18:55:47 -07:00
teknium1 c2b93ae5ac fix(platforms): QR fallback tip uses uv against the running interpreter
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.

Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:55:47 -07:00
Kevin Rajan 505b36b7af fix(platforms): make QR fallback install tip target the active interpreter
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.

Fixes #111695
2026-09-15 18:55:47 -07:00
teknium1 8017dfa4a8 fix(web): drop dead subprocess import from git router 2026-09-15 18:54:51 -07:00
teknium1 ac829e8dee fix(web): gh auth refresh waits out a probe that started before it was asked for
With the shared single-flight probe (previous commit) a `refresh=true`
request that landed while a probe was already running simply joined it. That
probe may have started before `gh auth login` completed, so the refresh
returned "not authenticated" and cached it for the full 5-minute TTL — the
composer pill kept offering /github-auth right after a successful login.

A refresh now accepts only a probe that started at or after the refresh was
requested: it awaits the in-flight one, then starts (or joins) the next.
Still only one `gh` runs at a time. A finished task whose done-callback has
not run yet is treated as absent so the loop cannot spin on it.

Live probe (real route, fake `gh` reading login state at start, 1 s answer):
PR head: refresh=true right after login -> authenticated False, cached False
fixed:   refresh=true right after login -> authenticated True,  cached True;
         5 concurrent requests -> peak concurrent probes 1

Co-authored-by: aron-intframe <aron-intframe@users.noreply.github.com>
2026-09-15 18:54:51 -07:00
KoNit-K 460768056d fix(web): bound and deduplicate gh auth probes 2026-09-15 18:54:51 -07:00
KoNit-K 68e4833134 fix(desktop): bound background profile hydration retries 2026-09-15 18:54:28 -07:00
teknium1 22716fab2a fix(plugins): one unreadable plugin dir no longer aborts loader or list discovery
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.

Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.

Part of #111804
2026-09-15 18:48:59 -07:00
KoNit-K 9af62e7ef0 fix(dashboard): skip unreadable plugin manifests 2026-09-15 18:48:59 -07:00
teknium1 5f9042bec6 fix(plugins): validate accepts model-provider plugins and mirrors the real PluginContext surface
Two false positives in `hermes plugins validate` that block catalog admission
for plugins that are correct at runtime:

- `kind: model-provider` plugins register at import via
  providers.register_provider(ProviderProfile); the PluginManager never calls a
  register(ctx) on them (plugins_discovery skips the kind). The probe demanded
  register() anyway, so every provider plugin -- including the in-tree
  plugins/model-providers/* -- failed with "no register() function". The probe
  now records register_provider calls for that kind and fails only when the
  import registers nothing.
- RecordingContext returned a no-op callable for ANY attribute, so
  `getattr(ctx, "profile_path", None)` was truthy under validation alone and
  register() crashed with an error the real PluginContext never produces. The
  parent now passes the real PluginContext method names into the probe; other
  names raise AttributeError exactly like the real object.

Surfaced by the 2026-09-15 catalog sweep (Gondola provider, hermes-persona).
2026-09-15 18:39:54 -07:00
KoNit-K addcf7d898 fix(cli): accept native npm paths under mnt 2026-09-15 18:35:59 -07:00
teknium1 bd63866253 fix(kanban): give a finished worker a grace window before the terminal reaper signals it
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.

Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
2026-09-15 18:35:32 -07:00
teknium1 aa5817d9be fix(kanban): reap workers that outlive their finished run
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).

Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.

Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.

Fixes #111791
2026-09-15 18:35:32 -07:00
teknium1 4465b8d7d5 fix(kanban): BLOB cells in comment/event/run rows degrade like a BLOB task body
Only Task.from_row coerced BLOB-typed cells; a BLOB task_comments.body (or
event payload / run summary) still came back as bytes and crashed
`hermes kanban show <id> --json` with "Object of type bytes is not JSON
serializable". Apply _lossy_text in the other from_row constructors.

Test: BLOB comment body and event payload -> str with U+FFFD and
JSON-serialisable (red before).
2026-09-15 18:35:09 -07:00
teknium1 a647c6cb2d fix(kanban): one undecodable task field no longer breaks the whole board listing
A tasks row whose TEXT body holds invalid UTF-8 made sqlite3 raise
"Could not decode to UTF-8 column 'body'" inside fetchall, so `hermes kanban
list` (and `show`, and every other reader of that row) failed board-wide
until the row was deleted by hand; a BLOB-typed body came back as bytes and
crashed `--json` (#111743).

Fix it once at the connection: every board connection (`_open_configured`
and the read-only descendant path in `connect`) installs a lossy
text_factory that substitutes U+FFFD, and `Task.from_row` runs BLOB cells
through the same helper so a corrupt row renders with replacement
characters instead of taking its neighbours down.

Fixes #111743
2026-09-15 18:35:09 -07:00
teknium1 85d4415bed fix(kanban): request_review shares complete_task's live-worker fence
The PR docstring said complete_task applies "the same fence request_review
applies", but the two disagreed: complete_task keyed on a live worker
process while request_review still refused any running task with a
claim_lock, so a claim whose worker is gone (or a CLI/library claim that
never spawned one) could be completed but not sent to review without
force. Factor the liveness test into _claim_is_live and use it in both.

TTL expiry is deliberately not part of "live": reclaim_stale_tasks extends
(not reclaims) the claim of a live worker, so the process stays the
liveness authority.

Test: claim -> request_review without a worker PID now succeeds (red
before); a live worker's claim is still refused without expected_run_id.
2026-09-15 18:34:40 -07:00
teknium1 72916de360 fix(kanban): live-claim guard keys on a live worker process, not on any claim
The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
2026-09-15 18:34:40 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1 c7f4bc5bd7 fix(kanban): text dispatch output and both "dispatcher stuck" warnings name the hold reason
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.

- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
  task plus rate_limited / skipped_locked / memory_pressure for one or more
  DispatchResults, so the CLI daemon and gateway warnings share one wording:
  `Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
  tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.

Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>

Part of #111910
2026-09-15 18:34:11 -07:00
KoNit-K ec64ec0d24 fix(kanban): dispatch --json reports respawn_guarded and other suppression reasons
`hermes kanban dispatch --json` only emitted `spawned` and the skip buckets it
already knew about, so a ready card held by the respawn guard (`active_pr`,
`recent_success`, ...), a quota-released worker, a lost dispatch lock or a
memory-pressure hold all looked like `spawned: []` with no reason. Emit
`respawn_guarded`, `rate_limited`, `skipped_locked` and `memory_pressure`
from the DispatchResult the tick already returns.

Salvaged from #111917 by @KoNit-K. Dropped hunk: the `_ACTIVE_PR_RECOVERY_LANES
= frozenset({"review"})` rename in kanban_db_dispatch.py, which is behaviour-
identical to the existing `lane == "review"` check and does not implement the
role-aware exemption the issue asks for.

Part of #111910
2026-09-15 18:34:11 -07:00
teknium1 058e620bef fix(kanban): headless specify/decompose aux calls carry a relay session key
`hermes kanban specify|decompose`, the dashboard specify/decompose routes and
the gateway auto-decomposer all reach the LLM through
hermes_cli/kanban_specify.py::_call_aux outside any agent turn. No
conversation affinity scope is bound there, so agent/opencode_affinity.py
resolved an empty key and sent no `x-opencode-session`; the OpenCode Go relay
rejects such requests with 400 MissingSessionID and the user sees
"Specify failed: LLM error: BadRequestError".

Declare a per-task affinity scope (`kanban:<task_id>`) around the call —
the same host-declared scope the main turn, compression and the
OpenRouter/Portal sticky keys already resolve first — but only when no scope
is bound, so an in-turn caller keeps its conversation's key. Reset in a
finally so nothing leaks past the call.

Live: real httpx transport capture against auxiliary.triage_specifier
provider=opencode-go — before: no x-opencode-session header; after:
`kanban:t_45567533` on specify and decompose, stable per task, distinct per
task; a pre-declared scope is preserved; an openai route gets no header.

Fixes #112043
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:33:43 -07:00
teknium1 7c6b8c7a8d fix(update): Desktop-owned serve no longer fails the update or holds fleet_restart_pending
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.

- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
  reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
  be untrue — the process provably runs old code and the Desktop app does not respawn it after
  a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
  `unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
  that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
  hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
  excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
  gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
  in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
  serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
  origin/main.

Fixes #111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:30:32 -07:00
teknium1 336227bf00 refactor(gateway): move the on-demand s6 slot helper into service_manager, trim tests
Reshape of the cherry-picked fix from #104194:

- The SOUL.md gate + `register_profile_gateway(start_now=False)` now live in
  `hermes_cli/service_manager.py::register_unregistered_profile_gateway`, next to the s6
  manager and `_profile_dir_for_gateway_service` it needs, instead of a private reach-in
  from the 6.5k-line `hermes_cli/gateway.py` facade. The facade only decides "start
  repairs, stop/restart re-raise" and keeps ONE error handler (S6Error is a RuntimeError;
  register's ValueError/RuntimeError/OSError surface as the same `✗` + exit 1).
- Tests trimmed from four to two invariants: start on a real profile registers `down`
  and then starts; stop on an unregistered profile / start on a directory without
  SOUL.md keep the original error and mint nothing (parametrized). Dropped: the
  registration-failure traceback test (covered by the single except clause) and the
  duplicate no-marker/stop split. The test now resolves the profile dir through the real
  HERMES_HOME mapping instead of monkeypatching `_profile_dir_for_gateway_service`.
- Docs: docker.md multi-profile section says `gateway start` inside the container
  registers a slot for a profile created from the host.
2026-09-15 18:30:05 -07:00
John Paul Soliva 9c36cafcec fix(gateway): in-container gateway start registers a missing s6 slot
`profile create` registers an s6 gateway slot only when the creating process is
itself inside the container: `detect_service_manager()` reads `/proc/1/comm` in
the CALLER's PID namespace, so on a host whose `~/.hermes` is bind-mounted into
the container the hook is a silent no-op. The profile directory lands exactly where
the container reads it, but no slot exists, and `hermes -p <name> gateway start`
inside the container fails with "not registered" until the operator restarts the
container so the boot reconciler notices. Same symptom as #54174 reaches `profile
install` by the same route.

Register the slot on demand instead. When `start` hits `GatewayNotRegisteredError`
and the profile directory carries `SOUL.md` — the boot reconciler's own "real
profile" marker — create the slot and start it. That makes
`_maybe_register_gateway_service`'s promise true without a restart.

Deliberately narrow:

- Only `start` self-heals. `stop` and `restart` on an unregistered profile keep the
  original error; registering a slot in order to stop it would be absurd.
- `SOUL.md` gates it, so a mistyped `-p` name or a stray directory cannot mint a
  phantom slot for a profile that does not exist.
- Registers with `start_now=False` and then goes through the ordinary `start` path,
  so the `desired_state` write that lets boot reconciliation restore want-up after
  a container restart keeps a single owner.
- `ValueError` (slot appeared underneath us) and `RuntimeError` (s6-svscanctl
  failed) surface as the existing actionable error, never a traceback.

Refs #54174. PR #54182 fixes the narrower `profile install` case by adding the same
registration call; this closes the symptom for any existing profile directory and
without a container restart.
2026-09-15 18:30:05 -07:00