_set_worker_pid on a fake PID (54321) now persists the 'unverified' marker, under which
the archive path correctly refuses to signal; the test pins the verified-spawn behaviour,
so it stubs the fingerprint capture.
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):
- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
a reboot, and that counter does not, so an unrelated process on a later boot with the
same PID and the same tick value passed _start_times_agree(). The fingerprint is now
"<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
the witness the drain marker already uses); both halves must match. Integer values on
rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
row and falls back to bare PID existence - a new spawn silently recreated the #89614/
#99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
and must not outlive its pid.
Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.
Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
Three P1 findings from the review of #111062 (fixes#110850 remainder):
- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
installed-but-dead unit; `systemd_install` writes the unit before the start that
can still be killed, so an apply interrupted at default/start (or mid-restart of
a stopped default) was reported as "already multiplexed" and never recovered.
`interrupted` is now derived from the manifest vs the LIVE default's recorded
served_profiles: flag on + manifest + not every migrated profile served = resume.
- the compensation `try` began after `_remove_secondary_gateways` and the flag
write, so a later secondary's stop, a unit's daemon-reload or the config write
failing left the first secondary removed with no rollback. The flag write,
removals and default bring-up now all sit inside one boundary that rolls back
through the manifest; `_preflight_apply` refuses failures knowable from the plan
(root system unit without a recorded User=, unresolvable recorded user, an
unwritable config.yaml) before any working gateway is stopped.
- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
`default_uid is None`; a default system unit whose User= this host cannot
resolve is the same boundary from the other side and now blocks the update
hook (same known uid still folds, different known uid still refuses).
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.
Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.
- `hermes profile migrate-identity <old> <new>`: retries the migration —
delegates to the gateway control verb while a multiplexer is live, performs the
durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
naming the offending database on a collision, a lock, or a partial failure. Only
the name format and the existence of the new profile are checked; the old profile
directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
`click`: a failed second database raised NameError instead of printing its warning.
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.
Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.
- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
into config.yaml (--force included); the env writer's denylist
(HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
bridging the value into os.environ; a name neither registered nor in the
environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
and lowercase bare keys are unchanged.
Fixes#111848 (first half landed in #112250).
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.
After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.
Fixes#111804 (its discovery half landed in #112293).
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.
`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.
`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.
This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.
Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.
Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
Follow-up to the ported status fix:
- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
closed wire enum; `mcp.servers.status` would raise `ContractViolation`
on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
contract files.
- `ui-tui` session panel: an unknown status fell through to the red
`failed` branch; render `lazy` with its cached tool count (inline
branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
yields `status: lazy` with the cached tool count and a summary without
`failed` (eager control stays `configured`, live control stays
`connected`); a lazy-only run neither warns nor re-arms the startup
retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
`cli-config.yaml.example`, the MCP config reference and the MCP guide.
Follow-up to the cherry-picked #111584 (@chelsealong):
- website/docs/user-guide/docker.md: new warning block next to the existing
"do not override the entrypoint" note explaining WHY (with `/init` gone the
hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
children), the Compose `init: true` / `docker run --init` remedy, and that
supervision is still lost on that path; plus a Troubleshooting entry for
`<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
check and drops the `platform.system()` gate and the blanket
`try/except Exception: pass` — a user process is never PID 1 on any host OS
(PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).
Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
A deployment that overrides the image's `entrypoint:` to invoke hermes
directly skips docker/entrypoint-dispatch.sh entirely, so hermes itself
becomes PID 1 with no s6-overlay /init (or any other init) above it.
Nothing then reaps orphaned grandchildren (browser tooling, MCP
subprocesses, shell-tool children) reparented to PID 1, and they
accumulate as zombies without bound.
entrypoint-dispatch.sh already warns on its own non-PID-1 fallback
path, but that script never runs in the entrypoint-override case, so
there was no signal at all. Add the same style of warning inside
hermes_cli.main, gated on being PID 1 on Linux, pointing users at the
image's default ENTRYPOINT or `docker run --init` / `init: true`.
Fixes#111577
Review finding: hermes_cli/kanban_db_dispatch.py::_resolve_hermes_argv still resolved which('hermes') before sys.executable -m hermes_cli.main while claiming to mirror gateway.run._resolve_hermes_bin, which this PR made module-first (#111569). Keep the explicit HERMES_BIN override first, then the module argv whenever hermes_cli is importable, PATH only as fallback; docstring updated.
Two diagnostics diverged after a session-only `/model` switch (#111436).
/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).
Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).
Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.
Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.
Fixes#111695
Keep: overlapping cache misses share one probe (never two `gh` at once);
a refresh issued mid-probe after a login flip gets the fresh answer and
still never overlaps. Dropped: the bounded-helper call-shape detector (the
descendant cleanup is proven by the live wrapper probe, not by asserting
the helper's name), and the two tests of pre-existing behaviour (fresh
cache short-circuit, missing `gh`).
`scan_directory` (the PluginManager sweep every CLI/gateway/Desktop backend
runs at startup) and `plugins_cmd._scan_level` (`hermes plugins list`, the
dashboard plugins hub, the TUI plugin picker) probed `plugin.yaml` with
`Path.exists()` outside any error handling. `stat()` raises instead of
returning False when the plugin directory itself is unsearchable — Windows
ACLs (WinError 5, the #111804 report) or a POSIX mode-000 folder — so a
single bad plugin folder took every other plugin down with it and the
Desktop backend exited before announcing its port.
Both scans now warn and skip that one directory, matching the dashboard
manifest scan fixed in the preceding (salvaged) commit.
Part of #111804
Once the launchd file carried the macos_only marker, the first real macOS
lane run showed TestIncompleteWarningMentionsLaunchctl asserting the
systemd-host branch (kickstart hint, systemctl line) that never runs on
macOS. The two Linux-arm tests move to their own linux_only file; the
launchd file keeps the macOS arm (bootstrap/list hint, no systemctl).
With the macos_only marker the win32 skipif is redundant (the marker already
skips every non-darwin host) and the is_macos() fake in _wire contradicts the
repo rule that host-specific behaviour is tested on that host, not by making
the interpreter believe it is elsewhere. Drop both.
_get_service_pids(all_profiles=True) also runs a bare `launchctl list` prefix
scan, so on a macOS box with a live ai.hermes.gateway* fleet the scoping
assertions picked up real PIDs (the one failure the reporter could clear by
booting the gateway out). Stub subprocess.run in _wire so the tests assert on
routing, not on the developer machine.
The cherry-picked tests/ci change-detector (asserting one specific file is in
the macos_only list) is dropped: the selector test suite already covers the
mechanism, and the marker is now the file's declaration.
Two false positives in `hermes plugins validate` that block catalog admission
for plugins that are correct at runtime:
- `kind: model-provider` plugins register at import via
providers.register_provider(ProviderProfile); the PluginManager never calls a
register(ctx) on them (plugins_discovery skips the kind). The probe demanded
register() anyway, so every provider plugin -- including the in-tree
plugins/model-providers/* -- failed with "no register() function". The probe
now records register_provider calls for that kind and fails only when the
import registers nothing.
- RecordingContext returned a no-op callable for ANY attribute, so
`getattr(ctx, "profile_path", None)` was truthy under validation alone and
register() crashed with an error the real PluginContext never produces. The
parent now passes the real PluginContext method names into the probe; other
names raise AttributeError exactly like the real object.
Surfaced by the 2026-09-15 catalog sweep (Gondola provider, hermes-persona).
Fold the four salvaged cases into one predicate truth table (WSL drive
mounts and .cmd shims refused, /mnt/data and /usr/bin accepted) and one
resolver test that walks the real re-scan: PATH interop hands back a
Windows npm first, the scan skips the /mnt/c entry and accepts the native
npm under /mnt/data. Drops the dead `hermes_cli.main._is_windows` patch —
the resolver calls main_install_repair's own `_is_windows`, which is
already False on the POSIX hosts these tests run on.
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.
Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).
Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.
Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.
Fixes#111791
Only Task.from_row coerced BLOB-typed cells; a BLOB task_comments.body (or
event payload / run summary) still came back as bytes and crashed
`hermes kanban show <id> --json` with "Object of type bytes is not JSON
serializable". Apply _lossy_text in the other from_row constructors.
Test: BLOB comment body and event payload -> str with U+FFFD and
JSON-serialisable (red before).
A tasks row whose TEXT body holds invalid UTF-8 made sqlite3 raise
"Could not decode to UTF-8 column 'body'" inside fetchall, so `hermes kanban
list` (and `show`, and every other reader of that row) failed board-wide
until the row was deleted by hand; a BLOB-typed body came back as bytes and
crashed `--json` (#111743).
Fix it once at the connection: every board connection (`_open_configured`
and the read-only descendant path in `connect`) installs a lossy
text_factory that substitutes U+FFFD, and `Task.from_row` runs BLOB cells
through the same helper so a corrupt row renders with replacement
characters instead of taking its neighbours down.
Fixes#111743
The PR docstring said complete_task applies "the same fence request_review
applies", but the two disagreed: complete_task keyed on a live worker
process while request_review still refused any running task with a
claim_lock, so a claim whose worker is gone (or a CLI/library claim that
never spawned one) could be completed but not sent to review without
force. Factor the liveness test into _claim_is_live and use it in both.
TTL expiry is deliberately not part of "live": reclaim_stale_tasks extends
(not reclaims) the claim of a live worker, so the process stays the
liveness authority.
Test: claim -> request_review without a worker PID now succeeds (red
before); a live worker's claim is still refused without expected_run_id.
The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).
Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.
Fixes#111764
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.
- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
task plus rate_limited / skipped_locked / memory_pressure for one or more
DispatchResults, so the CLI daemon and gateway warnings share one wording:
`Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.
Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>
Part of #111910