Commit Graph

35797 Commits

Author SHA1 Message Date
teknium1 e86b6e55ab test(approval): neutralize a host --yolo env in the cancelled-attribution tests
_YOLO_MODE_FROZEN is read from HERMES_YOLO_MODE at import; a host shell running yolo
auto-approved the gate under test so the prompt was never enqueued.
2026-09-15 18:44:46 -07:00
teknium1 6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1 54d7f75590 fix(bot-mode): a pending command approval is not a failed DM delivery
terminal_tool's approval gate answers `status: pending_approval` with an
EMPTY `error` (#28323) and no `session_id`, so _spawn_delivery's specific
branch (`if parsed.get("error")`) was skipped and every unanswered
approval fell through to "Delivery to X failed to start: no process id
returned" — blaming the spawn for an approval nobody in a non-interactive
turn (api_server, `hermes peer dm`, cron) could grant.

- _spawn_delivery: the pending shape gets its own message (the runner
  command needs terminal approval nobody in this turn can grant); a
  local/peer DM adds "nothing was sent — approve it or add it to
  command_allowlist and send again". Ownership is never transferred, so
  the existing finally still reclaims the plaintext DM file.
- _try_relay_delivery: the envelope is queued on disk BEFORE the reply
  waiter spawns and the Desktop drains it independently, so ANY waiter
  spawn failure is a lost wake-up, not a failed delivery; reporting it as
  an error made the sender resend and deliver the message twice. The
  relay path now returns the shape _start_delivery's live-owner branch
  already uses (status queued + notification_error + "Do NOT resend")
  instead of inventing a new status value nothing reads.

Slimmer redo of #92971 by @jonpol01 (same diagnosis, same relay/local
split on `dm_file is None`); the source-text contract test and the
`sent_no_reply_wake` status were dropped.

Fixes #111716

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-15 18:44:23 -07:00
fangliquanflq 27e3fc51ff fix(tools): keep async delegation results past a failed durable write and never prune live records
Two defects in tools/async_delegation.py:

- _push_completion_event called _persist_completion unguarded before
  publishing onto completion_queue. One sqlite3 error (locked/full
  state.db) dropped the completion event, left the record parked on
  "finalizing" (a permanently leaked max_concurrent_children slot) and
  let recover_abandoned_delegations later rewrite a succeeded unit as
  "unknown". The write is now try/except: the failure is logged and the
  event is still delivered, so _finalize flips the status and frees the
  slot. A lost durable row is acceptable degradation; a lost result and
  a leaked slot are not.

- _prune_completed_locked treated anything != "running" as finished,
  while the module's own _LIVE_STATES also names stalling/finalizing.
  A stalling record has no completed_at, so it sorted oldest and was the
  first eviction candidate once the retained cap overflowed; its late
  runner return then hit the missing-record path and the real result was
  dropped. The predicate is now `status not in _LIVE_STATES`.

Slim redo of #76606 (earliest fix) and #112031: the converge/shield/
delete-row machinery both PRs built around the write is dropped as
defense-in-depth; the two core hunks are ported as-is.

Fixes #76605
Fixes #112030
Co-authored-by: luckystar2026 <1393268817@qq.com>
2026-09-15 18:43:56 -07:00
teknium1 4b00842536 fix(web): keyless MCP honours SSE CR terminators and declared charsets
Review follow-up on the keyless Exa UTF-8 fix:

- `_parse_mcp_body` split SSE frames on `\n` only, so a body using bare CR
  line terminators (permitted by the SSE spec, and parsed fine before via
  `splitlines()`) failed with "Unrecognized MCP response shape". Split on
  the SSE terminators CRLF / CR / LF instead.
- `mcp_call` decoded the body as UTF-8 unconditionally, ignoring a charset
  the server did declare. Decode with the declared charset when the
  Content-Type carries one and fall back to UTF-8 only when it is absent.
- The HTTP >= 400 branch and `_keenable_request` still surfaced
  `response.text`, so a non-ASCII error body from a charset-less text/*
  response reached the user as ISO-8859-1 mojibake. Both now go through the
  same decode helper as the success path.
2026-09-15 18:43:28 -07:00
teknium1 28cdd72815 fix(web): keyless Exa decodes its SSE body as UTF-8 so CJK searches work
Exa's MCP endpoint answers `text/event-stream` without a charset, so
`requests` decoded `.text` as ISO-8859-1: every non-ASCII character came back
as mojibake, and for CJK results the UTF-8 continuation byte 0x85 became
U+0085, which `str.splitlines()` treats as a line break — the `data:` JSON
line was cut in two and the perfectly valid result surfaced as "Unrecognized
MCP response shape", after which the ring fell through to keyless Firecrawl
(403). Decode the bytes as UTF-8 (JSON-RPC and SSE are UTF-8 by spec) and
split SSE frames on newlines only. A parsed envelope that really carries no
text is now reported as "no text content" instead of an unrecognized shape.
2026-09-15 18:43:28 -07:00
teknium1 80f76edfaf fix(web): rescue eligibility asks the provider whether the ring was walked
Follow-up to the cherry-picked gateway fix: instead of re-inferring "keyless
mode" from the key env var (wrong for Firecrawl, whose managed-gateway and
self-hosted routes bypass the ring without a key), `_rescue_eligible` asks the
ring vendor's own predicate — `_use_keyless_ring()` for Firecrawl, `use_keyless`
for the others. That covers the persisted `nous` selection the contributor fix
handled AND the legacy never-configured fallback onto a ready gateway, plus
`FIRECRAWL_API_URL`. A ring vendor that actually walked the ring stays
ineligible (its failure means the ring already failed). Docs mention the
gateway route is rescued.
2026-09-15 18:43:05 -07:00
KoNit-K 911eea567f fix(web): rescue failed Nous gateway searches 2026-09-15 18:43:05 -07:00
teknium1 793529467d test(web): make the redirect cache-key test bite on positional pairing
The redirect test used a single requested URL, so main's positional
zip(fetch_urls, results) happened to key it correctly and the test only
pinned. Omit the first requested URL and redirect the second: positional
pairing now files the page under the wrong requested URL, while the
metadata.sourceURL key still lands on the right one.
2026-09-15 18:42:38 -07:00
teknium1 45a4db2225 fix(web): key extract cache on metadata.sourceURL too, pin redirect case
Follow-up to the cherry-picked "cache extracts by returned URL": Keenable and
Firecrawl report the post-redirect address in `url` and the REQUESTED URL in
`metadata.sourceURL`, so matching on `url` alone left every redirected page
uncached. Accept either field, as long as it names a URL from this batch;
anything else is served but never cached (a miss re-fetches, a mis-key poisons
the cache for the whole TTL). Docs: say the cache key is the requested URL the
provider reports, not the batch position.

Co-authored-by: nemofq <5635994+nemofq@users.noreply.github.com>
Co-authored-by: wooyongbin3-cpu <256294002+wooyongbin3-cpu@users.noreply.github.com>
2026-09-15 18:42:38 -07:00
KoNit-K ffd02f37e8 fix(web): cache extracts by returned URL 2026-09-15 18:42:38 -07:00
teknium1 1d14418ea2 fix(agent): concurrent worker survives a dict error result
_detect_tool_failure now classifies dict results as failures, so the
concurrent worker's failure log line sliced result[:200] on a dict and
raised TypeError; the worker died and the model saw "thread did not
return a result" instead of the tool's own error payload. Stringify the
preview like the sequential path does.
2026-09-15 18:42:10 -07:00
teknium1 339fa6d918 fix(gateway): bounded redacted result preview on tool.completed run events
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.

The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.

Salvages #111821 (@KoNit-K), part of #111815.

Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
2026-09-15 18:42:10 -07:00
KoNit-K c89209d564 fix(gateway): report structured tool failures in run events 2026-09-15 18:42:10 -07:00
teknium1 4432bccea8 chore(contributors): map co-author emails for TheNeuralVault and gongyi
Required by scripts/contributor_audit.py --strict for the Co-authored-by
trailers crediting #65182 and #57459 on the daemon-pool 3.14 fix.
2026-09-15 18:41:46 -07:00
teknium1 0262b376b5 test(tools): pin daemon-pool worker arg shape for both stdlib contracts
Replace the cherry-picked 3.14-shape test with two interpreter-agnostic
invariants. The original test did `del pool._initializer` (raises on 3.14,
where the attribute never exists) and drove `_WorkItem.run()` with no
`ctx` (TypeError on 3.14), so it could only ever pass on 3.11-3.13 — the
exact interpreters where the bug does not occur.

The fake worker now records the args tuple and resolves the work item's
future directly, so the tests assert only on the shape the executor picks:
`(ref, ctx, queue)` when `_create_worker_context` exists (3.14+),
`(ref, queue, initializer, initargs)` when the legacy fields do (3.11-3.13).
Both run green on 3.11 and 3.14; the 3.14-shape test is red on main under
either interpreter (`AttributeError: ... no attribute '_initializer'`).

Refs #58596, #111813.
2026-09-15 18:41:46 -07:00
KoNit-K c9fa191334 fix(tools): support daemon pool workers on Python 3.14
Cherry-picked from #111814. The same feature-detected fix was proposed earlier in
#58699, #65182 and #57459 (final form).

Co-authored-by: nankingjing <76432572+nankingjing@users.noreply.github.com>
Co-authored-by: TheNeuralVault <jdkabattles@gmail.com>
Co-authored-by: gongyi <yigongsyl@gmail.com>
2026-09-15 18:41:46 -07:00
teknium1 4c4d54554d docs(kanban): stop-nudge scope covers delegate children and in-process cron
State that the turn-end guard fires only for the dispatcher-owned worker, not
for delegate_task children or cron runs that inherit HERMES_KANBAN_TASK.
2026-09-15 18:41:19 -07:00
teknium1 8e16bde491 test: fold the schema-retry context test into the output-schema module
Reuse tests/tools/test_delegate_output_schema.py's _StubChild instead of a
new one-test file with its own double; the invariant (the retry turn sees
is_delegated_child_context() True and the flag is restored afterwards) is
unchanged. Trim the source comment to the WHY.
2026-09-15 18:41:19 -07:00
moep90 e002cdb92d fix(delegation): run the schema-retry turn in the delegated-child context
_validate_child_output_schema issues a second run_conversation on the child when
the first answer fails the declared output_schema. The main child turn is wrapped
in delegated_child_context; this one was not. It runs on the parent worker's
thread, where HERMES_KANBAN_TASK is set and nothing marks the execution as a
child, so every identity gate keyed on is_delegated_child_context() fails open.

The visible effect is the kanban stop guard: it nudges the child to call
kanban_complete or kanban_block. A child owns no board task and carries no kanban
toolset, so it cannot, and the nudge text ("do not narrate intent", "finish any
remaining deliverable") displaces the structured answer the retry exists to
produce. The retry then fails the same schema and delegate_task reports an error
for a child whose work was already complete.

Observed with four children, each nudged during its retry:

  [subagent-0] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-2] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-3] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-1] Kanban worker tried to exit without kanban_complete/kanban_block
  4/4 - Final answer does not satisfy the declared output_schema (after 1 retry)

Wrap the retry the same way the main turn is wrapped. The context is entered and
exited around the single call, so nothing outside the retry sees it.

Signed-off-by: moep90 <volleyballlive@googlemail.com>
2026-09-15 18:41:19 -07:00
jerryhjones 2bd0f1c59e fix(kanban): scope the stop nudge to the dispatcher-owned worker
agent/kanban_stop.py::kanban_stop_nudge_enabled tested only HERMES_KANBAN_TASK,
which in-process delegate_task children (and cron runs fired inside a worker)
inherit from the worker's process environment. Those executions own no board
task and have the kanban toolset withheld, so the turn-end nudge ordered them to
call a tool they cannot reach — burning attempts, and in production driving
children to complete the parent's card through the CLI.

Gate on agent/delegation_context.py::is_dispatcher_owned_worker_context, the
predicate every other HERMES_KANBAN_* identity gate already uses. The real
worker and the HERMES_KANBAN_STOP_NUDGE opt-out are unchanged.

Salvaged from PR #84656 by @jerryhjones (re-applied onto the current facade
shape); the same gate was first proposed in PR #80023 by @webdevfrancisco
using the narrower delegated-child predicate.

Co-authored-by: webdevfrancisco <franciscombautista2015@gmail.com>
2026-09-15 18:41:19 -07:00
teknium1 fa28f43285 fix(contributors): map dresvyanskiydenis@gmail.com to DresvyanskiyDenis
The salvage commit carries a Co-authored-by trailer for this non-noreply
address but no contributors/emails mapping existed, so release attribution
could not resolve it.
2026-09-15 18:40:52 -07:00
teknium1 b121a02416 fix(delegate): composite parents can grant their included toolsets to children
_expand_parent_toolsets built the parent's tool surface from each
toolset's declared `tools` only, so a composite parent's `includes` were
invisible: a child of a `debugging` parent (terminal/process_manage +
includes web/file) asking for `file` or `web` was refused, and `safe` /
`hermes-gateway` parents could grant nothing but their own name. Same
root cause as the `_strip_blocked_tools` fix in the previous commit
(#111700, "Related" section).

Both sides of the subset check now use the resolved static surface
(`resolve_toolset(name, include_registry=False)`), so a child may request
any toolset whose real tools the parent genuinely holds, and still never
gains a tool the parent lacks. Candidates that resolve to nothing are not
expanded into (they cannot be a meaningful subset).

Co-authored-by: DresvyanskiyDenis <dresvyanskiydenis@gmail.com>
2026-09-15 18:40:52 -07:00
KoNit-K 3d4c4cc24d fix(delegate): retain composite child toolsets 2026-09-15 18:40:52 -07:00
KoNit-K ae5666f7fc fix(tools): quiet expected unavailable toolsets 2026-09-15 18:40:29 -07:00
teknium1 efdf766cad test: trim concurrent multimodal log test to one invariant, reuse the stub
Keep only the invariant the fix owns (an envelope dict logs its serialized
size on the concurrent path); the plain-string control was already covered
by the pre-existing behaviour and doubled the file. Import the AIAgent stub
and fake tool-call shapes from test_start_order_gate.py instead of copying
them, so the concurrent-executor stub has one home.

Part of #112095. Salvage of #112104 (@kokhlo).
2026-09-15 18:40:03 -07:00
Konstantin Khlopkov 699037176c fix(agent): log the serialized size of multimodal results in the concurrent executor
The concurrent completion line logged len(result) directly, so a native-path
vision_analyze envelope dict reported "4 chars" — its key count — while the
sequential path already logs the serialized length. Mirror the sequential
measurement so parallel multimodal calls stop looking truncated in logs.
2026-09-15 18:40:03 -07:00
teknium1 5f9042bec6 fix(plugins): validate accepts model-provider plugins and mirrors the real PluginContext surface
Two false positives in `hermes plugins validate` that block catalog admission
for plugins that are correct at runtime:

- `kind: model-provider` plugins register at import via
  providers.register_provider(ProviderProfile); the PluginManager never calls a
  register(ctx) on them (plugins_discovery skips the kind). The probe demanded
  register() anyway, so every provider plugin -- including the in-tree
  plugins/model-providers/* -- failed with "no register() function". The probe
  now records register_provider calls for that kind and fails only when the
  import registers nothing.
- RecordingContext returned a no-op callable for ANY attribute, so
  `getattr(ctx, "profile_path", None)` was truthy under validation alone and
  register() crashed with an error the real PluginContext never produces. The
  parent now passes the real PluginContext method names into the probe; other
  names raise AttributeError exactly like the real object.

Surfaced by the 2026-09-15 catalog sweep (Gondola provider, hermes-persona).
2026-09-15 18:39:54 -07:00
teknium1 996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
teknium1 40b47b84e7 chore(contributors): map Cr4ckMe email for the #20151 co-author credit 2026-09-15 18:39:09 -07:00
teknium1 decf8e3f26 fix(tools): keep an empty required: [] in sanitized tool schemas
`_sanitize_node` deleted the `required` key whenever the pruned list came
out empty. Four built-in tools (skills_list, todo, delegate_task,
session_search) declare `required: []`, so they left the sanitizer with no
key at all. Strict OpenAI-compatible proxies read the missing key as
`null` and 400 the whole request ("null is not of type array"), which is
non-retryable and kills the session on its first call.

An empty array is valid for every backend; the pruning was added (34c3e67)
to drop names that are not in `properties`, not to delete the key. Keep the
key with the filtered list, even when that list is empty.

Fixes #111684
Fixes #59386
Co-authored-by: Cr4ckMe <jiqing.liu@whu.edu.cn>
2026-09-15 18:39:09 -07:00
Konstantin Khlopkov 0724a6a0fd fix(tools): persist the kill outcome when the reader thread finalises first 2026-09-15 18:38:44 -07:00
teknium1 0f38867a1b docs(dashboard): say lifetime-capping proxies still close the chat socket
The keepalive only defeats idle timeouts; proxies that cap total socket
lifetime (some tunnels) still close it and the chat reattaches on its own.
2026-09-15 18:38:16 -07:00
teknium1 0dba105b9d fix(web): clear stale banner when hidden-tab reconnect is deferred
scheduleReconnect parks the PTY as "closed" while the tab is hidden or the
chat route is inactive, but left any earlier non-rejection banner (e.g. a
failed image upload) in place. maybeReconnectOnPageResume refuses to
reconnect while a banner sits on a closed PTY and the reconnect overlay
hides behind a banner too, so the tab came back disconnected with no way
to recover. Clear the banner in the deferral branch like the normal
reconnect path does.

Test: hidden-tab 1001 close after a failed upload now opens a second
socket on visibilitychange (1 -> 2), red before the fix.
2026-09-15 18:38:16 -07:00
teknium1 a8a36c461b docs(dashboard): describe the chat PTY keepalive and hidden-tab reconnect pause 2026-09-15 18:38:16 -07:00
teknium1 115706aa53 fix(web): release xterm WebGL contexts on reconnect instead of dropping the renderer
Follow-up to the salvaged #111918. That commit fixed the WebGL context
pile-up by removing the WebGL renderer altogether. With @xterm/xterm 6 the
fallback is the DOM renderer, so wide layouts would have lost the crisp,
fast rendering the renderer split in 63975aa deliberately gave them, and
the salvaged comment ("default canvas renderer") described a renderer that
no longer exists.

The leak itself is one missing call: @xterm/addon-webgl 0.19 removes its
canvas on dispose but never calls WEBGL_lose_context.loseContext(), so
every PTY reconnect (which rebuilds the Terminal) leaves a live GL context
until GC. Browsers cap live contexts at ~16 and force-lose the oldest, so a
reconnect storm eventually blanks the terminal the user is looking at.
`loseWebglContexts(host)` runs in the terminal effect's cleanup before
`term.dispose()` and releases every WebGL context under the host; WebGL
stays on for wide layouts exactly as before.

The keepalive is no longer gated on tab visibility / chat activity: an open
PTY socket on a hidden or backgrounded tab still owns its PTY and is the one
most likely to sit idle through a proxy timeout, and the frame is ~20 bytes.
That removes the `shouldSendPtyKeepalive` helper and its two tests; the
socket-open check already lives in `sendTerminalResize`. The ChatPage test
drops its "no WebglAddon constructed" assertion, which passed on base too
(jsdom hosts are 0px wide, so the wide-layout WebGL branch never ran).
2026-09-15 18:38:16 -07:00
KoNit-K b23c4fad36 fix(web): keep chat PTY connections alive 2026-09-15 18:38:16 -07:00
teknium1 af0953985f test(cli): trim npm mount-path coverage to two invariants
Fold the four salvaged cases into one predicate truth table (WSL drive
mounts and .cmd shims refused, /mnt/data and /usr/bin accepted) and one
resolver test that walks the real re-scan: PATH interop hands back a
Windows npm first, the scan skips the /mnt/c entry and accepts the native
npm under /mnt/data. Drops the dead `hermes_cli.main._is_windows` patch —
the resolver calls main_install_repair's own `_is_windows`, which is
already False on the POSIX hosts these tests run on.
2026-09-15 18:35:59 -07:00
KoNit-K addcf7d898 fix(cli): accept native npm paths under mnt 2026-09-15 18:35:59 -07:00
teknium1 bd63866253 fix(kanban): give a finished worker a grace window before the terminal reaper signals it
reap_terminal_workers signalled any retained worker on the first tick after
its run closed, but a healthy worker is still alive for a moment after
kanban_complete / kanban_request_review returns (final assistant turn,
session persistence), so slow-but-healthy workers were killed mid-finalisation
and logged as terminal_worker_reaped. Reap only runs whose ended_at is at
least TERMINAL_WORKER_REAP_GRACE_SECONDS (120 s, two default ticks) old;
the fingerprint check is unchanged. Each row is now handled on its own so a
signal or /proc failure on one run is logged and skips only that run.

Tests: a just-closed run is not signalled and keeps its evidence, then is
reaped once the grace has passed (red before); one raising row no longer
aborts the sweep for the others (red before).
2026-09-15 18:35:32 -07:00
teknium1 aa5817d9be fix(kanban): reap workers that outlive their finished run
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).

Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.

Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.

Fixes #111791
2026-09-15 18:35:32 -07:00
teknium1 4465b8d7d5 fix(kanban): BLOB cells in comment/event/run rows degrade like a BLOB task body
Only Task.from_row coerced BLOB-typed cells; a BLOB task_comments.body (or
event payload / run summary) still came back as bytes and crashed
`hermes kanban show <id> --json` with "Object of type bytes is not JSON
serializable". Apply _lossy_text in the other from_row constructors.

Test: BLOB comment body and event payload -> str with U+FFFD and
JSON-serialisable (red before).
2026-09-15 18:35:09 -07:00
teknium1 a647c6cb2d fix(kanban): one undecodable task field no longer breaks the whole board listing
A tasks row whose TEXT body holds invalid UTF-8 made sqlite3 raise
"Could not decode to UTF-8 column 'body'" inside fetchall, so `hermes kanban
list` (and `show`, and every other reader of that row) failed board-wide
until the row was deleted by hand; a BLOB-typed body came back as bytes and
crashed `--json` (#111743).

Fix it once at the connection: every board connection (`_open_configured`
and the read-only descendant path in `connect`) installs a lossy
text_factory that substitutes U+FFFD, and `Task.from_row` runs BLOB cells
through the same helper so a corrupt row renders with replacement
characters instead of taking its neighbours down.

Fixes #111743
2026-09-15 18:35:09 -07:00
teknium1 85d4415bed fix(kanban): request_review shares complete_task's live-worker fence
The PR docstring said complete_task applies "the same fence request_review
applies", but the two disagreed: complete_task keyed on a live worker
process while request_review still refused any running task with a
claim_lock, so a claim whose worker is gone (or a CLI/library claim that
never spawned one) could be completed but not sent to review without
force. Factor the liveness test into _claim_is_live and use it in both.

TTL expiry is deliberately not part of "live": reclaim_stale_tasks extends
(not reclaims) the claim of a live worker, so the process stays the
liveness authority.

Test: claim -> request_review without a worker PID now succeeds (red
before); a live worker's claim is still refused without expected_run_id.
2026-09-15 18:34:40 -07:00
teknium1 72916de360 fix(kanban): live-claim guard keys on a live worker process, not on any claim
The first cut refused every claim-less complete of a running+claimed card,
which also refused the flows that have no worker to protect: a library or
CLI claim that never spawned a worker, and a worker whose process is gone
(12 sibling tests exercise exactly that shape). The guard now fires only
when tasks.worker_pid names a process that is still alive under its spawn
fingerprint (_worker_alive), which is the run the issue asked us to keep
open. Test updated to stand in as the live worker via _set_worker_pid.
2026-09-15 18:34:40 -07:00
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1 c7f4bc5bd7 fix(kanban): text dispatch output and both "dispatcher stuck" warnings name the hold reason
`hermes kanban dispatch` (plain output), the standalone daemon's stuck warning
and the gateway's embedded dispatcher stuck warning all reported a bare
`Spawned: 0` / "0 workers spawned" while the respawn guard held every ready
card — the reason existed only as a `respawn_guarded` task event visible via
`hermes kanban tail`. Operators watching the gateway health warning for 73+
ticks (#111910) had nothing to act on.

- `kanban_db_dispatch.describe_suppression()` renders the guard reasons per
  task plus rate_limited / skipped_locked / memory_pressure for one or more
  DispatchResults, so the CLI daemon and gateway warnings share one wording:
  `Last tick held back: active_pr=1, memory_pressure=elevated.`
- plain `dispatch` output prints `Guarded (<reason>): <task id>` and the
  tick-level holds, mirroring the JSON fields.
- kanban docs: how to see why a ready card is not spawning.

Co-authored-by: Steven Saehrig <trac3r726@users.noreply.github.com>

Part of #111910
2026-09-15 18:34:11 -07:00
KoNit-K ec64ec0d24 fix(kanban): dispatch --json reports respawn_guarded and other suppression reasons
`hermes kanban dispatch --json` only emitted `spawned` and the skip buckets it
already knew about, so a ready card held by the respawn guard (`active_pr`,
`recent_success`, ...), a quota-released worker, a lost dispatch lock or a
memory-pressure hold all looked like `spawned: []` with no reason. Emit
`respawn_guarded`, `rate_limited`, `skipped_locked` and `memory_pressure`
from the DispatchResult the tick already returns.

Salvaged from #111917 by @KoNit-K. Dropped hunk: the `_ACTIVE_PR_RECOVERY_LANES
= frozenset({"review"})` rename in kanban_db_dispatch.py, which is behaviour-
identical to the existing `lane == "review"` check and does not implement the
role-aware exemption the issue asks for.

Part of #111910
2026-09-15 18:34:11 -07:00
teknium1 f7ea39481a fix(kanban): dashboard estimate calls declare a relay-affinity key too
The dashboard's estimate endpoints make the same headless auxiliary call
as specify/decompose but never bound an affinity scope, so they still sent
no x-opencode-session and the OpenCode Go relay answered 400
MissingSessionID (#112043). Declare kanban:<task_id> for an existing task
and a stable kanban:estimate key for the create dialog (no task yet),
unless a scope is already bound.

Test: _run_estimate captured header None before; now kanban:t_1 /
kanban:estimate and nothing leaks past the call.
2026-09-15 18:33:43 -07:00
teknium1 1a8d922003 docs(providers): note the per-task x-opencode-session key for headless Kanban aux calls 2026-09-15 18:33:43 -07:00