98 Commits

Author SHA1 Message Date
teknium1 0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
Konstantin Khlopkov 35b1609fc3 fix(kanban): surface the link-time demotion of a ready child to todo
A ready child linked under an unfinished parent drops to todo with no
event and no operator signal; the only trace used to be claim_rejected
after a forced promote. Record a dependency_wait event when the demotion
fires, return the gate from link_tasks, warn in the CLI link command,
report gated in the kanban_link tool, and document the gate.
2026-09-15 06:25:42 -07:00
teknium1 8a8c3634e8 fix(kanban): scope the delegated-child write fence to the lineage's board root
HERMES_DELEGATED_CHILD_CONTEXT=1 is deliberately carried into every shell/
execute_code subprocess a delegate_task child spawns (the fence must survive
exec so a grandchild `hermes kanban complete` cannot promote itself). But the
readers treated the bare flag as "fence every Kanban DB": kanban_db_connect
opened ANY board ?mode=ro and write_txn refused ANY mutation. A subagent
running a Kanban reproduction against a scratch HERMES_HOME therefore got a
silently read-only board with a misleading "descendants require an
initialized board" error; only one lane in the retrospective ever discovered
why (deleg_15dac332), every earlier kanban repro ran degraded.

The marker's value is now the fenced board ROOT (kanban_home() at spawn) and
readers deny only paths under that root or the dispatcher-pinned
HERMES_KANBAN_DB (kanban_path_is_fenced). In-process children and a legacy
"1" marker still fence everything; an inherited path marker is never
re-derived, so a grandchild that moved HERMES_HOME cannot unfence the real
board. Owner-gate tests (test_kanban_descendant_scope, cron env isolation,
kanban CLI exit status) are unchanged and green.
2026-09-15 03:45:41 -07:00
teknium1 8b9df066e2 test(gateway): trim block-loop wording coverage to two invariants; CLI keeps the needs_input distinction
The salvaged parametrized set collapsed to one neutral case (transient) plus
the needs_input positive control. hermes kanban block now mirrors the
notifier: 'needs a human decision' only when the block was typed
needs_input, 'orchestration attention needed' otherwise.
2026-09-15 03:42:00 -07:00
KoNit-K 5c970d9745 fix(kanban): neutral block-loop wording on the CLI, Desktop toast, wake text and docs
The sibling surfaces of the gateway ping rendered the same false claim:
`hermes kanban block` said "needs a human decision", the Desktop toast title
said "needs a decision", the wake status line (locales/*.yaml
gateway.kanban.wake.block_loop_detected) said "needs a decision" and the
docs described the triage route as "for a human decision". A repeated-block
circuit breaker only establishes that orchestration attention is needed.

Surface sweep from PR #111131 (notifier/test hunks dropped in favour of the
typed-kind formatter from PR #111132).
2026-09-15 03:42:00 -07:00
KoNit-K 38bda39562 fix(kanban): preserve Discord thread route anchors 2026-09-14 16:14:33 -07:00
Teknium b7bef04861 fix(kanban): an explicit scratch workspace never inherits the board's project
Move the "explicit scratch means no project" decision into the one resolver
every surface funnels through, `kanban_db.create_task`: board-project
inheritance now runs only when the caller left `workspace_kind` open
(`None`), and `workspace_kind` defaults to scratch after that check. The
tool handler keeps the #106347 fix for `project=""` (no `or` collapse) and
`board=` scoping but drops its handler-local sentinel logic, since the
resolver now owns the rule; the `self_task` project inheritance for
dispatcher-owned workers is unchanged.

Sibling surfaces had the same bug through the same line and are fixed by
the same change:
- CLI `hermes kanban create --workspace scratch` on a project-scoped board
  produced a project worktree; `--workspace` no longer defaults in
  argparse so the resolver can tell "omitted" from "scratch".
- Dashboard `POST /tasks` with `workspace_kind: "scratch"` did the same;
  `CreateTaskBody.workspace_kind` defaults to `None` for the same reason.
- `kanban_swarm.create_swarm` threads `None` through for consistency.

Tests: one resolver invariant in test_kanban_board_project.py (explicit
scratch stays scratch, omitted still inherits) and the salvaged tool test
folded into a single parametrized matrix over scoped/unscoped target boards.
2026-09-09 12:18:58 -07:00
teknium1 02005cfe20 fix(kanban): promote refuses undone parents instead of a false --force success
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.

The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.

- drop `--force` from `promote` (parser, CLI handler, `promote_task`
  kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
  the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
  fake `ready`; the flag no longer parses

Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
2026-09-09 09:21:29 -07:00
Teknium 3b7ff435fd fix(kanban): preserve durable origins for worker-created tasks
Carry the owning task's notification subscriptions independently of dependency
edges, within the creation transaction. Prefer its durable session over worker
and request-local sessions while preserving explicit overrides. Cover worker
CLI create and built-in decomposition, and retain conversation route anchors.
Auto-subscribe no longer upgrades an inherited passive subscription.

Slim adaptation of Christopher-Schulze's session-precedence fix in #85687,
expanded to durable subscription provenance and sibling creation paths.
Related: #85575, #85687

Validation: strict RED/GREEN (7 failing cases before; 7 passing after), then
58 Kanban test files: 383 passed, 2 skipped. Real dispatcher-spawn subprocess
probe covers direct, linked, unlinked, explicit-session, worker CLI, built-in
children and a plain CLI negative control, with recording transport only.

Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
2026-09-07 14:16:57 -07:00
Teknium b578261584 fix: keep Kanban worker scope out of descendant processes
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.

Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.

Refs #103974, #104058, #104904
2026-09-07 07:10:28 -07:00
Teknium ac07da2674 fix(kanban): enforce declared PR acceptance at completion boundary 2026-09-07 04:46:54 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium e3ab65fe80 simplify(compat): kanban_db — drop 73 re-exports/aliases, repoint 794 callers 2026-09-03 13:48:14 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium f0b1d49d0f refactor(hermes_cli): join wrapped kanban call/arg lists; tighten env_loader guard comments (code AST-identical) 2026-09-02 21:55:16 -07:00
Teknium 36a74a3408 refactor(hermes_cli): fold kanban triage-sweep id validation into its driver, flatten small helpers 2026-09-02 20:55:57 -07:00
Teknium ce0eb1c8a2 refactor(hermes_cli): compact kanban CLI handlers (shared field/section printers, commented-op wrapper), no behavior change 2026-09-02 20:28:09 -07:00
Teknium 85c32cfd4e refactor(kanban-cli): compact docstrings/comments in kanban.py keeping every rationale 2026-09-02 16:34:31 -07:00
Teknium 55b8f0e0ab refactor(kanban-cli): extract dispatcher/maintenance verbs to kanban_ops; unify single-mutation ok/err reporting 2026-09-02 16:32:06 -07:00
Teknium 9ebda92cca refactor(kanban-cli): extract boards cluster to kanban_boards; unify goal-gate, id/reason, run-state and config helpers across handlers 2026-09-02 16:29:40 -07:00
Teknium 5cc3651b3b refactor(kanban-cli): extract data-driven argparse tree to kanban_parser and output helpers to kanban_output 2026-09-02 16:25:20 -07:00
Teknium ff660354f3 refactor(kanban): CLI micro-helpers, action dispatch tables, shared triage helpers; active_sessions dedupe
hermes_cli/kanban.py 3,565 -> 2,912; kanban_diagnostics 1,216 -> 996;
kanban_decompose 468 -> 393; kanban_transfer 478 -> 443; kanban_specify
264 -> 229; kanban_swarm 390 -> 378; active_sessions 871 -> 775. `hermes kanban
[sub] --help` byte-identical for all 55 parsers.

- kanban.py: _err / _print_json / _json_out / _fmt_counts / _bulk_apply /
  _obj_dict field tuples replace repeated print/JSON/exit-code blocks; action
  and board subcommand routing via dict dispatch; shared run-state and
  triage-sweep argparse blocks; argparse declarations re-packed (AST-identical).
- specify/decompose: one _run_triage_sweep driver, shared _extract_json_blob /
  _truncate / _profile_author / _title_body / _resolve_profile_from_cfg.
- diagnostics: rule helpers (_first_field / _latest_event_ts / _log_hint_action
  / _error_snippet), _rows_by_task fleet fetch; unreferenced DIAGNOSTIC_KINDS dropped.
- swarm: graph nodes share one create_task kwarg set.
- active_sessions: one _flock per platform, _pid_alive via _pid_liveness,
  shared _read_live_entries / _without_lease / _clean_metadata, table-driven
  strict registry validation.
- Docstrings/comments hand-compacted (AST-identical), invariants kept.
2026-09-02 13:32:14 -07:00
Teknium 5c6cbbc1be fix(loops): pause /loop --until on a blocked verdict; trim redundant gate condition and duplicate test
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
2026-09-02 05:32:01 -07:00
itsflownium 1bd9fce6cb fix(kanban): judge unachievable goals as blocked, never done 2026-09-02 05:32:01 -07:00
Brooklyn Nicholson 3150e444b2 feat(kanban): export and import a whole board as a portable archive
`hermes kanban boards export|import` moves a board between machines:
tasks, comments, links, history, and attachments in one .tar.gz.

Two things make this more than a tar of the board directory. The
database is live — kanban runs in WAL mode, so a filesystem copy loses
whatever still sits in the -wal sidecar and tears if the dispatcher
commits mid-copy; export goes through SQLite's online-backup API
instead. And rows carry machine-local state: claims, worker PIDs,
absolute workspace and attachment paths, session ids, and the gateway
chat ids subscribed to task events. Shipping those verbatim is how an
imported board arrives holding a claim owned by a process on someone
else's laptop, or starts pushing task events into a stranger's Telegram
thread. Everything machine-local is stripped on export and re-stripped
on import, since an archive is untrusted input.

Imports always land as a NEW board, auto-suffixing the slug on
collision, so an import can never merge into or overwrite a board that
is already there. Tasks whose workspace was a directory or git worktree
on the source machine are parked in triage rather than left for the
dispatcher to claim and burn into the failure breaker.
2026-08-28 22:38:01 -05:00
Shannon Sands 4beca7a943 feat(kanban): memory-aware dispatch guard + memory-derived default concurrency cap (OOF-30)
Two production incidents (OOF-77 "larrikin-lollies", OOF-30
"synclare-task-manager") followed the same shape: no
kanban.max_in_progress configured, a busy board, and a 1 GiB hosted VM.
The dispatcher fanned out 26-31 concurrent workers, the host went into
swap-thrash/OOM, and the whole machine — dashboard included — became
unreachable. NAS restart loops then masked the problem: each restart
"recovered" briefly before the kanban dispatcher immediately respawned
unbounded workers.

Building on the cherry-picked max_in_progress-across-both-lanes fix
(PR #28695, credit @Dusk1e), this adds two complementary safeguards to
hermes_cli/kanban_db.py:

1. Memory-DERIVED default concurrency cap. When kanban.max_in_progress
   is unset, resolve_max_in_progress() derives a default of
   clamp(MemTotal / 512 MiB, 2, 8) — e.g. 2 workers on a 1 GiB VM,
   8 on 4 GiB+. Explicit config always wins in either direction. On
   hosts where total memory can't be read (macOS/Windows dev machines),
   the default stays None (no cap — unchanged behaviour). Wired into
   both dispatch entry points (gateway/kanban_watchers.py and
   hermes kanban dispatch) so behaviour matches regardless of path.

2. Live memory-PRESSURE guard inside dispatch_once. A static cap can't
   see the host's actual memory state (other tenants, bloated
   long-lived workers). The dispatcher now samples system memory each
   tick via gateway.lifecycle_ledger.sample_memory() and classifies it
   with gateway.memory_status.classify_pressure() (same thresholds as
   the dashboard memory banner and OOM-suspicion heuristics from
   NS-608/NS-656): critical -> spawn nothing this tick; elevated ->
   at most one new worker; unknown -> no restriction (fail-open).
   Reclaim/promotion bookkeeping still runs under pressure, and
   deferred tasks stay queued — nothing is dropped. Restriction is
   surfaced on DispatchResult.memory_pressure and logged.

Tests: tests/hermes_cli/test_kanban_memory_guard.py (14 tests) covers
the derived cap (floor/ceiling/fail-open/explicit-config-wins), the
pressure classifier, and dispatch behaviour under critical/elevated/
unknown pressure including defer-not-drop and bookkeeping-still-runs.
An autouse fixture in tests/conftest.py pins the memory sample to
"no data" suite-wide so existing dispatch tests don't depend on the
CI runner's live memory state (opt-out marker: real_memory_guard).
2026-08-17 14:16:21 +05:30
Aleksei Razsadin 00c184170b fix(kanban): reap worktree workspaces at task completion and archive
Kanban worktree workspaces were never removed by anything: _cleanup_workspace
preserved them by design, the CLI startup pruner explicitly defers t_* trees
to 'hermes kanban gc', and gc only sweeps scratch — so every worktree task
leaked its checkout forever (measured ~130GB on one estate).

- _cleanup_workspace now dispatches worktree workspaces to a new
  _cleanup_worktree_workspace, which removes the worktree and its
  auto-generated wt/<task-id> branch only when the tree is clean AND every
  commit is reachable from a remote-tracking ref (reusing cli.py's
  _worktree_is_dirty / _worktree_has_unpushed_commits predicates). Any
  doubt preserves the worktree. dir workspaces stay untouched.
- The #33774 active-children deferral now covers worktree parents, and
  _try_cleanup_parent_workspaces reaps deferred worktree parents when the
  last child reaches a terminal state.
- archive_task reaps workspaces too; tasks archived without completing
  previously leaked forever.
- 'hermes kanban gc' gains a backstop sweep for archived worktree tasks
  that predate these hooks.

Tests: tests/hermes_cli/test_kanban_worktree_teardown.py (10 cases: removal,
dirty/unpushed/custom-branch/main-checkout/non-git preservation, complete/
archive integration, deferred-parent handoff).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 21:56:30 -07:00
verybigdog 6e81ce273c feat(kanban): explicit notify/wake delivery modes with faithful wake session routing
Salvage of #37865 by @verybigdog. Adds delivery_mode (notify / notify+wake / wake)
on kanban notify subscriptions, persists chat_type + user_id_alt so a woken turn
reconstructs the creator's real session key, inherits the return path to child
tasks, and keeps wake out of the model-exposed send_message schema.

Original commits were authored under a local placeholder identity
(hermes-agent@users.noreply.local); re-attributed to the contributor's
public email.
2026-08-13 10:47:40 -07:00
cmoiccool 5b4c03fa4b fix(kanban): query show graph before closing database 2026-08-12 13:20:38 +05:30
Teknium 1810cfc8dd fix(kanban): guard request_review against live-claim theft
request_review on a running task under a live claim now requires the
caller to prove ownership (expected_run_id, the unchanged worker path)
or pass an explicit force=True override (CLI --force; dashboard human
actions pass force=True) instead of silently clearing claim_lock /
worker_pid of a live run.

Failures now carry distinct diagnostic reasons via with_reason=True
(mirroring request_changes' tuple pattern): live-claim refusal,
malformed re-review provenance, unsatisfied parents, unknown task, and
CAS miss. Tool/CLI handlers surface the specific reason instead of the
generic 'unknown id or not in running/ready'.

Regression tests: live-claim refusal + force/worker paths; malformed
provenance gets a distinct reason and explicit reviewer= recovers.
2026-08-10 12:43:46 -07:00
Jakub Wolniewicz 0acf49b16f fix(kanban): isolate review handoff ownership 2026-08-10 12:43:46 -07:00
Jakub Wolniewicz 6d7e86c262 fix(kanban): enforce review lifecycle invariants 2026-08-10 12:43:46 -07:00
Jakub Wolniewicz ae23b1f676 fix: complete kanban review lifecycle
Close the autonomous implement-review-rework loop, preserve parent gating and implementer provenance, distinguish downstream review cards, and surface legacy review dependency deadlocks immediately.

Co-authored-by: kaishi00 <6590895+kaishi00@users.noreply.github.com>
2026-08-10 12:43:46 -07:00
Nikita Barkov 16accefd2f feat(kanban): add first-class "review" handoff lifecycle
Add a non-terminal "review" status so a worker that finished implementation
can hand off for human review without abusing kanban_block. The old
kanban_block(reason="review-required: ...") convention routed the handoff
through the unblock-loop breaker, so a normal review -> changes -> review
cycle was falsely escalated to triage.

- kanban_db: request_review (running/ready -> review, non-block, emits
  review_requested), reopen_review_task (review -> ready/todo, review_reopened),
  complete_task accepts review -> done, and a review_dispatch gate (default off,
  shared by the dispatcher loop and the gateway health probe).
- kanban_request_review worker tool + `request-review` / `reopen-review` CLI
  verbs; tool wired through toolsets, EXPOSED_TOOLS, _POLISHED_TOOLS.
- Gateway notifier wakes the origin subscriber on review_requested and
  block_loop_detected; the subscription survives until done/archived, so every
  review cycle re-notifies.
- Dashboard PATCH + bulk route the review transitions (request_review /
  reopen_review_task) and render the review column.
- goals.py goal-loop and KANBAN_GUIDANCE recognize review as a terminator.
- Docs (reference tables, user guide, AGENTS.md, zh-Hans mirrors) + tests.

needs_input / failed are unchanged: they still route through kanban_block,
still count toward block_recurrences, and still escalate to triage.
2026-08-10 12:43:46 -07:00
teknium1 5b751dc0ad chore: remove unused imports and dead locals (ruff F401/F841 sweep)
Cleans F401 unused imports and F841 dead local assignments across
root *.py, agent/, hermes_cli/, tools/, gateway/, cron/, tui_gateway/
(tests/, plugins/, skills/ excluded).

Intentionally KEPT (false positives / test-patch surfaces):
- agent/transports/__init__.py package re-exports
- cli.py browser_connect re-exports (DEFAULT_BROWSER_CDP_URL area,
  used by tests/cli/test_cli_browser_connect.py)
- hermes_cli/main.py _prompt_auth_credentials_choice /
  _model_flow_bedrock_api_key (accessed via main_mod attr in tests)
- gateway/run.py aliased replay_cleanup + whatsapp_identity re-exports
  and _PORT_BINDING_PLATFORM_VALUES (test-referenced)
- hermes_cli/web_server.py get_running_pid (tests monkeypatch it) and
  _OAUTH_TOKEN_URL availability probe
- hermes_cli/config.py get_process_hermes_home re-export (noqa'd F811
  chain) and yaml availability-probe import
- hermes_cli/nous_subscription.py managed_nous_tools_enabled
  (tests patch hermes_cli.nous_subscription.managed_nous_tools_enabled)
- try/except ImportError availability probes (env_loader, tts_tool,
  mcp_tool, web_server anthropic OAuth block)
- tools/web_tools.py noqa F401 re-exports
- hermes_cli/setup_whatsapp_cloud.py:263 'proceed' skipped: possible
  missing-guard bug, flagged for separate review
- unused function parameters (signature changes out of scope)

Side-effect RHS calls preserved where only the binding was dead
(e.g. web_server proc = _spawn_hermes_action -> bare call).
2026-07-29 11:53:39 -07:00
张满良 c03a06b8d9 fix(kanban): cover remaining add_notify_sub call sites for chat_type (#56580)
Follow-up to the main fix in this PR. rodriguez46p-ui's review on the
equivalent #56632 (closed stale) flagged that only the auto-subscribe
path in tools/kanban_tools.py was covered; the same gap existed in two
more call sites:

- gateway/slash_commands.py: the `/kanban create` slash command auto-
  subscribes the calling session but didn't pass chat_type. Read it
  from source.chat_type (already available on SessionSource).
- hermes_cli/kanban.py: the `kanban notify-subscribe` CLI command now
  accepts --chat-type and threads it through.

The dashboard plugin API (plugins/kanban/dashboard/plugin_api.py) still
has the gap because the home_channel config schema doesn't carry
chat_type — that's a follow-up that needs a config schema change.

Verified: 258 tests pass on the kanban + session_context suites.
2026-07-26 13:42:33 -07:00
teknium1 6179da5496 fix(dashboard): one gateway liveness ladder for status + channels
The sidebar strip and the Channels page could contradict each other on
the same page load — "Gateway running" next to "The gateway is not
running." /api/status and /api/messaging/platforms each open-coded their
own liveness ladder: status probed GATEWAY_HEALTH_URL and scoped its
PID/state reads to the requested profile, messaging did neither and used
the uncached raw PID probe.

Three deployments hit the split: a cross-container gateway (no local PID,
only the health probe can see it), a profile-scoped dashboard (messaging
borrowed a DIFFERENT profile's runtime state, reporting a false
"connected" that hides a real outage — #71211), and a launch-service
managed gateway with no PID file.

Adds resolve_gateway_liveness() in gateway/status.py as the single ladder
(cached PID -> HTTP health probe -> runtime-status PID with
expected_home) and routes both endpoints, /api/messaging/platforms/{id}/test,
and the kanban dispatcher-presence probe through it. Probe callables are
injectable so the existing monkeypatch seams keep working, and
GatewayLiveness.probe_error distinguishes "down" from "couldn't tell" so
the kanban warning keeps failing OPEN instead of crying wolf.

Closes #71211.
2026-07-26 08:08:38 -07:00
trkim a7dcf9787b fix(kanban): harden delegated-child mutation boundary 2026-07-23 07:33:36 -07:00
Teknium c1b0f6f3c1 feat(kanban): per-task model dropdown — set/override worker model+provider from the board (#69876)
Adds the missing write path for the per-task model_override column (which
was previously only settable via manual SQL) and pairs it with a
provider_override so cross-provider switches resolve correctly:

- kanban_db: provider_override column (+migration), set_model_override()
  with model_override_set event, create_task(model_override=,
  provider_override=), dispatcher spawns worker with -m <model>
  [--provider <name>]
- dashboard: Model row in the task drawer — dropdown fed by a new
  /model-options endpoint (build_models_payload substrate, provider-grouped,
  free-text fallback), PATCH + bulk model override support
- CLI: kanban create --model/--provider, new kanban set-model subcommand,
  show prints the provider
- agent tools: kanban_create accepts model/provider; show/list expose
  provider_override

Rate-limit recovery flow: override is settable on running tasks and takes
effect on the next dispatch, without touching the worker profile's config.
2026-07-22 22:23:24 -07:00
Teknium 60cfa11136 feat(kanban): add hermes kanban repair CLI verb
Adds kanban_db.repair_db() — a structured, non-raising wrapper around
the same narrow repair policy as the connect-time guard: probe with
PRAGMA integrity_check under the board's cross-process init flock;
quarantine the corrupt bytes FIRST via the content-addressed backup;
REINDEX only when every integrity message is index-scoped; re-check;
report ok / repaired / corrupt / missing. Locked/busy OperationalError
still propagates raw (a locked healthy DB is not corruption and gets
no quarantine), and a repair invalidates the per-process healthy-path
cache so the next connect() re-probes.

The CLI verb reports status human-readably (or --json), exits 0 for
ok/repaired/missing and 1 when the DB is still corrupt (non-index
corruption stays fail-closed with manual-recovery guidance). It
dispatches BEFORE kanban_command's auto-init: init_db() raises
KanbanDbCorruptError on a corrupt board, which previously would have
made a repair verb unreachable on exactly the boards that need it.

CLI tests drive the real argparse surface (build_parser +
kanban_command) against real corrupted SQLite fixtures.
2026-07-21 12:41:14 -07:00
Teknium 369afc60be fix: migrate CLI kanban gate + remaining mocks to 5-value judge contract
Follow-up to the salvaged transport-failure auto-pause (#54387): the PR
branch predates the CLI completion gate merged in #67985, so that new
judge_goal consumer (hermes_cli/kanban.py) needed the 5-value unpack too
— otherwise it would fail open again via the swallowed ValueError.

Also migrates the two remaining 4-value mocks in
tests/cli/test_cli_goal_interrupt.py flagged on the earlier PR #27760.
2026-07-20 05:38:25 -07:00
Teknium 34a304abb3 fix: unpack judge_goal 4-tuple in salvaged CLI gate; harden tests
Follow-up to the cherry-picked CLI judge gate (#55854): the gate carried
the same 3-value unpack bug just fixed on the tool path in PR #67973 —
the ValueError would have been swallowed by the fail-open handler,
silently disabling the gate.

Also: test mock now returns the real 4-value judge contract (the old
3-value mock masked the bug), tests track complete_task invocations and
assert the rejection path never writes, and the unused _make_goal_task
helper is dropped.
2026-07-20 03:36:44 -07:00
srojk34 aa32154e4b fix(kanban): apply goal_mode judge gate to CLI complete command
The three-commit hardening series (Issue #38367, PR #55408) added a
pre-completion judge gate to `tools/kanban_tools.py:_handle_complete`
(the kanban_complete tool used by agent tool-calls).  The structurally
identical `hermes_cli/kanban.py:_cmd_complete` (the `hermes kanban
complete` CLI subcommand) was left unguarded.

A goal_mode worker with terminal tool access — the overwhelming default
for coding agents — can bypass the judge entirely by running:
    hermes kanban complete <task_id>
This transitions the task to `done` status with no judge verdict, making
the acceptance-criteria enforcement worthless on that path.

Fix: apply the same gate in _cmd_complete before calling kb.complete_task.
When a judge is reachable and returns anything other than "done", the
command prints an actionable rejection message and exits non-zero without
modifying the task.  The fail-open policy (no judge configured → allowed)
is preserved to match the tool-call path.
2026-07-20 03:36:44 -07:00
Gille f29c28d6d0 docs(kanban): clarify unblock status routing 2026-07-17 15:47:39 -07:00
otsune 3fccd698fd feat(kanban): attachment toolset + CLI to match the dashboard surface
The kanban board has had full attachment storage and a dashboard HTTP
API (upload/list/download/delete) since #35338, but there was no agent
toolset tool and no `hermes kanban` CLI verb for attachments. Agents and
scripts that don't go through the dashboard server (or can't touch the DB
directly) had no way to create or read real attachments — only links in
comments.

Close that gap by mirroring the existing comment surface:

- `kanban_db.store_attachment_bytes()` — one shared write path (validate
  name, enforce the 25 MB cap, write the blob under the per-task dir with
  collision-free naming, insert the metadata row, clean up an orphan blob
  if the insert fails). `_MAX_ATTACHMENT_BYTES`, `_safe_attachment_name`,
  and a new `_collision_free_path` move here so the dashboard, the tool,
  and the CLI all share one implementation and can't drift.
- Tools (`tools/kanban_tools.py`): `kanban_attach` (inline base64),
  `kanban_attach_url` (server-side http/https fetch with the same cap),
  `kanban_attachments` (list). Write tools respect worker task-ownership;
  list is read-only. Registered in the `kanban` toolset.
- CLI (`hermes_cli/kanban.py`): `attach <id> <path>`, `attachments <id>`,
  `attach-rm <attachment_id>`.
- Dashboard `upload_task_attachment` now imports the shared helpers and
  uses `_collision_free_path` — behavior identical (still streams to disk
  with the cap, still 413 on overflow).
- Docs (AGENTS.md, kanban-worker skill) and toolset membership updated.

Tests: tool round-trip + oversize + bad base64 + ownership; attach_url
against a local HTTP fixture incl. oversize-mid-stream and non-http
scheme rejection; CLI attach/attachments/attach-rm; shared-helper unit
tests; dashboard parity preserved.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 07:33:14 -07:00
Teknium 5b5c79a8ef feat(kanban): typed block reasons + unblock-loop breaker (#52848)
* feat(kanban): typed block reasons + unblock-loop breaker

Stops the kanban blocked-task loop: a worker blocks a task, a cron
unblocks it, the worker re-blocks for the same reason, repeat forever.

block_task now takes a typed kind and a persistent block_recurrences
counter on the tasks table:

- kind=dependency routes to todo (parent-gated, auto-resumed), never
  the human 'blocked' bucket a cron would keep unblocking.
- needs_input/capability/transient/untyped land in blocked; each
  same-cause re-block after an unblock increments block_recurrences,
  and at BLOCK_RECURRENCE_LIMIT (default 2) the task routes to triage
  for a human instead of blocked.
- unblock_task no longer resets block_recurrences (the amnesia that
  let the loop run unbounded); complete_task clears it on success.

Wired through the worker kanban_block tool (new kind arg) and the
hermes kanban block --kind CLI flag, both reporting where the task
actually landed. Docs + 11 new tests; 536 existing kanban tests green.

* test(kanban): make second-block notify test use a distinct block cause

test_notifier_second_blocked_delivers blocked the same task twice with
the same (untyped) reason, which now trips the new unblock-loop breaker
and routes the second block to triage instead of blocked — so only one
'blocked' notification fired. The test's actual intent is that TWO
distinct block cycles each notify; give the two cycles different kinds
(needs_input then capability) so they're genuinely separate blocks. The
same-cause loop→triage path is covered by test_kanban_block_kinds.py.
2026-06-25 21:46:58 -07:00
Brooklyn Nicholson e7811345c1 feat(kanban): link tasks to project worktrees 2026-06-25 16:40:26 -05:00
Teknium 84e1d31e54 refactor(kanban): fold worker/orchestrator skills into injected guidance (#50473)
The kanban-worker and kanban-orchestrator bundled skills existed only to
be force-loaded into dispatcher-spawned workers, gated by
environments:[kanban] so they wouldn't leak into normal CLI listings.
That gating was fragile (the leak that #50443 patched) and the
--skills auto-load was already best-effort — most workers ran without it
because the bundled skill isn't present in profile-scoped skills dirs.

Remove the skills entirely and promote their load-bearing content
(workspace kinds, deliverable artifacts, created-card integrity, profile
discovery) into KANBAN_GUIDANCE, which is already injected into every
kanban worker's system prompt. Net result: every worker reliably gets
the guidance, nothing can leak into a CLI/blank-slate session, and the
gating machinery is gone.

- agent/prompt_builder.py: promote the 4 load-bearing rules into KANBAN_GUIDANCE
- hermes_cli/kanban_db.py: drop --skills kanban-worker auto-injection + _kanban_worker_skill_available probe
- hermes_cli/kanban_swarm.py: drop skills=[kanban-orchestrator] on the root card
- hermes_cli/kanban.py: drop kanban-init skill seeding; fix help text
- delete skills/devops/kanban-{worker,orchestrator}
- docs: delete the two skill pages (EN+zh), fix sidebars/catalog/kanban.md/kanban-worker-lanes.md and the video-orchestrator + codex-lane references
- tests: update spawn-argv expectations; re-bound the guidance-size guard

Supersedes the skill-leak half of #50443 (credit @helix4u for flagging the area).
2026-06-21 17:06:48 -07:00