Files
hermes-agent/tests
Shannon Sands 4beca7a943 feat(kanban): memory-aware dispatch guard + memory-derived default concurrency cap (OOF-30)
Two production incidents (OOF-77 "larrikin-lollies", OOF-30
"synclare-task-manager") followed the same shape: no
kanban.max_in_progress configured, a busy board, and a 1 GiB hosted VM.
The dispatcher fanned out 26-31 concurrent workers, the host went into
swap-thrash/OOM, and the whole machine — dashboard included — became
unreachable. NAS restart loops then masked the problem: each restart
"recovered" briefly before the kanban dispatcher immediately respawned
unbounded workers.

Building on the cherry-picked max_in_progress-across-both-lanes fix
(PR #28695, credit @Dusk1e), this adds two complementary safeguards to
hermes_cli/kanban_db.py:

1. Memory-DERIVED default concurrency cap. When kanban.max_in_progress
   is unset, resolve_max_in_progress() derives a default of
   clamp(MemTotal / 512 MiB, 2, 8) — e.g. 2 workers on a 1 GiB VM,
   8 on 4 GiB+. Explicit config always wins in either direction. On
   hosts where total memory can't be read (macOS/Windows dev machines),
   the default stays None (no cap — unchanged behaviour). Wired into
   both dispatch entry points (gateway/kanban_watchers.py and
   hermes kanban dispatch) so behaviour matches regardless of path.

2. Live memory-PRESSURE guard inside dispatch_once. A static cap can't
   see the host's actual memory state (other tenants, bloated
   long-lived workers). The dispatcher now samples system memory each
   tick via gateway.lifecycle_ledger.sample_memory() and classifies it
   with gateway.memory_status.classify_pressure() (same thresholds as
   the dashboard memory banner and OOM-suspicion heuristics from
   NS-608/NS-656): critical -> spawn nothing this tick; elevated ->
   at most one new worker; unknown -> no restriction (fail-open).
   Reclaim/promotion bookkeeping still runs under pressure, and
   deferred tasks stay queued — nothing is dropped. Restriction is
   surfaced on DispatchResult.memory_pressure and logged.

Tests: tests/hermes_cli/test_kanban_memory_guard.py (14 tests) covers
the derived cap (floor/ceiling/fail-open/explicit-config-wins), the
pressure classifier, and dispatch behaviour under critical/elevated/
unknown pressure including defer-not-drop and bookkeeping-still-runs.
An autouse fixture in tests/conftest.py pins the memory sample to
"no data" suite-wide so existing dispatch tests don't depend on the
CI runner's live memory state (opt-out marker: real_memory_guard).
2026-08-17 14:16:21 +05:30
..