feat(kanban): memory-aware dispatch guard + memory-derived default concurrency cap (OOF-30)
Two production incidents (OOF-77 "larrikin-lollies", OOF-30 "synclare-task-manager") followed the same shape: no kanban.max_in_progress configured, a busy board, and a 1 GiB hosted VM. The dispatcher fanned out 26-31 concurrent workers, the host went into swap-thrash/OOM, and the whole machine — dashboard included — became unreachable. NAS restart loops then masked the problem: each restart "recovered" briefly before the kanban dispatcher immediately respawned unbounded workers. Building on the cherry-picked max_in_progress-across-both-lanes fix (PR #28695, credit @Dusk1e), this adds two complementary safeguards to hermes_cli/kanban_db.py: 1. Memory-DERIVED default concurrency cap. When kanban.max_in_progress is unset, resolve_max_in_progress() derives a default of clamp(MemTotal / 512 MiB, 2, 8) — e.g. 2 workers on a 1 GiB VM, 8 on 4 GiB+. Explicit config always wins in either direction. On hosts where total memory can't be read (macOS/Windows dev machines), the default stays None (no cap — unchanged behaviour). Wired into both dispatch entry points (gateway/kanban_watchers.py and hermes kanban dispatch) so behaviour matches regardless of path. 2. Live memory-PRESSURE guard inside dispatch_once. A static cap can't see the host's actual memory state (other tenants, bloated long-lived workers). The dispatcher now samples system memory each tick via gateway.lifecycle_ledger.sample_memory() and classifies it with gateway.memory_status.classify_pressure() (same thresholds as the dashboard memory banner and OOM-suspicion heuristics from NS-608/NS-656): critical -> spawn nothing this tick; elevated -> at most one new worker; unknown -> no restriction (fail-open). Reclaim/promotion bookkeeping still runs under pressure, and deferred tasks stay queued — nothing is dropped. Restriction is surfaced on DispatchResult.memory_pressure and logged. Tests: tests/hermes_cli/test_kanban_memory_guard.py (14 tests) covers the derived cap (floor/ceiling/fail-open/explicit-config-wins), the pressure classifier, and dispatch behaviour under critical/elevated/ unknown pressure including defer-not-drop and bookkeeping-still-runs. An autouse fixture in tests/conftest.py pins the memory sample to "no data" suite-wide so existing dispatch tests don't depend on the CI runner's live memory state (opt-out marker: real_memory_guard).
This commit is contained in:
@@ -2643,6 +2643,10 @@ def _cmd_dispatch(args: argparse.Namespace) -> int:
|
||||
_kanban_cfg.get("max_in_progress_per_profile")
|
||||
)
|
||||
max_in_progress = _coerce_positive_int(_kanban_cfg.get("max_in_progress"))
|
||||
# Memory-derived default when unset (OOF-30/OOF-77) — same
|
||||
# fallback the gateway-embedded dispatcher applies, so behaviour
|
||||
# matches regardless of which path runs the tick.
|
||||
max_in_progress = kb.resolve_max_in_progress(max_in_progress)
|
||||
# CLI --max overrides config kanban.max_spawn when both are present;
|
||||
# CLI is the more explicit signal so it wins.
|
||||
cli_max = getattr(args, "max", None)
|
||||
|
||||
Reference in New Issue
Block a user