Commit Graph

8 Commits

Author SHA1 Message Date
Teknium ff660354f3 refactor(kanban): CLI micro-helpers, action dispatch tables, shared triage helpers; active_sessions dedupe
hermes_cli/kanban.py 3,565 -> 2,912; kanban_diagnostics 1,216 -> 996;
kanban_decompose 468 -> 393; kanban_transfer 478 -> 443; kanban_specify
264 -> 229; kanban_swarm 390 -> 378; active_sessions 871 -> 775. `hermes kanban
[sub] --help` byte-identical for all 55 parsers.

- kanban.py: _err / _print_json / _json_out / _fmt_counts / _bulk_apply /
  _obj_dict field tuples replace repeated print/JSON/exit-code blocks; action
  and board subcommand routing via dict dispatch; shared run-state and
  triage-sweep argparse blocks; argparse declarations re-packed (AST-identical).
- specify/decompose: one _run_triage_sweep driver, shared _extract_json_blob /
  _truncate / _profile_author / _title_body / _resolve_profile_from_cfg.
- diagnostics: rule helpers (_first_field / _latest_event_ts / _log_hint_action
  / _error_snippet), _rows_by_task fleet fetch; unreferenced DIAGNOSTIC_KINDS dropped.
- swarm: graph nodes share one create_task kwarg set.
- active_sessions: one _flock per platform, _pid_alive via _pid_liveness,
  shared _read_live_entries / _without_lease / _clean_metadata, table-driven
  strict registry validation.
- Docstrings/comments hand-compacted (AST-identical), invariants kept.
2026-09-02 13:32:14 -07:00
Teknium 5505042f40 fix(sessions): fail closed on ownership uncertainty and fence every turn source
Follow-ups on top of the #94595 cherry-pick, implementing the maintainer
review's two blockers:

Blocker 1 (turn-admission chokepoint): _run_prompt_submit itself now runs
the ownership admission, so synthesized turns that never pass through the
prompt.submit RPC handler (crash auto-continue from cold session.resume,
wake-ups) are fenced too. Auto-continue additionally checks ownership
BEFORE emitting message.start and leaves the marker in place, closing the
#94778 shape where backend B resumed a session backend A was actively
running and auto-continued A's fresh interrupted-turn marker into a
duplicate concurrent turn.

Blocker 2 (fail-closed registry semantics): try_acquire_active_session no
longer converts an unreadable/corrupt registry into an untracked go-ahead.
Ownership uncertainty is a distinct typed refusal —
SESSION_COORDINATION_UNAVAILABLE — because "could not prove ownership" must
never be collapsed into "no owner exists". The TUI gateway claim helper
fails closed on claim exceptions for every surface, not just desktop.

Also: empty session ids short-circuit to a no-op lease (nothing to fence,
and the strict registry schema rejects empty ids), and the existing
fail-open tests were updated to assert the new fail-closed contract.
2026-08-31 12:36:33 -07:00
Futahua a5f0fbb262 fix(sessions): per-session exclusivity is correctness, not a capacity policy
Cherry-picked from PR #94595 (author: Futahua) onto current main, with the
maintainer-review revision points folded in during the rebase:

- the lease engages UNCONDITIONALLY: try_acquire_active_session no longer
  returns a disabled no-op lease when max_concurrent_sessions is unset;
  the concurrency cap stays an orthogonal, optional policy checked second
- ownership uncertainty fails CLOSED (SESSION_COORDINATION_UNAVAILABLE)
  instead of degrading to an untracked go-ahead: a corrupt/unreadable
  registry must not be collapsed into 'no owner exists' (review blocker 2)
- the ownership admission sits at the _run_prompt_submit chokepoint that
  EVERY fresh turn source crosses, and crash auto-continue acquires (or
  bails) BEFORE emitting message.start — closing the #94778 bypass where
  backend B's auto-continue ran a duplicate turn while backend A was live
  (review blocker 1)
- the TUI gateway claim helper fails closed on claim exceptions for every
  surface, not just desktop
- CLI and messaging-gateway call sites pass live_session_id metadata so
  the (pid, live id) re-entrancy identity protects them from self-fencing
  on a leaked lease

Co-authored-by: teknium1 <teknium1@users.noreply.github.com>
2026-08-31 12:36:33 -07:00
Brooklyn Nicholson 51e67babca fix(cli): keep Desktop liveness leases when the session cap is off
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.

Co-authored-by: metamindedu <metamind@kakao.com>
2026-08-27 11:50:05 -05:00
Teknium 94ce8396e8 fix(sessions): release active-session leases against their acquisition registry
A gateway active-session lease is acquired against the root HERMES_HOME,
but release_active_session()/transfer_active_session() re-resolved the
registry path from the *current* HERMES_HOME. Under native multiplex a
routed turn runs agent cleanup inside _profile_runtime_scope, so the
release looked under the named profile while the root entry stayed
alive — after max_concurrent_sessions routed turns every new session was
rejected with 'Hermes is at the active session limit' (#85431).

Pin state/lock paths on the lease at acquisition time and prefer them on
release and transfer. Fixes #85431.
2026-08-15 00:33:01 -07:00
Brooklyn Nicholson e35c2f6049 fix(sessions): claim the cap slot on first turn, not on open
An open chat window took a session-cap slot at session.create/resume time.
Every desktop tile paint and every background reconnect-resume opens one, so
on a websocket-flappy host they accumulated: five parked desktop tabs filled a
5-slot cap and locked the messaging gateway (which shares the cap) out for
fourteen minutes while running no agents at all.

A slot held that way is invisible everywhere. An unprompted draft has no DB row
and the sidebar filters it out with min_messages=1, so the only way to diagnose
it was reading runtime/active_sessions.json by hand.

Claim on the first turn instead, mirroring the lazy contract
_ensure_session_db_row already uses for the row itself. Capacity now means an
agent can run rather than that a window exists, and anything holding a slot is
something the user can see.

Also reclaim leases whose session skipped teardown. _prune_dead only fires when
the owning pid dies, and a dashboard/serve backend runs for days, so a leaked
lease was held until restart. The owning process reconciles against the leases
it still holds, which is exact and needs no heartbeat write on the turn path.
2026-07-28 12:11:52 -05:00
konsisumer 02050859f3 fix(tui): preserve live session identity across compression (#49041)
When a session rotates id on compression, _sync_session_key_after_compress()
re-anchored the session_key, approval-notify routing, yolo state, and slash
worker — but never moved the active-session lease, which stayed keyed to the
pre-compression id. And _find_live_session_by_key() matched live sessions on
the stale session_key, not the live agent's current agent.session_id. After
compression a resume/create path failed to recognize the existing live agent
and could build a SECOND live agent against the same DB continuation -> forked
lineage / cross-session message mixing.

- active_sessions.transfer_active_session(): move a lease in place to the new
  id under the exclusive file lock (no slot drop).
- gateway _transfer_active_session_slot(): call it inside
  _sync_session_key_after_compress(); on the rare fallback (entry pruned)
  RESERVE the new slot before releasing the old lease (reserve-before-release),
  so a concurrent gateway at the session cap cannot grab the freed slot in a
  release-then-reacquire window and leave this session with no lease; if the
  reserve fails, keep the existing lease (review fix).
- _session_lookup_key(): make live-session lookup authoritative on
  agent.session_id, wired into all stale-session_key consumers
  (_find_live_session_by_key, _session_live_item, _live_session_payload) —
  fixes the whole lookup class.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
2026-06-24 00:54:18 +05:30
Robin Fernandes 639c1e3636 feat(sessions): add optional max session cap 2026-06-08 15:12:12 -07:00