Commit Graph

5 Commits

Author SHA1 Message Date
Teknium 9d86ac62b4 test: shrink the goal-DB timing tests to sub-second wall time
Review follow-up on the salvaged #88965 work:

- test_goal_command_slow_db_init_still_persists: drop the 4s slow-init
  loop-gap harness (wall-clock gap assertions on shared CI runners are
  their own flake class; loop-freeze bounds are already covered in
  test_goals_db_bootstrap_off_loop.py). The persistence contract keeps
  its discriminating power by shrinking the monkeypatched init window
  (0.2s) under a 0.8s slow init — the window-only path still fails it.
- test_slow_construction_does_not_block_the_loop: monkeypatch both
  bootstrap windows down (0.3s/0.05s) and shrink the blocking init from
  6s to 1.5s; the two-window contract is what's under test, not the
  production constants. Adds an elapsed ordering assertion so the
  kick-vs-in-flight window distinction stays pinned.

Combined wall time for the pair: ~10.5s -> ~1.4s.
2026-08-18 16:08:29 -07:00
ethernet 246477a809 fix(gateway): /goal no longer lies when state.db init is slow
A fresh state.db init (schema DDL, FTS tables, first config import)
measures ~300ms warm on a fast machine. The gateway constructs
GoalManager on the event-loop thread, and a cold cache ran that init
behind a 0.25s bootstrap grace window: on a slow CI box the /goal set
path's waits expired and save_goal silently no-oped — the reply said
"Goal set (7-turn budget)..." but nothing persisted, and a fresh
GoalManager read back no state (first assertion passes, second fails).

Two changes, one per caller shape:

- Async callers (_get_goal_manager_for_event,
  _get_heartbeat_manager_for_event, _post_turn_goal_continuation, and
  the heartbeat poller) warm the SessionDB cache off-loop through the
  context-preserving executor before constructing the manager (shared
  _warm_goals_session_db helper). The loop never blocks and the first
  write lands at any init duration. A bare to_thread would lose the
  per-turn profile home override under multiplex; the executor hop
  keeps it (same pattern as the goal judge path).
- Sync callers (heartbeat persistence, _goal_still_active_for_session)
  cannot await, so the bootstrap windows stay: the call that starts the
  bootstrap waits a one-time init window (1.5s) instead of the short
  per-call window (0.25s), giving healthy cold inits room to land while
  a contended migration still degrades to None with only a bounded
  one-time stall. The bootstrap thread binds the caller's home as a
  contextvar override so a multiplexed worker cannot cache the default
  profile's DB under another profile's key.

save_goal and heartbeat save_state now log at WARNING when they drop a
write, because the reply has already told the user the state was set.
Regression test pins the contract: init past the window, write
persists, loop gap under 2s (the flake-policy floor for wall-clock
bounds; the slow-init margin grew to match, so the test still tells
on-loop from off-loop).

Independent diagnosis + measurement by jackulau (#88965 review); the
off-loop warm-up shape follows their harness table. Simplify-code
review (4-agent) contributed the helper extraction and the poller
warm-up.
2026-08-18 16:08:29 -07:00
Victor Kyriazakos 8b03e65804 chore: generalize field-report attribution in code comments 2026-08-17 17:20:06 -07:00
Victor Kyriazakos c9dbdbcea7 fix(gateway): grace window keeps first-call goal persistence on healthy DBs
Review finding on the off-loop bootstrap: returning None on every cold-
cache loop-thread call silently dropped the first goal/heartbeat
persistence op even when the DB was perfectly healthy. The loop-thread
path now waits up to 250ms on the bootstrap event - a healthy init
(tens of ms) completes inside the window and the caller gets the real
DB; a contended init (the crash-loop scenario) exceeds it and degrades
to None with a bounded, watchdog-safe stall.
2026-08-17 17:20:06 -07:00
Victor Kyriazakos 8e81e2aaae fix(gateway): never construct SessionDB on the event-loop thread
SessionDB.__init__ runs schema init, and a migration against a contended
state.db blocks for seconds. The goal/heartbeat path reached it
synchronously on the gateway's event-loop thread (GoalManager() ->
load_goal -> _get_session_db -> SessionDB()), so a contended DB starved
the loop-liveness watchdog, which hard-exited with code 75 and the
supervisor restarted straight back into the same state - an unbounded
crash loop reported from an enterprise fleet.

_get_session_db now detects a running loop on the calling thread: on a
cache miss it kicks a one-shot background bootstrap thread and returns
None immediately (every caller already degrades gracefully on None);
the cached instance serves all later calls. Worker threads construct
inline as before, with a lock-guarded cache so a bootstrap race keeps
one instance and closes the loser. The heartbeat module shares this
boundary via the same _get_session_db.
2026-08-17 17:20:06 -07:00