Review follow-up on the salvaged #88965 work:
- test_goal_command_slow_db_init_still_persists: drop the 4s slow-init
loop-gap harness (wall-clock gap assertions on shared CI runners are
their own flake class; loop-freeze bounds are already covered in
test_goals_db_bootstrap_off_loop.py). The persistence contract keeps
its discriminating power by shrinking the monkeypatched init window
(0.2s) under a 0.8s slow init — the window-only path still fails it.
- test_slow_construction_does_not_block_the_loop: monkeypatch both
bootstrap windows down (0.3s/0.05s) and shrink the blocking init from
6s to 1.5s; the two-window contract is what's under test, not the
production constants. Adds an elapsed ordering assertion so the
kick-vs-in-flight window distinction stays pinned.
Combined wall time for the pair: ~10.5s -> ~1.4s.
A fresh state.db init (schema DDL, FTS tables, first config import)
measures ~300ms warm on a fast machine. The gateway constructs
GoalManager on the event-loop thread, and a cold cache ran that init
behind a 0.25s bootstrap grace window: on a slow CI box the /goal set
path's waits expired and save_goal silently no-oped — the reply said
"Goal set (7-turn budget)..." but nothing persisted, and a fresh
GoalManager read back no state (first assertion passes, second fails).
Two changes, one per caller shape:
- Async callers (_get_goal_manager_for_event,
_get_heartbeat_manager_for_event, _post_turn_goal_continuation, and
the heartbeat poller) warm the SessionDB cache off-loop through the
context-preserving executor before constructing the manager (shared
_warm_goals_session_db helper). The loop never blocks and the first
write lands at any init duration. A bare to_thread would lose the
per-turn profile home override under multiplex; the executor hop
keeps it (same pattern as the goal judge path).
- Sync callers (heartbeat persistence, _goal_still_active_for_session)
cannot await, so the bootstrap windows stay: the call that starts the
bootstrap waits a one-time init window (1.5s) instead of the short
per-call window (0.25s), giving healthy cold inits room to land while
a contended migration still degrades to None with only a bounded
one-time stall. The bootstrap thread binds the caller's home as a
contextvar override so a multiplexed worker cannot cache the default
profile's DB under another profile's key.
save_goal and heartbeat save_state now log at WARNING when they drop a
write, because the reply has already told the user the state was set.
Regression test pins the contract: init past the window, write
persists, loop gap under 2s (the flake-policy floor for wall-clock
bounds; the slow-init margin grew to match, so the test still tells
on-loop from off-loop).
Independent diagnosis + measurement by jackulau (#88965 review); the
off-loop warm-up shape follows their harness table. Simplify-code
review (4-agent) contributed the helper extraction and the poller
warm-up.
Review finding on the off-loop bootstrap: returning None on every cold-
cache loop-thread call silently dropped the first goal/heartbeat
persistence op even when the DB was perfectly healthy. The loop-thread
path now waits up to 250ms on the bootstrap event - a healthy init
(tens of ms) completes inside the window and the caller gets the real
DB; a contended init (the crash-loop scenario) exceeds it and degrades
to None with a bounded, watchdog-safe stall.
SessionDB.__init__ runs schema init, and a migration against a contended
state.db blocks for seconds. The goal/heartbeat path reached it
synchronously on the gateway's event-loop thread (GoalManager() ->
load_goal -> _get_session_db -> SessionDB()), so a contended DB starved
the loop-liveness watchdog, which hard-exited with code 75 and the
supervisor restarted straight back into the same state - an unbounded
crash loop reported from an enterprise fleet.
_get_session_db now detects a running loop on the calling thread: on a
cache miss it kicks a one-shot background bootstrap thread and returns
None immediately (every caller already degrades gracefully on None);
the cached instance serves all later calls. Worker threads construct
inline as before, with a lock-guarded cache so a bootstrap race keeps
one instance and closes the loser. The heartbeat module shares this
boundary via the same _get_session_db.