Review fixups for #88965. The goal, loop, and heartbeat managers each
had a copy of the same WARNING text. The shared _warn_dropped_write
helper in goals.py keeps the three logs identical and greppable as one
bug class. The _warm_goals_session_db parameter is now label. The old
name ctx said context, but the value is a log label.
A fresh state.db init (schema DDL, FTS tables, first config import)
measures ~300ms warm on a fast machine. The gateway constructs
GoalManager on the event-loop thread, and a cold cache ran that init
behind a 0.25s bootstrap grace window: on a slow CI box the /goal set
path's waits expired and save_goal silently no-oped — the reply said
"Goal set (7-turn budget)..." but nothing persisted, and a fresh
GoalManager read back no state (first assertion passes, second fails).
Two changes, one per caller shape:
- Async callers (_get_goal_manager_for_event,
_get_heartbeat_manager_for_event, _post_turn_goal_continuation, and
the heartbeat poller) warm the SessionDB cache off-loop through the
context-preserving executor before constructing the manager (shared
_warm_goals_session_db helper). The loop never blocks and the first
write lands at any init duration. A bare to_thread would lose the
per-turn profile home override under multiplex; the executor hop
keeps it (same pattern as the goal judge path).
- Sync callers (heartbeat persistence, _goal_still_active_for_session)
cannot await, so the bootstrap windows stay: the call that starts the
bootstrap waits a one-time init window (1.5s) instead of the short
per-call window (0.25s), giving healthy cold inits room to land while
a contended migration still degrades to None with only a bounded
one-time stall. The bootstrap thread binds the caller's home as a
contextvar override so a multiplexed worker cannot cache the default
profile's DB under another profile's key.
save_goal and heartbeat save_state now log at WARNING when they drop a
write, because the reply has already told the user the state was set.
Regression test pins the contract: init past the window, write
persists, loop gap under 2s (the flake-policy floor for wall-clock
bounds; the slow-init margin grew to match, so the test still tells
on-loop from off-loop).
Independent diagnosis + measurement by jackulau (#88965 review); the
off-loop warm-up shape follows their harness table. Simplify-code
review (4-agent) contributed the helper extraction and the poller
warm-up.
/heartbeat every <interval> <prompt> gives the current session one
recurring instruction. When the session is idle and the interval has
elapsed, the prompt is injected as a plain user turn — same
conversation, same context, prompt cache and role alternation
untouched.
- CLI: idle-poll watchdog thread (wake-word watchdog pattern) feeding
_pending_input; gateway: single gateway-wide async poller injecting
through the adapter FIFO. Busy sessions coalesce their tick to the
next idle poll.
- Missed ticks coalesce (anchor resets on fire) — a busy hour yields
ONE heartbeat turn, never a backlog. Real user messages always win.
- 60s interval floor; injected prompt carries a don't-invent-work
guard so idle heartbeats don't generate busywork.
- State persists in SessionDB.state_meta (heartbeat:<session_id>),
survives /resume, migrates across compression session rotations
alongside /goal state.
- Session-scoped and in-process by design — durable cross-process
schedules remain the cron subsystem's job (docs draw the boundary).
- Slack stays under the 50-slash cap via /hermes heartbeat; ghost-text
suggester now prefers the shortest prefix match so /he still
suggests /help.
Adapted from the session-heartbeat concept in Prime Intellect's
Prime-Agent (/heartbeat).