578f85cfb0
The initial /btw implementation (#97937) answered from a rendered plain-text transcript digest — truncated context, cold-written tokens on every question. Teknium's call: reuse the self-improvement review fork instead, which keeps the entire prompt cache stable for the fork and gives it the complete conversation for very cheap. - agent/background_review.py: extract the review-fork construction into build_cache_parity_fork() — same runtime/credentials as the parent, byte-identical system prompt / tools[] / reasoning config on the same-model path, shared session_id for prefix warmth, full persistence detachment (no state.db writes, no rotation, no external memory, in-place-only compaction). The review thread now calls the helper; behavior unchanged (full review test suite green). - agent/side_question.py: /btw prefers the fork when a live parent AIAgent exists — replays the untruncated snapshot as warm cache reads, denies every tool at dispatch via an empty thread whitelist (tools[] stays byte-identical for cache parity), attributes usage to the parent, and trims a mid-turn snapshot tail so role alternation holds. The one-shot digest remains as fallback (no live agent = cold cache anyway, and any fork failure degrades gracefully). - CLI passes self.agent, TUI passes the session agent, gateway looks up the chat's cached agent (parity with how turns reuse it). Live-verified: /btw on the worktree runs the fork path (agent.log shows the side question as a forked conversation turn on the parent session_id with the full history replayed), answers correctly from context.