feat: /btw rides the background-review cache-parity fork for full-context answers

The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.

- agent/background_review.py: extract the review-fork construction into
  build_cache_parity_fork() — same runtime/credentials as the parent,
  byte-identical system prompt / tools[] / reasoning config on the
  same-model path, shared session_id for prefix warmth, full persistence
  detachment (no state.db writes, no rotation, no external memory,
  in-place-only compaction). The review thread now calls the helper;
  behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
  AIAgent exists — replays the untruncated snapshot as warm cache reads,
  denies every tool at dispatch via an empty thread whitelist (tools[]
  stays byte-identical for cache parity), attributes usage to the parent,
  and trims a mid-turn snapshot tail so role alternation holds. The
  one-shot digest remains as fallback (no live agent = cold cache anyway,
  and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
  the chat's cached agent (parity with how turns reuse it).

Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
This commit is contained in:
Teknium
2026-08-29 08:14:34 -07:00
parent 360761c8cf
commit 578f85cfb0
6 changed files with 628 additions and 292 deletions
+201 -31
View File
@@ -1,19 +1,31 @@
"""Context-aware side questions (``/btw``).
``/btw <question>`` answers a quick question ABOUT the current conversation
without interrupting it: a one-shot auxiliary LLM call receives a read-only
transcript snapshot plus the question, and the answer is delivered alongside
the live session. The live conversation history is never touched — no
without interrupting it. The live conversation history is never touched — no
synthetic turns, no role-alternation risk, no prompt-cache invalidation.
This is deliberately different from ``/bg`` (``/background``'s successor),
which spawns a fresh, contextless agent session for independent work.
Two execution paths, picked automatically:
Model selection rides the standard auxiliary plumbing
(:func:`agent.auxiliary_client.call_llm` via :func:`agent.oneshot.run_oneshot`):
pass ``main_runtime`` to inherit the live session's provider/model; users can
override per-task via ``auxiliary.side_question.provider`` / ``.model`` in
config.yaml.
* **Cache-parity fork (preferred).** When a live parent ``AIAgent`` is
available, the answer comes from a detached fork built by
:func:`agent.background_review.build_cache_parity_fork` — the exact
mechanism the self-improvement background review uses. The fork inherits
the parent's runtime, byte-identical system prompt / ``tools[]`` /
reasoning config, and shared ``session_id``, then replays the parent's
message snapshot verbatim. The provider prefix cache is already warm for
that entire replay, so the fork sees the FULL untruncated conversation at
cache-read prices. Tool calls are denied at dispatch (thread whitelist),
persistence is fully detached, and usage is attributed to the parent.
* **One-shot digest (fallback).** When no live parent exists (e.g. the
gateway evicted the session's cached agent — the provider cache is cold
there anyway), a rendered plain-text transcript snapshot is sent through
one auxiliary :func:`agent.oneshot.run_oneshot` call.
Model selection rides the standard auxiliary plumbing: main model by
default; users can override per-task via ``auxiliary.side_question.provider``
/ ``.model`` in config.yaml (an override routes the fork to that model and
replays a compact digest, since the cache is cold on a different model).
"""
import logging
@@ -25,14 +37,29 @@ logger = logging.getLogger(__name__)
# config.yaml, falls back main-model-first like every other aux task.
SIDE_QUESTION_TASK = "side_question"
# Per-message and total character budgets for the transcript snapshot. The
# snapshot is rendered to plain text (never replayed as raw provider messages)
# so assistant tool_calls entries can't trip provider-side validation on a
# tools-less one-shot request.
# Fork path: the model may waste an iteration attempting a (denied) tool
# call before answering in text; give it a little headroom.
_FORK_MAX_ITERATIONS = 3
# Fallback one-shot path: per-message and total character budgets for the
# rendered transcript snapshot.
_PER_MESSAGE_CHAR_CAP = 2000
_TRANSCRIPT_CHAR_BUDGET = 24000
_INSTRUCTIONS = (
_FORK_PROMPT = (
"The user asked a quick SIDE question with /btw while the main work "
"continues in the original session.\n"
"Rules:\n"
"- Answer ONLY the side question, using the conversation above as "
"context. Do not continue, redo, or critique the main task.\n"
"- Do NOT call any tools — they are disabled for this side question. "
"Answer directly in text.\n"
"- If the conversation does not contain enough information to answer, "
"say so plainly instead of guessing.\n"
"- Be concise and direct."
)
_ONESHOT_INSTRUCTIONS = (
"You are the same AI assistant that is currently working inside the "
"conversation transcribed below. The user has asked a quick SIDE question "
"with /btw while the main work continues.\n"
@@ -63,16 +90,40 @@ def _msg_text(msg: Dict[str, Any]) -> str:
return ""
def trim_snapshot_for_fork(history: Optional[List[Dict[str, Any]]]) -> List[Dict[str, Any]]:
"""Trim a possibly mid-turn snapshot so appending a user message is valid.
A /btw issued while a turn is running can snapshot the transcript in the
middle of a tool loop — ending on an assistant message with unresolved
``tool_calls``, a tool result, or the in-flight user message. Appending
the side question after any of those would violate role alternation on
strict providers. Drop trailing messages until the snapshot ends with a
completed assistant text message. Trimming only the TAIL preserves the
warm prefix-cache property of everything kept.
"""
msgs = list(history or [])
while msgs:
last = msgs[-1]
if not isinstance(last, dict):
msgs.pop()
continue
role = last.get("role")
if role == "assistant" and not last.get("tool_calls"):
break
msgs.pop()
return msgs
def render_history_for_side_question(
history: Optional[List[Dict[str, Any]]],
char_budget: int = _TRANSCRIPT_CHAR_BUDGET,
) -> str:
"""Render a conversation snapshot as a plain-text transcript.
Keeps the most recent messages that fit ``char_budget``, newest-biased
(older context is what gets dropped). Tool calls are summarized by name;
tool results are included truncated so "what did that command output"
style questions remain answerable.
Fallback path only. Keeps the most recent messages that fit
``char_budget``, newest-biased (older context is what gets dropped).
Tool calls are summarized by name; tool results are included truncated
so "what did that command output" style questions remain answerable.
"""
lines: List[str] = []
for msg in history or []:
@@ -119,7 +170,90 @@ def render_history_for_side_question(
return prefix + "\n".join(kept)
def answer_side_question(
def _side_question_task_config() -> Dict[str, Any]:
"""Return ``auxiliary.side_question`` from config (or ``{}``)."""
try:
from hermes_cli.config import load_config_readonly
cfg = load_config_readonly()
except Exception:
return {}
aux = cfg.get("auxiliary", {}) if isinstance(cfg.get("auxiliary"), dict) else {}
task = aux.get(SIDE_QUESTION_TASK, {})
return task if isinstance(task, dict) else {}
def _answer_via_fork(
parent_agent: Any,
question: str,
history: Optional[List[Dict[str, Any]]],
) -> str:
"""Answer via a cache-parity fork of ``parent_agent``.
Runs synchronously on the CALLING thread (all /btw surfaces invoke this
from a worker thread). The thread-scoped tool whitelist is emptied so
any tool call the fork attempts is denied at dispatch — the request's
``tools[]`` stays byte-identical to the parent's for cache parity, but
the side question can never mutate anything.
"""
from agent.background_review import (
_digest_history,
_record_review_usage_to_parent,
_snapshot_review_usage,
build_cache_parity_fork,
)
from hermes_cli.plugins import (
clear_thread_tool_whitelist,
set_thread_tool_whitelist,
)
task_cfg = _side_question_task_config()
fork, _rt, routed = build_cache_parity_fork(
parent_agent,
task_cfg,
max_iterations=_FORK_MAX_ITERATIONS,
write_origin="side_question",
)
try:
set_thread_tool_whitelist(
set(),
deny_msg_fmt=(
"Side question (/btw) denied tool call: {tool_name}. "
"Tools are disabled here — answer directly from the "
"conversation context."
),
)
snapshot = trim_snapshot_for_fork(history)
replay = _digest_history(snapshot) if routed else snapshot
result = fork.run_conversation(
user_message=f"{_FORK_PROMPT}\n\nSide question: {question}",
conversation_history=replay,
)
answer = (result or {}).get("final_response", "") or ""
if not answer and result and result.get("error"):
raise RuntimeError(str(result["error"]))
return answer.strip()
finally:
clear_thread_tool_whitelist()
# Attribute the fork's token usage to the parent session (same
# pattern as the background review, issue #87250). Best-effort.
try:
_record_review_usage_to_parent(
parent_agent, _snapshot_review_usage(fork)
)
except Exception:
pass
try:
fork.shutdown_memory_provider()
except Exception:
pass
try:
fork.close()
except Exception:
pass
def _answer_via_oneshot(
question: str,
history: Optional[List[Dict[str, Any]]],
*,
@@ -128,18 +262,9 @@ def answer_side_question(
temperature: Optional[float] = 0.3,
timeout: float = 180.0,
) -> str:
"""Answer ``question`` against a snapshot of ``history``.
Returns the model's text answer. Raises whatever the auxiliary client
raises (RuntimeError on no provider, etc.) — callers surface the error
on their own UI.
"""
"""Fallback: answer from a rendered transcript digest in one aux call."""
from agent.oneshot import run_oneshot
question = (question or "").strip()
if not question:
raise ValueError("answer_side_question requires a non-empty question")
transcript = render_history_for_side_question(history)
user_input = (
"Conversation transcript (snapshot):\n"
@@ -149,7 +274,7 @@ def answer_side_question(
f"Side question: {question}"
)
return run_oneshot(
instructions=_INSTRUCTIONS,
instructions=_ONESHOT_INSTRUCTIONS,
user_input=user_input,
task=SIDE_QUESTION_TASK,
max_tokens=max_tokens,
@@ -157,3 +282,48 @@ def answer_side_question(
timeout=timeout,
main_runtime=main_runtime,
)
def answer_side_question(
question: str,
history: Optional[List[Dict[str, Any]]],
*,
parent_agent: Any = None,
main_runtime: Optional[Dict[str, Any]] = None,
max_tokens: int = 2048,
temperature: Optional[float] = 0.3,
timeout: float = 180.0,
) -> str:
"""Answer ``question`` against a snapshot of ``history``.
When ``parent_agent`` is a live ``AIAgent``, the answer comes from a
cache-parity fork replaying the full snapshot against the warm provider
prefix cache (see module docstring). Otherwise a one-shot digest call is
used. Raises on failure — callers surface the error on their own UI.
"""
question = (question or "").strip()
if not question:
raise ValueError("answer_side_question requires a non-empty question")
if parent_agent is not None:
try:
answer = _answer_via_fork(parent_agent, question, history)
if answer:
return answer
logger.warning(
"/btw fork returned an empty answer; falling back to one-shot"
)
except Exception:
logger.warning(
"/btw cache-parity fork failed; falling back to one-shot",
exc_info=True,
)
return _answer_via_oneshot(
question,
history,
main_runtime=main_runtime,
max_tokens=max_tokens,
temperature=temperature,
timeout=timeout,
)