refactor(agent/runtime): hooks/guardrails/dispatch — unify shell-hook and webhook plumbing, table-driven guardrail thresholds

- shell_hooks is the shared home: _ToolMatcherMixin (matcher compile + matches_tool),
  _payload_fields, _forget_home_registrations, _home_key, _utc_now_iso now serve
  outbound_webhooks too (copies deleted; every log string byte-identical).
- shell_hooks: response parsing is a per-event dispatch table; _spawn diagnostic
  dict + _evaluate_result shared by the live callback and run_once;
  _locked_update_approvals POSIX/non-POSIX bodies merged via ExitStack.
- tool_guardrails: ToolCallGuardrailConfig thresholds from a _THRESHOLD_SOURCES
  table (nested-wins-over-flat preserved); _int_at_least replaces
  _positive_int/_non_negative_int; observe_identical_call (0 refs) folded into
  observe_call; _halt helper for hard-stop decisions.
- tool_dispatch_helpers: _plan_tool_batch_segments split into _batch_admission +
  close/extend helpers with the post-hoc normalization merged in.
- Comment/docstring compaction keeping every stated rule.
This commit is contained in:
Teknium
2026-09-02 10:36:07 -07:00
parent 2ad284b473
commit 4bfc55fbf0
5 changed files with 847 additions and 1901 deletions
+205 -457
View File
@@ -1,26 +1,12 @@
"""Tool-dispatch helpers — parallelism gating, multimodal envelopes, mutation tracking.
Pure module-level utilities extracted from ``run_agent.py``:
* ``_is_destructive_command`` — terminal-command heuristic used to gate
parallel batch dispatch.
* ``_should_parallelize_tool_batch`` / ``_extract_parallel_scope_paths`` /
``_extract_parallel_scope_path`` / ``_paths_overlap`` — the rules engine
deciding when a multi-tool batch can run concurrently (V4A patch scope
uses patch-body file headers, not a decoy ``path=``).
* ``_is_multimodal_tool_result`` / ``_multimodal_text_summary`` /
``_append_subdir_hint_to_multimodal`` — envelope helpers for the
``{"_multimodal": True, "content": [...], "text_summary": ...}`` dict
shape returned by tools like ``computer_use``.
* ``_extract_file_mutation_targets`` / ``_extract_landed_file_mutation_paths`` /
``_extract_error_preview`` —
per-turn file-mutation verifier inputs.
* ``_trajectory_normalize_msg`` — strip image blobs from a message for
trajectory saving.
All helpers are stateless. ``run_agent`` re-exports each name so existing
``from run_agent import ...`` imports in tests and other modules keep
working unchanged.
Stateless module-level utilities extracted from ``run_agent.py``, which
re-exports each name so existing ``from run_agent import ...`` imports keep
working. Groups: batch-parallelism planner (path-overlap admission; V4A patch
scope comes from patch-body headers, not a decoy ``path=``), multimodal
``{"_multimodal": True, "content": [...], "text_summary": ...}`` envelope
helpers, per-turn file-mutation verifier inputs, trajectory normalisation, and
the tool-result message constructor with its untrusted-content wrapping.
"""
from __future__ import annotations
@@ -40,8 +26,7 @@ from tools.threat_patterns import scan_for_threats
logger = logging.getLogger(__name__)
# Tools that must never run concurrently (interactive / user-facing).
# When any of these appear in a batch, we fall back to sequential execution.
# Interactive / user-facing tools never run concurrently: any of these in a batch is a barrier.
_NEVER_PARALLEL_TOOLS = frozenset({"clarify"})
# Read-only tools with no shared mutable session state.
@@ -60,16 +45,12 @@ _PARALLEL_SAFE_TOOLS = frozenset({
"web_search",
})
# Filesystem tools whose parallel admission is decided by path overlap.
# Readers may share a subtree with other readers; a writer conflicts with
# ANY overlapping reservation (reader or writer). This is what keeps a
# batched ``search_files``/``read_file`` from observing pre-mutation file
# state when the model batches it alongside the ``patch``/``write_file``
# it depends on (the classic same-block write→read race).
# Filesystem tools admitted by path overlap. Readers may share a subtree; a
# writer conflicts with ANY overlapping reservation. This keeps a batched
# read_file/search_files from observing pre-mutation state when the model
# batches it alongside the patch/write_file it depends on.
_PATH_SCOPED_READERS = frozenset({"read_file", "search_files"})
_PATH_SCOPED_WRITERS = frozenset({"write_file", "patch"})
# File tools can run concurrently when they target independent paths.
_PATH_SCOPED_TOOLS = _PATH_SCOPED_READERS | _PATH_SCOPED_WRITERS
# Patterns that indicate a terminal command may modify/delete files.
@@ -92,47 +73,29 @@ _REDIRECT_OVERWRITE = re.compile(r'[^>]>[^>]|^>[^>]')
def _is_destructive_command(cmd: str) -> bool:
"""Heuristic: does this terminal command look like it modifies/deletes files?"""
if not cmd:
return False
if _DESTRUCTIVE_PATTERNS.search(cmd):
return True
if _REDIRECT_OVERWRITE.search(cmd):
return True
return False
return bool(cmd) and bool(_DESTRUCTIVE_PATTERNS.search(cmd) or _REDIRECT_OVERWRITE.search(cmd))
def _is_mcp_tool_parallel_safe(tool_name: str) -> bool:
"""Check if an MCP tool comes from a server with parallel tool calls enabled.
Lazy-imports from ``tools.mcp_tool`` to avoid circular dependencies.
Returns False if the MCP module is not available.
"""
"""Whether an MCP tool's server opted into parallel calls; False if MCP is unavailable."""
try:
from tools.mcp_tool import is_mcp_tool_parallel_safe
from tools.mcp_tool import is_mcp_tool_parallel_safe # lazy: avoids import cycle
return is_mcp_tool_parallel_safe(tool_name)
except Exception:
return False
# Read-only bridge lookups: dispatch_tool_search / dispatch_tool_describe are
# stateless catalog reads (the catalog is rebuilt from the current tool-defs
# list on every call), so a batch of them can run concurrently.
# Stateless catalog reads (rebuilt from the current tool-defs on every call) — parallel-safe.
_PARALLEL_SAFE_BRIDGE_LOOKUPS = frozenset({"tool_search", "tool_describe"})
def _peel_bridge_call(tool_name: str, function_args: dict) -> tuple[str, dict]:
"""Resolve a ``tool_call`` bridge invocation to its underlying tool.
"""Resolve a ``tool_call`` bridge invocation to ``(underlying_name, underlying_args)``.
The batch planner admits calls to a parallel run by tool NAME, but when
tool search is active the model emits the literal name ``tool_call`` for
every deferred tool — so a server opted in via
``supports_parallel_tool_calls: true`` silently lost concurrency the
moment the bridge activated. Peel the wrapper here so admission is
decided on the underlying tool, exactly like the executors' unwrap.
Returns ``(underlying_name, underlying_args)`` when the wrapper parses
cleanly, else ``(tool_name, function_args)`` unchanged — an unparseable
bridge call stays a sequential barrier and fails at dispatch as before.
With tool search active the model emits the literal name ``tool_call`` for
every deferred tool, so admission must be decided on the underlying tool
(as the executors' unwrap does). An unparseable bridge call is returned
unchanged: it stays a sequential barrier and fails at dispatch as before.
"""
try:
from tools.tool_search import TOOL_CALL_NAME, resolve_underlying_call
@@ -146,147 +109,97 @@ def _peel_bridge_call(tool_name: str, function_args: dict) -> tuple[str, dict]:
return tool_name, function_args
def _batch_admission(tool_call, execution_cwd: Optional[Path]) -> tuple[str, List[Path], bool] | None:
"""Classify one call for the planner: ``None`` = sequential barrier, else
``(effective_name, scoped_paths, is_writer)`` (empty paths = unscoped parallel-safe)."""
tool_name = tool_call.function.name
if tool_name in _NEVER_PARALLEL_TOOLS:
return None
try:
function_args = json.loads(tool_call.function.arguments)
except Exception:
_raw = tool_call.function.arguments
logging.debug(
"Could not parse args for %s — treating as sequential barrier; raw=%s",
tool_name,
_raw[:200] if isinstance(_raw, str) else repr(_raw)[:200],
)
return None
if not isinstance(function_args, dict):
logging.debug(
"Non-dict args for %s (%s) — treating as sequential barrier",
tool_name,
type(function_args).__name__,
)
return None
name, args = _peel_bridge_call(tool_name, function_args)
if name in _NEVER_PARALLEL_TOOLS:
return None
if name in _PATH_SCOPED_TOOLS:
scoped = _extract_parallel_scope_paths(name, args, execution_cwd=execution_cwd)
return (name, scoped, name in _PATH_SCOPED_WRITERS) if scoped else None
if name in _PARALLEL_SAFE_TOOLS or name in _PARALLEL_SAFE_BRIDGE_LOOKUPS or _is_mcp_tool_parallel_safe(name):
return name, [], False
return None
def _plan_tool_batch_segments(tool_calls, *, execution_cwd: Optional[Path] = None) -> List[tuple]:
"""Split a tool-call batch into ordered ``(kind, calls)`` segments.
"""Split a tool-call batch into ordered ``("parallel"|"sequential", calls)`` segments.
``kind`` is ``"parallel"`` (a maximal contiguous run of parallel-safe
calls) or ``"sequential"`` (one or more barrier calls that must run
in-order on the sequential path). Segments preserve the model's
original call order exactly — a later call never crosses an earlier
barrier — so tool-result ordering and side-effect boundaries are
identical to fully-sequential execution. The per-call safety rules
are the same ones the old all-or-nothing gate applied to the whole
batch:
* ``_NEVER_PARALLEL_TOOLS`` (interactive tools) → barrier.
* Unparseable / non-dict arguments → barrier.
* Path-scoped tools (``read_file``/``search_files``/``write_file``/
``patch``) join a parallel run only when their target path(s) do not
CONFLICT with a path already reserved in the same run. Reservations
carry a reader/writer role: reader↔reader overlap is harmless (two
reads of the same file commute) and stays parallel; any overlap
involving a writer closes the run so the conflicting call starts a
NEW run after the first completes. ``search_files`` reserves its
search root (default ``.``) as a reader — a search batched after a
write into the searched subtree is ordered behind that write instead
of racing it. For V4A ``patch(mode="patch")`` the reserved paths are
the file headers in the patch body, not a possibly-stale ``path=``
argument.
* Anything not in ``_PARALLEL_SAFE_TOOLS`` and not an opted-in MCP
tool → barrier.
Parallel runs shorter than two calls are demoted to sequential (no
concurrency win, and the sequential executor owns the richer inline
dispatch), and adjacent sequential segments are merged.
Segments preserve the model's call order exactly — a later call never
crosses an earlier barrier — so result ordering and side-effect boundaries
match fully-sequential execution. Barriers: ``_NEVER_PARALLEL_TOOLS``,
unparseable/non-dict args, and anything not parallel-safe (built-in list,
bridge lookups, opted-in MCP tools). Path-scoped tools join a run only when
their paths don't conflict with the run's reservations: reader↔reader
overlap commutes and stays parallel; any overlap involving a writer closes
the run so the call starts a NEW run after the conflicting one lands.
``search_files`` reserves its root (default ``.``) as a reader. Parallel
runs shorter than two calls demote to sequential (the sequential executor
owns the richer inline dispatch); adjacent sequential segments merge.
"""
segments: list[list] = [] # [kind, calls] pairs, normalized to tuples on return
segments: List[tuple] = []
current: list = []
# (canonical_path, is_writer) reservations for the current parallel run.
reserved_paths: list[tuple[Path, bool]] = []
reserved_paths: list[tuple[Path, bool]] = [] # (canonical_path, is_writer) for the current run
def _close_parallel() -> None:
nonlocal current, reserved_paths
if current:
segments.append(["parallel", current])
current = []
reserved_paths = []
if len(current) >= 2:
segments.append(("parallel", current))
elif current:
_extend_sequential(current)
current = []
reserved_paths = []
def _add_sequential(tc) -> None:
_close_parallel()
def _extend_sequential(calls: list) -> None:
if segments and segments[-1][0] == "sequential":
segments[-1][1].append(tc)
segments[-1][1].extend(calls)
else:
segments.append(["sequential", [tc]])
segments.append(("sequential", list(calls)))
for tool_call in tool_calls:
tool_name = tool_call.function.name
if tool_name in _NEVER_PARALLEL_TOOLS:
_add_sequential(tool_call)
admission = _batch_admission(tool_call, execution_cwd)
if admission is None:
_close_parallel()
_extend_sequential([tool_call])
continue
try:
function_args = json.loads(tool_call.function.arguments)
except Exception:
_raw = tool_call.function.arguments
logging.debug(
"Could not parse args for %s — treating as sequential barrier; raw=%s",
tool_name,
_raw[:200] if isinstance(_raw, str) else repr(_raw)[:200],
)
_add_sequential(tool_call)
continue
if not isinstance(function_args, dict):
logging.debug(
"Non-dict args for %s (%s) — treating as sequential barrier",
tool_name,
type(function_args).__name__,
)
_add_sequential(tool_call)
continue
# Bridge unwrap: admission is decided on the UNDERLYING tool, not on
# the literal wrapper name the model emitted. Read-only bridge
# lookups (tool_search / tool_describe) are parallel-safe as-is.
effective_name, effective_args = _peel_bridge_call(tool_name, function_args)
if effective_name in _NEVER_PARALLEL_TOOLS:
_add_sequential(tool_call)
continue
if effective_name in _PATH_SCOPED_TOOLS:
scoped_paths = _extract_parallel_scope_paths(
effective_name, effective_args, execution_cwd=execution_cwd
)
if not scoped_paths:
_add_sequential(tool_call)
continue
is_writer = effective_name in _PATH_SCOPED_WRITERS
if any(
(is_writer or existing_is_writer)
and _paths_overlap(scoped_path, existing)
for scoped_path in scoped_paths
for existing, existing_is_writer in reserved_paths
):
# Same-subtree conflict inside this run: close it so this
# call starts a fresh run AFTER the conflicting one lands.
# Reader↔reader overlap never conflicts — concurrent reads
# of the same subtree commute.
_close_parallel()
reserved_paths.extend((p, is_writer) for p in scoped_paths)
current.append(tool_call)
continue
if (
effective_name in _PARALLEL_SAFE_TOOLS
or effective_name in _PARALLEL_SAFE_BRIDGE_LOOKUPS
or _is_mcp_tool_parallel_safe(effective_name)
_name, scoped_paths, is_writer = admission
if any(
(is_writer or existing_is_writer) and _paths_overlap(scoped_path, existing)
for scoped_path in scoped_paths
for existing, existing_is_writer in reserved_paths
):
current.append(tool_call)
continue
_add_sequential(tool_call)
_close_parallel()
reserved_paths.extend((p, is_writer) for p in scoped_paths)
current.append(tool_call)
_close_parallel()
normalized: list[list] = []
for kind, calls in segments:
if kind == "parallel" and len(calls) < 2:
kind = "sequential"
if normalized and normalized[-1][0] == "sequential" and kind == "sequential":
normalized[-1][1].extend(calls)
else:
normalized.append([kind, calls])
return [(kind, calls) for kind, calls in normalized]
return segments
def _should_parallelize_tool_batch(tool_calls) -> bool:
"""Return True when the WHOLE tool-call batch is safe to run concurrently.
Thin view over ``_plan_tool_batch_segments`` kept for callers/tests that
only care about the homogeneous case: True iff the planner produces a
single all-parallel segment.
"""
"""True iff the planner yields a single all-parallel segment for the WHOLE batch."""
if len(tool_calls) <= 1:
return False
segments = _plan_tool_batch_segments(tool_calls)
@@ -294,19 +207,13 @@ def _should_parallelize_tool_batch(tool_calls) -> bool:
def _canonical_path(raw_path: str, execution_cwd: Optional[Path] = None) -> Path:
"""Return a canonical, OS-aware path for overlap detection.
Uses ``os.path.realpath`` to resolve symlinks on existing path components
and ``os.path.normcase`` for case-insensitive platforms (Windows).
Falls back to ``Path.cwd()`` when *execution_cwd* is not supplied.
"""
"""Canonical, OS-aware path for overlap detection (realpath for symlinks on
existing components, normcase for case-insensitive platforms); relative
paths resolve against *execution_cwd* or ``Path.cwd()``."""
expanded = Path(raw_path).expanduser()
base = execution_cwd if execution_cwd is not None else Path.cwd()
candidate = expanded if expanded.is_absolute() else base / expanded
# realpath resolves symlinks on path components that exist; for
# not-yet-created files it canonicalises as far as possible.
resolved = os.path.normcase(os.path.realpath(os.path.abspath(str(candidate))))
return Path(resolved)
return Path(os.path.normcase(os.path.realpath(os.path.abspath(str(candidate)))))
def _extract_parallel_scope_paths(
@@ -314,17 +221,12 @@ def _extract_parallel_scope_paths(
function_args: dict,
execution_cwd: Optional[Path] = None,
) -> List[Path]:
"""Return every canonical path this call reserves for overlap checks.
"""Every canonical path this call reserves for overlap checks.
*execution_cwd* should be the working directory that the tool will
actually use at runtime. When omitted the process cwd is used,
which may differ from the tool execution environment on some
platforms (e.g. WSL, sandboxed sub-processes).
For ``patch`` in V4A ``mode=patch``, scope comes from patch-body
``*** Update/Add/Delete/Move File:`` headers (not a possibly-decoy
``path=``). An empty result means the planner cannot determine the
scope and must treat the call as a sequential barrier.
*execution_cwd* should be the cwd the tool will actually use (may differ
from the process cwd on WSL / sandboxed backends). For V4A ``patch`` the
scope comes from patch-body file headers. An empty result means the scope
is unknown and the planner must treat the call as a sequential barrier.
"""
if tool_name not in _PATH_SCOPED_TOOLS:
return []
@@ -337,25 +239,15 @@ def _extract_parallel_scope_paths(
if isinstance(raw_path, str) and raw_path.strip():
raw_paths.append(raw_path)
elif tool_name == "search_files":
# ``search_files`` defaults its search root to the cwd when
# ``path`` is omitted — reserve that root rather than falling
# back to a sequential barrier (an empty result here would
# demote every bare search to a barrier and destroy read
# parallelism).
# search_files defaults its root to the cwd; reserve that rather than
# demoting every bare search to a barrier.
raw_paths.append(".")
scoped: List[Path] = []
seen: set[str] = set()
for raw in raw_paths:
if not isinstance(raw, str) or not raw.strip():
continue
canonical = _canonical_path(raw, execution_cwd)
key = str(canonical)
if key in seen:
continue
seen.add(key)
scoped.append(canonical)
return scoped
# dict.fromkeys dedupes while preserving first-seen order.
return list(dict.fromkeys(
_canonical_path(raw, execution_cwd)
for raw in raw_paths if isinstance(raw, str) and raw.strip()
))
def _extract_parallel_scope_path(
@@ -363,41 +255,24 @@ def _extract_parallel_scope_path(
function_args: dict,
execution_cwd: Optional[Path] = None,
) -> Optional[Path]:
"""Return the primary canonical file target for path-scoped tools.
Thin view over ``_extract_parallel_scope_paths`` kept for callers/tests
that only need a single representative path. For multi-file V4A
patches this is the first header target.
"""
scoped = _extract_parallel_scope_paths(
tool_name, function_args, execution_cwd=execution_cwd
)
"""Primary canonical target (first header target for multi-file V4A patches), or None."""
scoped = _extract_parallel_scope_paths(tool_name, function_args, execution_cwd=execution_cwd)
return scoped[0] if scoped else None
def _paths_overlap(left: Path, right: Path) -> bool:
"""Return True when two paths may refer to the same subtree.
Both *left* and *right* must already be canonical (as returned by
``_extract_parallel_scope_paths`` / ``_canonical_path``) so that
symlink aliases and case differences are already normalised.
"""
"""True when two already-canonical paths may refer to the same subtree."""
left_parts = left.parts
right_parts = right.parts
if not left_parts or not right_parts:
# Empty paths shouldn't reach here (guarded upstream), but be safe.
# Empty paths are guarded upstream; only two non-empty equal prefixes overlap.
return bool(left_parts) == bool(right_parts) and bool(left_parts)
common_len = min(len(left_parts), len(right_parts))
return left_parts[:common_len] == right_parts[:common_len]
def _is_multimodal_tool_result(value: Any) -> bool:
"""True if the value is a multimodal tool result envelope.
Multimodal handlers (e.g. tools/computer_use) return a dict with
`_multimodal=True`, a `content` key holding OpenAI-style content
parts, and an optional `text_summary` for string-only fallbacks.
"""
"""True for the multimodal envelope: dict with ``_multimodal=True`` and a ``content`` list."""
return (
isinstance(value, dict)
and value.get("_multimodal") is True
@@ -405,23 +280,17 @@ def _is_multimodal_tool_result(value: Any) -> bool:
)
def _multimodal_text_summary(value: Any) -> str:
"""Extract a plain text view of a multimodal tool result.
def _is_text_part(p: Any) -> bool:
return isinstance(p, dict) and p.get("type") == "text"
Used wherever downstream code needs a string — logging, previews,
persistence size heuristics, fall-back content for providers that
don't support multipart tool messages.
"""
def _multimodal_text_summary(value: Any) -> str:
"""Plain-text view of a tool result (logging, previews, string-only providers)."""
if _is_multimodal_tool_result(value):
if value.get("text_summary"):
return str(value["text_summary"])
parts = []
for p in value.get("content") or []:
if isinstance(p, dict) and p.get("type") == "text":
parts.append(str(p.get("text", "")))
if parts:
return "\n".join(parts)
return "[multimodal tool result]"
parts = [str(p.get("text", "")) for p in value.get("content") or [] if _is_text_part(p)]
return "\n".join(parts) if parts else "[multimodal tool result]"
if isinstance(value, str):
return value
try:
@@ -431,17 +300,12 @@ def _multimodal_text_summary(value: Any) -> str:
def _append_subdir_hint_to_multimodal(value: Dict[str, Any], hint: str) -> None:
"""Mutate a multimodal tool-result envelope to append a subdir hint.
The hint is added to the first text part so the model sees it; image
parts are left untouched. `text_summary` is also updated for
string-fallback callers.
"""
"""Append a subdir hint to the envelope's first text part (and ``text_summary``) in place."""
if not _is_multimodal_tool_result(value):
return
parts = value.get("content") or []
for p in parts:
if isinstance(p, dict) and p.get("type") == "text":
if _is_text_part(p):
p["text"] = str(p.get("text", "")) + hint
break
else:
@@ -451,52 +315,34 @@ def _append_subdir_hint_to_multimodal(value: Dict[str, Any], hint: str) -> None:
value["text_summary"] = value["text_summary"] + hint
def _extract_file_mutation_targets(tool_name: str, args: Dict[str, Any]) -> List[str]:
"""Return the file paths a ``write_file`` or ``patch`` call is targeting.
# ``\s*`` (not ``\s+``) after ``***`` matches patch_parser / file_tools, which
# accept ``***Update File:`` with no space.
_V4A_FILE_HEADER = re.compile(r'^\*\*\*\s*(?:Update|Add|Delete)\s+File:\s*(.+)$', re.MULTILINE)
_V4A_MOVE_HEADER = re.compile(r'^\*\*\*\s*Move\s+File:\s*(.+?)\s*->\s*(.+)$', re.MULTILINE)
For ``write_file`` and ``patch`` in replace mode this is just ``args["path"]``.
For ``patch`` in V4A patch mode we parse the patch content for
``*** Update File:`` / ``*** Add File:`` / ``*** Delete File:`` headers so
the verifier can track each file in a multi-file patch separately.
def _extract_file_mutation_targets(tool_name: str, args: Dict[str, Any]) -> List[str]:
"""File paths a ``write_file`` / ``patch`` call targets.
Replace mode uses ``args["path"]``; V4A patch mode parses the
``*** Update/Add/Delete/Move File:`` headers so each file in a multi-file
patch is tracked separately.
"""
if tool_name not in _FILE_MUTATING_TOOLS:
return []
if tool_name == "write_file":
p = args.get("path")
return [str(p)] if p else []
# tool_name == "patch"
mode = args.get("mode") or "replace"
mode = "replace" if tool_name == "write_file" else (args.get("mode") or "replace")
if mode == "replace":
p = args.get("path")
return [str(p)] if p else []
if mode == "patch":
body = args.get("patch") or ""
if not isinstance(body, str) or not body:
return []
paths: List[str] = []
# ``\s*`` (not ``\s+``) after ``***`` matches patch_parser / file_tools:
# they accept ``***Update File:`` with no space after the asterisks.
for _m in re.finditer(
r'^\*\*\*\s*(?:Update|Add|Delete)\s+File:\s*(.+)$',
body,
re.MULTILINE,
):
p = _m.group(1).strip()
if p:
paths.append(p)
for _m in re.finditer(
r'^\*\*\*\s*Move\s+File:\s*(.+?)\s*->\s*(.+)$',
body,
re.MULTILINE,
):
src = _m.group(1).strip()
dst = _m.group(2).strip()
if src:
paths.append(src)
if dst:
paths.append(dst)
return paths
return []
if mode != "patch":
return []
body = args.get("patch") or ""
if not isinstance(body, str) or not body:
return []
paths = [m.group(1).strip() for m in _V4A_FILE_HEADER.finditer(body)]
for m in _V4A_MOVE_HEADER.finditer(body):
paths.extend((m.group(1).strip(), m.group(2).strip()))
return [p for p in paths if p]
def _extract_landed_file_mutation_paths(
@@ -504,7 +350,8 @@ def _extract_landed_file_mutation_paths(
args: Dict[str, Any],
result: Any,
) -> List[str]:
"""Return the concrete file paths a successful mutation reports."""
"""Concrete file paths a successful mutation reports (``files_modified`` /
``resolved_path`` in the JSON result), falling back to the declared targets."""
targets = _extract_file_mutation_targets(tool_name, args)
if tool_name not in _FILE_MUTATING_TOOLS or not isinstance(result, str):
return targets
@@ -516,28 +363,17 @@ def _extract_landed_file_mutation_paths(
return targets
files = data.get("files_modified")
if isinstance(files, list):
landed = [str(p) for p in files if p]
if landed:
return landed
landed = [str(p) for p in files if p] if isinstance(files, list) else []
if landed:
return landed
resolved = data.get("resolved_path")
if resolved:
return [str(resolved)]
return targets
return [str(resolved)] if resolved else targets
def _extract_error_preview(result: Any, max_len: int = 180) -> str:
"""Pull a one-line error summary out of a tool result for footer display."""
"""One-line error summary of a tool result for footer display."""
text = _multimodal_text_summary(result) if result is not None else ""
if not isinstance(text, str):
try:
text = str(text)
except Exception:
return ""
# Try to parse JSON and pull the ``error`` field — tool handlers return
# ``{"success": false, "error": "..."}``; raw string wins if parse fails.
# Handlers return {"success": false, "error": "..."}; the raw string wins if parse fails.
stripped = text.strip()
if stripped.startswith("{"):
try:
@@ -546,7 +382,6 @@ def _extract_error_preview(result: Any, max_len: int = 180) -> str:
text = data["error"]
except Exception:
pass
# Collapse whitespace, trim to max_len.
text = " ".join(text.split())
if len(text) > max_len:
text = text[: max_len - 1] + "…"
@@ -554,25 +389,20 @@ def _extract_error_preview(result: Any, max_len: int = 180) -> str:
def _trajectory_normalize_msg(msg: Dict[str, Any]) -> Dict[str, Any]:
"""Strip image blobs from a message for trajectory saving.
Returns a shallow copy with multimodal tool results replaced by their
text_summary, and image parts in content lists replaced by
`[screenshot]` placeholders. Keeps the message schema otherwise intact.
"""
"""Shallow copy with image blobs stripped for trajectory saving: multimodal
results become their text summary, image parts become ``[screenshot]``."""
if not isinstance(msg, dict):
return msg
content = msg.get("content")
if _is_multimodal_tool_result(content):
return {**msg, "content": _multimodal_text_summary(content)}
if isinstance(content, list):
cleaned = []
for p in content:
if isinstance(p, dict) and p.get("type") in {"image", "image_url", "input_image"}:
cleaned.append({"type": "text", "text": "[screenshot]"})
else:
cleaned.append(p)
return {**msg, "content": cleaned}
return {**msg, "content": [
{"type": "text", "text": "[screenshot]"}
if isinstance(p, dict) and p.get("type") in {"image", "image_url", "input_image"}
else p
for p in content
]}
return msg
@@ -590,33 +420,19 @@ def make_tool_result_message(
*,
effect_disposition: str | None = None,
) -> dict:
"""Build a tool-result message dict with both the OpenAI-format ``name``
field (required by the wire format and provider adapters) and the internal
``tool_name`` field (written to the session DB messages table).
"""Build a tool-result message with the OpenAI ``name`` field (wire format)
and the internal ``tool_name`` field (session DB).
Content from high-risk tools (``web_extract``, ``web_search``, ``browser_*``,
``mcp_*``) gets wrapped in semantic delimiters telling the model the content
is untrusted data, not instructions. This is the architectural defense
against indirect prompt injection from poisoned web pages, GitHub issues,
and MCP responses — it changes how the model interprets the content rather
than relying on regex pattern matching catching every payload.
Wrapping applies to plain string content and to multimodal content
lists (``[{"type": "text", "text": "..."}, {"type": "image_url", ...}]``):
each text-type part is wrapped individually using the same rules as plain
string content (short text passes through unchanged; longer text is
neutralized and framed). Non-text parts (e.g. image_url) are preserved.
The outer list itself is rebuilt rather than returned by identity, so
callers should compare by value, not by ``is``.
Content from high-risk tools (web_extract, web_search, browser_*, mcp_*) is
wrapped in untrusted-data delimiters — the architectural defense against
indirect prompt injection; see ``_maybe_wrap_untrusted``.
"""
# Keep the constructor safe for every caller, including replay recovery
# paths that do not go through the live executor's canonical-id helper.
# Replay-recovery callers bypass the executor's canonical-id helper, so normalize here too.
tool_call_id = _normalize_tool_call_id(tool_call_id)
# Order matters: detect provider-side elision on the RAW content and
# append the notice first, THEN wrap — so the notice lives inside the
# untrusted block next to the data it describes, appended exactly once
# at construction time (cache-safe).
# Order matters: detect elision on the RAW content and append the notice
# first, THEN wrap, so the notice sits inside the untrusted block next to
# the data it describes — once, at construction time (cache-safe).
wrapped = _maybe_wrap_untrusted(name, _maybe_append_elision_notice(name, content))
message = stamp_message_timestamp({
"role": "tool",
@@ -637,63 +453,41 @@ def make_tool_result_message(
return message
# Tools whose results carry attacker-controllable content. Wrapping their
# string output in ``<untrusted_tool_result>`` delimiters tells the model the
# payload is data, not instructions — the architectural piece of the
# promptware defense. Skipped for short outputs (under 32 chars) where the
# overhead of the wrapper outweighs any indirect-injection risk.
_UNTRUSTED_TOOL_NAMES = frozenset({
"web_extract",
"web_search",
})
_UNTRUSTED_TOOL_PREFIXES = (
"browser_",
"mcp_",
)
# Tools whose results carry attacker-controllable content. Short outputs
# (under 32 chars) skip wrapping: the overhead outweighs any injection risk.
_UNTRUSTED_TOOL_NAMES = frozenset({"web_extract", "web_search"})
_UNTRUSTED_TOOL_PREFIXES = ("browser_", "mcp_")
_UNTRUSTED_WRAP_MIN_CHARS = 32
# Matches the delimiter token in any case so attacker content can't forge or
# prematurely close the boundary with a differently-cased variant the model
# would still read as a tag (e.g. ``</UNTRUSTED_TOOL_RESULT>``).
# Case-insensitive so attacker content can't forge or prematurely close the
# boundary with a differently-cased tag the model would still read as one.
_DELIMITER_TOKEN_RE = re.compile(r"untrusted_tool_result", re.IGNORECASE)
def _is_untrusted_tool(name: Optional[str]) -> bool:
if not name:
return False
if name in _UNTRUSTED_TOOL_NAMES:
return True
return any(name.startswith(p) for p in _UNTRUSTED_TOOL_PREFIXES)
return bool(name) and (name in _UNTRUSTED_TOOL_NAMES or name.startswith(_UNTRUSTED_TOOL_PREFIXES))
# --- Upstream-elision detection --------------------------------------------
#
# Some MCP servers elide data SERVER-SIDE and mark the elision inside the
# payload itself (e.g. Composio: '...13 more items' inside a JSON array,
# '"has_more": true', 'Complete response was large (N tokens). Full data
# saved to sandbox in /mnt/files/...', 'data_preview' envelopes). Because the
# result looks structurally complete, models treat the visible slice as the
# whole dataset and falsely claim completeness. When one of these markers is
# present, we append ONE compact notice at result-construction time — before
# the message enters history, never mutated later, so prompt caching is safe.
def _is_text_item(item: Any) -> bool:
return _is_text_part(item) and isinstance(item.get("text"), str)
# Conservative patterns only: each one is an explicit provider-side "there is
# more data than what you can see" signal, not a generic truncation heuristic.
# --- Upstream-elision detection ---
# Some MCP servers elide data SERVER-SIDE and mark it inside the payload
# ('...13 more items', '"has_more": true', 'saved to sandbox', 'data_preview'
# envelopes). The result looks structurally complete, so models treat the
# visible slice as the whole dataset. Conservative, explicit markers only —
# not a generic truncation heuristic. One notice is appended at construction
# time, never mutated later (prompt-cache safe).
_UPSTREAM_ELISION_PATTERNS = (
re.compile(r"\.\.\.\s*\d+\s+more\s+items?", re.IGNORECASE),
re.compile(r'"has_more"\s*:\s*true', re.IGNORECASE),
re.compile(r"saved to sandbox", re.IGNORECASE),
re.compile(r"data_preview", re.IGNORECASE),
)
# Results smaller than this can't meaningfully hide an elided enumeration —
# skip the scan entirely so tiny results pay nothing.
# Tiny results can't hide an elided enumeration; markers for the sizes that
# matter (20-50K) are always inside the first 64KB.
_ELISION_SCAN_MIN_CHARS = 1_000
# Bound the regex scan: markers appear near the elided structure, which for
# the payload sizes that matter (20-50K) is always inside the first 64KB.
_ELISION_SCAN_MAX_CHARS = 65_536
_UPSTREAM_ELISION_NOTICE = (
@@ -704,53 +498,31 @@ _UPSTREAM_ELISION_NOTICE = (
def _detect_upstream_elision(content: Any) -> bool:
"""True when a string tool result carries provider-side elision markers.
Cheap and safe by construction: non-string content is never scanned,
results under ``_ELISION_SCAN_MIN_CHARS`` short-circuit, and the regex
scan is capped at the first ``_ELISION_SCAN_MAX_CHARS`` chars.
"""
if not isinstance(content, str):
return False
if len(content) < _ELISION_SCAN_MIN_CHARS:
"""True when a string result carries provider-side elision markers (bounded scan)."""
if not isinstance(content, str) or len(content) < _ELISION_SCAN_MIN_CHARS:
return False
window = content[:_ELISION_SCAN_MAX_CHARS]
return any(p.search(window) for p in _UPSTREAM_ELISION_PATTERNS)
def _maybe_append_elision_notice(name: str, content: Any) -> Any:
"""Append the incompleteness notice to untrusted string results that
embed upstream elision markers. Returns ``content`` unchanged otherwise.
Runs on the RAW result before untrusted-wrapping so the notice sits with
the data it describes, and only at result-construction time (cache-safe).
"""
if not _is_untrusted_tool(name):
return content
if _detect_upstream_elision(content):
"""Append the incompleteness notice to untrusted string results with elision markers."""
if _is_untrusted_tool(name) and _detect_upstream_elision(content):
return content + _UPSTREAM_ELISION_NOTICE
return content
def _tool_output_risk_metadata(name: str, content: Any) -> Optional[Dict[str, Any]]:
"""Classify textual attacker-controlled output without retaining a copy.
"""Internal-only advisory classification of attacker-controlled output.
The advisory metadata is internal-only. It records deterministic finding
identifiers, never blocks or redacts the normal result, and deliberately
omits raw scanned text.
Records deterministic finding ids, never blocks or redacts, and omits the scanned text.
"""
if not _is_untrusted_tool(name):
return None
if isinstance(content, str):
text_parts = [content]
elif isinstance(content, list):
text_parts = [
item["text"]
for item in content
if isinstance(item, dict)
and item.get("type") == "text"
and isinstance(item.get("text"), str)
]
text_parts = [item["text"] for item in content if _is_text_item(item)]
if not text_parts:
return None
else:
@@ -769,40 +541,20 @@ def _tool_output_risk_metadata(name: str, content: Any) -> Optional[Dict[str, An
def _neutralize_delimiters(content: str) -> str:
"""Defang any literal ``untrusted_tool_result`` delimiter embedded in
attacker-controlled content so it can't break out of the wrapper.
Without this, a poisoned web page / GitHub issue / MCP response that
contains ``</untrusted_tool_result>`` would close the trust boundary early
— everything the attacker writes after it then reads as trusted instructions
outside the block. Replacing the underscores with hyphens leaves the text
readable but means it no longer matches the real (underscore) delimiter.
"""
"""Defang embedded ``untrusted_tool_result`` tokens so poisoned content
can't close the trust boundary early (hyphens keep it readable but non-matching)."""
return _DELIMITER_TOKEN_RE.sub("untrusted-tool-result", content)
def _maybe_wrap_untrusted(name: str, content: Any) -> Any:
"""Wrap content from high-risk tools in untrusted-data delimiters.
"""Wrap high-risk tool content in untrusted-data delimiters.
Handles plain string content and multimodal content lists
(``[{"type": "text", "text": "..."}, {"type": "image_url", ...}]``).
Text parts inside a multimodal list are wrapped individually — the same
rules as plain string content — so vision-capable adapters still receive
a valid content list while an injection payload embedded in a text chunk
is still marked as untrusted data. Non-text parts (image_url, etc.) are
preserved unchanged. The outer list is rebuilt rather than returned by
identity, so callers must compare by value, not by ``is``.
Returns ``content`` unchanged when:
- the tool is not in the high-risk set
- the content is neither a string nor a list (dict, None, …)
- (string) the content is too short to be worth wrapping
Wrapped string content is always neutralized (any embedded delimiter token
is defanged) and wrapped in exactly one well-formed block. There is no
"already wrapped" fast-path: such a check is attacker-forgeable — content
that merely starts with the opening tag would be returned with no data
framing at all — so re-wrapping (harmlessly) is the safe choice.
Strings are neutralized and wrapped in exactly one block; text parts of a
multimodal list are wrapped individually (non-text parts preserved, outer
list rebuilt — compare by value, not ``is``). Unchanged when the tool is
not high-risk, the content is neither str nor list, or a string is too
short. There is deliberately no "already wrapped" fast-path: it would be
attacker-forgeable, so harmless re-wrapping is the safe choice.
"""
if not _is_untrusted_tool(name):
return content
@@ -821,11 +573,7 @@ def _maybe_wrap_untrusted(name: str, content: Any) -> Any:
)
if isinstance(content, list):
return [
{**item, "text": _maybe_wrap_untrusted(name, item["text"])}
if isinstance(item, dict)
and item.get("type") == "text"
and isinstance(item.get("text"), str)
else item
{**item, "text": _maybe_wrap_untrusted(name, item["text"])} if _is_text_item(item) else item
for item in content
]
return content