refactor(agent/prompt): remove dead code, unify duplicated helpers, compact docstrings across prompt/skill/redaction modules
Dead (zero refs): coding_system_blocks, get_friendly_tool_labels, get_scan_ordered_skills_dirs, _project_quarantine_cache_clear, clear_stable_prefixes, _redact_http_request_target_query_params, _has_http_method_substring, PromptCachePlan.marker_count, display _diff_* colour thunks (-> _diff_ansi), pass-through RedactingFormatter.__init__. Unified: _slugify -> slugify_skill_name; reload diff -> diff_command_snapshots; _is_summary_item -> is_compaction_summary_message alias; sanitizer walkers -> _sanitize_messages/_sanitize_structure; assignment redaction passes -> _redact_assignments/_should_redact_assignment; quiet-mode tool lines -> _CUTE_LINES table.
This commit is contained in:
+138
-312
@@ -1,52 +1,29 @@
|
||||
"""Coding-context awareness — base Hermes, every interactive surface.
|
||||
|
||||
When the user runs Hermes inside a code workspace (CLI, TUI, desktop app, or an
|
||||
editor over ACP), Hermes shifts into a **coding posture**. This module is the
|
||||
single place that decides whether we're in that posture and what it implies,
|
||||
so the rest of the codebase never re-derives "are we coding?" on its own.
|
||||
When Hermes runs inside a code workspace (CLI, TUI, desktop, ACP editor) it
|
||||
shifts into a **coding posture**. This module is the single place that decides
|
||||
whether we're in that posture and what it implies, so nothing else re-derives
|
||||
"are we coding?". The posture is a frozen :class:`RuntimeMode` selected from a
|
||||
small :class:`ContextProfile` registry (``coding`` / ``general``); a profile is
|
||||
*data* (toolset, operating brief, skill-index hints) that every domain reads:
|
||||
|
||||
Architecture — one seam, many consumers
|
||||
----------------------------------------
|
||||
The posture is modelled as a frozen :class:`RuntimeMode` selected from a small
|
||||
:class:`ContextProfile` registry (today: ``coding`` and ``general``). A profile
|
||||
is *data* — it declares the toolset to collapse to, the operating brief to
|
||||
inject, and hints for other domains (model routing, memory, subagents). Every
|
||||
domain reads the same resolved object instead of probing git/config itself:
|
||||
* System prompt — ``RuntimeMode.system_prompt_parts()`` → operating brief +
|
||||
live git/workspace snapshot (``agent/system_prompt.py``).
|
||||
* Toolset — ``RuntimeMode.toolset_selection()`` → ``coding`` toolset + enabled
|
||||
MCP servers, ONLY under the opt-in ``focus`` mode. The default posture is
|
||||
prompt-only and never strips a toolset the user explicitly enabled.
|
||||
* Delegation — subagents inherit the toolset and prompt builder, so the
|
||||
posture propagates for free.
|
||||
|
||||
* **System prompt** — ``RuntimeMode.system_blocks()`` → the operating brief +
|
||||
a live git/workspace snapshot (``agent/system_prompt.py``).
|
||||
* **Toolset** — ``RuntimeMode.toolset_selection()`` → the ``coding`` toolset
|
||||
plus the user's enabled MCP servers (``cli.py`` / ``tui_gateway``). Only
|
||||
under the opt-in ``focus`` mode: the default posture is prompt-only and
|
||||
never touches the user's configured toolsets (toolsets like messaging /
|
||||
smart-home / music are off-by-default anyway, and someone who explicitly
|
||||
enabled image-gen or Spotify shouldn't lose it for being in a git repo).
|
||||
* **Delegation** — subagents inherit the parent's toolset and run through the
|
||||
same prompt builder, so the coding posture propagates to children for free.
|
||||
* **Model / memory / compression** — declared on the profile
|
||||
(``model_hint``, ``memory_policy``) as the extension seam; consumers read
|
||||
``mode.profile`` rather than re-deciding.
|
||||
Cache safety: the mode is resolved once and immutable; the workspace snapshot
|
||||
is built once at prompt-build time and never re-probed per turn (the brief
|
||||
tells the model to re-check with ``git``). A ``/coding`` flip takes effect next
|
||||
session.
|
||||
|
||||
Cache safety
|
||||
------------
|
||||
The mode is resolved **once** and is immutable. The workspace snapshot is built
|
||||
once at prompt-build time and baked into the *stable* system-prompt tier — never
|
||||
re-probed per turn (that would shatter the prompt cache). Branch and dirty state
|
||||
drift mid-session, so the brief tells the model to re-check with ``git`` before
|
||||
acting on the snapshot. A ``/coding`` flip therefore only takes effect next
|
||||
session (deferred), the same contract as ``/skills install`` vs ``--now``.
|
||||
|
||||
Activation (config ``agent.coding_context``):
|
||||
|
||||
* ``auto`` (default) — posture (brief + snapshot) on an interactive coding
|
||||
surface sitting in a code workspace (git repo or recognised project root).
|
||||
Prompt-only; toolsets and the skill index untouched.
|
||||
* ``focus`` — like ``auto``, but additionally collapses the toolset to the
|
||||
``coding`` set + enabled MCP servers and demotes non-coding skill
|
||||
categories to names-only in the prompt's skill index (no skill is ever
|
||||
hidden). Explicit opt-in for a lean schema.
|
||||
* ``on`` — force the posture anywhere (incl. non-workspaces). Prompt-only.
|
||||
* ``off`` — disable entirely.
|
||||
Activation (config ``agent.coding_context``): ``auto`` (default) — posture on an
|
||||
interactive surface in a code workspace, prompt-only; ``focus`` — also collapse
|
||||
the toolset and demote non-coding skill categories to names-only (never
|
||||
hidden); ``on`` — force the posture anywhere; ``off`` — disable.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -67,8 +44,7 @@ logger = logging.getLogger("hermes.coding_context")
|
||||
CODING_TOOLSET = "coding"
|
||||
|
||||
# Surfaces where a coding posture makes sense under ``auto``. Messaging
|
||||
# platforms (telegram, discord, slack, …) are intentionally absent — a chat bot
|
||||
# in a group is not pair-programming.
|
||||
# platforms are intentionally absent — a chat bot in a group is not pairing.
|
||||
INTERACTIVE_CODING_PLATFORMS = {"cli", "tui", "acp", "desktop", ""}
|
||||
|
||||
# Project-root signals that mark a directory as a code workspace even when it
|
||||
@@ -85,10 +61,8 @@ _PROJECT_MARKERS = (
|
||||
# Agent-instruction files surfaced separately from manifests in the snapshot.
|
||||
_CONTEXT_FILES = ("AGENTS.md", "CLAUDE.md", ".cursorrules")
|
||||
|
||||
# Source-file extensions that make a git repo a *code* workspace even with no
|
||||
# manifest. Without this, `git init` on a notes/writing/research folder (a huge
|
||||
# non-coding use case) would flip the whole session into the coding posture just
|
||||
# for having a `.git`. A manifest still wins on its own (see `_PROJECT_MARKERS`).
|
||||
# Source extensions that make a manifest-less git repo a *code* workspace, so
|
||||
# `git init` on a notes/writing folder does not flip the session into coding.
|
||||
_CODE_EXTENSIONS = frozenset({
|
||||
".py", ".pyi", ".ipynb", ".js", ".jsx", ".ts", ".tsx", ".mjs", ".cjs",
|
||||
".go", ".rs", ".java", ".kt", ".kts", ".scala", ".rb", ".php", ".c", ".h",
|
||||
@@ -97,24 +71,16 @@ _CODE_EXTENSIONS = frozenset({
|
||||
".hs", ".clj", ".erl", ".pl",
|
||||
})
|
||||
|
||||
# Dirs never worth scanning for the code check (deps/build/vcs/venv noise).
|
||||
_CODE_SCAN_SKIP_DIRS = frozenset({
|
||||
".git", "node_modules", "venv", ".venv", "__pycache__", "dist", "build",
|
||||
"target", ".next", ".turbo", "vendor",
|
||||
})
|
||||
|
||||
# Bounded sweep: a code workspace reveals itself in the first handful of entries.
|
||||
_CODE_SCAN_MAX_ENTRIES = 500
|
||||
|
||||
|
||||
def _has_code_files(root: Path) -> bool:
|
||||
"""Cheap, bounded check for source files in a repo's top two levels.
|
||||
|
||||
Lets a git repo of loose scripts (no manifest) still read as a code
|
||||
workspace while a bare notes/writing repo does not. Scans the root and its
|
||||
immediate subdirectories only, capped at ``_CODE_SCAN_MAX_ENTRIES`` stats —
|
||||
a handful of readdirs at session start, not a full walk.
|
||||
"""
|
||||
"""Bounded check for source files in the root and its immediate subdirs."""
|
||||
seen = 0
|
||||
stack = [(root, True)]
|
||||
while stack:
|
||||
@@ -138,6 +104,7 @@ def _has_code_files(root: Path) -> bool:
|
||||
continue
|
||||
return False
|
||||
|
||||
|
||||
# Lockfile → package manager, checked in priority order.
|
||||
_PY_LOCKFILES = (("uv.lock", "uv"), ("poetry.lock", "poetry"), ("Pipfile.lock", "pipenv"))
|
||||
_JS_LOCKFILES = (
|
||||
@@ -153,21 +120,13 @@ _MAX_FACT_FILE_BYTES = 256 * 1024
|
||||
_GIT_TIMEOUT = 2.5
|
||||
|
||||
|
||||
# Per-model edit-format steering. Matching the edit tool format to how a model
|
||||
# was trained reduces mistakes and wasted reasoning (OpenAI/Codex handle
|
||||
# patch-style diffs best; Anthropic models — and most open-weight coding
|
||||
# models, whose RL scaffolds use str_replace-style editors — do best with
|
||||
# string-replacement). Our `patch` tool exposes both: mode="patch" (V4A
|
||||
# multi-file) and mode="replace" (find-and-swap). We nudge each family toward
|
||||
# its native format. Unknown families get nothing (the brief's neutral wording
|
||||
# stands). Substrings match the model id; aligned with TOOL_USE_ENFORCEMENT_MODELS.
|
||||
#
|
||||
# GPT/Codex get V4A for ALL edits, single-file included: in codex-rs,
|
||||
# apply_patch (V4A — apply_patch.lark) is the ONLY file editor, no
|
||||
# str_replace-style tool exists, and the shipped model prompts say to use
|
||||
# apply_patch even "for single file edits" — so a replace-mode nudge would
|
||||
# steer those models toward a format their first-party harness never taught
|
||||
# them.
|
||||
# Per-model edit-format steering: nudge each family toward the `patch` mode it
|
||||
# was trained on (unknown families get nothing). GPT/Codex get V4A for ALL
|
||||
# edits incl. single-file — codex-rs ships apply_patch as its ONLY editor and
|
||||
# its prompts say to use it even for single files, so a replace-mode nudge
|
||||
# would steer them toward a format their first-party harness never taught.
|
||||
# Anthropic and most open-weight coding models were RL'd on str_replace-style
|
||||
# editors. Substrings match the model id; aligned with TOOL_USE_ENFORCEMENT_MODELS.
|
||||
_EDIT_FORMAT_GUIDANCE: dict[str, tuple[tuple[str, ...], str]] = {
|
||||
"patch": (
|
||||
("gpt", "codex"),
|
||||
@@ -188,12 +147,7 @@ _EDIT_FORMAT_GUIDANCE: dict[str, tuple[tuple[str, ...], str]] = {
|
||||
|
||||
|
||||
def _model_family(model: Optional[str]) -> Optional[str]:
|
||||
"""Classify a model id into an edit-format family key, or ``None``.
|
||||
|
||||
Used to steer the coding posture toward the edit tool format a model was
|
||||
trained on. Family-agnostic by design: an unrecognised model gets ``None``
|
||||
and the operating brief's neutral edit wording applies.
|
||||
"""
|
||||
"""Edit-format family key for a model id, or ``None`` (neutral wording applies)."""
|
||||
if not model:
|
||||
return None
|
||||
lowered = model.lower()
|
||||
@@ -206,14 +160,11 @@ def _model_family(model: Optional[str]) -> Optional[str]:
|
||||
def _edit_format_line(model: Optional[str]) -> str:
|
||||
"""The edit-format guidance line for this model's family (``""`` if none)."""
|
||||
family = _model_family(model)
|
||||
if family is None:
|
||||
return ""
|
||||
return _EDIT_FORMAT_GUIDANCE[family][1]
|
||||
return "" if family is None else _EDIT_FORMAT_GUIDANCE[family][1]
|
||||
|
||||
|
||||
# Operating brief for the coding posture. Tool names referenced here (read_file,
|
||||
# search_files, patch, write_file, terminal, todo) are in the coding toolset and
|
||||
# in _HERMES_CORE_TOOLS, so they're present on every surface this fires on.
|
||||
# Operating brief for the coding posture. Tool names referenced here are in the
|
||||
# coding toolset and in _HERMES_CORE_TOOLS, so they exist on every surface this fires on.
|
||||
CODING_AGENT_GUIDANCE = (
|
||||
"You are a coding agent pairing with the user inside their codebase. "
|
||||
"Operate like a careful senior engineer.\n"
|
||||
@@ -264,6 +215,12 @@ CODING_AGENT_GUIDANCE = (
|
||||
"answer, not a preamble."
|
||||
)
|
||||
|
||||
_TODO_SENTENCE = (
|
||||
"- Track multi-step work with `todo_list`. Reference code as "
|
||||
"`path:line` instead of pasting whole files."
|
||||
)
|
||||
_NO_TODO_SENTENCE = "- Reference code as `path:line` instead of pasting whole files."
|
||||
|
||||
|
||||
# ── Context profiles (declarative posture definitions) ──────────────────────
|
||||
|
||||
@@ -272,35 +229,22 @@ CODING_AGENT_GUIDANCE = (
|
||||
class ContextProfile:
|
||||
"""A named operating posture. Pure data — consumers read these fields.
|
||||
|
||||
``toolset`` — collapse to this toolset (+ enabled MCP) when no explicit
|
||||
selection is pinned; ``None`` keeps the platform default.
|
||||
``guidance`` — operating brief injected into the stable system prompt;
|
||||
``""`` injects nothing.
|
||||
``model_hint`` — routing preference key for smart model routing
|
||||
(extension seam; not yet consumed by the router).
|
||||
``memory_policy``— memory namespace/weighting hint (extension seam).
|
||||
``compact_skill_categories`` — skill categories DEMOTED to names-only in
|
||||
the system-prompt skill index under the opt-in ``focus``
|
||||
mode. Never hidden: every skill name stays visible
|
||||
(so memory-anchored recall keeps working) — only the
|
||||
descriptions are dropped to cut index noise. Deny-list
|
||||
semantics so unknown/custom categories keep full
|
||||
entries.
|
||||
``toolset``: collapse to this toolset (+ enabled MCP) under ``focus``;
|
||||
``None`` keeps the platform default. ``guidance``: operating brief for the
|
||||
stable system prompt. ``model_hint``: routing preference (extension seam).
|
||||
``compact_skill_categories``: categories DEMOTED to names-only in the skill
|
||||
index under ``focus`` — deny-list, never hidden, so recall keeps working.
|
||||
"""
|
||||
|
||||
name: str
|
||||
toolset: Optional[str] = None
|
||||
guidance: str = ""
|
||||
model_hint: Optional[str] = None
|
||||
memory_policy: str = "default"
|
||||
compact_skill_categories: tuple[str, ...] = ()
|
||||
|
||||
|
||||
# Skill categories that are clearly not part of a coding workflow. Demoted to
|
||||
# names-only in the prompt's skill index under the opt-in ``focus`` mode only
|
||||
# (deny-list — anything not listed here, incl. custom user categories, keeps
|
||||
# full entries). Coding-adjacent categories (devops, github, mcp,
|
||||
# data-science, diagramming, research, security, …) are intentionally absent.
|
||||
# Clearly non-coding skill categories (deny-list: custom categories keep full
|
||||
# entries). Coding-adjacent ones (devops, github, mcp, research, …) are absent.
|
||||
_NON_CODING_SKILL_CATEGORIES = (
|
||||
"apple", "communication", "cooking", "creative", "email", "finance",
|
||||
"gaming", "gifs", "health", "media", "music", "note-taking",
|
||||
@@ -315,7 +259,6 @@ CODING_PROFILE = ContextProfile(
|
||||
toolset=CODING_TOOLSET,
|
||||
guidance=CODING_AGENT_GUIDANCE,
|
||||
model_hint="coding",
|
||||
memory_policy="project",
|
||||
compact_skill_categories=_NON_CODING_SKILL_CATEGORIES,
|
||||
)
|
||||
|
||||
@@ -332,45 +275,38 @@ def get_profile(name: str) -> ContextProfile:
|
||||
|
||||
# ── Helpers ─────────────────────────────────────────────────────────────────
|
||||
|
||||
_MODE_ALIASES = {
|
||||
**dict.fromkeys(("focus", "strict", "lean"), "focus"),
|
||||
**dict.fromkeys(("on", "true", "yes", "1", "always"), "on"),
|
||||
**dict.fromkeys(("off", "false", "no", "0", "never"), "off"),
|
||||
}
|
||||
|
||||
def _coding_mode(config: Optional[dict[str, Any]]) -> str:
|
||||
"""Return the normalized ``agent.coding_context`` mode (auto/focus/on/off)."""
|
||||
|
||||
def _agent_config_value(config: Optional[dict[str, Any]], key: str, default: Any, *, readonly: bool) -> Any:
|
||||
"""``config["agent"][key]``, loading config when none was passed."""
|
||||
if config is None:
|
||||
try:
|
||||
from hermes_cli.config import load_config_readonly
|
||||
from hermes_cli.config import load_config, load_config_readonly
|
||||
|
||||
config = load_config_readonly()
|
||||
config = load_config_readonly() if readonly else load_config()
|
||||
except Exception:
|
||||
config = {}
|
||||
raw = ((config or {}).get("agent", {}) or {}).get("coding_context", "auto")
|
||||
mode = str(raw).strip().lower()
|
||||
if mode in {"focus", "strict", "lean"}:
|
||||
return "focus"
|
||||
if mode in {"on", "true", "yes", "1", "always"}:
|
||||
return "on"
|
||||
if mode in {"off", "false", "no", "0", "never"}:
|
||||
return "off"
|
||||
return "auto"
|
||||
return ((config or {}).get("agent", {}) or {}).get(key, default)
|
||||
|
||||
|
||||
def _coding_mode(config: Optional[dict[str, Any]]) -> str:
|
||||
"""Normalized ``agent.coding_context`` mode (auto/focus/on/off)."""
|
||||
raw = _agent_config_value(config, "coding_context", "auto", readonly=True)
|
||||
return _MODE_ALIASES.get(str(raw).strip().lower(), "auto")
|
||||
|
||||
|
||||
def _coding_instructions(config: Optional[dict[str, Any]]) -> str:
|
||||
"""Standing operator instructions for the coding posture (config).
|
||||
"""Standing operator instructions (``agent.coding_instructions``: str or list).
|
||||
|
||||
``agent.coding_instructions`` — a string or list of strings appended to the
|
||||
coding brief as an extra stable system block, so a user can pin project-wide
|
||||
coding-workflow rules (e.g. "for UI work don't run tsc/lint until I approve;
|
||||
clean the diff before committing") without editing the shipped brief.
|
||||
Cache-safe: resolved once per session into the stable system-prompt tier,
|
||||
like the rest of the posture.
|
||||
Appended to the brief as an extra stable block so a user can pin
|
||||
project-wide workflow rules without editing the shipped brief.
|
||||
"""
|
||||
if config is None:
|
||||
try:
|
||||
from hermes_cli.config import load_config
|
||||
|
||||
config = load_config()
|
||||
except Exception:
|
||||
config = {}
|
||||
raw = ((config or {}).get("agent", {}) or {}).get("coding_instructions", "")
|
||||
raw = _agent_config_value(config, "coding_instructions", "", readonly=False)
|
||||
if isinstance(raw, (list, tuple)):
|
||||
return "\n".join(str(item).strip() for item in raw if str(item).strip())
|
||||
return str(raw or "").strip()
|
||||
@@ -403,19 +339,14 @@ def _home() -> Optional[Path]:
|
||||
|
||||
|
||||
def _marker_root(cwd: Path) -> Optional[Path]:
|
||||
"""Nearest ancestor that looks like a project root, or ``None``.
|
||||
"""Nearest ancestor (≤6 levels) that looks like a project root, or ``None``.
|
||||
|
||||
Walks up at most a few levels so a manifest in the workspace root counts
|
||||
even when the user is in a subdirectory. ``$HOME`` itself is skipped — a
|
||||
Makefile or AGENTS.md sitting in the home directory is global user config,
|
||||
not a project-root signal.
|
||||
``$HOME`` and the shared temp root are skipped: a Makefile/AGENTS.md in the
|
||||
home dir is global user config, and a stray manifest in /tmp must not flip
|
||||
every session whose cwd lives under it into the coding posture.
|
||||
"""
|
||||
current = cwd.resolve()
|
||||
home = _home()
|
||||
# Shared world-writable temp roots are never project roots: a stray
|
||||
# manifest in /tmp (left by any process) must not flip every session
|
||||
# whose cwd lives under the temp dir into the coding posture. Same
|
||||
# reasoning as the $HOME skip below.
|
||||
try:
|
||||
temp_root = Path(tempfile.gettempdir()).resolve()
|
||||
except Exception:
|
||||
@@ -435,17 +366,10 @@ def _detect_profile_name(mode: str, platform: str, cwd_str: str) -> str:
|
||||
"""Resolve which profile applies.
|
||||
|
||||
``auto``/``focus``: coding when the surface is interactive AND the cwd is a
|
||||
code workspace (a git repo or a recognised project root). ``on``: always
|
||||
coding. ``off``: always general.
|
||||
|
||||
A git repo rooted at ``$HOME`` (the dotfiles pattern) is NOT a workspace
|
||||
signal — without the guard, every session anywhere under a dotfiles-managed
|
||||
home directory would silently flip to the coding posture.
|
||||
|
||||
Detection is intentionally not memoized: it's a handful of ``stat`` calls,
|
||||
and callers resolve the mode once per session anyway. Caching here would
|
||||
risk a stale posture if a long-lived process (gateway/TUI) serves sessions
|
||||
from different working directories.
|
||||
code workspace (project root, or a git repo that actually holds code).
|
||||
``on``: always coding. ``off``: always general. A git repo rooted at
|
||||
``$HOME`` (dotfiles) is NOT a workspace signal. Deliberately not memoized:
|
||||
a long-lived gateway/TUI process serves sessions from different cwds.
|
||||
"""
|
||||
if mode == "off":
|
||||
return GENERAL_PROFILE.name
|
||||
@@ -454,16 +378,10 @@ def _detect_profile_name(mode: str, platform: str, cwd_str: str) -> str:
|
||||
if platform and platform.strip().lower() not in INTERACTIVE_CODING_PLATFORMS:
|
||||
return GENERAL_PROFILE.name
|
||||
cwd = Path(cwd_str)
|
||||
# A recognized project root (manifest / AGENTS.md / .cursorrules) is a code
|
||||
# workspace on its own — cheap stat checks, no scan.
|
||||
if _marker_root(cwd) is not None:
|
||||
return CODING_PROFILE.name
|
||||
git_root = _git_root(cwd)
|
||||
if git_root is not None and git_root == _home():
|
||||
git_root = None # dotfiles repo at $HOME — not a code workspace
|
||||
# A bare git repo only counts when it actually holds code, so `git init` on a
|
||||
# notes/writing/research folder stays in the general posture.
|
||||
if git_root is not None and _has_code_files(git_root):
|
||||
if git_root is not None and git_root != _home() and _has_code_files(git_root):
|
||||
return CODING_PROFILE.name
|
||||
return GENERAL_PROFILE.name
|
||||
|
||||
@@ -475,23 +393,18 @@ def _detect_profile_name(mode: str, platform: str, cwd_str: str) -> str:
|
||||
class RuntimeMode:
|
||||
"""The resolved operating posture for a session. Immutable by construction.
|
||||
|
||||
Built once via :func:`resolve_runtime_mode` and consumed by every domain
|
||||
that cares about the coding/general distinction. Never mutate or re-resolve
|
||||
mid-session — that would break the prompt cache.
|
||||
Built once via :func:`resolve_runtime_mode`; never re-resolved mid-session
|
||||
(that would break the prompt cache).
|
||||
"""
|
||||
|
||||
profile: ContextProfile
|
||||
surface: str
|
||||
cwd: Path
|
||||
# The normalized ``agent.coding_context`` mode this posture was resolved
|
||||
# under (auto/focus/on/off). Toolset collapse is gated on ``focus``.
|
||||
# Normalized ``agent.coding_context`` mode; toolset collapse is gated on ``focus``.
|
||||
config_mode: str = "auto"
|
||||
# The model id this session runs (e.g. "anthropic/claude-opus-4.8"). Used
|
||||
# only to steer edit-format guidance toward the model's family — see
|
||||
# ``_edit_format_line``. Fixed for the session, so cache-safe.
|
||||
# Model id, used only to steer edit-format guidance (fixed per session).
|
||||
model: Optional[str] = None
|
||||
# Standing operator instructions (``agent.coding_instructions``), appended
|
||||
# as an extra stable system block. Empty unless the user configures it.
|
||||
# ``agent.coding_instructions``, appended as an extra stable block.
|
||||
instructions: str = ""
|
||||
|
||||
@property
|
||||
@@ -505,94 +418,58 @@ class RuntimeMode:
|
||||
def toolset_selection(self, config: Optional[dict[str, Any]] = None) -> Optional[list[str]]:
|
||||
"""Toolset list for this posture, or ``None`` to keep the platform default.
|
||||
|
||||
Non-``None`` only under the opt-in ``focus`` mode. The default posture
|
||||
is prompt-only: most strippable toolsets are off-by-default anyway, and
|
||||
a user who explicitly enabled one (image-gen for frontend/game assets,
|
||||
messaging for build notifications, …) keeps it while coding.
|
||||
|
||||
Callers apply this only when the user hasn't pinned an explicit
|
||||
selection (``--toolsets``, ``HERMES_TUI_TOOLSETS``, …); they never
|
||||
override a pin. Returns the profile's toolset plus enabled MCP servers.
|
||||
Non-``None`` only under ``focus``. Callers apply it only when the user
|
||||
hasn't pinned an explicit selection (``--toolsets``, ``HERMES_TUI_TOOLSETS``).
|
||||
"""
|
||||
if self.config_mode != "focus":
|
||||
return None
|
||||
if self.profile.toolset is None:
|
||||
if self.config_mode != "focus" or self.profile.toolset is None:
|
||||
return None
|
||||
return [self.profile.toolset, *_enabled_mcp_servers(config)]
|
||||
|
||||
def system_prompt_parts(
|
||||
self, valid_tool_names=None
|
||||
) -> tuple[list[str], list[str], list[str]]:
|
||||
"""Return prefix, workspace, and trailing posture blocks separately.
|
||||
"""Return (prefix, workspace, trailing) posture blocks.
|
||||
|
||||
The operating brief carries a model-family edit-format nudge appended
|
||||
to it (one cached string, not a separate block) so the model is steered
|
||||
toward the `patch` mode it handles best — see ``_edit_format_line``.
|
||||
|
||||
``valid_tool_names`` (when provided) tailors the brief to the session's
|
||||
toolset: the ``todo`` tracking sentence is dropped when the todo tool
|
||||
isn't loaded (e.g. Blank Slate), so the brief never references a tool
|
||||
the model can't call. The toolset is fixed at session construction,
|
||||
so the rendered brief is deterministic per session — cache-safe.
|
||||
|
||||
The three lists preserve the historical flat prompt order: the brief,
|
||||
the live workspace snapshot, then configured operator instructions.
|
||||
Prompt assembly can therefore put a cache boundary before the snapshot
|
||||
without changing the persisted system-prompt bytes.
|
||||
The brief carries the model-family edit-format nudge appended to it
|
||||
(one cached string). ``valid_tool_names`` drops the ``todo_list``
|
||||
sentence when that tool isn't loaded (e.g. Blank Slate). The three
|
||||
lists preserve the historical flat order — brief, workspace snapshot,
|
||||
operator instructions — so prompt assembly can put a cache boundary
|
||||
before the snapshot without changing the persisted bytes.
|
||||
"""
|
||||
if not self.is_coding:
|
||||
return [], [], []
|
||||
prefix: list[str] = []
|
||||
workspace_parts: list[str] = []
|
||||
trailing: list[str] = []
|
||||
if self.profile.guidance:
|
||||
brief = self.profile.guidance
|
||||
if valid_tool_names is not None and "todo_list" not in valid_tool_names:
|
||||
brief = brief.replace(
|
||||
"- Track multi-step work with `todo_list`. Reference code as "
|
||||
"`path:line` instead of pasting whole files.",
|
||||
"- Reference code as `path:line` instead of pasting "
|
||||
"whole files.",
|
||||
)
|
||||
brief = brief.replace(_TODO_SENTENCE, _NO_TODO_SENTENCE)
|
||||
edit_line = _edit_format_line(self.model)
|
||||
if edit_line:
|
||||
brief = f"{brief}\n{edit_line}"
|
||||
prefix.append(brief)
|
||||
workspace = build_coding_workspace_block(self.cwd)
|
||||
if workspace:
|
||||
workspace_parts.append(workspace)
|
||||
# Operator instructions ride their own block so the brief (block 0) stays
|
||||
# byte-stable and cache-keyed independently of user config.
|
||||
if self.instructions:
|
||||
trailing.append(f"Operator instructions (from config):\n{self.instructions}")
|
||||
workspace_parts = [workspace] if workspace else []
|
||||
# Operator instructions ride their own block so the brief stays
|
||||
# byte-stable independently of user config.
|
||||
trailing = (
|
||||
[f"Operator instructions (from config):\n{self.instructions}"]
|
||||
if self.instructions else []
|
||||
)
|
||||
return prefix, workspace_parts, trailing
|
||||
|
||||
def system_blocks(self) -> list[str]:
|
||||
"""Return posture blocks in their historical display order.
|
||||
|
||||
``system_prompt_parts`` is the cache-aware API. This compatibility
|
||||
helper retains the public flat list for callers outside prompt assembly.
|
||||
"""
|
||||
"""Posture blocks as one flat list in historical order (compat helper)."""
|
||||
prefix, workspace, trailing = self.system_prompt_parts()
|
||||
return [*prefix, *workspace, *trailing]
|
||||
|
||||
def compact_skill_categories(self) -> frozenset[str]:
|
||||
"""Skill categories to demote to names-only in the prompt's skill index.
|
||||
"""Skill categories to demote to names-only in the skill index.
|
||||
|
||||
Gated on the opt-in ``focus`` mode, like the toolset collapse: the
|
||||
default posture leaves the skill index untouched. Users who didn't ask
|
||||
for a lean prompt keep full entries for every category — index changes
|
||||
under ``auto`` proved too surprising in practice, even names-only ones
|
||||
(a demoted description is information the model no longer weighs when
|
||||
deciding what to load).
|
||||
|
||||
Demoted — never hidden — even under ``focus``. An earlier revision
|
||||
fully pruned these categories from the index, which caused silent
|
||||
capability loss in a real workflow: agent-created skills are the
|
||||
model's accumulated project memory (server-ops runbooks, learned
|
||||
pitfalls, …), and models do not reliably reach for ``skills_list`` to
|
||||
rediscover what the index stopped showing them. Names-only keeps every
|
||||
skill loadable on recall while still cutting the description noise.
|
||||
Gated on ``focus`` like the toolset collapse — index changes under
|
||||
``auto`` proved too surprising. Demoted, never hidden: fully pruning
|
||||
them caused silent capability loss (agent-created skills are the
|
||||
model's project memory and models don't reliably re-run ``skills_list``).
|
||||
"""
|
||||
if not self.is_coding or self.config_mode != "focus":
|
||||
return frozenset()
|
||||
@@ -606,14 +483,10 @@ def resolve_runtime_mode(
|
||||
config: Optional[dict[str, Any]] = None,
|
||||
model: Optional[str] = None,
|
||||
) -> RuntimeMode:
|
||||
"""Resolve the operating posture once. Cheap — a handful of ``stat`` calls.
|
||||
"""Resolve the operating posture once (a handful of ``stat`` calls).
|
||||
|
||||
This is the single entry point every domain should call. The returned
|
||||
object is immutable and safe to cache for the session. Detection itself is
|
||||
intentionally *not* memoized (see ``_detect_profile_name``) so a long-lived
|
||||
process can't pin a stale posture; callers resolve once per session and
|
||||
hold the result. ``model`` is recorded only to steer edit-format guidance;
|
||||
it never affects detection.
|
||||
The single entry point every domain should call; the result is immutable
|
||||
and safe to hold for the session. ``model`` only steers edit-format guidance.
|
||||
"""
|
||||
resolved_cwd = _resolve_cwd(cwd)
|
||||
mode = _coding_mode(config)
|
||||
@@ -649,32 +522,12 @@ def coding_selection(
|
||||
cwd: Optional[str | Path] = None,
|
||||
config: Optional[dict[str, Any]] = None,
|
||||
) -> Optional[list[str]]:
|
||||
"""Toolset selection for the coding posture.
|
||||
|
||||
``None`` unless the user opted into ``focus`` mode AND the posture is
|
||||
active — the default coding posture never overrides configured toolsets.
|
||||
"""
|
||||
"""Toolset selection for the coding posture (``None`` unless ``focus`` and active)."""
|
||||
return resolve_runtime_mode(
|
||||
platform=platform, cwd=cwd, config=config
|
||||
).toolset_selection(config)
|
||||
|
||||
|
||||
def coding_system_blocks(
|
||||
*,
|
||||
platform: Optional[str] = None,
|
||||
cwd: Optional[str | Path] = None,
|
||||
config: Optional[dict[str, Any]] = None,
|
||||
model: Optional[str] = None,
|
||||
) -> list[str]:
|
||||
"""Stable system-prompt blocks for the current posture (empty when general).
|
||||
|
||||
``model`` steers the brief's edit-format nudge toward the model's family.
|
||||
"""
|
||||
return resolve_runtime_mode(
|
||||
platform=platform, cwd=cwd, config=config, model=model
|
||||
).system_blocks()
|
||||
|
||||
|
||||
def coding_system_prompt_parts(
|
||||
*,
|
||||
platform: Optional[str] = None,
|
||||
@@ -695,25 +548,14 @@ def coding_compact_skill_categories(
|
||||
cwd: Optional[str | Path] = None,
|
||||
config: Optional[dict[str, Any]] = None,
|
||||
) -> frozenset[str]:
|
||||
"""Skill categories the active posture demotes to names-only in the index.
|
||||
|
||||
Empty outside the coding posture and outside the opt-in ``focus`` mode —
|
||||
the default posture never touches the skill index. Under ``focus``,
|
||||
demoted — never hidden: every skill name stays in the index and remains
|
||||
loadable via ``skill_view`` / ``skills_list``; only descriptions are
|
||||
dropped.
|
||||
"""
|
||||
"""Skill categories the active posture demotes to names-only (empty outside ``focus``)."""
|
||||
return resolve_runtime_mode(
|
||||
platform=platform, cwd=cwd, config=config
|
||||
).compact_skill_categories()
|
||||
|
||||
|
||||
def _enabled_mcp_servers(config: Optional[dict[str, Any]]) -> list[str]:
|
||||
"""Names of MCP servers the user has enabled — kept in the coding posture.
|
||||
|
||||
MCP servers (figma, browser, tophat, …) are explicitly configured and part
|
||||
of the coding workflow, not noise to strip.
|
||||
"""
|
||||
"""Names of MCP servers the user has enabled — kept in the coding posture."""
|
||||
try:
|
||||
from hermes_cli.config import read_raw_config
|
||||
from hermes_cli.tools_config import _parse_enabled_flag
|
||||
@@ -735,10 +577,9 @@ def _enabled_mcp_servers(config: Optional[dict[str, Any]]) -> list[str]:
|
||||
def _git(cwd: Path, *args: str) -> str:
|
||||
"""``git -C <cwd> <args>`` → stripped stdout, or ``""`` on any failure.
|
||||
|
||||
Uses the shared :func:`bounded_git_probe` so the post-kill cleanup is bounded
|
||||
on Windows — a plain ``subprocess.run(timeout=...)`` here deadlocked the agent
|
||||
turn inside ``build_coding_workspace_block`` when a killed git left a suspended
|
||||
descendant holding the pipe handles (issue #66037).
|
||||
:func:`bounded_git_probe` bounds the post-kill cleanup on Windows — a plain
|
||||
``subprocess.run(timeout=...)`` deadlocked when a killed git left a
|
||||
suspended descendant holding the pipe handles.
|
||||
"""
|
||||
return bounded_git_probe(["git", "-C", str(cwd), *args], timeout=_GIT_TIMEOUT)
|
||||
|
||||
@@ -780,12 +621,7 @@ def _read_small(path: Path) -> str:
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ProjectFacts:
|
||||
"""Structured project facts — the model's verify loop, detected once.
|
||||
|
||||
The same data that feeds the workspace snapshot, exposed structurally so
|
||||
non-prompt consumers (e.g. the desktop verify UI) read it instead of
|
||||
re-detecting and drifting from the prompt.
|
||||
"""
|
||||
"""Structured project facts — exposed so non-prompt consumers (desktop verify UI) don't re-detect."""
|
||||
|
||||
manifests: list[str]
|
||||
package_managers: list[str]
|
||||
@@ -796,9 +632,8 @@ class ProjectFacts:
|
||||
def detect_project_facts(root: Path) -> ProjectFacts:
|
||||
"""Detect manifests, package manager(s), verify commands, and context files.
|
||||
|
||||
Cheap: stat calls plus reads of a couple of small files. The single source
|
||||
of truth for both the prompt snapshot (:func:`_project_facts`) and the
|
||||
gateway's ``project.facts`` — so the UI never re-sniffs verify commands.
|
||||
Single source of truth for the prompt snapshot and the gateway's
|
||||
``project.facts``. Cheap: stat calls plus a couple of small file reads.
|
||||
"""
|
||||
manifests = [m for m in _PROJECT_MARKERS if m not in _CONTEXT_FILES and (root / m).is_file()]
|
||||
package_managers = list(
|
||||
@@ -833,16 +668,9 @@ def detect_project_facts(root: Path) -> ProjectFacts:
|
||||
|
||||
|
||||
def _project_facts(root: Path) -> list[str]:
|
||||
"""Render :func:`detect_project_facts` as workspace-snapshot lines.
|
||||
|
||||
Hands the model its *verify loop* up front — which manifest, which package
|
||||
manager, and the exact test/lint/build commands — instead of making it
|
||||
rediscover them every session. Built once at prompt-build time; the string
|
||||
output must stay byte-stable to preserve the prompt cache.
|
||||
"""
|
||||
"""Render :func:`detect_project_facts` as workspace-snapshot lines (byte-stable)."""
|
||||
f = detect_project_facts(root)
|
||||
facts: list[str] = []
|
||||
|
||||
if f.manifests:
|
||||
line = f"- Project: {', '.join(f.manifests[:6])}"
|
||||
if f.package_managers:
|
||||
@@ -852,22 +680,25 @@ def _project_facts(root: Path) -> list[str]:
|
||||
facts.append(f"- Verify: {'; '.join(f.verify_commands)}")
|
||||
if f.context_files:
|
||||
facts.append(f"- Context files: {', '.join(f.context_files)}")
|
||||
|
||||
return facts
|
||||
|
||||
|
||||
def _workspace_roots(cwd: Optional[str | Path]) -> tuple[Optional[Path], Optional[Path]]:
|
||||
"""(git_root, workspace_root) for *cwd*; workspace root is git root else marker root."""
|
||||
resolved = _resolve_cwd(cwd)
|
||||
git_root = _git_root(resolved)
|
||||
return git_root, git_root or _marker_root(resolved)
|
||||
|
||||
|
||||
def project_facts_for(cwd: Optional[str | Path] = None) -> Optional[dict[str, Any]]:
|
||||
"""Structured project facts for ``cwd`` — ``None`` outside a workspace.
|
||||
|
||||
Same detection the system-prompt snapshot uses (git root, else marker root),
|
||||
exposed for non-prompt consumers (the desktop verify UI) so they never
|
||||
re-derive "are we coding?" or duplicate the verify-command sniffing.
|
||||
Same detection the system-prompt snapshot uses, exposed for non-prompt
|
||||
consumers (the desktop verify UI).
|
||||
"""
|
||||
resolved = _resolve_cwd(cwd)
|
||||
root = _git_root(resolved) or _marker_root(resolved)
|
||||
_, root = _workspace_roots(cwd)
|
||||
if root is None:
|
||||
return None
|
||||
|
||||
f = detect_project_facts(root)
|
||||
return {
|
||||
"root": str(root),
|
||||
@@ -881,13 +712,10 @@ def project_facts_for(cwd: Optional[str | Path] = None) -> Optional[dict[str, An
|
||||
def build_coding_workspace_block(cwd: Optional[str | Path] = None) -> str:
|
||||
"""Workspace snapshot for the system prompt (empty outside a workspace).
|
||||
|
||||
Git state (branch/status/commits) when the cwd is in a repo, plus detected
|
||||
project facts (manifest, package manager, verify commands, context files)
|
||||
— so marker-only (non-git) projects still get a snapshot.
|
||||
Git state when the cwd is in a repo, plus detected project facts — so
|
||||
marker-only (non-git) projects still get a snapshot.
|
||||
"""
|
||||
resolved = _resolve_cwd(cwd)
|
||||
git_root = _git_root(resolved)
|
||||
root = git_root or _marker_root(resolved)
|
||||
git_root, root = _workspace_roots(cwd)
|
||||
if root is None:
|
||||
return ""
|
||||
|
||||
@@ -908,11 +736,9 @@ def build_coding_workspace_block(cwd: Optional[str | Path] = None) -> str:
|
||||
elif head == "(detached)":
|
||||
lines.append("- Branch: (detached HEAD)")
|
||||
|
||||
# Linked worktree: the per-worktree git dir differs from the shared common dir.
|
||||
# We surface the fact that it's a worktree (so the model knows branches/stashes
|
||||
# are shared state) but deliberately do NOT expose the primary tree path —
|
||||
# giving the model a second absolute path causes it to sometimes run commands
|
||||
# in the wrong directory.
|
||||
# Linked worktree: say so (branches/stashes are shared state) but do
|
||||
# NOT expose the primary tree path — a second absolute path makes the
|
||||
# model run commands in the wrong directory.
|
||||
git_dir, common_dir = _git(root, "rev-parse", "--git-dir"), _git(root, "rev-parse", "--git-common-dir")
|
||||
if git_dir and common_dir and Path(git_dir).resolve() != Path(common_dir).resolve():
|
||||
lines.append("- Worktree: linked (git state shared with primary tree)")
|
||||
|
||||
+98
-161
@@ -1,3 +1,5 @@
|
||||
"""@-reference expansion (``@file:``, ``@folder:``, ``@diff``, ``@git:``, ``@url:`` + plugin prefixes)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
@@ -7,6 +9,7 @@ import mimetypes
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
from abc import ABC, abstractmethod
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Awaitable, Callable
|
||||
@@ -20,10 +23,8 @@ from hermes_cli._subprocess_compat import (
|
||||
)
|
||||
from hermes_cli.sizefmt import format_bytes
|
||||
|
||||
from abc import ABC, abstractmethod
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Plugin context-reference provider API (Issue #26193)
|
||||
# Plugin context-reference provider API
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
BUILTIN_PREFIXES = frozenset({"diff", "staged", "file", "folder", "git", "url"})
|
||||
@@ -43,11 +44,7 @@ class ContextCompletionItem:
|
||||
|
||||
|
||||
class ContextReferenceProvider(ABC):
|
||||
"""Base class for plugin-registered @-prefix context reference providers.
|
||||
|
||||
Plugins subclass this and register via
|
||||
``PluginContext.register_context_reference()``.
|
||||
"""
|
||||
"""Base class for plugin @-prefix providers, registered via ``PluginContext.register_context_reference()``."""
|
||||
|
||||
prefix: str = "" # e.g. "issue", "channel", "doc"
|
||||
description: str = "" # shown in autocomplete meta column
|
||||
@@ -86,8 +83,7 @@ _QUOTED_REFERENCE_VALUE = r'(?:`[^`\n]+`|"[^"\n]+"|\'[^\'\n]+\')'
|
||||
REFERENCE_PATTERN = re.compile(
|
||||
rf"(?<![\w/])@(?:(?P<simple>diff|staged)\b|(?P<kind>file|folder|git|url):(?P<value>{_QUOTED_REFERENCE_VALUE}(?::\d+(?:-\d+)?)?|\S+))"
|
||||
)
|
||||
# Plugin fallback pattern – catches any @<word>:<value> not handled by the
|
||||
# built-in regex so that plugin-registered prefixes can be resolved.
|
||||
# Plugin fallback: any @<word>:<value> the built-in regex did not claim.
|
||||
_PLUGIN_REFERENCE_PATTERN = re.compile(
|
||||
rf"(?<![\w/])@(?P<kind>[a-zA-Z][a-zA-Z0-9_-]*):(?P<value>{_QUOTED_REFERENCE_VALUE}(?::\d+(?:-\d+)?)?|\S+)"
|
||||
)
|
||||
@@ -138,9 +134,8 @@ class ContextReferenceResult:
|
||||
def format_reference_value(value: str) -> str:
|
||||
"""Quote a reference value so ``REFERENCE_PATTERN`` reads it back whole.
|
||||
|
||||
The unquoted alternative in the pattern is ``\\S+``, so a path containing a
|
||||
space parses as a truncated ref with the tail left behind as loose text.
|
||||
Mirrors ``formatRefValue`` in the desktop's directive-text.tsx.
|
||||
The unquoted alternative is ``\\S+``, so a path with a space would parse as a
|
||||
truncated ref. Mirrors ``formatRefValue`` in the desktop's directive-text.tsx.
|
||||
"""
|
||||
if not _NEEDS_QUOTING.search(value):
|
||||
return value
|
||||
@@ -158,26 +153,14 @@ def parse_context_references(message: str) -> list[ContextReference]:
|
||||
for match in REFERENCE_PATTERN.finditer(message):
|
||||
simple = match.group("simple")
|
||||
if simple:
|
||||
refs.append(
|
||||
ContextReference(
|
||||
raw=match.group(0),
|
||||
kind=simple,
|
||||
target="",
|
||||
start=match.start(),
|
||||
end=match.end(),
|
||||
)
|
||||
)
|
||||
refs.append(ContextReference(raw=match.group(0), kind=simple, target="", start=match.start(), end=match.end()))
|
||||
continue
|
||||
|
||||
kind = match.group("kind")
|
||||
value = _strip_trailing_punctuation(match.group("value") or "")
|
||||
line_start = None
|
||||
line_end = None
|
||||
target = _strip_reference_wrappers(value)
|
||||
|
||||
if kind == "file":
|
||||
target, line_start, line_end = _parse_file_reference_value(value)
|
||||
|
||||
else:
|
||||
target, line_start, line_end = _strip_reference_wrappers(value), None, None
|
||||
refs.append(
|
||||
ContextReference(
|
||||
raw=match.group(0),
|
||||
@@ -190,26 +173,24 @@ def parse_context_references(message: str) -> list[ContextReference]:
|
||||
)
|
||||
)
|
||||
|
||||
# Second pass: resolve plugin-registered prefixes the built-in pattern missed
|
||||
# Second pass: plugin-registered prefixes the built-in pattern missed.
|
||||
if _context_reference_providers:
|
||||
for match in _PLUGIN_REFERENCE_PATTERN.finditer(message):
|
||||
kind = match.group("kind")
|
||||
if kind in BUILTIN_PREFIXES:
|
||||
if kind in BUILTIN_PREFIXES or kind not in _context_reference_providers:
|
||||
continue
|
||||
# Skip if already captured by the built-in pattern
|
||||
if any(r.kind == kind and r.start == match.start() for r in refs):
|
||||
continue
|
||||
if kind in _context_reference_providers:
|
||||
value = _strip_trailing_punctuation(match.group("value") or "")
|
||||
refs.append(
|
||||
ContextReference(
|
||||
raw=match.group(0),
|
||||
kind=kind,
|
||||
target=_strip_reference_wrappers(value),
|
||||
start=match.start(),
|
||||
end=match.end(),
|
||||
)
|
||||
value = _strip_trailing_punctuation(match.group("value") or "")
|
||||
refs.append(
|
||||
ContextReference(
|
||||
raw=match.group(0),
|
||||
kind=kind,
|
||||
target=_strip_reference_wrappers(value),
|
||||
start=match.start(),
|
||||
end=match.end(),
|
||||
)
|
||||
)
|
||||
|
||||
return refs
|
||||
|
||||
@@ -222,6 +203,7 @@ def preprocess_context_references(
|
||||
url_fetcher: Callable[[str], str | Awaitable[str]] | None = None,
|
||||
allowed_root: str | Path | None = None,
|
||||
) -> ContextReferenceResult:
|
||||
"""Sync wrapper; safe both without a loop (CLI) and inside a running loop (gateway)."""
|
||||
coro = preprocess_context_references_async(
|
||||
message,
|
||||
cwd=cwd,
|
||||
@@ -229,7 +211,6 @@ def preprocess_context_references(
|
||||
url_fetcher=url_fetcher,
|
||||
allowed_root=allowed_root,
|
||||
)
|
||||
# Safe for both CLI (no loop) and gateway (loop already running).
|
||||
try:
|
||||
loop = asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
@@ -254,31 +235,17 @@ async def preprocess_context_references_async(
|
||||
return ContextReferenceResult(message=message, original_message=message)
|
||||
|
||||
cwd_path = Path(cwd).expanduser().resolve()
|
||||
# Default to the current working directory so @ references cannot escape
|
||||
# the active workspace unless a caller explicitly widens the root.
|
||||
allowed_root_path = (
|
||||
Path(allowed_root).expanduser().resolve() if allowed_root is not None else cwd_path
|
||||
)
|
||||
# Default root = cwd so @ references cannot escape the workspace unless a caller widens it.
|
||||
allowed_root_path = Path(allowed_root).expanduser().resolve() if allowed_root is not None else cwd_path
|
||||
warnings: list[str] = []
|
||||
blocks: list[str] = []
|
||||
injected_tokens = 0
|
||||
|
||||
# Expand all references concurrently. Each _expand_reference is independent
|
||||
# (no shared state during expansion) — a message with several @url: refs
|
||||
# would otherwise pay one full web_extract round-trip per ref in series.
|
||||
# gather preserves positional order, so we reassemble warnings/blocks in the
|
||||
# original ref order exactly as the prior serial loop did; the token-budget
|
||||
# check below is unchanged (it runs once, after all refs are expanded).
|
||||
# Expand concurrently (each ref is independent; several @url: refs would otherwise
|
||||
# serialize web_extract round-trips). gather preserves order, so warnings/blocks
|
||||
# are assembled in ref order; the token-budget check runs once afterwards.
|
||||
expanded = await asyncio.gather(
|
||||
*(
|
||||
_expand_reference(
|
||||
ref,
|
||||
cwd_path,
|
||||
url_fetcher=url_fetcher,
|
||||
allowed_root=allowed_root_path,
|
||||
)
|
||||
for ref in refs
|
||||
)
|
||||
*(_expand_reference(ref, cwd_path, url_fetcher=url_fetcher, allowed_root=allowed_root_path) for ref in refs)
|
||||
)
|
||||
for warning, block in expanded:
|
||||
if warning:
|
||||
@@ -302,17 +269,14 @@ async def preprocess_context_references_async(
|
||||
expanded=False,
|
||||
blocked=True,
|
||||
)
|
||||
|
||||
if injected_tokens > soft_limit:
|
||||
warnings.append(
|
||||
f"@ context injection warning: {injected_tokens} tokens exceeds the 25% soft limit ({soft_limit})."
|
||||
)
|
||||
|
||||
# Leave the `@file:`/`@folder:` tokens where the user typed them. The token
|
||||
# IS the reference, not scaffolding around it: clients render each one as an
|
||||
# inline chip, so stripping them left a sentence with a hole in it ("review
|
||||
# and ship") and made the desktop re-derive the refs from the attached block
|
||||
# to show them as a detached list above the prose.
|
||||
# The `@file:`/`@folder:` tokens stay where the user typed them: the token IS the
|
||||
# reference (clients render it as an inline chip); stripping it left a hole in the
|
||||
# sentence and forced the desktop to re-derive refs from the attached block.
|
||||
final = message
|
||||
if warnings:
|
||||
final = f"{final}\n\n--- Context Warnings ---\n" + "\n".join(f"- {warning}" for warning in warnings)
|
||||
@@ -337,6 +301,7 @@ async def _expand_reference(
|
||||
url_fetcher: Callable[[str], str | Awaitable[str]] | None = None,
|
||||
allowed_root: Path | None = None,
|
||||
) -> tuple[str | None, str | None]:
|
||||
"""Return ``(warning, block)`` for one reference; exactly one side is set."""
|
||||
try:
|
||||
if ref.kind == "file":
|
||||
return _expand_file_reference(ref, cwd, allowed_root=allowed_root)
|
||||
@@ -357,7 +322,6 @@ async def _expand_reference(
|
||||
except Exception as exc:
|
||||
return f"{ref.raw}: {exc}", None
|
||||
|
||||
# Plugin-provided context references
|
||||
provider = _context_reference_providers.get(ref.kind)
|
||||
if provider is not None:
|
||||
try:
|
||||
@@ -383,13 +347,8 @@ def _expand_file_reference(
|
||||
if not path.is_file():
|
||||
return f"{ref.raw}: path is not a file", None
|
||||
if _is_binary_file(path):
|
||||
# A binary file can't be inlined as text, but it IS on disk (the agent's
|
||||
# tools run where this resolves — the local cwd, or the staged copy in a
|
||||
# remote session workspace). Returning a bare "not supported" warning
|
||||
# with no content was a dead end: the model saw a failure and gave up
|
||||
# (told the user the file type wasn't supported). Instead, hand it an
|
||||
# actionable block — the path, type, size, and a nudge to use its tools —
|
||||
# so it can read/convert/view the file itself.
|
||||
# A bare "not supported" warning was a dead end (the model gave up); the file IS
|
||||
# on disk where the agent's tools run, so hand it an actionable block instead.
|
||||
return None, _binary_reference_block(ref, path)
|
||||
|
||||
text = path.read_text(encoding="utf-8")
|
||||
@@ -400,8 +359,7 @@ def _expand_file_reference(
|
||||
text = "\n".join(lines[start_idx:end_idx])
|
||||
|
||||
lang = _code_fence_language(path)
|
||||
label = ref.raw
|
||||
return None, f"📄 {label} ({estimate_tokens_rough(text)} tokens)\n```{lang}\n{text}\n```"
|
||||
return None, f"📄 {ref.raw} ({estimate_tokens_rough(text)} tokens)\n```{lang}\n{text}\n```"
|
||||
|
||||
|
||||
def _expand_folder_reference(
|
||||
@@ -416,37 +374,45 @@ def _expand_folder_reference(
|
||||
return f"{ref.raw}: folder not found", None
|
||||
if not path.is_dir():
|
||||
return f"{ref.raw}: path is not a folder", None
|
||||
|
||||
listing = _build_folder_listing(path, cwd)
|
||||
return None, f"📁 {ref.raw} ({estimate_tokens_rough(listing)} tokens)\n{listing}"
|
||||
|
||||
|
||||
def _run_quiet(
|
||||
cmd: list[str], cwd: Path, timeout: int, env: dict | None = None
|
||||
) -> subprocess.CompletedProcess:
|
||||
"""subprocess.run with captured text output, no stdin, and no console flash on Windows."""
|
||||
popen_kwargs: dict = {"creationflags": windows_hide_flags()} if IS_WINDOWS else {}
|
||||
if env is not None:
|
||||
popen_kwargs["env"] = env
|
||||
return subprocess.run(
|
||||
cmd,
|
||||
cwd=cwd,
|
||||
capture_output=True,
|
||||
text=True, encoding='utf-8', errors='replace',
|
||||
timeout=timeout,
|
||||
stdin=subprocess.DEVNULL,
|
||||
**popen_kwargs,
|
||||
)
|
||||
|
||||
|
||||
def _expand_git_reference(
|
||||
ref: ContextReference,
|
||||
cwd: Path,
|
||||
args: list[str],
|
||||
label: str,
|
||||
) -> tuple[str | None, str | None]:
|
||||
_popen_kwargs = {"creationflags": windows_hide_flags()} if IS_WINDOWS else {}
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["git", *harden_git_argv(args)],
|
||||
cwd=cwd,
|
||||
capture_output=True,
|
||||
text=True, encoding='utf-8', errors='replace',
|
||||
timeout=30,
|
||||
stdin=subprocess.DEVNULL,
|
||||
env=noninteractive_git_env(),
|
||||
**_popen_kwargs,
|
||||
# Repo-supplied config/attributes must never execute code (GHSA-7x36-8jrh-v4pw).
|
||||
result = _run_quiet(
|
||||
["git", *harden_git_argv(args)], cwd, 30, env=noninteractive_git_env()
|
||||
)
|
||||
except subprocess.TimeoutExpired:
|
||||
return f"{ref.raw}: git command timed out (30s)", None
|
||||
if result.returncode != 0:
|
||||
stderr = (result.stderr or "").strip() or "git command failed"
|
||||
return f"{ref.raw}: {stderr}", None
|
||||
content = result.stdout.strip()
|
||||
if not content:
|
||||
content = "(no output)"
|
||||
content = result.stdout.strip() or "(no output)"
|
||||
return None, f"🧾 {label} ({estimate_tokens_rough(content)} tokens)\n```diff\n{content}\n```"
|
||||
|
||||
|
||||
@@ -466,8 +432,7 @@ async def _default_url_fetcher(url: str) -> str:
|
||||
from tools.web_tools import web_extract_tool
|
||||
|
||||
raw = await web_extract_tool([url], format="markdown")
|
||||
payload = json.loads(raw)
|
||||
docs = payload.get("results", [])
|
||||
docs = json.loads(raw).get("results", [])
|
||||
if not docs:
|
||||
return ""
|
||||
doc = docs[0]
|
||||
@@ -488,6 +453,7 @@ def _resolve_path(cwd: Path, target: str, *, allowed_root: Path | None = None) -
|
||||
|
||||
|
||||
def _ensure_reference_path_allowed(path: Path) -> None:
|
||||
"""Refuse credential/internal paths. Fails CLOSED: the gateway feeds untrusted remote text here."""
|
||||
from hermes_constants import get_hermes_home
|
||||
home = Path(os.path.expanduser("~")).resolve()
|
||||
hermes_home = get_hermes_home().resolve()
|
||||
@@ -499,7 +465,6 @@ def _ensure_reference_path_allowed(path: Path) -> None:
|
||||
|
||||
if path in blocked_exact:
|
||||
raise ValueError("path is a sensitive credential file and cannot be attached")
|
||||
|
||||
for blocked_dir in blocked_dirs:
|
||||
try:
|
||||
path.relative_to(blocked_dir)
|
||||
@@ -507,16 +472,9 @@ def _ensure_reference_path_allowed(path: Path) -> None:
|
||||
continue
|
||||
raise ValueError("path is a sensitive credential or internal Hermes path and cannot be attached")
|
||||
|
||||
# Anchor to the canonical read deny-list (agent/file_safety.get_read_block_error),
|
||||
# the single source of truth used by the file/terminal read path. The narrow
|
||||
# list above predates that guard and never caught the real credential stores:
|
||||
# provider keys (auth.json), Anthropic OAuth tokens (.anthropic_oauth.json),
|
||||
# MCP OAuth material (mcp-tokens/), webhook HMAC secrets, and project-local
|
||||
# .env files. That gap matters because the gateway feeds UNTRUSTED remote
|
||||
# message text into reference expansion, so `@file:~/.hermes/auth.json` from a
|
||||
# chat peer would otherwise read the operator's keys straight into context.
|
||||
# Routing through the canonical guard closes the gap today and keeps this path
|
||||
# protected automatically whenever that deny-list grows.
|
||||
# Anchor to the canonical read deny-list (agent/file_safety.get_read_block_error): the
|
||||
# narrow list above never caught auth.json, .anthropic_oauth.json, mcp-tokens/, webhook
|
||||
# secrets or project .env files, and it grows automatically with that deny-list.
|
||||
try:
|
||||
from agent.file_safety import get_read_block_error
|
||||
|
||||
@@ -527,13 +485,8 @@ def _ensure_reference_path_allowed(path: Path) -> None:
|
||||
except ValueError:
|
||||
raise
|
||||
except Exception:
|
||||
# Fail CLOSED on the security path. This guard exists specifically to
|
||||
# cover credential stores the narrow list above misses (auth.json,
|
||||
# .anthropic_oauth.json, mcp-tokens/, ...). If the canonical lookup
|
||||
# ever fails, silently falling through would re-open that exact hole —
|
||||
# the gateway feeds untrusted remote text here, so a probe could then
|
||||
# attach the operator's keys. Refuse instead: a spurious block on a
|
||||
# legitimate file is a recoverable annoyance; a leaked credential is not.
|
||||
# If the canonical lookup fails, falling through would re-open the exact hole this
|
||||
# guard closes; a spurious block is recoverable, a leaked credential is not.
|
||||
raise ValueError(
|
||||
"path could not be verified against the credential deny-list and cannot be attached"
|
||||
)
|
||||
@@ -583,27 +536,26 @@ def _parse_file_reference_value(value: str) -> tuple[str, int | None, int | None
|
||||
return _strip_reference_wrappers(value), None, None
|
||||
|
||||
|
||||
_TEXT_EXTENSIONS = (".py", ".md", ".txt", ".json", ".yaml", ".yml", ".toml", ".js", ".ts")
|
||||
|
||||
|
||||
def _is_binary_file(path: Path) -> bool:
|
||||
mime, _ = mimetypes.guess_type(path.name)
|
||||
if mime and not mime.startswith("text/") and not any(
|
||||
path.name.endswith(ext) for ext in (".py", ".md", ".txt", ".json", ".yaml", ".yml", ".toml", ".js", ".ts")
|
||||
):
|
||||
if mime and not mime.startswith("text/") and not path.name.endswith(_TEXT_EXTENSIONS):
|
||||
return True
|
||||
chunk = path.read_bytes()[:4096]
|
||||
return b"\x00" in chunk
|
||||
return b"\x00" in path.read_bytes()[:4096]
|
||||
|
||||
|
||||
def _build_folder_listing(path: Path, cwd: Path, limit: int = 200) -> str:
|
||||
lines = [f"{path.relative_to(cwd)}/"]
|
||||
entries = _iter_visible_entries(path, cwd, limit=limit)
|
||||
base_depth = len(path.relative_to(cwd).parts)
|
||||
for entry in entries:
|
||||
rel = entry.relative_to(cwd)
|
||||
indent = " " * max(len(rel.parts) - len(path.relative_to(cwd).parts) - 1, 0)
|
||||
indent = " " * max(len(entry.relative_to(cwd).parts) - base_depth - 1, 0)
|
||||
if entry.is_dir():
|
||||
lines.append(f"{indent}- {entry.name}/")
|
||||
else:
|
||||
meta = _file_metadata(entry)
|
||||
lines.append(f"{indent}- {entry.name} ({meta})")
|
||||
lines.append(f"{indent}- {entry.name} ({_file_metadata(entry)})")
|
||||
if len(entries) >= limit:
|
||||
lines.append("- ...")
|
||||
return "\n".join(lines)
|
||||
@@ -629,29 +581,16 @@ def _iter_visible_entries(path: Path, cwd: Path, limit: int) -> list[Path]:
|
||||
dirs[:] = sorted(d for d in dirs if not d.startswith(".") and d != "__pycache__")
|
||||
files = sorted(f for f in files if not f.startswith("."))
|
||||
root_path = Path(root)
|
||||
for d in dirs:
|
||||
output.append(root_path / d)
|
||||
if len(output) >= limit:
|
||||
return output
|
||||
for f in files:
|
||||
output.append(root_path / f)
|
||||
for name in dirs + files:
|
||||
output.append(root_path / name)
|
||||
if len(output) >= limit:
|
||||
return output
|
||||
return output
|
||||
|
||||
|
||||
def _rg_files(path: Path, cwd: Path, limit: int) -> list[Path] | None:
|
||||
_popen_kwargs = {"creationflags": windows_hide_flags()} if IS_WINDOWS else {}
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["rg", "--files", str(path.relative_to(cwd))],
|
||||
cwd=cwd,
|
||||
capture_output=True,
|
||||
text=True, encoding='utf-8', errors='replace',
|
||||
timeout=10,
|
||||
stdin=subprocess.DEVNULL,
|
||||
**_popen_kwargs,
|
||||
)
|
||||
result = _run_quiet(["rg", "--files", str(path.relative_to(cwd))], cwd, 10)
|
||||
except (FileNotFoundError, OSError, subprocess.TimeoutExpired):
|
||||
return None
|
||||
if result.returncode != 0:
|
||||
@@ -661,19 +600,15 @@ def _rg_files(path: Path, cwd: Path, limit: int) -> list[Path] | None:
|
||||
|
||||
|
||||
def _agent_visible_path(path: Path) -> str:
|
||||
"""Map a host path to the path the agent's tools can read in the active backend.
|
||||
"""Map a host path to what the agent's tools can read in the active backend.
|
||||
|
||||
Under a container backend (docker) the gateway host path dangles inside the
|
||||
sandbox — the container has its own filesystem and the host path is not
|
||||
mounted. Files staged into an auto-mounted cache dir (``images/``,
|
||||
``attachments/``, ...) are translated to their in-container path via the
|
||||
existing ``tools.credential_files`` machinery (#76577). Falls back to the
|
||||
host path when the backend is local or translation is unavailable.
|
||||
Under a container backend the host path dangles inside the sandbox; files staged
|
||||
into an auto-mounted cache dir are translated via ``tools.credential_files``.
|
||||
Falls back to the host path when the backend is local or translation fails.
|
||||
"""
|
||||
try:
|
||||
# Desktop/in-process gateways may not have bridged ``terminal.*``
|
||||
# config into ``TERMINAL_ENV`` at startup; run the idempotent bridge so
|
||||
# the credential_files translation gate sees the active backend.
|
||||
# In-process gateways may not have bridged terminal.* config into TERMINAL_ENV
|
||||
# yet; run the idempotent bridge so the translation gate sees the active backend.
|
||||
from tools.terminal_tool import _ensure_terminal_env_bridged
|
||||
|
||||
_ensure_terminal_env_bridged()
|
||||
@@ -709,18 +644,20 @@ def _file_metadata(path: Path) -> str:
|
||||
return f"{line_count} lines"
|
||||
|
||||
|
||||
_FENCE_LANGUAGES = {
|
||||
".py": "python",
|
||||
".js": "javascript",
|
||||
".ts": "typescript",
|
||||
".tsx": "tsx",
|
||||
".jsx": "jsx",
|
||||
".json": "json",
|
||||
".md": "markdown",
|
||||
".sh": "bash",
|
||||
".yml": "yaml",
|
||||
".yaml": "yaml",
|
||||
".toml": "toml",
|
||||
}
|
||||
|
||||
|
||||
def _code_fence_language(path: Path) -> str:
|
||||
mapping = {
|
||||
".py": "python",
|
||||
".js": "javascript",
|
||||
".ts": "typescript",
|
||||
".tsx": "tsx",
|
||||
".jsx": "jsx",
|
||||
".json": "json",
|
||||
".md": "markdown",
|
||||
".sh": "bash",
|
||||
".yml": "yaml",
|
||||
".yaml": "yaml",
|
||||
".toml": "toml",
|
||||
}
|
||||
return mapping.get(path.suffix.lower(), "")
|
||||
return _FENCE_LANGUAGES.get(path.suffix.lower(), "")
|
||||
|
||||
+342
-608
File diff suppressed because it is too large
Load Diff
+33
-119
@@ -1,32 +1,8 @@
|
||||
"""Lightweight internationalization (i18n) for Hermes static user-facing messages.
|
||||
"""Lightweight i18n for Hermes' static user-facing strings (approval prompts, a few gateway replies).
|
||||
|
||||
Scope (thin slice, by design): only the highest-impact static strings shown
|
||||
to the user by Hermes itself -- approval prompts, a handful of gateway slash
|
||||
command replies, restart-drain notices. Agent-generated output, log lines,
|
||||
error tracebacks, tool outputs, and slash-command descriptions all stay in
|
||||
English.
|
||||
|
||||
Catalog files live under ``locales/<lang>.yaml`` at the repo root. Each
|
||||
catalog is a flat dict keyed by dotted paths (e.g. ``approval.choose`` or
|
||||
``gateway.approval_expired``). Missing keys fall back to English; if English
|
||||
is missing too, the key path itself is returned so a broken catalog never
|
||||
crashes the agent.
|
||||
|
||||
Usage::
|
||||
|
||||
from agent.i18n import t
|
||||
print(t("approval.choose_long")) # current lang
|
||||
print(t("gateway.draining", count=3)) # {count} formatted
|
||||
print(t("approval.choose_long", lang="zh")) # explicit override
|
||||
|
||||
Language resolution order:
|
||||
1. Explicit ``lang=`` argument passed to :func:`t`
|
||||
2. ``HERMES_LANGUAGE`` environment variable (for tests / quick override)
|
||||
3. ``display.language`` from config.yaml
|
||||
4. ``"en"`` (baseline)
|
||||
|
||||
Supported languages: en, zh, zh-hant, ja, de, es, fr, tr, uk, af, ko, it, ga,
|
||||
pt, ru, hu, ar. Unknown values fall back to en.
|
||||
Catalogs are ``locales/<lang>.yaml`` flattened to dotted keys. Missing keys
|
||||
fall back to English, then to the key itself, so a broken catalog never crashes.
|
||||
Language resolution: explicit ``lang=`` > ``HERMES_LANGUAGE`` > ``display.language`` > ``en``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -46,15 +22,13 @@ SUPPORTED_LANGUAGES: tuple[str, ...] = (
|
||||
)
|
||||
DEFAULT_LANGUAGE = "en"
|
||||
|
||||
# Accept a few natural aliases so users who type "chinese" / "zh-CN" / "jp"
|
||||
# get the right catalog instead of silently falling back to English.
|
||||
# Natural aliases so "chinese" / "zh-CN" / "jp" hit the right catalog instead of
|
||||
# silently falling back to English. Bare "chinese" defaults to Simplified;
|
||||
# Taiwan/HK/Macau tags route to the distinct Traditional catalog. pt-br shares
|
||||
# the pt catalog (no separate br one).
|
||||
_LANGUAGE_ALIASES: dict[str, str] = {
|
||||
"english": "en", "en-us": "en", "en-gb": "en",
|
||||
# Simplified Chinese — explicit codes route here; bare "chinese" / "mandarin"
|
||||
# also default to Simplified since that's the larger user base.
|
||||
"chinese": "zh", "mandarin": "zh", "zh-cn": "zh", "zh-hans": "zh", "zh-sg": "zh",
|
||||
# Traditional Chinese — distinct catalog. Cover Taiwan / Hong Kong / Macau
|
||||
# locale tags plus the common "traditional" alias.
|
||||
"traditional-chinese": "zh-hant", "traditional_chinese": "zh-hant",
|
||||
"zh-tw": "zh-hant", "zh-hk": "zh-hant", "zh-mo": "zh-hant",
|
||||
"japanese": "ja", "jp": "ja", "ja-jp": "ja",
|
||||
@@ -63,23 +37,14 @@ _LANGUAGE_ALIASES: dict[str, str] = {
|
||||
"french": "fr", "français": "fr", "france": "fr", "fr-fr": "fr", "fr-be": "fr", "fr-ca": "fr", "fr-ch": "fr",
|
||||
"ukrainian": "uk", "ukrainisch": "uk", "українська": "uk", "uk-ua": "uk", "ua": "uk",
|
||||
"turkish": "tr", "türkçe": "tr", "tr-tr": "tr",
|
||||
# Afrikaans — South African Dutch-derived language; "af-ZA" is the common BCP-47 tag.
|
||||
"afrikaans": "af", "af-za": "af",
|
||||
# Korean
|
||||
"korean": "ko", "한국어": "ko", "ko-kr": "ko",
|
||||
# Italian
|
||||
"italian": "it", "italiano": "it", "it-it": "it", "it-ch": "it",
|
||||
# Irish (Gaeilge) — ga is the BCP-47 code
|
||||
"irish": "ga", "gaeilge": "ga", "ga-ie": "ga",
|
||||
# Portuguese — bare "portuguese" routes to European Portuguese; pt-br
|
||||
# is in the same family but rendered identically here (no separate br catalog).
|
||||
"portuguese": "pt", "português": "pt", "portugues": "pt",
|
||||
"pt-pt": "pt", "pt-br": "pt", "brazilian": "pt", "brasileiro": "pt",
|
||||
# Russian
|
||||
"russian": "ru", "русский": "ru", "ru-ru": "ru",
|
||||
# Hungarian
|
||||
"hungarian": "hu", "magyar": "hu", "hu-hu": "hu",
|
||||
# Arabic — bare "arabic"/endonym plus the common regional BCP-47 tags.
|
||||
"arabic": "ar", "العربية": "ar",
|
||||
"ar-sa": "ar", "ar-eg": "ar", "ar-ae": "ar", "ar-ma": "ar", "ar-dz": "ar",
|
||||
}
|
||||
@@ -89,18 +54,10 @@ _catalog_lock = threading.Lock()
|
||||
|
||||
|
||||
def _locales_dir() -> Path:
|
||||
"""Return the directory containing locale YAML files.
|
||||
"""Locale dir: ``HERMES_BUNDLED_LOCALES`` (sealed packaging, e.g. Nix) if it exists, else ``<repo-root>/locales``.
|
||||
|
||||
Resolution order, first existing wins:
|
||||
|
||||
1. ``HERMES_BUNDLED_LOCALES`` env var -- set by the Nix wrapper (or any
|
||||
sealed-packaging system) to point at the installed catalog directory.
|
||||
2. ``<repo-root>/locales`` -- source checkouts and editable installs,
|
||||
where the working tree sits next to ``agent/``.
|
||||
|
||||
Falling through to the source-style path (even when missing) keeps
|
||||
``_load_catalog`` error messages informative -- it logs the path it
|
||||
looked at -- rather than raising.
|
||||
The source path is returned even when missing so ``_load_catalog`` can log
|
||||
the path it looked at rather than raise.
|
||||
"""
|
||||
override = os.getenv("HERMES_BUNDLED_LOCALES", "").strip()
|
||||
if override:
|
||||
@@ -112,19 +69,11 @@ def _locales_dir() -> Path:
|
||||
"falling back to bundled/source locale resolution",
|
||||
override,
|
||||
)
|
||||
|
||||
# agent/i18n.py -> agent/ -> repo root (source checkout, editable install)
|
||||
source_dir = Path(__file__).resolve().parent.parent / "locales"
|
||||
return source_dir
|
||||
return Path(__file__).resolve().parent.parent / "locales"
|
||||
|
||||
|
||||
def _normalize_lang(value: Any) -> str:
|
||||
"""Normalize a user-supplied language value to a supported code.
|
||||
|
||||
Accepts supported codes directly, common aliases (``chinese`` -> ``zh``),
|
||||
and case-insensitive regional tags (``zh-CN`` -> ``zh``). Returns the
|
||||
default language for unknown values.
|
||||
"""
|
||||
"""Map a user-supplied value (code, alias, or regional tag like ``zh-CN``) to a supported code, else default."""
|
||||
if not isinstance(value, str):
|
||||
return DEFAULT_LANGUAGE
|
||||
key = value.strip().lower()
|
||||
@@ -134,20 +83,20 @@ def _normalize_lang(value: Any) -> str:
|
||||
return key
|
||||
if key in _LANGUAGE_ALIASES:
|
||||
return _LANGUAGE_ALIASES[key]
|
||||
# Try stripping a region suffix (e.g. "pt-br" -> "pt" won't be supported,
|
||||
# but "zh-CN" -> "zh" will).
|
||||
base = key.split("-", 1)[0]
|
||||
base = key.split("-", 1)[0] # strip region suffix
|
||||
if base in SUPPORTED_LANGUAGES:
|
||||
return base
|
||||
return DEFAULT_LANGUAGE
|
||||
|
||||
|
||||
def _load_catalog(lang: str) -> dict[str, str]:
|
||||
"""Load and flatten one locale YAML file into a dotted-key dict.
|
||||
def _cache_catalog(lang: str, flat: dict[str, str]) -> dict[str, str]:
|
||||
with _catalog_lock:
|
||||
_catalog_cache[lang] = flat
|
||||
return flat
|
||||
|
||||
YAML files can be nested for human readability; this produces the flat
|
||||
key space :func:`t` expects. Cached per-language for the process.
|
||||
"""
|
||||
|
||||
def _load_catalog(lang: str) -> dict[str, str]:
|
||||
"""Load one locale YAML flattened to dotted keys; cached per language (empty dict on any failure)."""
|
||||
with _catalog_lock:
|
||||
cached = _catalog_cache.get(lang)
|
||||
if cached is not None:
|
||||
@@ -156,46 +105,34 @@ def _load_catalog(lang: str) -> dict[str, str]:
|
||||
path = _locales_dir() / f"{lang}.yaml"
|
||||
if not path.is_file():
|
||||
logger.debug("i18n catalog missing for %s at %s", lang, path)
|
||||
with _catalog_lock:
|
||||
_catalog_cache[lang] = {}
|
||||
return {}
|
||||
return _cache_catalog(lang, {})
|
||||
|
||||
try:
|
||||
import yaml # PyYAML is already a hermes dependency
|
||||
import yaml
|
||||
with path.open("r", encoding="utf-8") as f:
|
||||
raw = yaml.safe_load(f) or {}
|
||||
except Exception as exc:
|
||||
logger.warning("Failed to load i18n catalog %s: %s", path, exc)
|
||||
with _catalog_lock:
|
||||
_catalog_cache[lang] = {}
|
||||
return {}
|
||||
return _cache_catalog(lang, {})
|
||||
|
||||
flat: dict[str, str] = {}
|
||||
_flatten_into(raw, "", flat)
|
||||
with _catalog_lock:
|
||||
_catalog_cache[lang] = flat
|
||||
return flat
|
||||
return _cache_catalog(lang, flat)
|
||||
|
||||
|
||||
def _flatten_into(node: Any, prefix: str, out: dict[str, str]) -> None:
|
||||
# Non-string, non-dict leaves are ignored -- catalogs are text-only.
|
||||
if isinstance(node, dict):
|
||||
for key, value in node.items():
|
||||
child_key = f"{prefix}.{key}" if prefix else str(key)
|
||||
_flatten_into(value, child_key, out)
|
||||
elif isinstance(node, str):
|
||||
out[prefix] = node
|
||||
# Non-string, non-dict leaves are ignored -- catalogs are text-only.
|
||||
|
||||
|
||||
@lru_cache(maxsize=1)
|
||||
def _config_language_cached() -> str | None:
|
||||
"""Read ``display.language`` from config.yaml once per process.
|
||||
|
||||
Cached because ``t()`` is called in hot paths (every approval prompt,
|
||||
every gateway reply) and re-reading YAML each call would be wasteful.
|
||||
``reset_language_cache()`` clears this when config changes at runtime
|
||||
(e.g. after the setup wizard).
|
||||
"""
|
||||
"""``display.language`` from config.yaml, read once per process (``t()`` is a hot path)."""
|
||||
try:
|
||||
from hermes_cli.config import load_config_readonly
|
||||
cfg = load_config_readonly()
|
||||
@@ -208,11 +145,7 @@ def _config_language_cached() -> str | None:
|
||||
|
||||
|
||||
def reset_language_cache() -> None:
|
||||
"""Invalidate cached language resolution and catalogs.
|
||||
|
||||
Call after :func:`hermes_cli.config.save_config` if a running process
|
||||
needs to pick up a changed ``display.language`` without restart.
|
||||
"""
|
||||
"""Invalidate cached language resolution and catalogs (call after ``save_config`` changes ``display.language``)."""
|
||||
_config_language_cached.cache_clear()
|
||||
with _catalog_lock:
|
||||
_catalog_cache.clear()
|
||||
@@ -223,41 +156,22 @@ def get_language() -> str:
|
||||
env_lang = os.environ.get("HERMES_LANGUAGE")
|
||||
if env_lang:
|
||||
return _normalize_lang(env_lang)
|
||||
cfg_lang = _config_language_cached()
|
||||
if cfg_lang:
|
||||
return cfg_lang
|
||||
return DEFAULT_LANGUAGE
|
||||
return _config_language_cached() or DEFAULT_LANGUAGE
|
||||
|
||||
|
||||
def t(key: str, lang: str | None = None, **format_kwargs: Any) -> str:
|
||||
"""Translate a dotted key to the active language.
|
||||
"""Translate a dotted catalog key to the active (or explicit ``lang``) language.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
key
|
||||
Dotted path into the catalog, e.g. ``"approval.choose_long"``.
|
||||
lang
|
||||
Explicit language override. Takes precedence over env + config.
|
||||
**format_kwargs
|
||||
``str.format`` substitution arguments (``t("gateway.drain", count=3)``
|
||||
expects a catalog entry with a ``{count}`` placeholder).
|
||||
|
||||
Returns
|
||||
-------
|
||||
The translated string, or the English fallback if the key is missing in
|
||||
the target language, or the bare key if English is also missing.
|
||||
``format_kwargs`` are applied with ``str.format``. Falls back to English,
|
||||
then to the bare key; a format failure returns the unformatted string.
|
||||
"""
|
||||
target = _normalize_lang(lang) if lang else get_language()
|
||||
catalog = _load_catalog(target)
|
||||
value = catalog.get(key)
|
||||
value = _load_catalog(target).get(key)
|
||||
|
||||
if value is None and target != DEFAULT_LANGUAGE:
|
||||
# Fall through to English rather than showing a key path to the user.
|
||||
value = _load_catalog(DEFAULT_LANGUAGE).get(key)
|
||||
|
||||
if value is None:
|
||||
# Last-ditch: return the key itself. A broken catalog should not
|
||||
# crash anything; it just looks ugly until someone fixes it.
|
||||
logger.debug("i18n miss: key=%r lang=%r", key, target)
|
||||
value = key
|
||||
|
||||
|
||||
+72
-155
@@ -1,30 +1,15 @@
|
||||
"""CJK/wide-character-aware re-alignment of model-emitted markdown tables.
|
||||
|
||||
Models pad markdown tables assuming each character occupies one terminal
|
||||
cell. CJK glyphs and most emoji render as two cells, so the model's
|
||||
spacing collapses into drift the moment a table reaches a real terminal —
|
||||
header pipes line up, every body row drifts right by N cells per CJK
|
||||
char.
|
||||
Models pad tables assuming one cell per character; CJK glyphs and most emoji
|
||||
take two, so body rows drift right on real terminals. This rebuilds padding
|
||||
with ``wcwidth.wcswidth`` while preserving pipes/dashes so the table still reads
|
||||
as plain text in ``strip``/unrendered modes (Rich already aligns CJK itself).
|
||||
|
||||
This module rebuilds row padding using ``wcwidth.wcswidth`` (display
|
||||
columns), preserving the table's pipes and dashes so it still reads as a
|
||||
plain-text table in ``strip`` / unrendered display modes. Standard Rich
|
||||
markdown rendering already aligns CJK correctly inside a wide enough
|
||||
panel; this helper is for the paths that print the model's text more or
|
||||
less verbatim.
|
||||
|
||||
The helper is deliberately conservative:
|
||||
|
||||
* Only contiguous ``| ... |`` blocks with a divider line are rewritten.
|
||||
* Anything that does not look like a table is passed through unchanged.
|
||||
* Single-line / mid-stream fragments are left alone — callers buffer
|
||||
table rows and flush them once the block is complete.
|
||||
|
||||
There is a small, intentional caveat: ``wcwidth`` returns ``-1`` for some
|
||||
emoji-with-variation-selector sequences (e.g. ``⚠️``); we clamp those to
|
||||
0 so they do not corrupt the column width math. The 1-cell drift on
|
||||
those specific glyphs is preferable to silently widening every table
|
||||
that contains one.
|
||||
Deliberately conservative: only contiguous ``| ... |`` blocks with a divider are
|
||||
rewritten; everything else passes through; single-line/mid-stream fragments are
|
||||
left alone (callers buffer rows and flush complete blocks). ``wcwidth`` returns
|
||||
``-1`` for some emoji+variation-selector sequences (``⚠️``); those clamp to 0 —
|
||||
a 1-cell drift on that glyph beats widening every table that contains one.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -47,13 +32,7 @@ _MIN_COL_WIDTH = 3 # matches the divider's minimum dash run.
|
||||
|
||||
|
||||
def _disp_width(s: str) -> int:
|
||||
"""``wcswidth`` clamped to a non-negative integer.
|
||||
|
||||
``wcswidth`` returns ``-1`` when it encounters a control char or an
|
||||
unknown sequence; treat those as zero-width rather than letting a
|
||||
negative number flow into ``max`` and break the column-width math.
|
||||
"""
|
||||
|
||||
"""``wcswidth`` clamped to >= 0 (it returns -1 for control/unknown sequences)."""
|
||||
w = wcswidth(s)
|
||||
return w if w > 0 else 0
|
||||
|
||||
@@ -64,7 +43,6 @@ def _pad_to_width(s: str, target: int) -> str:
|
||||
|
||||
def split_table_row(row: str) -> List[str]:
|
||||
"""Split ``| a | b | c |`` into ``["a", "b", "c"]`` with trims."""
|
||||
|
||||
s = row.strip()
|
||||
if s.startswith("|"):
|
||||
s = s[1:]
|
||||
@@ -75,7 +53,6 @@ def split_table_row(row: str) -> List[str]:
|
||||
|
||||
def is_table_divider(row: str) -> bool:
|
||||
"""True when ``row`` is a markdown table separator line."""
|
||||
|
||||
cells = split_table_row(row)
|
||||
return len(cells) > 1 and all(_DIVIDER_CELL_RE.match(c) for c in cells)
|
||||
|
||||
@@ -83,76 +60,70 @@ def is_table_divider(row: str) -> bool:
|
||||
def looks_like_table_row(row: str) -> bool:
|
||||
"""True when ``row`` could plausibly be a markdown table row.
|
||||
|
||||
Used by streaming callers to decide whether to buffer an in-flight
|
||||
line. We are intentionally permissive here — the realigner itself
|
||||
only rewrites blocks that are accompanied by a divider, so a false
|
||||
positive here at most delays the print of one line.
|
||||
Intentionally permissive for streaming callers deciding whether to buffer a
|
||||
line: the realigner only rewrites divider-backed blocks, so a false positive
|
||||
at most delays printing one line. A leading pipe is the strongest signal;
|
||||
without it we accept >= 2 pipes so models that omit the leading pipe still match.
|
||||
"""
|
||||
|
||||
if "|" not in row:
|
||||
return False
|
||||
stripped = row.strip()
|
||||
if not stripped:
|
||||
return False
|
||||
# A leading pipe is the strongest signal; without it we still allow
|
||||
# rows with at least two pipes so models that omit the leading pipe
|
||||
# don't slip past us.
|
||||
if stripped.startswith("|"):
|
||||
return True
|
||||
return stripped.count("|") >= 2
|
||||
return stripped.startswith("|") or stripped.count("|") >= 2
|
||||
|
||||
|
||||
def _render_block(rows: List[List[str]], available_width: int | None = None) -> List[str]:
|
||||
"""Render ``rows`` (header + body, divider implied) at uniform widths.
|
||||
|
||||
If ``available_width`` is given and the rebuilt horizontal table
|
||||
would exceed it, fall back to a vertical key-value rendering so
|
||||
rows do not soft-wrap mid-cell — terminal soft-wrap destroys
|
||||
column alignment visually even when the underlying bytes are
|
||||
perfectly padded, which is exactly the "tables look broken"
|
||||
user report this code path is meant to address.
|
||||
When the horizontal table would exceed ``available_width`` fall back to a
|
||||
vertical key-value rendering: terminal soft-wrap mid-cell destroys alignment
|
||||
visually even when the bytes are perfectly padded.
|
||||
"""
|
||||
|
||||
ncols = max(len(r) for r in rows)
|
||||
rows = [r + [""] * (ncols - len(r)) for r in rows]
|
||||
widths = [max(_MIN_COL_WIDTH, *(_disp_width(r[c]) for r in rows)) for c in range(ncols)]
|
||||
|
||||
widths = [
|
||||
max(_MIN_COL_WIDTH, *(_disp_width(r[c]) for r in rows))
|
||||
for c in range(ncols)
|
||||
]
|
||||
|
||||
# Total horizontal width for the rendered row:
|
||||
# `| ` + cell + ` ` for each column, plus the final closing `|`.
|
||||
# `| ` + cell + ` ` per column, plus the closing `|`.
|
||||
horizontal_width = sum(widths) + 3 * ncols + 1
|
||||
|
||||
if available_width is not None and horizontal_width > max(available_width, 20):
|
||||
return _render_vertical(rows, ncols, available_width)
|
||||
|
||||
def _row(cells: List[str]) -> str:
|
||||
return (
|
||||
"| "
|
||||
+ " | ".join(_pad_to_width(c, widths[k]) for k, c in enumerate(cells))
|
||||
+ " |"
|
||||
)
|
||||
return "| " + " | ".join(_pad_to_width(c, widths[k]) for k, c in enumerate(cells)) + " |"
|
||||
|
||||
out = [_row(rows[0])]
|
||||
out.append("|" + "|".join("-" * (w + 2) for w in widths) + "|")
|
||||
for r in rows[1:]:
|
||||
out.append(_row(r))
|
||||
out = [_row(rows[0]), "|" + "|".join("-" * (w + 2) for w in widths) + "|"]
|
||||
out.extend(_row(r) for r in rows[1:])
|
||||
return out
|
||||
|
||||
|
||||
def _hard_break(word: str, w: int) -> List[str]:
|
||||
"""Split a single over-wide word into display-width-``w`` chunks."""
|
||||
out: List[str] = []
|
||||
buf = ""
|
||||
bw = 0
|
||||
for ch in word:
|
||||
cw = _disp_width(ch) or 1
|
||||
if bw + cw > w and buf:
|
||||
out.append(buf)
|
||||
buf = ch
|
||||
bw = cw
|
||||
else:
|
||||
buf += ch
|
||||
bw += cw
|
||||
if buf:
|
||||
out.append(buf)
|
||||
return out
|
||||
|
||||
|
||||
def _wrap_to_width(text: str, width: int) -> List[str]:
|
||||
"""Soft-wrap ``text`` at word boundaries to fit ``width`` display cells.
|
||||
"""Soft-wrap ``text`` at word boundaries to ``width`` display cells.
|
||||
|
||||
Falls back to hard-breaking the longest word if a single token is
|
||||
wider than ``width``. Empty input yields a single empty string so
|
||||
the caller's row count stays predictable.
|
||||
Words wider than ``width`` are hard-broken. Empty input yields a single
|
||||
empty string so the caller's row count stays predictable.
|
||||
"""
|
||||
|
||||
if width <= 0 or not text:
|
||||
return [text]
|
||||
|
||||
words = text.split()
|
||||
if not words:
|
||||
return [""]
|
||||
@@ -161,100 +132,61 @@ def _wrap_to_width(text: str, width: int) -> List[str]:
|
||||
current = ""
|
||||
current_w = 0
|
||||
|
||||
def _hard_break(word: str, w: int) -> List[str]:
|
||||
out: List[str] = []
|
||||
buf = ""
|
||||
bw = 0
|
||||
for ch in word:
|
||||
cw = _disp_width(ch) or 1
|
||||
if bw + cw > w and buf:
|
||||
out.append(buf)
|
||||
buf = ch
|
||||
bw = cw
|
||||
else:
|
||||
buf += ch
|
||||
bw += cw
|
||||
if buf:
|
||||
out.append(buf)
|
||||
return out
|
||||
def _start(word: str, ww: int) -> None:
|
||||
nonlocal current, current_w
|
||||
if ww <= width:
|
||||
current, current_w = word, ww
|
||||
else:
|
||||
pieces = _hard_break(word, width)
|
||||
lines.extend(pieces[:-1])
|
||||
current = pieces[-1] if pieces else ""
|
||||
current_w = _disp_width(current)
|
||||
|
||||
for word in words:
|
||||
ww = _disp_width(word)
|
||||
if not current:
|
||||
if ww <= width:
|
||||
current = word
|
||||
current_w = ww
|
||||
else:
|
||||
pieces = _hard_break(word, width)
|
||||
lines.extend(pieces[:-1])
|
||||
current = pieces[-1] if pieces else ""
|
||||
current_w = _disp_width(current)
|
||||
continue
|
||||
if current_w + 1 + ww <= width:
|
||||
_start(word, ww)
|
||||
elif current_w + 1 + ww <= width:
|
||||
current += " " + word
|
||||
current_w += 1 + ww
|
||||
else:
|
||||
lines.append(current)
|
||||
if ww <= width:
|
||||
current = word
|
||||
current_w = ww
|
||||
else:
|
||||
pieces = _hard_break(word, width)
|
||||
lines.extend(pieces[:-1])
|
||||
current = pieces[-1] if pieces else ""
|
||||
current_w = _disp_width(current)
|
||||
_start(word, ww)
|
||||
if current:
|
||||
lines.append(current)
|
||||
return lines or [""]
|
||||
|
||||
|
||||
def _render_vertical(
|
||||
rows: List[List[str]], ncols: int, available_width: int
|
||||
) -> List[str]:
|
||||
"""Render a too-wide table as vertical ``Header: value`` rows.
|
||||
def _render_vertical(rows: List[List[str]], ncols: int, available_width: int) -> List[str]:
|
||||
"""Render a too-wide table as ``Header: value`` blocks (Claude Code's narrow fallback).
|
||||
|
||||
Mirrors Claude Code's narrow-terminal fallback in
|
||||
``MarkdownTable.tsx``: each body row becomes a small block of
|
||||
``Header: cell-value`` lines (continuation lines indented two
|
||||
spaces) separated by a thin ``─`` divider between rows. Keeps
|
||||
every line narrower than ``available_width`` so the terminal does
|
||||
not soft-wrap mid-cell.
|
||||
Each body row becomes one block with continuation lines indented two spaces,
|
||||
blocks separated by a thin ``─`` rule; every line stays under ``available_width``.
|
||||
"""
|
||||
|
||||
if not rows:
|
||||
return []
|
||||
|
||||
headers = rows[0] + [""] * (ncols - len(rows[0]))
|
||||
body = rows[1:]
|
||||
|
||||
labels = [h or f"Column {i + 1}" for i, h in enumerate(headers)]
|
||||
|
||||
sep_width = max(20, min(40, available_width - 2)) if available_width else 30
|
||||
separator = "─" * sep_width
|
||||
indent = " "
|
||||
indent_w = _disp_width(indent)
|
||||
cont_budget = max(10, available_width - _disp_width(indent))
|
||||
|
||||
out: List[str] = []
|
||||
for ri, row in enumerate(body):
|
||||
for ri, row in enumerate(rows[1:]):
|
||||
if ri > 0:
|
||||
out.append(separator)
|
||||
for ci in range(ncols):
|
||||
label = labels[ci]
|
||||
value = row[ci] if ci < len(row) else ""
|
||||
label_w = _disp_width(label)
|
||||
first_budget = max(10, available_width - label_w - 2)
|
||||
cont_budget = max(10, available_width - indent_w)
|
||||
if not value:
|
||||
out.append(f"{label}:")
|
||||
continue
|
||||
wrapped = _wrap_to_width(value, first_budget)
|
||||
wrapped = _wrap_to_width(value, max(10, available_width - _disp_width(label) - 2))
|
||||
out.append(f"{label}: {wrapped[0]}")
|
||||
if len(wrapped) > 1:
|
||||
# Re-flow continuation text at the wider continuation
|
||||
# budget — words split across the narrower first-line
|
||||
# budget should re-pack greedily for the rest.
|
||||
cont_text = " ".join(wrapped[1:])
|
||||
for cl in _wrap_to_width(cont_text, cont_budget):
|
||||
# Re-flow continuation text at the wider continuation budget.
|
||||
for cl in _wrap_to_width(" ".join(wrapped[1:]), cont_budget):
|
||||
if cl.strip():
|
||||
out.append(f"{indent}{cl}")
|
||||
return out
|
||||
@@ -263,16 +195,10 @@ def _render_vertical(
|
||||
def realign_markdown_tables(text: str, available_width: int | None = None) -> str:
|
||||
"""Rewrite every ``| ... |`` + divider block with wcwidth-aware padding.
|
||||
|
||||
Lines that are not part of a recognised table are returned verbatim,
|
||||
so this is safe to apply to arbitrary assistant prose.
|
||||
|
||||
If ``available_width`` is given (terminal cells available for the
|
||||
rendered table), tables wider than that are rendered as vertical
|
||||
key-value pairs instead of a horizontal pipe-bordered grid. This
|
||||
avoids the terminal soft-wrapping mid-cell, which destroys column
|
||||
alignment visually even when the bytes are perfectly padded.
|
||||
Non-table lines are returned verbatim, so this is safe on arbitrary prose.
|
||||
With ``available_width`` (terminal cells), tables wider than that render as
|
||||
vertical key-value pairs instead of soft-wrapping mid-cell.
|
||||
"""
|
||||
|
||||
if "|" not in text:
|
||||
return text
|
||||
|
||||
@@ -280,30 +206,21 @@ def realign_markdown_tables(text: str, available_width: int | None = None) -> st
|
||||
out: List[str] = []
|
||||
i = 0
|
||||
n = len(lines)
|
||||
|
||||
while i < n:
|
||||
line = lines[i]
|
||||
# A table starts with a header row whose next line is a divider.
|
||||
if (
|
||||
"|" in line
|
||||
and i + 1 < n
|
||||
and is_table_divider(lines[i + 1])
|
||||
):
|
||||
if "|" in line and i + 1 < n and is_table_divider(lines[i + 1]):
|
||||
header = split_table_row(line)
|
||||
body: List[List[str]] = []
|
||||
j = i + 2
|
||||
while j < n and "|" in lines[j] and lines[j].strip():
|
||||
if is_table_divider(lines[j]):
|
||||
j += 1
|
||||
continue
|
||||
body.append(split_table_row(lines[j]))
|
||||
if not is_table_divider(lines[j]):
|
||||
body.append(split_table_row(lines[j]))
|
||||
j += 1
|
||||
|
||||
if any(c for c in header) or body:
|
||||
out.extend(_render_block([header] + body, available_width))
|
||||
i = j
|
||||
continue
|
||||
out.append(line)
|
||||
i += 1
|
||||
|
||||
return "\n".join(out)
|
||||
|
||||
+236
-569
File diff suppressed because it is too large
Load Diff
+99
-203
@@ -1,43 +1,24 @@
|
||||
"""Native OpenAI Responses server-side compaction — gpt-5.6 on direct OpenAI routes only.
|
||||
|
||||
OpenAI's Responses API supports server-side compaction: include
|
||||
``context_management=[{"type": "compaction", "compact_threshold": N}]`` in a
|
||||
``/v1/responses`` request and, when the rendered input crosses N tokens, the
|
||||
server summarizes older context into an opaque ``compaction`` output item
|
||||
(``encrypted_content``, sealed to the issuing endpoint). Replaying that item
|
||||
as an input item on later requests stands in for the pruned history, so the
|
||||
model keeps long-horizon recall without the client ever seeing a summary.
|
||||
Docs: https://developers.openai.com/api/docs/guides/compaction
|
||||
Including ``context_management=[{"type": "compaction", "compact_threshold": N}]``
|
||||
in a ``/v1/responses`` request makes the server summarize older context into an
|
||||
opaque ``compaction`` item (``encrypted_content``, sealed to the issuing
|
||||
endpoint) once the input crosses N tokens; replaying that item stands in for
|
||||
the pruned history. Docs: https://developers.openai.com/api/docs/guides/compaction
|
||||
|
||||
Hermes' support is deliberately narrow (live verification, Aug 2026):
|
||||
Support is deliberately narrow (live-verified):
|
||||
* gpt-5.6 family only — gpt-5.1/5.2 fail server-side (HTTP 500 blocking, a
|
||||
permanent stall streaming) with no structured "unsupported" rejection, so an
|
||||
explicit model-family check is the only safe gate.
|
||||
* Direct OpenAI routes only (api.openai.com or the ChatGPT Codex backend) —
|
||||
other Responses surfaces would 400 on the field and cannot mint/decrypt the blob.
|
||||
|
||||
* **gpt-5.6 family only.** gpt-5.6 and its variants compact correctly.
|
||||
Sending the field to gpt-5.1 / gpt-5.2 reliably fails server-side —
|
||||
HTTP 500 on the blocking path and a permanent stall on the streaming
|
||||
path (90s watchdog x 3 retries = a dead turn). There is no structured
|
||||
"unsupported" rejection to downgrade on, so the only safe gate is an
|
||||
explicit model-family check.
|
||||
* **Direct OpenAI routes only:** api.openai.com (API key) or the ChatGPT
|
||||
Codex backend (subscription OAuth). Every other Responses surface
|
||||
(xAI, GitHub/Copilot, relays, local servers) never sees the field —
|
||||
most would 400 on the unknown parameter, and none can mint or decrypt
|
||||
the compaction blob.
|
||||
|
||||
Ownership model: Hermes' local compression stays fully armed as the
|
||||
fallback owner. The native threshold is clamped safely below the local
|
||||
compressor's trigger so the server compacts first; if it doesn't (native
|
||||
disabled mid-session, provider hiccup, non-eligible route), the local
|
||||
summarizer fires exactly as before. There is no new custody state — the
|
||||
captured compaction items ride the existing ``codex_reasoning_items``
|
||||
sidecar, which already handles persistence (state.db), gateway session
|
||||
replay, cross-issuer stamping, and the encrypted-replay kill switch.
|
||||
|
||||
This module stays free of transport/adapter dependencies so the transport,
|
||||
adapter, and conversation loop can share the gate without import cycles. The
|
||||
two exceptions — ``agent.context_compressor`` and ``agent.message_content`` —
|
||||
sit below this module in the dependency graph (neither imports
|
||||
``native_compaction``), so importing their provenance/text primitives here
|
||||
introduces no cycle.
|
||||
Hermes' local compressor stays armed as fallback owner: the native threshold is
|
||||
clamped below the local trigger so the server compacts first, and captured
|
||||
compaction items ride the existing ``codex_reasoning_items`` sidecar (persistence,
|
||||
replay, cross-issuer stamping, kill switch). This module stays free of
|
||||
transport/adapter imports so transport, adapter, and loop share the gate
|
||||
without cycles; ``context_compressor`` and ``message_content`` sit below it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -52,14 +33,11 @@ from agent.message_content import flatten_message_text
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Native compaction fires this many tokens below the local compressor's
|
||||
# trigger so the server always gets the first shot at compaction.
|
||||
# trigger so the server always gets the first shot.
|
||||
LOCAL_TRIGGER_SAFETY_MARGIN = 8_192
|
||||
|
||||
# Deterministic fallback when automatic mode cannot inspect a local trigger.
|
||||
# Fallback when automatic mode has no local trigger to follow.
|
||||
DEFAULT_COMPACT_THRESHOLD = 200_000
|
||||
|
||||
# Model-family gate. Substring match on the lowercased model id so dated
|
||||
# snapshots (gpt-5.6-2026-07-xx) and variants (gpt-5.6-mini) stay eligible.
|
||||
# Substring match so dated snapshots and variants (gpt-5.6-mini) stay eligible.
|
||||
_ELIGIBLE_MODEL_MARKER = "gpt-5.6"
|
||||
|
||||
|
||||
@@ -77,11 +55,10 @@ def resolve_native_compaction_capabilities(
|
||||
) -> Dict[str, bool]:
|
||||
"""Resolve the native-compaction capability for a runtime destination.
|
||||
|
||||
The result is deliberately explicit: a resolved ``False`` is different
|
||||
from an unresolved capability and must survive model switches unchanged.
|
||||
A resolved ``False`` is distinct from "unresolved" and must survive model
|
||||
switches unchanged.
|
||||
"""
|
||||
normalized_provider = (provider or "").strip().lower()
|
||||
direct_default = normalized_provider == "openai" and not base_url
|
||||
direct_default = (provider or "").strip().lower() == "openai" and not base_url
|
||||
eligible = is_native_compaction_model(model) and (
|
||||
direct_default
|
||||
or is_direct_openai_route(base_url, is_codex_backend=is_codex_backend)
|
||||
@@ -110,10 +87,10 @@ def resolve_compact_threshold(
|
||||
) -> int:
|
||||
"""Resolve automatic mode or clamp an explicit native threshold.
|
||||
|
||||
An omitted or invalid setting follows the resolved local compressor trigger.
|
||||
An explicit positive integer remains absolute unless it must be clamped so
|
||||
native compaction fires first. ``local_trigger_tokens`` is
|
||||
``ContextCompressor.threshold_tokens`` when a compressor is attached.
|
||||
An omitted/invalid setting follows the local compressor trigger
|
||||
(``ContextCompressor.threshold_tokens``) minus the safety margin. An
|
||||
explicit positive integer is absolute unless it must be clamped so native
|
||||
compaction fires first. Booleans are never thresholds.
|
||||
"""
|
||||
local = None
|
||||
try:
|
||||
@@ -139,7 +116,7 @@ def resolve_compact_threshold(
|
||||
)
|
||||
except (TypeError, ValueError):
|
||||
configured = None
|
||||
if isinstance(configured_threshold, bool) or configured is None or configured <= 0:
|
||||
if configured is None or configured <= 0:
|
||||
return upper if upper is not None else DEFAULT_COMPACT_THRESHOLD
|
||||
if upper is None:
|
||||
return configured
|
||||
@@ -150,11 +127,7 @@ _checkpoint_suppression_logged = False
|
||||
|
||||
|
||||
def _warn_native_compaction_suppressed_by_checkpoint_gate() -> None:
|
||||
"""Log once per process that the checkpoint gate suppresses native compaction.
|
||||
|
||||
The suppression itself is re-evaluated per request; only the log line is
|
||||
deduplicated so a long session does not repeat it on every API call.
|
||||
"""
|
||||
"""Log once per process; the suppression itself is re-evaluated per request."""
|
||||
global _checkpoint_suppression_logged
|
||||
if _checkpoint_suppression_logged:
|
||||
return
|
||||
@@ -175,27 +148,22 @@ def native_compaction_context_management(
|
||||
) -> Optional[List[Dict[str, Any]]]:
|
||||
"""Return the ``context_management`` payload for this request, or None.
|
||||
|
||||
None means "do not send the field" — the request is byte-identical to
|
||||
pre-feature behavior. All gates are re-checked per request so a
|
||||
mid-session model switch or the in-session kill switch
|
||||
(``agent.codex_responses_native_compaction = False``, set by the
|
||||
conversation loop's rejection recovery) takes effect on the next call.
|
||||
None means "do not send the field" (request byte-identical to pre-feature).
|
||||
Every gate is re-checked per request so a mid-session model switch or the
|
||||
in-session kill switch (``agent.codex_responses_native_compaction = False``,
|
||||
set by rejection recovery) takes effect on the next call.
|
||||
"""
|
||||
capabilities = getattr(agent, "runtime_capabilities", None)
|
||||
if isinstance(capabilities, dict):
|
||||
if not bool(capabilities.get("native_compaction", False)):
|
||||
return None
|
||||
if not bool(getattr(agent, "codex_responses_native_compaction", False)):
|
||||
if isinstance(capabilities, dict) and not capabilities.get("native_compaction", False):
|
||||
return None
|
||||
# compression.enabled: false disables ALL automatic compaction, native
|
||||
# included — mirrors the codex_app_server_auto contract.
|
||||
if not bool(getattr(agent, "compression_enabled", True)):
|
||||
if not getattr(agent, "codex_responses_native_compaction", False):
|
||||
return None
|
||||
# compression.checkpoint_required: server-side compaction is a lossy
|
||||
# boundary the provider owns — no pre-compress checkpoint can run before
|
||||
# the server replaces older context. Keep the checkpoint-aware Hermes
|
||||
# compressor authoritative instead of silently letting the server
|
||||
# compact. Explicit-True check matches the compress_context() gate.
|
||||
# compression.enabled: false disables ALL automatic compaction, native included.
|
||||
if not getattr(agent, "compression_enabled", True):
|
||||
return None
|
||||
# Server-side compaction is a lossy boundary the provider owns — no
|
||||
# pre-compress checkpoint can run first — so the checkpoint-aware Hermes
|
||||
# compressor stays authoritative. Explicit-True matches compress_context().
|
||||
if getattr(agent, "compression_checkpoint_required", False) is True:
|
||||
_warn_native_compaction_suppressed_by_checkpoint_gate()
|
||||
return None
|
||||
@@ -219,13 +187,10 @@ def native_compaction_context_management(
|
||||
return [{"type": "compaction", "compact_threshold": threshold}]
|
||||
|
||||
|
||||
# Retention budget for plaintext user messages carried across a native
|
||||
# compaction boundary (mirrors Codex CLI's RETAINED_MESSAGE_TOKEN_BUDGET).
|
||||
# Live verification (Aug 2026, gpt-5.6 @ api.openai.com): the server renders
|
||||
# Retention budgets for plaintext user messages / local compression summaries
|
||||
# carried across a native compaction boundary (mirrors Codex CLI's
|
||||
# RETAINED_MESSAGE_TOKEN_BUDGET; the summary budget prevents summary inflation).
|
||||
RETAINED_USER_MESSAGE_TOKEN_BUDGET = 64_000
|
||||
|
||||
# Retention budget for local compression summary messages carried across a native
|
||||
# compaction boundary to prevent summary token inflation.
|
||||
RETAINED_SUMMARY_TOKEN_BUDGET = 32_000
|
||||
|
||||
|
||||
@@ -235,12 +200,7 @@ def _approx_tokens(text: str) -> int:
|
||||
|
||||
|
||||
def _extract_item_text(item: Any) -> Optional[str]:
|
||||
"""Extract measurable text from message content and fallback fields.
|
||||
|
||||
Returns None when the item carries no measurable text. Handles string
|
||||
content, multipart lists (input_text/text/output_text), and nested
|
||||
metadata text.
|
||||
"""
|
||||
"""Measurable text from a Responses item (string/multipart/metadata), or None."""
|
||||
if not isinstance(item, dict):
|
||||
return None
|
||||
|
||||
@@ -272,12 +232,10 @@ def _extract_item_text(item: Any) -> Optional[str]:
|
||||
|
||||
|
||||
def _has_retainable_image_content(item: Any) -> bool:
|
||||
"""Return True for a converted Responses message with a valid image part.
|
||||
"""True for a converted Responses message with a valid ``input_image`` part.
|
||||
|
||||
The pruning boundary receives normalized Responses items, so only the
|
||||
adapter-owned ``input_image`` shape is authority here. Unknown, malformed,
|
||||
or empty multipart placeholders must not become durable history merely
|
||||
because their list is non-empty.
|
||||
Only the adapter-owned ``input_image`` shape counts: unknown or empty
|
||||
multipart placeholders must not become durable history for being non-empty.
|
||||
"""
|
||||
if not isinstance(item, dict):
|
||||
return False
|
||||
@@ -295,26 +253,11 @@ def _has_retainable_image_content(item: Any) -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _is_summary_item(item: Any) -> bool:
|
||||
"""True when *item* is a canonical Hermes compression-summary message.
|
||||
|
||||
Delegates entirely to
|
||||
``agent.context_compressor.is_compaction_summary_message`` — the single
|
||||
authoritative provenance check already used by every other summary
|
||||
consumer (memory providers, frontends, the compactor itself). It prefers
|
||||
the exact, truthy ``COMPRESSED_SUMMARY_METADATA_KEY`` marker and falls
|
||||
back to the canonical prefix classifier (``SUMMARY_PREFIX`` /
|
||||
``LEGACY_SUMMARY_PREFIX`` / historical prefixes, including the
|
||||
merge-into-tail shape) for the case where the underscore-prefixed key
|
||||
was already stripped by a wire sanitizer.
|
||||
|
||||
Deliberately NOT a second heuristic: no arbitrary underscore-key scan, no
|
||||
inference from a falsy or unrelated metadata key, and no matching on
|
||||
ad-hoc content headings like ``"## Summary"`` in ordinary text — any of
|
||||
those can promote a normal user/assistant message (or adversarial
|
||||
content) to durable retained history (#90975 review).
|
||||
"""
|
||||
return is_compaction_summary_message(item)
|
||||
# Canonical provenance check (metadata marker, then canonical prefix classifier).
|
||||
# Deliberately NOT a second heuristic: no underscore-key scan, no matching on
|
||||
# ad-hoc headings — either could promote ordinary or adversarial content to
|
||||
# durable retained history.
|
||||
_is_summary_item = is_compaction_summary_message
|
||||
|
||||
|
||||
def prune_pre_checkpoint_items(
|
||||
@@ -326,48 +269,29 @@ def prune_pre_checkpoint_items(
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Restructure Responses input around the newest compaction checkpoint.
|
||||
|
||||
The server drops every input item that precedes a replayed ``compaction``
|
||||
item (live-verified Aug 2026), so sending pre-checkpoint history is dead
|
||||
weight AND silently erases the user's plaintext asks — including any
|
||||
local-compression summary the agent already produced, which previously
|
||||
vanished here because it carries ``role="assistant"``, not ``"user"``
|
||||
(#90975). When a checkpoint is present, rebuild the wire as::
|
||||
The server drops every input item preceding a replayed ``compaction`` item,
|
||||
which silently erases the user's plaintext asks and any local-compression
|
||||
summary (``role="assistant"``). With a checkpoint present, rebuild as::
|
||||
|
||||
[checkpoint run] + [retained user & summary messages (newest-first budget)] + [post]
|
||||
|
||||
- The NEWEST contiguous run of checkpoints wins.
|
||||
- Retained user messages are kept verbatim within
|
||||
``retained_user_token_budget``; the boundary message is head-truncated
|
||||
when it only partially fits (string content only) — goals are usually
|
||||
stated up front, so the head is the valuable end. A recognized
|
||||
- User messages are kept verbatim within ``retained_user_token_budget``;
|
||||
the boundary message is head-truncated when it only partially fits
|
||||
(string content only — goals are stated up front). A recognized
|
||||
image-only user message is retained whole at one-token cost.
|
||||
- Compression summary messages (``_is_summary_item``, the canonical
|
||||
``agent.context_compressor`` provenance check) are retained whole
|
||||
within ``retained_summary_token_budget``. A summary is never
|
||||
byte/character-sliced: Hermes summaries carry structural framing
|
||||
(handoff prefix, end marker, merge-into-tail delimiters) that a blind
|
||||
slice can corrupt, so one that doesn't fit whole is dropped instead.
|
||||
A summary already retained once (identical text) is never duplicated,
|
||||
so repeated checkpoints stay idempotent.
|
||||
- ``enable_summary_retention`` is a function-level override (used by
|
||||
tests and callers that need the pre-#90975 behavior back); it is not
|
||||
wired to a user-facing config surface.
|
||||
- Original relative chronological order between user messages and
|
||||
summaries is preserved.
|
||||
- ``item_sources`` (optional, parallel to ``items``) is the raw chat
|
||||
message each Responses item was converted from. By the time a summary
|
||||
reaches this function as a converted ``item`` it can already be lossy:
|
||||
a merge-into-tail tool-result carrier becomes a typed
|
||||
``function_call_output`` (no ``content``/``role`` survives the
|
||||
conversion at all), and a merge-into-tail assistant carrier can be
|
||||
shadowed by a stale exact ``codex_message_items`` replay captured
|
||||
before the merge rewrote its content. When a source is provided and is
|
||||
itself a canonical summary carrier (``is_compaction_summary_message``),
|
||||
its content is read directly from the source — never from the
|
||||
converted item — and it is retained as a synthesized
|
||||
``role="assistant"`` message regardless of what shape the original
|
||||
item took. Without ``item_sources`` (default), retention only sees
|
||||
what survived conversion, matching pre-#90976 behavior (#90976).
|
||||
- Summaries are retained whole within ``retained_summary_token_budget`` and
|
||||
never sliced (their structural framing would corrupt); one that doesn't
|
||||
fit is dropped. Identical summary text is never retained twice.
|
||||
- Relative order between user messages and summaries is preserved.
|
||||
- ``item_sources`` (parallel to ``items``) is the raw chat message each item
|
||||
was converted from. Conversion can be lossy for summaries (a
|
||||
merge-into-tail carrier becomes a typed ``function_call_output``, or an
|
||||
assistant carrier is shadowed by a stale exact replay), so when a source
|
||||
is itself a canonical summary carrier its content is read from the
|
||||
SOURCE and retained as a synthesized ``role="assistant"`` message.
|
||||
- ``enable_summary_retention`` is a function-level override for tests, not
|
||||
a config surface.
|
||||
"""
|
||||
if not isinstance(items, list) or not items:
|
||||
return items
|
||||
@@ -379,7 +303,6 @@ def prune_pre_checkpoint_items(
|
||||
if last_cp is None:
|
||||
return items
|
||||
|
||||
# Extend backwards over the contiguous run ending at last_cp.
|
||||
first_cp = last_cp
|
||||
while (
|
||||
first_cp > 0
|
||||
@@ -403,14 +326,12 @@ def prune_pre_checkpoint_items(
|
||||
seen_summary_texts: set = set()
|
||||
|
||||
def _try_retain_summary(text: Optional[str]) -> Optional[Dict[str, Any]]:
|
||||
"""Check budget/dedup/cost for a summary; return cost info or None."""
|
||||
"""Budget/dedup check for a summary; return its cost or None."""
|
||||
if not text or summary_remaining <= 0 or text in seen_summary_texts:
|
||||
return None
|
||||
cost = _approx_tokens(text)
|
||||
if cost > summary_remaining:
|
||||
# Never byte-slice a summary's structural framing — drop it
|
||||
# whole rather than corrupt the handoff prefix / end marker.
|
||||
return None
|
||||
return None # never slice a summary's structural framing
|
||||
seen_summary_texts.add(text)
|
||||
return {"cost": cost}
|
||||
|
||||
@@ -418,15 +339,10 @@ def prune_pre_checkpoint_items(
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
|
||||
# Canonical source-based summary detection: reads the ORIGINAL chat
|
||||
# message's own content, so it sees past a lossy conversion (a
|
||||
# typed `function_call_output` wrapper, or a stale exact-replay
|
||||
# message) that erased the summary from `item` itself (#90976).
|
||||
# This is never a heuristic promotion of arbitrary item content —
|
||||
# it only fires when the source message itself is a canonical,
|
||||
# provenance-tagged summary carrier.
|
||||
# Source-based detection sees past a lossy conversion; it only fires
|
||||
# when the source itself is a provenance-tagged summary carrier.
|
||||
if enable_summary_retention and isinstance(source, dict) and _is_summary_item(source):
|
||||
text = flatten_message_text(source.get("content")) if isinstance(source, dict) else ""
|
||||
text = flatten_message_text(source.get("content"))
|
||||
text = text if text.strip() else None
|
||||
result = _try_retain_summary(text)
|
||||
if result:
|
||||
@@ -438,9 +354,7 @@ def prune_pre_checkpoint_items(
|
||||
summary_remaining -= result["cost"]
|
||||
continue
|
||||
|
||||
# Skip typed non-message items (function_call_output etc. never
|
||||
# carry role=user or a summary flag, but stay defensive about
|
||||
# future shapes).
|
||||
# Typed non-message items never carry role=user or a summary flag.
|
||||
if "type" in item and item.get("type") != "message":
|
||||
continue
|
||||
|
||||
@@ -476,8 +390,7 @@ def prune_pre_checkpoint_items(
|
||||
retained_reversed.append(truncated)
|
||||
user_remaining = 0
|
||||
|
||||
retained_ordered = list(reversed(retained_reversed))
|
||||
result = checkpoint_run + retained_ordered + post
|
||||
result = checkpoint_run + list(reversed(retained_reversed)) + post
|
||||
|
||||
logger.debug(
|
||||
"Pruned pre-checkpoint items: %d input -> %d retained (user_rem=%d, summary_rem=%d)",
|
||||
@@ -490,25 +403,21 @@ def prune_pre_checkpoint_items(
|
||||
return result
|
||||
|
||||
|
||||
_REJECTION_MARKERS = (
|
||||
"unknown", "unsupported", "invalid", "unexpected", "not permitted",
|
||||
"not allowed", "unrecognized", "extra field", "no such", "bad request",
|
||||
"not supported",
|
||||
)
|
||||
|
||||
|
||||
def is_native_compaction_rejection(error: Any, status_code: Any = None) -> bool:
|
||||
"""True when a provider error is a STRUCTURED rejection of the
|
||||
context_management field.
|
||||
"""True when a provider error is a STRUCTURED rejection of ``context_management``.
|
||||
|
||||
Used by the conversation loop's one-shot recovery: strip the field,
|
||||
disable native compaction for the rest of the session, retry. Matching
|
||||
is deliberately narrow — a transient 5xx/timeout whose body merely
|
||||
ECHOES the request (and therefore contains the field name) must NOT
|
||||
permanently downgrade native compaction for the session (#82777).
|
||||
|
||||
Two conditions, both required when a status is known:
|
||||
|
||||
* ``status_code`` is 400 (or unknown/None — some transports surface
|
||||
only a message string; field-name matching alone is then the best
|
||||
available signal, preserving pre-#82777 behavior for them), and
|
||||
* the error text names ``context_management`` / ``compact_threshold``
|
||||
alongside rejection language ("unknown", "unsupported", "invalid",
|
||||
"unexpected", "not permitted"...). A bare field-name echo without
|
||||
rejection language does not match.
|
||||
Drives the loop's one-shot recovery (strip the field, disable for the
|
||||
session, retry), so matching is narrow: a transient 5xx whose body merely
|
||||
ECHOES the request must not permanently downgrade native compaction. Requires
|
||||
``status_code`` 400 (or unknown — some transports surface only a message)
|
||||
AND the field name alongside rejection language.
|
||||
"""
|
||||
text = str(error or "").lower()
|
||||
if "context_management" not in text and "compact_threshold" not in text:
|
||||
@@ -519,23 +428,15 @@ def is_native_compaction_rejection(error: Any, status_code: Any = None) -> bool:
|
||||
return False
|
||||
except (TypeError, ValueError):
|
||||
pass
|
||||
rejection_markers = (
|
||||
"unknown", "unsupported", "invalid", "unexpected", "not permitted",
|
||||
"not allowed", "unrecognized", "extra field", "no such", "bad request",
|
||||
"not supported",
|
||||
)
|
||||
return any(marker in text for marker in rejection_markers)
|
||||
return any(marker in text for marker in _REJECTION_MARKERS)
|
||||
|
||||
|
||||
def has_compaction_checkpoint(items: Any) -> bool:
|
||||
"""Does this ``codex_reasoning_items`` sidecar carry a compaction checkpoint?
|
||||
|
||||
A ``type: "compaction"`` item is the server-side stand-in for history that
|
||||
has already been pruned — cumulative context, not per-turn reasoning. It
|
||||
rides the same sidecar as ordinary reasoning items, so anything that
|
||||
rewrites or discards that sidecar (or the message carrying it) has to ask
|
||||
this question first: the checkpoint exists in exactly one place, and the
|
||||
request that loses it loses the compacted history with it.
|
||||
A ``type: "compaction"`` item is cumulative context, not per-turn
|
||||
reasoning, and exists in exactly one place: anything that rewrites or
|
||||
discards the sidecar must ask this first or lose the compacted history.
|
||||
"""
|
||||
return any(
|
||||
isinstance(item, dict) and item.get("type") == "compaction"
|
||||
@@ -547,16 +448,11 @@ def merge_interim_reasoning_items(
|
||||
prior_items: Any,
|
||||
new_items: Any,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Merge ``codex_reasoning_items`` across Codex incomplete-continuation
|
||||
dedup, preserving native compaction checkpoints.
|
||||
"""Merge ``codex_reasoning_items`` across Codex incomplete-continuation dedup.
|
||||
|
||||
The incomplete-retry path updates a visually-duplicate interim assistant
|
||||
message in place with the newer response's replay payload. A checkpoint
|
||||
captured on the EARLIER response is a cumulative context carrier the
|
||||
continuation won't re-emit (the replayed checkpoint keeps the server
|
||||
render under threshold), so a blind overwrite drops the only copy and the
|
||||
next request balloons back to full history. Rule: newer items win, but
|
||||
prior checkpoints are prepended unless the newer payload carries its own.
|
||||
A checkpoint captured on the EARLIER response is not re-emitted by the
|
||||
continuation, so a blind overwrite drops the only copy. Rule: newer items
|
||||
win, but prior checkpoints are prepended unless the newer payload has its own.
|
||||
"""
|
||||
kept_checkpoints = [
|
||||
item
|
||||
|
||||
+74
-110
@@ -1,13 +1,9 @@
|
||||
"""
|
||||
Contextual first-touch onboarding hints.
|
||||
"""Contextual first-touch onboarding hints.
|
||||
|
||||
Instead of blocking first-run questionnaires, show a one-time hint the *first*
|
||||
time a user hits a behavior fork — message-while-running, first long-running
|
||||
tool, etc. Each hint is shown once per install (tracked in ``config.yaml`` under
|
||||
``onboarding.seen.<flag>``) and then never again.
|
||||
|
||||
Keep this module tiny and dependency-free so both the CLI and gateway can import
|
||||
it without pulling in heavy modules.
|
||||
Each hint is shown once per install the *first* time a user hits a behavior
|
||||
fork (message-while-running, first long tool, ...), tracked in ``config.yaml``
|
||||
under ``onboarding.seen.<flag>``. Kept tiny and dependency-free so both the CLI
|
||||
and gateway can import it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -19,79 +15,75 @@ from typing import Any, Mapping, Optional
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
# Flag names (stable — used as config.yaml keys under onboarding.seen)
|
||||
# -------------------------------------------------------------------------
|
||||
|
||||
BUSY_INPUT_FLAG = "busy_input_prompt"
|
||||
TOOL_PROGRESS_FLAG = "tool_progress_prompt"
|
||||
OPENCLAW_RESIDUE_FLAG = "openclaw_residue_cleanup"
|
||||
PROFILE_BUILD_FLAG = "profile_build_offered"
|
||||
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
# Hint content
|
||||
# -------------------------------------------------------------------------
|
||||
# ── Hint content ──────────────────────────────────────────────────────────
|
||||
# Busy-input hints are keyed by the effective busy_input_mode that was just
|
||||
# applied so the message matches reality; "interrupt" is the default branch.
|
||||
|
||||
_BUSY_INPUT_HINTS_GATEWAY = {
|
||||
"queue": (
|
||||
"💡 First-time tip — I queued your message instead of interrupting. "
|
||||
"Send `/busy interrupt` to make new messages stop the current task "
|
||||
"immediately, or `/busy status` to check. This notice won't appear again."
|
||||
),
|
||||
"steer": (
|
||||
"💡 First-time tip — I steered your message into the current run; "
|
||||
"it will arrive after the next tool call instead of interrupting. "
|
||||
"Send `/busy interrupt` or `/busy queue` to change this, or "
|
||||
"`/busy status` to check. This notice won't appear again."
|
||||
),
|
||||
"redirect": (
|
||||
"💡 First-time tip — I redirected the current run using your message. "
|
||||
"Completed work stays in context, and `/stop` still cancels the task. "
|
||||
"Send `/busy queue` to wait for a separate turn, or `/busy status` "
|
||||
"to check. This notice won't appear again."
|
||||
),
|
||||
}
|
||||
_BUSY_INPUT_HINT_GATEWAY_DEFAULT = (
|
||||
"💡 First-time tip — I just interrupted my current task to answer you. "
|
||||
"Send `/busy queue` to queue follow-ups for after the current task instead, "
|
||||
"`/busy steer` to inject them mid-run without interrupting, or "
|
||||
"`/busy status` to check. This notice won't appear again."
|
||||
)
|
||||
|
||||
_BUSY_INPUT_HINTS_CLI = {
|
||||
"queue": (
|
||||
"(tip) Your message was queued for the next turn. "
|
||||
"Use /busy interrupt to make Enter stop the current run instead, "
|
||||
"or /busy steer to inject mid-run. This tip only shows once."
|
||||
),
|
||||
"steer": (
|
||||
"(tip) Your message was steered into the current run; it arrives "
|
||||
"after the next tool call. Use /busy interrupt or /busy queue to "
|
||||
"change this. This tip only shows once."
|
||||
),
|
||||
"redirect": (
|
||||
"(tip) Your correction redirected the current run without discarding "
|
||||
"completed work. Use /stop to cancel or /busy queue to wait for a "
|
||||
"separate turn. This tip only shows once."
|
||||
),
|
||||
}
|
||||
_BUSY_INPUT_HINT_CLI_DEFAULT = (
|
||||
"(tip) Your message interrupted the current run. "
|
||||
"Use /busy queue to queue messages for the next turn instead, "
|
||||
"or /busy steer to inject mid-run. This tip only shows once."
|
||||
)
|
||||
|
||||
|
||||
def busy_input_hint_gateway(mode: str) -> str:
|
||||
"""Hint shown the first time a user messages while the agent is busy.
|
||||
|
||||
``mode`` is the effective busy_input_mode that was just applied, so the
|
||||
message matches reality ("I just interrupted…" vs "I just queued…").
|
||||
"""
|
||||
if mode == "queue":
|
||||
return (
|
||||
"💡 First-time tip — I queued your message instead of interrupting. "
|
||||
"Send `/busy interrupt` to make new messages stop the current task "
|
||||
"immediately, or `/busy status` to check. This notice won't appear again."
|
||||
)
|
||||
if mode == "steer":
|
||||
return (
|
||||
"💡 First-time tip — I steered your message into the current run; "
|
||||
"it will arrive after the next tool call instead of interrupting. "
|
||||
"Send `/busy interrupt` or `/busy queue` to change this, or "
|
||||
"`/busy status` to check. This notice won't appear again."
|
||||
)
|
||||
if mode == "redirect":
|
||||
return (
|
||||
"💡 First-time tip — I redirected the current run using your message. "
|
||||
"Completed work stays in context, and `/stop` still cancels the task. "
|
||||
"Send `/busy queue` to wait for a separate turn, or `/busy status` "
|
||||
"to check. This notice won't appear again."
|
||||
)
|
||||
return (
|
||||
"💡 First-time tip — I just interrupted my current task to answer you. "
|
||||
"Send `/busy queue` to queue follow-ups for after the current task instead, "
|
||||
"`/busy steer` to inject them mid-run without interrupting, or "
|
||||
"`/busy status` to check. This notice won't appear again."
|
||||
)
|
||||
"""Hint shown the first time a user messages while the agent is busy (markdown)."""
|
||||
return _BUSY_INPUT_HINTS_GATEWAY.get(mode, _BUSY_INPUT_HINT_GATEWAY_DEFAULT)
|
||||
|
||||
|
||||
def busy_input_hint_cli(mode: str) -> str:
|
||||
"""CLI version of the busy-input hint (plain text, no markdown)."""
|
||||
if mode == "queue":
|
||||
return (
|
||||
"(tip) Your message was queued for the next turn. "
|
||||
"Use /busy interrupt to make Enter stop the current run instead, "
|
||||
"or /busy steer to inject mid-run. This tip only shows once."
|
||||
)
|
||||
if mode == "steer":
|
||||
return (
|
||||
"(tip) Your message was steered into the current run; it arrives "
|
||||
"after the next tool call. Use /busy interrupt or /busy queue to "
|
||||
"change this. This tip only shows once."
|
||||
)
|
||||
if mode == "redirect":
|
||||
return (
|
||||
"(tip) Your correction redirected the current run without discarding "
|
||||
"completed work. Use /stop to cancel or /busy queue to wait for a "
|
||||
"separate turn. This tip only shows once."
|
||||
)
|
||||
return (
|
||||
"(tip) Your message interrupted the current run. "
|
||||
"Use /busy queue to queue messages for the next turn instead, "
|
||||
"or /busy steer to inject mid-run. This tip only shows once."
|
||||
)
|
||||
return _BUSY_INPUT_HINTS_CLI.get(mode, _BUSY_INPUT_HINT_CLI_DEFAULT)
|
||||
|
||||
|
||||
def tool_progress_hint_gateway() -> str:
|
||||
@@ -110,13 +102,7 @@ def tool_progress_hint_cli() -> str:
|
||||
|
||||
|
||||
def openclaw_residue_hint_cli() -> str:
|
||||
"""Banner shown the first time Hermes starts and finds ``~/.openclaw/``.
|
||||
|
||||
Points users at ``hermes claw migrate`` (non-destructive port of config,
|
||||
memory, and skills) first. ``hermes claw cleanup`` is mentioned as the
|
||||
follow-up step for users who have already migrated and want to archive
|
||||
the old directory — with a warning that archiving breaks OpenClaw.
|
||||
"""
|
||||
"""Banner shown the first time Hermes finds ``~/.openclaw/``: migrate first, cleanup (which breaks OpenClaw) after."""
|
||||
return (
|
||||
"A legacy OpenClaw directory was detected at ~/.openclaw/.\n"
|
||||
"To port your config, memory, and skills over to Hermes, run "
|
||||
@@ -129,10 +115,7 @@ def openclaw_residue_hint_cli() -> str:
|
||||
|
||||
|
||||
def detect_openclaw_residue(home: Optional[Path] = None) -> bool:
|
||||
"""Return True if an OpenClaw workspace directory is present in ``$HOME``.
|
||||
|
||||
Pure filesystem check — no side effects. ``home`` override exists for tests.
|
||||
"""
|
||||
"""True if ``$HOME/.openclaw`` is a directory (pure check; ``home`` override for tests)."""
|
||||
base = home or Path.home()
|
||||
try:
|
||||
return (base / ".openclaw").is_dir()
|
||||
@@ -140,25 +123,15 @@ def detect_openclaw_residue(home: Optional[Path] = None) -> bool:
|
||||
return False
|
||||
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
# Onboarding profile-build path (opt-in, consent-gated)
|
||||
# -------------------------------------------------------------------------
|
||||
# ── Onboarding profile-build path (opt-in, consent-gated) ─────────────────
|
||||
|
||||
def profile_build_mode(config: Mapping[str, Any]) -> str:
|
||||
"""Resolve the onboarding profile-build mode from config.
|
||||
"""``config.onboarding.profile_build``: ``"off"`` never offers; anything else -> ``"ask"`` (offer on first contact).
|
||||
|
||||
Returns one of:
|
||||
``"ask"`` — on first contact, OFFER to build a profile (default).
|
||||
``"off"`` — never offer; the first-message note stays a plain intro.
|
||||
|
||||
Read from ``config.onboarding.profile_build``. Unknown / missing values
|
||||
fall back to ``"ask"`` so the default experience offers the flow. Any
|
||||
network/account lookups inside the flow are separately consented to in
|
||||
conversation — this setting only governs whether the offer is made.
|
||||
This only governs whether the offer is made; lookups inside the flow are
|
||||
consented to separately in conversation.
|
||||
"""
|
||||
if not isinstance(config, Mapping):
|
||||
return "ask"
|
||||
onboarding = config.get("onboarding")
|
||||
onboarding = config.get("onboarding") if isinstance(config, Mapping) else None
|
||||
if not isinstance(onboarding, Mapping):
|
||||
return "ask"
|
||||
mode = onboarding.get("profile_build")
|
||||
@@ -170,11 +143,9 @@ def profile_build_mode(config: Mapping[str, Any]) -> str:
|
||||
def profile_build_directive() -> str:
|
||||
"""System-note directive appended to the very first message ever.
|
||||
|
||||
Instructs the agent to run a short, opt-in, consent-gated profile-build
|
||||
flow and persist confirmed facts to the user-profile memory store
|
||||
(``memory`` tool, ``target="user"``). Phrased so the agent ASKS before any
|
||||
lookup and never silently reads connected accounts — directly addressing
|
||||
the privacy concern that reading email/accounts unprompted feels invasive.
|
||||
Runs a short opt-in profile-build flow persisting to the user-profile memory
|
||||
store; phrased so the agent ASKS before any lookup and never silently reads
|
||||
connected accounts.
|
||||
"""
|
||||
return (
|
||||
"\n\n[System note: This is the user's very first message ever. "
|
||||
@@ -196,9 +167,7 @@ def profile_build_directive() -> str:
|
||||
)
|
||||
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
# State read / write
|
||||
# -------------------------------------------------------------------------
|
||||
# ── State read / write ────────────────────────────────────────────────────
|
||||
|
||||
def _get_seen_dict(config: Mapping[str, Any]) -> Mapping[str, Any]:
|
||||
onboarding = config.get("onboarding") if isinstance(config, Mapping) else None
|
||||
@@ -214,12 +183,7 @@ def is_seen(config: Mapping[str, Any], flag: str) -> bool:
|
||||
|
||||
|
||||
def mark_seen(config_path: Path, flag: str) -> bool:
|
||||
"""Persist ``onboarding.seen.<flag> = True`` to ``config_path``.
|
||||
|
||||
Uses the atomic YAML writer so a concurrent process can't observe a
|
||||
partially-written file. Returns True on success, False on any error
|
||||
(including the config file being absent — onboarding is best-effort).
|
||||
"""
|
||||
"""Persist ``onboarding.seen.<flag> = True`` atomically; False on any error (best-effort)."""
|
||||
try:
|
||||
import yaml
|
||||
from hermes_cli.config import atomic_config_write
|
||||
@@ -239,7 +203,7 @@ def mark_seen(config_path: Path, flag: str) -> bool:
|
||||
seen = {}
|
||||
cfg["onboarding"]["seen"] = seen
|
||||
if seen.get(flag) is True:
|
||||
return True # already marked — nothing to do
|
||||
return True
|
||||
seen[flag] = True
|
||||
atomic_config_write(config_path, cfg)
|
||||
return True
|
||||
|
||||
+8
-29
@@ -1,31 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""``/plan`` — build the plan-mode prompt that turns the user's request into a
|
||||
saved markdown implementation plan, with no execution.
|
||||
"""``/plan`` — build the plan-mode prompt: a saved markdown implementation plan, no execution.
|
||||
|
||||
``/plan`` used to be a bundled skill (``skills/software-development/plan``)
|
||||
whose auto-generated slash command fell off the capped Telegram/Discord command
|
||||
menus for most installs (skills are the only tier trimmed at the platform
|
||||
caps, alphabetically — ``plan`` sat past the cutoff). It is now a first-class
|
||||
built-in: this module builds ONE prompt that instructs the live agent to
|
||||
|
||||
1. Stay in planning mode for the turn — read-only inspection is allowed,
|
||||
but no implementation, no mutating commands, no side effects.
|
||||
2. Write a concrete, bite-sized, TDD-shaped markdown plan under
|
||||
``.hermes/plans/`` in the active workspace via ``write_file``.
|
||||
|
||||
There is no engine and no model-tool footprint: the agent does the work with
|
||||
its existing toolset, so this works identically on local, Docker, and remote
|
||||
terminal backends. Every surface (CLI ``/plan``, gateway ``/plan``, TUI
|
||||
``/plan``) calls :func:`build_plan_prompt` and feeds the result to the agent
|
||||
as a normal turn — same pattern as ``/learn`` and ``/init``, preserving
|
||||
prompt-cache invariants (no system-prompt or history mutation).
|
||||
A first-class built-in (the former bundled skill fell off capped Telegram/Discord
|
||||
command menus). No engine, no model-tool footprint: every surface feeds
|
||||
:func:`build_plan_prompt` to the agent as a normal turn, like ``/learn`` and
|
||||
``/init``, so system prompt and history stay untouched (prompt-cache safe).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
# The plan-mode ground rules + authoring craft, distilled from the retired
|
||||
# bundled skill (v2.0.0, writing-craft adapted from obra/superpowers).
|
||||
# Embedded in the prompt so the agent plans the way a maintainer would.
|
||||
# Plan-mode ground rules + authoring craft, distilled from the retired bundled
|
||||
# skill (writing-craft adapted from obra/superpowers).
|
||||
_PLAN_MODE_RULES = """\
|
||||
For this turn, you are in PLAN MODE — planning only.
|
||||
|
||||
@@ -76,13 +61,7 @@ Interaction style:
|
||||
|
||||
|
||||
def build_plan_prompt(task: str = "") -> str:
|
||||
"""Build the plan-mode prompt for the live agent.
|
||||
|
||||
Args:
|
||||
task: What to plan. Empty → infer the task from the current
|
||||
conversation context (mirrors the retired skill's behavior and
|
||||
issue #36821's "plan from context" expectation).
|
||||
"""
|
||||
"""Build the plan-mode prompt; empty *task* asks the agent to infer it from conversation context."""
|
||||
task = (task or "").strip()
|
||||
if task:
|
||||
task_block = f"Task to plan:\n{task}\n"
|
||||
|
||||
@@ -1,52 +1,30 @@
|
||||
"""Builder-declared stable prefixes for Anthropic prompt caching (#81867).
|
||||
"""Builder-declared stable prefixes for Anthropic prompt caching.
|
||||
|
||||
Skill, webhook, and cron builders concatenate a large static scaffold
|
||||
(activation note + expanded skill body) with a small volatile invocation
|
||||
tail (ticket payload, timestamps, run context) into one user-message
|
||||
string. Only the builder knows the exact byte where the volatile tail
|
||||
begins, so it registers the stable prefix here at construction time; the
|
||||
cache planner consults the registry to place a cache breakpoint at that
|
||||
boundary instead of caching the whole message as one atomic block.
|
||||
Skill/webhook/cron builders concatenate a large static scaffold with a small
|
||||
volatile invocation tail into one user-message string. Only the builder knows
|
||||
where the tail begins, so it registers the stable prefix here and the cache
|
||||
planner places a breakpoint at that boundary instead of caching the whole
|
||||
message. Re-parsing marker strings out of the message at request time is
|
||||
deliberately avoided: markers can legitimately appear inside skill bodies or
|
||||
event payloads, and any delimiter heuristic then shrinks the cached prefix or
|
||||
silently absorbs volatile bytes into it.
|
||||
|
||||
This deliberately avoids re-parsing scaffold marker strings out of the
|
||||
message at request time: markers can legitimately appear inside skill
|
||||
bodies or inside event payloads (e.g. a helpdesk ticket quoting an agent
|
||||
transcript), and any delimiter-search heuristic then either shrinks the
|
||||
cached prefix or — worse — silently absorbs volatile bytes into it,
|
||||
reintroducing the per-invocation cache miss this exists to fix.
|
||||
|
||||
The registry is process-local by design. A freshly fired webhook/cron
|
||||
invocation is always built and sent by the same process, which is the
|
||||
only window where the split pays off. Any miss (restart, eviction,
|
||||
historic message) falls back to the pre-existing whole-message policy.
|
||||
|
||||
Split-shape lifetime: the split is applied only while the skill message is
|
||||
one of the plan's marked endpoints (the last few cacheable messages). Once
|
||||
later turns rotate it out of that window it ships as a single string block
|
||||
again, which changes the block boundary once and re-ingests the prefix from
|
||||
that message onward exactly one time in a long-lived session. Webhook/cron
|
||||
invocations — the workload this exists for — send the skill turn as the
|
||||
newest message every time, so they always hit the split shape; the one-time
|
||||
re-ingest only affects long interactive sessions and nets out far below the
|
||||
per-invocation full rewrite this removes.
|
||||
Process-local by design: a webhook/cron fire is built and sent by the same
|
||||
process, and any miss (restart, eviction, historic message) falls back to the
|
||||
whole-message policy. The split only applies while the message is one of the
|
||||
plan's marked endpoints; once it rotates out it ships as one block again
|
||||
(one-time re-ingest in long interactive sessions, never for webhook/cron).
|
||||
"""
|
||||
|
||||
import threading
|
||||
from collections import OrderedDict
|
||||
from typing import Optional
|
||||
|
||||
# A couple dozen distinct active scaffolds (webhook routes x skills x cron
|
||||
# jobs) is generous for one gateway process; beyond that, oldest entries
|
||||
# fall back to whole-message caching rather than growing unboundedly.
|
||||
# A couple dozen active scaffolds is generous for one gateway process.
|
||||
_MAX_ENTRIES = 32
|
||||
|
||||
# Entries hold whole expanded skill bodies, so an entry count alone does not
|
||||
# bound memory — a handful of large skills can retain tens of MB in a
|
||||
# long-lived gateway process. Evict by total retained characters too (a
|
||||
# conservative proxy for bytes: actual memory is 1–4x depending on the
|
||||
# string's widest code point), always keeping the newest entry so a single
|
||||
# oversized scaffold still gets a boundary instead of silently disabling
|
||||
# the split.
|
||||
# Entries hold whole expanded skill bodies, so also bound total retained chars
|
||||
# (1-4x bytes). The newest entry is always kept so one oversized scaffold still
|
||||
# gets a boundary instead of silently disabling the split.
|
||||
_MAX_CHARS = 4 * 1024 * 1024
|
||||
|
||||
_lock = threading.Lock()
|
||||
@@ -67,29 +45,19 @@ def register_stable_prefix(prefix: str) -> None:
|
||||
|
||||
|
||||
def find_stable_prefix(content: str) -> Optional[str]:
|
||||
"""Longest registered prefix that is a *proper* prefix of ``content`` with non-whitespace tail.
|
||||
"""Longest registered *proper* prefix of ``content`` with a non-whitespace tail.
|
||||
|
||||
Proper with non-whitespace tail (``bool(content[len(prefix):].strip())``) so the
|
||||
split never produces an empty or whitespace-only volatile text block, which
|
||||
Anthropic rejects on the wire (HTTP 400).
|
||||
|
||||
A hit refreshes the entry's LRU position: a scaffold fired every minute
|
||||
by cron must not be evicted by a burst of one-off skill invocations,
|
||||
which would silently drop it back to whole-message caching.
|
||||
The tail must be non-whitespace so the split never yields an empty text
|
||||
block (Anthropic rejects it with HTTP 400). A hit refreshes the entry's LRU
|
||||
position so a scaffold fired every minute by cron is not evicted by a
|
||||
burst of one-off skill invocations.
|
||||
"""
|
||||
with _lock:
|
||||
best: Optional[str] = None
|
||||
for prefix in _prefixes:
|
||||
if content.startswith(prefix) and bool(content[len(prefix):].strip()):
|
||||
if content.startswith(prefix) and content[len(prefix):].strip():
|
||||
if best is None or len(prefix) > len(best):
|
||||
best = prefix
|
||||
if best is not None:
|
||||
# After the scan so the OrderedDict is never mutated mid-iteration.
|
||||
_prefixes.move_to_end(best)
|
||||
_prefixes.move_to_end(best) # after the scan: never mutate mid-iteration
|
||||
return best
|
||||
|
||||
|
||||
def clear_stable_prefixes() -> None:
|
||||
"""Test isolation helper."""
|
||||
with _lock:
|
||||
_prefixes.clear()
|
||||
|
||||
+68
-174
@@ -1,68 +1,30 @@
|
||||
"""Rotation-stable logical cache scope for prompt_cache_key derivation.
|
||||
|
||||
Context-compression rotation (legacy ``compression.in_place: false`` mode)
|
||||
mints a new physical ``session_id`` mid-conversation to segment the
|
||||
transcript. The prompt-cache scope introduced by #79161 was derived from that
|
||||
physical id, so every rotation moved the conversation into a fresh cache
|
||||
bucket even though it is logically the same conversation continuing
|
||||
(issue #79017).
|
||||
Legacy compression rotation (``compression.in_place: false``) mints a new
|
||||
physical ``session_id`` mid-conversation, which moved the conversation into a
|
||||
fresh cache bucket each time. ``resolve_prompt_cache_scope()`` instead maps
|
||||
the physical id to the ROOT of its compression lineage via
|
||||
``SessionDB.get_compression_lineage()`` — NOT ``get_conversation_root`` /
|
||||
``_conversation_root_id`` (the Portal-attribution walk), which follows
|
||||
``parent_session_id`` blindly and would collapse /branch children and delegate
|
||||
trees into one id. The two resolvers are intentionally different.
|
||||
|
||||
``resolve_prompt_cache_scope()`` maps the physical session id to the ROOT of
|
||||
its *compression lineage* — the pre-rotation session id — using
|
||||
``SessionDB.get_compression_lineage()``, whose fork-aware semantics
|
||||
(hardened in #79193) give exactly the scope boundaries the cache key needs.
|
||||
NOT ``SessionDB.get_conversation_root`` / ``run_agent._conversation_root_id``
|
||||
(the Portal-attribution walk): that one follows ``parent_session_id`` blindly,
|
||||
collapsing /branch children and whole delegate trees into one id, which would
|
||||
violate the #79161 isolation this scope must preserve. The two resolvers are
|
||||
intentionally different — do not "deduplicate" them.
|
||||
Scope boundaries: rotation children walk back to the original segment; ``/new``
|
||||
starts a fresh scope; ``/branch`` children, delegate subagents, and tool-tagged
|
||||
children are explicit fork children with their own isolated scope; cron fires
|
||||
keep their physical id (the per-fire timestamp is stripped later).
|
||||
|
||||
- compression-rotation children walk back to the original segment
|
||||
(rotation-stable scope — the fix);
|
||||
- ``/new`` starts a lineage-less session (fresh scope);
|
||||
- ``/branch`` children (``_branched_from``), delegate subagents
|
||||
(``_delegate_from``), and tool-tagged children (``source="tool"``) are
|
||||
explicit fork children and keep their own isolated scope, preserving the
|
||||
sibling/subagent isolation #79161 established;
|
||||
- cron fires keep their physical ``cron_<job>_<ts>`` id here — the per-fire
|
||||
timestamp is stripped later by ``_cache_scope_from_session_id`` exactly as
|
||||
before.
|
||||
Hosts that mint one physical id per RESPONSE (Studio group chat, ``/v1/responses``
|
||||
with client-managed history) carry no lineage, so the walk returns the physical
|
||||
id and the scope moves every reply. Hermes must not infer the conversation from
|
||||
id SYNTAX (that collides client-supplied ids); the host declares it via
|
||||
``gateway_session_key`` (``X-Hermes-Session-Key`` / ``build_session_key``),
|
||||
consumed by ``declared_conversation_scope()``, which wins over the lineage walk.
|
||||
The declared key is hashed to ``gwk_<sha256[:24]>`` because it embeds
|
||||
platform/chat/user identifiers and leaves the process as a provider routing key.
|
||||
|
||||
A host that mints one physical ``session_id`` per RESPONSE (Hermes Studio's
|
||||
group chat, and ``POST /v1/responses`` with client-managed history, which
|
||||
mints ``str(uuid4())`` per request) re-keys every conversation-affinity hint
|
||||
Hermes sends — ``prompt_cache_key`` on both OpenAI-wire transports, plus the
|
||||
OpenRouter/Nous sticky ``session_id`` and xAI's ``x-grok-conv-id`` through
|
||||
``portal_tags`` (issue #96811). Those rows carry no lineage, so the walk
|
||||
above correctly returns the physical id and the scope moves every reply.
|
||||
|
||||
Hermes must not infer the logical conversation from the id's SYNTAX (that
|
||||
rule collides independent client-supplied ids and merges Studio members
|
||||
truncated past its 96-character boundary — the #79017 failure class). The
|
||||
host has to declare it, and one carrier already means exactly that:
|
||||
``gateway_session_key`` — the "stable per-chat key" (``agent:main:telegram:
|
||||
dm:123``) built by ``gateway.session.build_session_key`` from the
|
||||
``X-Hermes-Session-Key`` header, which branching deliberately does NOT key
|
||||
off. ``declared_conversation_scope()`` consumes it, and it wins over the
|
||||
lineage walk because it is stable across rotation AND across per-response
|
||||
ids. Two boundaries it must not cross:
|
||||
|
||||
- explicit fork children (``/branch``, delegate subagents, tool children)
|
||||
share their parent's chat key but are separate conversations — the row's
|
||||
fork markers keep them on their own scope (#79161);
|
||||
- background-review forks run on a clone of the live runtime, so they are
|
||||
excluded by ``_persist_disabled`` for the same reason.
|
||||
|
||||
The declared key is hashed into ``gwk_<sha256[:24]>`` before it becomes a
|
||||
scope: unlike a session id it embeds platform/chat/user identifiers, and
|
||||
this value leaves the process verbatim as OpenRouter's sticky ``session_id``
|
||||
and xAI's ``x-grok-conv-id``.
|
||||
|
||||
The resolution is memoized per (agent, session_id): the lineage walk runs
|
||||
once per transcript segment — NOT per API call — and re-runs only when
|
||||
rotation actually changes ``agent.session_id`` (per the no-DB-on-the-hot-path
|
||||
constraint recorded on #79017). Default installs compact in place and never
|
||||
rotate, so they hit the memo forever and behave byte-identically to before.
|
||||
Resolution is memoized per (agent, session_id, db-present): the lineage walk
|
||||
runs once per transcript segment, never per API call.
|
||||
"""
|
||||
|
||||
import hashlib
|
||||
@@ -72,16 +34,13 @@ from typing import Any, Optional
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
_MEMO_ATTR = "_prompt_cache_scope_memo"
|
||||
# Namespace for a scope resolved from a host-declared conversation key.
|
||||
_DECLARED_SCOPE_PREFIX = "gwk_"
|
||||
|
||||
|
||||
def _lineage_root(session_id: str, session_db: Any) -> Optional[str]:
|
||||
"""Return the compression-lineage root of *session_id*, or None.
|
||||
"""Compression-lineage root of *session_id*, or None.
|
||||
|
||||
Defensive about the DB handle: test doubles and partially constructed
|
||||
agents can hand back non-list results — anything that is not a non-empty
|
||||
list/tuple whose first element is a non-empty string is ignored.
|
||||
Tolerates non-list results from test doubles / partially built agents.
|
||||
"""
|
||||
if session_db is None:
|
||||
return None
|
||||
@@ -102,24 +61,13 @@ def _agent_source(
|
||||
) -> str:
|
||||
"""The ``sessions.source`` this agent's conversation is recorded under.
|
||||
|
||||
Read from the agent's own row when it exists, because that is the value
|
||||
the peer queries below match on.
|
||||
|
||||
``row_source`` is that value when the caller already has it — the single
|
||||
identity read in :func:`declared_conversation_scope` — where ``""`` means
|
||||
"the row was read and carries no source". ``None`` means "not read yet"
|
||||
and keeps the original lookup, which is the path a ``SessionDB`` without
|
||||
:meth:`~hermes_state.SessionDB.declared_scope_identity` still takes.
|
||||
|
||||
Before the row lands — this module resolves the first scope ahead of
|
||||
``_ensure_db_session`` — it uses the SAME resolver persistence will use,
|
||||
``run_agent._session_source_for_agent``, not ``agent.platform``. The two
|
||||
diverge whenever ``HERMES_SESSION_SOURCE`` overrides the platform, and the
|
||||
divergence is not a cosmetic one: the declared scope is non-``None``
|
||||
immediately, so ``resolve_prompt_cache_scope`` memoizes it for this session
|
||||
id and never re-resolves once the authoritative row appears. Both sides of
|
||||
a ``/new`` would then read the platform domain, miss the boundary recorded
|
||||
under the override, and hash the same scope.
|
||||
``row_source`` is the row's value when the caller already read it (``""``
|
||||
= read, no source; ``None`` = not read yet, do the lookup). Before the row
|
||||
lands, use the SAME resolver persistence uses
|
||||
(``run_agent._session_source_for_agent``), not ``agent.platform``: they
|
||||
diverge under ``HERMES_SESSION_SOURCE``, and the declared scope is memoized
|
||||
immediately, so both sides of a ``/new`` would otherwise miss the boundary
|
||||
recorded under the override and hash the same scope.
|
||||
"""
|
||||
if row_source is None and session_id and session_db is not None:
|
||||
try:
|
||||
@@ -132,8 +80,7 @@ def _agent_source(
|
||||
return row_source
|
||||
platform = getattr(agent, "platform", None)
|
||||
try:
|
||||
# Imported lazily: run_agent imports this module, and this is the
|
||||
# single owner of the source a session row is created with.
|
||||
# Lazy: run_agent imports this module.
|
||||
from run_agent import _session_source_for_agent
|
||||
|
||||
source = str(_session_source_for_agent(platform) or "").strip()
|
||||
@@ -145,23 +92,14 @@ def _agent_source(
|
||||
|
||||
|
||||
def _conversation_generation(session_key: str, source: str, session_db: Any) -> str:
|
||||
"""Return the durable generation for *session_key*'s current conversation.
|
||||
"""Durable generation for *session_key*'s current conversation (``""`` if none).
|
||||
|
||||
The declared key names a chat and deliberately survives `/new` and policy
|
||||
resets. Hashing it alone would therefore reuse one affinity scope across
|
||||
distinct conversations, violating the #79017/#86733 contract: warm across
|
||||
compression, cold across a conversation boundary.
|
||||
|
||||
``SessionDB.latest_conversation_boundary`` reads the monotonic
|
||||
``conversation_generations`` counter for ``(source, session_key)``. The
|
||||
counter advances in the same transaction that records an
|
||||
``_RESET_END_REASONS`` boundary. It is independent of prunable session rows
|
||||
and wall-clock time, so deletion, bulk pruning, and clock rollback cannot
|
||||
reissue an old generation. Compression continues the current conversation
|
||||
and does not advance it.
|
||||
|
||||
This lookup runs on the memoized resolution path, not once per API call.
|
||||
Return ``""`` when the key has never reset or the DB exposes no generation.
|
||||
The declared key names a chat and survives ``/new`` and policy resets, so
|
||||
hashing it alone would reuse one scope across distinct conversations. The
|
||||
``conversation_generations`` counter advances in the same transaction that
|
||||
records a reset boundary and is independent of prunable rows and
|
||||
wall-clock, so pruning or clock rollback cannot reissue a generation.
|
||||
Compression does not advance it.
|
||||
"""
|
||||
reader = getattr(session_db, "latest_conversation_boundary", None)
|
||||
if not callable(reader):
|
||||
@@ -173,30 +111,17 @@ def _conversation_generation(session_key: str, source: str, session_db: Any) ->
|
||||
|
||||
|
||||
def declared_conversation_scope(agent: Any) -> Optional[str]:
|
||||
"""Return the host-declared logical conversation scope, or None.
|
||||
"""Host-declared logical conversation scope (``gwk_<sha256[:24]>``), or None.
|
||||
|
||||
Resolved from ``agent._gateway_session_key`` (the ``X-Hermes-Session-Key``
|
||||
/``build_session_key`` per-chat key) qualified by the conversation
|
||||
generation currently live on it (:func:`_conversation_generation`), hashed
|
||||
together into ``gwk_<sha256[:24]>`` so no platform/chat/user identifier
|
||||
reaches a provider on the wire and the value stays inside every caller's
|
||||
length/charset budget.
|
||||
|
||||
The key alone would outlive the conversation — it survives ``/new`` and the
|
||||
idle/daily policy resets by design — so the generation is what makes this
|
||||
carrier legal: stable across a host's per-response physical ids, and cold
|
||||
on every conversation replacement.
|
||||
|
||||
None — meaning "fall back to the physical-id scope" — when no key was
|
||||
declared, when this agent is a background-review fork (``_persist_disabled``:
|
||||
it clones the live runtime, including the key), when the session row is an
|
||||
explicit fork child (``/branch``, delegate, tool), and on any DB error
|
||||
during either lookup.
|
||||
Hashes ``(source, gateway_session_key, generation)`` so no platform/chat/
|
||||
user identifier reaches a provider. None — fall back to the physical-id
|
||||
scope — when no key is declared, when the agent is a background-review
|
||||
fork (``_persist_disabled`` clones the live runtime incl. the key), when
|
||||
the row is an explicit fork child, and on any DB error (fail closed rather
|
||||
than merge a fork onto its parent's key).
|
||||
"""
|
||||
key = str(getattr(agent, "_gateway_session_key", "") or "").strip()
|
||||
if not key:
|
||||
return None
|
||||
if getattr(agent, "_persist_disabled", False):
|
||||
if not key or getattr(agent, "_persist_disabled", False):
|
||||
return None
|
||||
sid = str(getattr(agent, "session_id", None) or "")
|
||||
db = getattr(agent, "_session_db", None)
|
||||
@@ -204,13 +129,9 @@ def declared_conversation_scope(agent: Any) -> Optional[str]:
|
||||
row_source: Optional[str] = None
|
||||
if sid and db is not None:
|
||||
try:
|
||||
# One read for both halves of the row's identity: the fork verdict
|
||||
# and the source the peer queries match on live on the same
|
||||
# ``sessions`` row, and asking for them separately read it twice
|
||||
# per resolution (@teknium1 on #98811). A SessionDB without the
|
||||
# combined view keeps the original call, so nothing that predates
|
||||
# it — including the doubles that certify the fail-closed contract
|
||||
# below — changes behaviour.
|
||||
# One read for both halves of the row identity (fork verdict +
|
||||
# source). A SessionDB without the combined view keeps the
|
||||
# original call.
|
||||
identity = getattr(db, "declared_scope_identity", None)
|
||||
if callable(identity):
|
||||
is_fork, row_source = identity(sid)
|
||||
@@ -219,11 +140,6 @@ def declared_conversation_scope(agent: Any) -> Optional[str]:
|
||||
if is_fork:
|
||||
return None
|
||||
except Exception:
|
||||
# Degrade to the physical-id scope rather than risk merging a
|
||||
# fork onto its parent's key on a transient DB failure. The
|
||||
# source read is inside this same guard for the same reason: it
|
||||
# was always the second half of a read that had already failed
|
||||
# closed here.
|
||||
logger.debug("declared-scope fork check failed", exc_info=True)
|
||||
return None
|
||||
source = _agent_source(agent, sid, db, row_source)
|
||||
@@ -231,65 +147,46 @@ def declared_conversation_scope(agent: Any) -> Optional[str]:
|
||||
try:
|
||||
generation = _conversation_generation(key, source, db)
|
||||
except Exception:
|
||||
# Same fail-closed rule as the fork check: an unqualified key
|
||||
# spans /new, so degrade to the physical-id scope instead.
|
||||
logger.debug("declared-scope generation read failed", exc_info=True)
|
||||
return None
|
||||
# The carrier is the SAME identity tuple the peer queries use: two hosts
|
||||
# may legally declare the same key string under different sources, and the
|
||||
# scope leaves this process as a routing key, so it must not collapse them.
|
||||
# Same identity tuple the peer queries use: two hosts may declare the
|
||||
# same key under different sources and must not collapse.
|
||||
carrier = f"{source}|{key}|{generation}"
|
||||
digest = hashlib.sha256(carrier.encode("utf-8", errors="replace")).hexdigest()[:24]
|
||||
return f"{_DECLARED_SCOPE_PREFIX}{digest}"
|
||||
|
||||
|
||||
def resolve_prompt_cache_scope(agent: Any) -> str:
|
||||
"""Resolve the rotation-stable cache-scope id for *agent*'s conversation.
|
||||
"""Rotation-stable cache-scope id for *agent*'s conversation.
|
||||
|
||||
Returns the host-declared conversation scope when one applies
|
||||
(``declared_conversation_scope``), else the compression-lineage ROOT of
|
||||
``agent.session_id`` (the physical id itself when the session has no
|
||||
compression ancestry, no DB is attached, or the walk fails). The result is memoized on the agent
|
||||
keyed by the current session id, so the DB walk happens once per
|
||||
transcript segment rather than once per API call.
|
||||
Declared scope when one applies, else the compression-lineage root of
|
||||
``agent.session_id`` (the physical id when there is no ancestry, no DB, or
|
||||
the walk fails). Memoized on the agent keyed by session id.
|
||||
"""
|
||||
sid = str(getattr(agent, "session_id", None) or "")
|
||||
if not sid:
|
||||
return ""
|
||||
db = getattr(agent, "_session_db", None)
|
||||
# Memo key includes DB presence: an agent that starts DB-less and gains a
|
||||
# handle later (run_agent._get_session_db_for_recall lazily attaches one)
|
||||
# DB presence is part of the key: an agent that gains a DB handle later
|
||||
# must re-resolve instead of staying pinned to the physical id.
|
||||
key = (sid, db is not None)
|
||||
memo = getattr(agent, _MEMO_ATTR, None)
|
||||
if isinstance(memo, tuple) and len(memo) == 2 and memo[0] == key:
|
||||
return memo[1]
|
||||
# A declared conversation key outranks the lineage walk: it is stable
|
||||
# across compression rotation AND across a host's per-response ids, which
|
||||
# the walk cannot see (#96811).
|
||||
root = declared_conversation_scope(agent) or (
|
||||
_lineage_root(sid, db) if db is not None else None
|
||||
)
|
||||
scope = root or sid
|
||||
# Memoize on a successful walk, or when there is no DB to consult at all,
|
||||
# or when the agent will never persist a row (background-review forks set
|
||||
# _persist_disabled but still hold a DB handle — without this, every API
|
||||
# call would re-run the lineage query forever).
|
||||
# A failed/empty walk on a persisting agent is NOT memoized: falling back
|
||||
# to the physical id is the correct degraded answer right now (row not
|
||||
# persisted yet, transient DB error), but pinning it for the whole segment
|
||||
# would keep the scope wrong after the session row lands.
|
||||
if (
|
||||
root is not None
|
||||
or db is None
|
||||
or getattr(agent, "_persist_disabled", False)
|
||||
):
|
||||
# Memoize on success, with no DB, or when the agent never persists a row
|
||||
# (background-review forks hold a DB handle but set _persist_disabled).
|
||||
# A failed/empty walk on a persisting agent is NOT memoized: the physical
|
||||
# id is right for now (row not yet persisted, transient error) but would
|
||||
# stay wrong for the whole segment once the row lands.
|
||||
if root is not None or db is None or getattr(agent, "_persist_disabled", False):
|
||||
try:
|
||||
setattr(agent, _MEMO_ATTR, (key, scope))
|
||||
except Exception:
|
||||
# Frozen/slotted test doubles — resolution still works, just
|
||||
# unmemoized.
|
||||
pass
|
||||
pass # frozen/slotted doubles: resolution works, just unmemoized
|
||||
return scope
|
||||
|
||||
|
||||
@@ -303,14 +200,11 @@ def declared_conversation_scope_safe(agent: Any) -> Optional[str]:
|
||||
|
||||
|
||||
def resolve_prompt_cache_scope_safe(agent: Any) -> Optional[str]:
|
||||
"""Never-raising variant of :func:`resolve_prompt_cache_scope`.
|
||||
"""Never-raising variant of :func:`resolve_prompt_cache_scope` (None on failure/empty).
|
||||
|
||||
Returns None on any failure (or when there is no scope). Consumers treat
|
||||
None/empty as "fall back to the physical session_id", so a resolution
|
||||
failure degrades to pre-#79017 behavior instead of blocking the caller —
|
||||
important at turn_context's call site, where an exception raised inside
|
||||
the ``set_runtime_main(...)`` argument list would otherwise skip the whole
|
||||
runtime binding, not just the cache scope.
|
||||
Consumers treat None as "use the physical session_id"; at turn_context's
|
||||
call site an exception inside the ``set_runtime_main(...)`` argument list
|
||||
would skip the whole runtime binding, not just the cache scope.
|
||||
"""
|
||||
try:
|
||||
return resolve_prompt_cache_scope(agent) or None
|
||||
|
||||
+88
-212
@@ -1,13 +1,10 @@
|
||||
"""Anthropic prompt caching strategy.
|
||||
"""Anthropic prompt caching strategy — pure functions, no AIAgent dependency.
|
||||
|
||||
The default layout uses 4 cache_control breakpoints: the static system
|
||||
prefix, the end of the system prompt, and the last 2 non-system messages.
|
||||
When a static system prefix is unavailable, it falls back to one system
|
||||
breakpoint plus the last 3 messages. All markers use the same TTL (5m or 1h).
|
||||
This preserves intra-session caching while allowing new sessions to reuse the
|
||||
stable system-prompt prefix.
|
||||
|
||||
Pure functions -- no class state, no AIAgent dependency.
|
||||
Default layout: 4 cache_control breakpoints — the static system prefix, the end
|
||||
of the system prompt, and the last 2 non-system messages. Without a static
|
||||
prefix: one system breakpoint plus the last 3 messages. All markers share one
|
||||
TTL (5m or 1h). This keeps intra-session caching while letting new sessions
|
||||
reuse the stable system-prompt prefix.
|
||||
"""
|
||||
|
||||
import copy
|
||||
@@ -24,30 +21,18 @@ class PromptCachePlan:
|
||||
messages: List[Dict[str, Any]]
|
||||
tools: List[Dict[str, Any]]
|
||||
|
||||
@property
|
||||
def marker_count(self) -> int:
|
||||
"""Wire-visible cache markers in this plan (computed on demand).
|
||||
|
||||
Only tests consume this; keeping it lazy avoids walking every
|
||||
message part and tool schema on the per-request hot path.
|
||||
"""
|
||||
return _count_cache_markers(self.messages, self.tools)
|
||||
|
||||
|
||||
def envelope_tool_part_cache_markers_supported(
|
||||
provider: str | None, base_url: str | None
|
||||
) -> bool:
|
||||
"""Whether the envelope-layout route honors part-level markers on role:tool.
|
||||
|
||||
OpenRouter (and Nous Portal, which proxies to it) relocate a
|
||||
``cache_control`` sitting on a tool message's content part onto the
|
||||
``tool_result`` block during their OpenAI→Anthropic translation, so the
|
||||
marker is honored there. LiteLLM-style OpenAI-wire proxies instead map
|
||||
content parts verbatim: the part-level marker lands at
|
||||
``tool_result.content[0]``, which the Anthropic Messages schema forbids —
|
||||
a non-retryable HTTP 400 that kills the whole turn (#89886). On those
|
||||
routes tool messages must not carry part-level markers at all; the
|
||||
breakpoint budget reallocates to the nearest eligible message instead.
|
||||
OpenRouter (and Nous Portal, which proxies to it) relocate a part-level
|
||||
``cache_control`` onto the ``tool_result`` block during OpenAI→Anthropic
|
||||
translation. LiteLLM-style proxies copy parts verbatim, so the marker lands
|
||||
at ``tool_result.content[0]`` — forbidden by the Anthropic schema, a
|
||||
non-retryable 400. On those routes tool messages carry no part markers and
|
||||
the breakpoint budget reallocates to the nearest eligible message.
|
||||
"""
|
||||
from agent.agent_runtime_helpers import _is_litellm_route
|
||||
|
||||
@@ -65,27 +50,19 @@ def _apply_cache_marker(
|
||||
content = msg.get("content")
|
||||
|
||||
if role == "tool" and native_anthropic:
|
||||
# Native Anthropic layout: top-level marker; the adapter moves it
|
||||
# inside the tool_result block.
|
||||
# Top-level marker; the native adapter moves it inside tool_result.
|
||||
msg["cache_control"] = cache_marker
|
||||
return
|
||||
|
||||
if role == "tool" and not tool_part_markers:
|
||||
# Envelope route whose OpenAI→Anthropic translation copies content
|
||||
# parts verbatim (LiteLLM et al.): a part-level marker becomes
|
||||
# tool_result.content[0].cache_control → non-retryable 400 (#89886).
|
||||
# LiteLLM-style envelope: a part marker becomes
|
||||
# tool_result.content[0].cache_control → non-retryable 400.
|
||||
return
|
||||
|
||||
if content is None or content == "":
|
||||
if role == "tool" and not native_anthropic:
|
||||
# OpenRouter rejects top-level cache_control on role:tool (silent
|
||||
# hang) and an empty message has no content part to carry the
|
||||
# marker — skip. Non-empty tool content falls through below and
|
||||
# gets the marker on a content part, which OpenRouter honors.
|
||||
return
|
||||
if role == "assistant" and not native_anthropic:
|
||||
# Empty assistant turns are pure tool_calls. A top-level marker
|
||||
# here is ignored on the envelope layout, so skip.
|
||||
# Envelope layout: OpenRouter rejects top-level cache_control on
|
||||
# role:tool (silent hang), and ignores it on empty assistant turns
|
||||
# (pure tool_calls) — neither has a content part to carry it.
|
||||
if role in ("tool", "assistant") and not native_anthropic:
|
||||
return
|
||||
msg["cache_control"] = cache_marker
|
||||
return
|
||||
@@ -96,17 +73,13 @@ def _apply_cache_marker(
|
||||
if stable_prefix is not None:
|
||||
suffix = content[len(stable_prefix):]
|
||||
if suffix.strip():
|
||||
# Builder-declared boundary (#81867): the scaffold carries the
|
||||
# breakpoint, the volatile invocation tail rides unmarked so a
|
||||
# changed ticket ID or timestamp no longer invalidates the
|
||||
# whole skill body. Request-local only — the canonical session
|
||||
# message stays a plain string.
|
||||
# Builder-declared boundary: the scaffold carries the
|
||||
# breakpoint and the volatile tail rides unmarked, so a
|
||||
# changed ticket ID/timestamp no longer invalidates the
|
||||
# skill body. Request-local only — the stored message
|
||||
# stays a plain string.
|
||||
msg["content"] = [
|
||||
{
|
||||
"type": "text",
|
||||
"text": stable_prefix,
|
||||
"cache_control": cache_marker,
|
||||
},
|
||||
{"type": "text", "text": stable_prefix, "cache_control": cache_marker},
|
||||
{"type": "text", "text": suffix},
|
||||
]
|
||||
return
|
||||
@@ -126,17 +99,12 @@ def _can_carry_marker(
|
||||
) -> bool:
|
||||
"""True if a marker on this message is actually honored by the provider.
|
||||
|
||||
On the native Anthropic layout every message works (top-level markers are
|
||||
relocated by the adapter). On the envelope layout (OpenRouter et al.) only
|
||||
markers inside content parts are honored: empty-content messages (e.g.
|
||||
assistant turns that are pure tool_calls) and empty tool messages would
|
||||
receive a top-level marker the provider ignores — wasting one of the four
|
||||
breakpoints. Skip those so the breakpoints land on messages that count.
|
||||
|
||||
``tool_part_markers=False`` (LiteLLM-style envelope routes, #89886)
|
||||
additionally excludes ALL role:tool messages: their part-level marker
|
||||
would be forwarded verbatim into ``tool_result.content[]`` and rejected
|
||||
with a non-retryable 400, so the breakpoint must reallocate instead.
|
||||
Native Anthropic honors every message (the adapter relocates top-level
|
||||
markers). The envelope layout only honors markers inside content parts, so
|
||||
empty-content messages would waste one of the four breakpoints; with
|
||||
``tool_part_markers=False`` (LiteLLM-style routes) every role:tool message
|
||||
is excluded too, since its part marker would be rejected with a 400.
|
||||
Must agree with :func:`_apply_cache_marker`, which marks only the LAST part.
|
||||
"""
|
||||
if native_anthropic:
|
||||
return True
|
||||
@@ -146,10 +114,6 @@ def _can_carry_marker(
|
||||
if content is None or content == "":
|
||||
return False
|
||||
if isinstance(content, list):
|
||||
# _apply_cache_marker only marks the LAST content part, so the carrier
|
||||
# predicate must agree: a list whose last element isn't a dict cannot
|
||||
# actually receive a marker and would waste a breakpoint. Mirror the
|
||||
# `content` truthiness + last-element-dict check in _apply_cache_marker.
|
||||
return bool(content) and isinstance(content[-1], dict)
|
||||
return isinstance(content, str)
|
||||
|
||||
@@ -162,66 +126,31 @@ def _build_marker(ttl: str) -> Dict[str, str]:
|
||||
return marker
|
||||
|
||||
|
||||
# Alibaba-family providers (Qwen routes). Their context cache documents a
|
||||
# five-minute window (renewed on hit) and rejects the Anthropic 1h tier.
|
||||
# Shared with agent_runtime_helpers.anthropic_prompt_cache_policy so the
|
||||
# cache-policy opt-in and the TTL clamp can never desync (#84733).
|
||||
# Alibaba-family providers (Qwen routes): documented five-minute context cache,
|
||||
# Anthropic 1h tier rejected. Shared with
|
||||
# agent_runtime_helpers.anthropic_prompt_cache_policy so the cache-policy
|
||||
# opt-in and the TTL clamp never desync. Do NOT narrow this set to extend a
|
||||
# TTL — it also drives the marker-layout opt-in, so narrowing DISABLES caching.
|
||||
ALIBABA_FAMILY_PROVIDERS = frozenset({
|
||||
"opencode",
|
||||
"opencode-zen",
|
||||
"opencode-go",
|
||||
"opencode-zen",
|
||||
"alibaba",
|
||||
})
|
||||
|
||||
|
||||
# --- 1h-tier membership: an ALLOW-list, deliberately minimal ----------------
|
||||
#
|
||||
# #84733 clamped 1h -> 5m for the whole alibaba/opencode family, reasoning from
|
||||
# Alibaba's PUBLISHED Qwen docs. Wire measurement on the opencode-go route
|
||||
# contradicts the docs. Controlled run: identical request, only the ttl flag
|
||||
# varying, read back after 11 minutes with no intervening call (a read renews
|
||||
# the window and would mask expiry):
|
||||
#
|
||||
# qwen3.8-max ttl=1h -> cache_read 2122 SURVIVED
|
||||
# qwen3.8-max ttl=- -> cache_read 0 EXPIRED <- control
|
||||
# glm-5.2 ttl=1h -> cache_read 2092 SURVIVED
|
||||
# minimax-m2.5 ttl=1h -> cache_read 0 EXPIRED
|
||||
#
|
||||
# Read the two non-qwen rows for what they are: evidence about the ROUTE, not
|
||||
# about traffic Hermes sends today. anthropic_prompt_cache_policy currently
|
||||
# opts opencode-go in only for qwen models, so glm-5.2 and minimax-m2.5 on
|
||||
# that route receive no cache_control marker at all and never reach this
|
||||
# clamp in production. They constrain the route-level rule; they are not
|
||||
# live paths.
|
||||
#
|
||||
# Only opencode-go is listed: it is the only route measured. Other opencode
|
||||
# routes stay clamped because they were NOT measured, not because they are
|
||||
# known bad. opencode-zen returns cache_creation.ephemeral_1h_input_tokens for
|
||||
# Claude models, so it is a candidate -- but qwen on zen is unmeasured, so
|
||||
# adding the provider wholesale would outrun the evidence.
|
||||
#
|
||||
# WARNING: opencode-go labels EVERY write `ephemeral_5m_input_tokens` whatever
|
||||
# ttl was requested. That label is NOT evidence of the retention window -- it
|
||||
# is what made the original docs-based reasoning look confirmed. Verify only
|
||||
# with a delayed read past 5 minutes and no intervening call.
|
||||
#
|
||||
# NOTE: kept separate from ALIBABA_FAMILY_PROVIDERS on purpose. That set also
|
||||
# drives the cache-marker-layout OPT-IN in
|
||||
# agent_runtime_helpers.anthropic_prompt_cache_policy; narrowing it would
|
||||
# silently DISABLE caching for qwen on opencode-go rather than extend its TTL.
|
||||
# 1h-tier ALLOW-list: only routes wire-measured to retain a 1h marker (delayed
|
||||
# read past 5 minutes with no intervening call — an intervening read renews the
|
||||
# window and masks expiry). Other opencode routes stay clamped because they are
|
||||
# UNMEASURED, not known-bad. Note opencode-go labels every write
|
||||
# `ephemeral_5m_input_tokens` regardless of requested ttl; that label is not
|
||||
# evidence of the retention window.
|
||||
MEASURED_1H_PROVIDERS = frozenset({
|
||||
"opencode-go",
|
||||
})
|
||||
|
||||
# Models measured to ignore the 1h tier even on a 1h-capable route.
|
||||
#
|
||||
# SCOPE: consulted only for providers already in MEASURED_1H_PROVIDERS. The
|
||||
# measurement was taken on the opencode-go route, so it says nothing about the
|
||||
# same model reached some other way -- and MiniMax on its own
|
||||
# Anthropic-compatible endpoint IS a separate, cache-eligible route
|
||||
# (anthropic_prompt_cache_policy opts it in by provider id / host match).
|
||||
# Checking this set globally would have silently regressed that unrelated
|
||||
# route's configured 1h to 5m off the back of an opencode-go observation.
|
||||
# Models measured to ignore the 1h tier on a MEASURED_1H_PROVIDERS route.
|
||||
# Consulted only there: the same model on its own Anthropic-compatible endpoint
|
||||
# is a separate cache-eligible route and must not inherit this clamp.
|
||||
NO_1H_TIER_MODELS = frozenset({
|
||||
"minimax-m2.5",
|
||||
})
|
||||
@@ -235,9 +164,8 @@ def _flat_model(model: str) -> str:
|
||||
def is_qwen_model(model: str) -> bool:
|
||||
"""True when ``model`` names a Qwen-family model (case-insensitive).
|
||||
|
||||
Shared by the TTL clamp below and
|
||||
``agent_runtime_helpers.anthropic_prompt_cache_policy`` so the
|
||||
cache-policy opt-in and the clamp can never desync (#84733).
|
||||
Shared with ``agent_runtime_helpers.anthropic_prompt_cache_policy`` so the
|
||||
cache-policy opt-in and the TTL clamp never desync.
|
||||
"""
|
||||
return "qwen" in (model or "").lower()
|
||||
|
||||
@@ -250,30 +178,18 @@ def effective_cache_ttl(
|
||||
) -> str:
|
||||
"""Clamp a requested cache TTL to what the destination route supports.
|
||||
|
||||
Qwen/Alibaba context caching documents an explicit five-minute window
|
||||
(renewed on hit); the Anthropic ``1h`` tier is ignored/rejected there,
|
||||
so a configured ``1h`` regresses to ``5m`` instead of shipping a marker
|
||||
the provider drops and creating a false 1h-cache expectation (#84733).
|
||||
Exception: routes in ``MEASURED_1H_PROVIDERS`` were wire-measured to
|
||||
honour the tier (delayed read past 5 minutes) and keep ``1h`` — minus
|
||||
any model in ``NO_1H_TIER_MODELS`` measured to ignore it on that route.
|
||||
All other caching routes keep the requested TTL.
|
||||
|
||||
``None`` (caching active with no explicit tier) resolves to ``5m``.
|
||||
Qwen/Alibaba routes document a five-minute window and drop the ``1h``
|
||||
tier, so a configured ``1h`` regresses to ``5m`` there instead of creating
|
||||
a false 1h-cache expectation — except on ``MEASURED_1H_PROVIDERS``, which
|
||||
keep ``1h`` minus any ``NO_1H_TIER_MODELS`` model. The measured-route check
|
||||
runs BEFORE the generic Qwen clamp, which would otherwise swallow every
|
||||
Qwen model on it. ``None`` resolves to ``5m``.
|
||||
"""
|
||||
if ttl != "1h":
|
||||
return ttl or "5m"
|
||||
if (provider or "").lower() in MEASURED_1H_PROVIDERS:
|
||||
# Route measured to honour the tier -- checked BEFORE the generic
|
||||
# is_qwen_model clamp below, which would otherwise swallow every Qwen
|
||||
# model on it. Within the route, a model measured to ignore the tier
|
||||
# still wins; the denial stays nested here so an opencode-go
|
||||
# observation cannot leak out and reclamp the same model on an
|
||||
# unrelated route.
|
||||
return "5m" if _flat_model(model) in NO_1H_TIER_MODELS else "1h"
|
||||
if is_qwen_model(model):
|
||||
return "5m"
|
||||
if (provider or "").lower() in ALIBABA_FAMILY_PROVIDERS:
|
||||
if is_qwen_model(model) or (provider or "").lower() in ALIBABA_FAMILY_PROVIDERS:
|
||||
return "5m"
|
||||
return "1h"
|
||||
|
||||
@@ -289,21 +205,13 @@ def _apply_system_cache_markers(
|
||||
) -> int:
|
||||
"""Mark the static system prefix (and optionally the full prompt).
|
||||
|
||||
The system prompt remains one stored string. Splitting it only in the
|
||||
outgoing request keeps session persistence and non-Anthropic transports
|
||||
unchanged while making the stable prefix independently cacheable.
|
||||
|
||||
``mark_suffix=False`` is the tool-cache-plan layout: only the static
|
||||
prefix carries a marker, the volatile suffix rides unmarked (its
|
||||
breakpoint budget is spent on the tools array instead).
|
||||
|
||||
``fallback_to_whole=False`` skips marking entirely when the prefix
|
||||
split is not possible (no prefix, mismatched prefix, non-string
|
||||
content) instead of marking the whole message.
|
||||
|
||||
When the prompt IS exactly the static prefix (empty suffix), the whole
|
||||
message is marked as a single block — never a two-part split with an
|
||||
empty text block, which Anthropic rejects.
|
||||
The system prompt stays one stored string; it is split only in the
|
||||
outgoing request so persistence and non-Anthropic transports are
|
||||
unchanged. ``mark_suffix=False`` is the tool-cache-plan layout (suffix
|
||||
unmarked, its budget spent on the tools array). ``fallback_to_whole=False``
|
||||
marks nothing when the prefix split is impossible. When the prompt IS the
|
||||
prefix (empty/whitespace suffix) the whole message is marked as one block —
|
||||
never a split with an empty text block, which Anthropic rejects.
|
||||
|
||||
Returns the number of markers applied (0, 1, or 2).
|
||||
"""
|
||||
@@ -320,17 +228,10 @@ def _apply_system_cache_markers(
|
||||
if mark_suffix:
|
||||
suffix_part["cache_control"] = cache_marker
|
||||
message["content"] = [
|
||||
{
|
||||
"type": "text",
|
||||
"text": static_system_prefix,
|
||||
"cache_control": cache_marker,
|
||||
},
|
||||
{"type": "text", "text": static_system_prefix, "cache_control": cache_marker},
|
||||
suffix_part,
|
||||
]
|
||||
return 2 if mark_suffix else 1
|
||||
# Empty/whitespace-only suffix: the stored prompt IS the static prefix. Mark it as
|
||||
# one whole block — a [marked-prefix, ""] split would put an empty
|
||||
# text block on the wire (HTTP 400 on native Anthropic).
|
||||
_apply_cache_marker(message, cache_marker, native_anthropic=native_anthropic)
|
||||
return 1
|
||||
|
||||
@@ -345,27 +246,17 @@ def strip_anthropic_cache_control(
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Remove ``cache_control`` markers and undo decoration-produced list shapes.
|
||||
|
||||
Used before re-applying decoration after a mid-turn provider failover so
|
||||
the mutated, undecorated shape (image shrink / ASCII cleanup / etc.) is
|
||||
preserved while markers match the *new* provider's cache policy (#72626).
|
||||
Used before re-decorating after a mid-turn provider failover, so the
|
||||
mutated undecorated shape is preserved while markers match the new
|
||||
provider's policy. Flattening back to a plain string is restricted to the
|
||||
exact shapes :func:`apply_anthropic_cache_control` produces from string
|
||||
content — a single text part, the two-part ``[static, volatile]`` system
|
||||
split, or the two-part skill split — so the ``""``-join is provably
|
||||
byte-exact; organic multi-part text and parts with extra keys keep their
|
||||
structure. Marker removal is copy-on-write on part dicts: parts can alias
|
||||
caller-held lists and stripping must never rewrite the stored transcript.
|
||||
|
||||
Flattening back to a plain string is restricted to the exact shapes
|
||||
:func:`apply_anthropic_cache_control` produces from string content —
|
||||
a single ``{"type": "text"}`` part, the two-part ``[static, volatile]``
|
||||
system split, or the two-part builder-declared skill split (recognised
|
||||
by its marker-on-the-first-part shape, so flattening never depends on
|
||||
the prefix registry still holding the entry) — so the ``""``-join is
|
||||
provably byte-exact. Organic
|
||||
multi-part text (merged user turns, imported transcripts) and parts
|
||||
carrying extra keys (``citations`` etc.) keep their structure; only
|
||||
per-part markers are removed. Marker removal is copy-on-write on the
|
||||
part dicts: content parts can alias caller-held message lists (the main
|
||||
send path now hands structurally-cloned copies via
|
||||
_clone_message_for_send, but other callers may pass shallow copies),
|
||||
and stripping must never rewrite the stored transcript.
|
||||
|
||||
Mutates the top-level message dicts of ``api_messages`` in place and
|
||||
returns the same list.
|
||||
Mutates the top-level message dicts in place and returns the same list.
|
||||
"""
|
||||
for msg in api_messages:
|
||||
if not isinstance(msg, dict):
|
||||
@@ -374,13 +265,11 @@ def strip_anthropic_cache_control(
|
||||
content = msg.get("content")
|
||||
if not isinstance(content, list):
|
||||
continue
|
||||
# Two-part skill-invocation split (#81867). The builder-declared
|
||||
# boundary is the only decoration that marks the *first* part of a
|
||||
# user message: list content otherwise receives its marker on the
|
||||
# last part, and the two-part [static, volatile] split is role-gated
|
||||
# to system. So the shape alone identifies it, and flattening stays
|
||||
# correct even when the prefix registry has since evicted the entry
|
||||
# (failover re-decorates a request built many messages ago, #72626).
|
||||
# The builder-declared skill split is the only decoration that marks
|
||||
# the FIRST part of a user message (list content is otherwise marked
|
||||
# on the last part; the [static, volatile] split is system-only), so
|
||||
# the shape alone identifies it even after the prefix registry has
|
||||
# evicted the entry.
|
||||
skill_split_shape = (
|
||||
msg.get("role") == "user"
|
||||
and len(content) == 2
|
||||
@@ -505,9 +394,9 @@ def build_prompt_cache_plan(
|
||||
) -> PromptCachePlan:
|
||||
"""Build isolated cache sections for one resolved request destination.
|
||||
|
||||
``tool_part_markers=False`` (LiteLLM-style envelope routes, #89886)
|
||||
keeps ``cache_control`` off role:tool content parts; breakpoints
|
||||
reallocate to the nearest eligible non-tool message.
|
||||
``tool_part_markers=False`` (LiteLLM-style envelope routes) keeps
|
||||
``cache_control`` off role:tool content parts; breakpoints reallocate to
|
||||
the nearest eligible non-tool message.
|
||||
"""
|
||||
messages = copy.deepcopy(api_messages or [])
|
||||
strip_anthropic_cache_control(messages)
|
||||
@@ -558,23 +447,13 @@ def apply_anthropic_cache_control(
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Apply Anthropic cache-control markers to API messages.
|
||||
|
||||
When ``static_system_prefix`` exactly matches the beginning of a string
|
||||
system prompt, it receives an early marker and the full system prompt gets
|
||||
a trailing marker. The remaining two markers target the latest cacheable
|
||||
non-system messages. Without that prefix, the legacy system-and-3 layout
|
||||
is retained.
|
||||
|
||||
Idempotent: pre-existing ``cache_control`` markers are stripped from a
|
||||
per-message copy before new ones are placed, so calling this twice (or
|
||||
handing it messages a prior call already marked) can never accumulate
|
||||
past 4 markers. Only messages that already carry a marker pay the copy
|
||||
cost — a shallow top-level copy suffices because
|
||||
:func:`strip_anthropic_cache_control` is copy-on-write on content parts —
|
||||
and the rest of the copy-on-write contract is unchanged (#90971).
|
||||
|
||||
``tool_part_markers=False`` (LiteLLM-style envelope routes, #89886)
|
||||
keeps markers off role:tool messages entirely; the breakpoint budget
|
||||
reallocates to the nearest eligible non-tool message.
|
||||
With a matching ``static_system_prefix`` the prefix gets an early marker
|
||||
and the full system prompt a trailing one; the remaining two markers go to
|
||||
the latest cacheable non-system messages. Without it, the legacy
|
||||
system-and-3 layout applies. Idempotent: pre-existing markers are stripped
|
||||
from a per-message copy first, so repeated calls never accumulate past 4
|
||||
markers; a shallow top-level copy suffices because
|
||||
:func:`strip_anthropic_cache_control` is copy-on-write on content parts.
|
||||
|
||||
Returns:
|
||||
Shallow copy of message list with selective deep copies of modified messages.
|
||||
@@ -594,9 +473,6 @@ def apply_anthropic_cache_control(
|
||||
and any(isinstance(part, dict) and "cache_control" in part for part in content)
|
||||
)
|
||||
if has_marker:
|
||||
# Shallow top-level copy is enough: strip pops the top-level key
|
||||
# and rebuilds content lists/part dicts copy-on-write, so the
|
||||
# caller's message (and any aliased parts) are never mutated.
|
||||
messages[i] = strip_anthropic_cache_control([dict(msg)])[0]
|
||||
|
||||
breakpoints_used = 0
|
||||
|
||||
+415
-766
File diff suppressed because it is too large
Load Diff
+71
-129
@@ -1,18 +1,11 @@
|
||||
"""Replay-history sanitization shared across resume code paths.
|
||||
|
||||
When a session's last turn dies mid-tool-loop — the process is killed by a
|
||||
restart/shutdown command, a stale-timeout fires, or an interrupt lands before
|
||||
the tool result is written — the persisted transcript can end with a dangling
|
||||
``assistant(tool_calls)`` (no matching ``tool`` answer) or an interrupted
|
||||
``assistant→tool`` block. On resume the model sees that broken tail and
|
||||
re-issues the unanswered call, producing an endless "thinking"/reboot loop
|
||||
(#49201, #29086).
|
||||
|
||||
These pure helpers strip those tails before the history is replayed to the
|
||||
model. They were originally local to ``gateway/run.py`` (which fixed the
|
||||
messaging-gateway path) and are extracted here so every resume surface — the
|
||||
messaging gateway AND the TUI/WebUI gateway — shares the same cleanup instead
|
||||
of the WebUI path silently skipping it.
|
||||
A session whose last turn died mid-tool-loop (process killed by a restart
|
||||
command, stale timeout, interrupt before the tool result was written) persists
|
||||
a dangling ``assistant(tool_calls)`` or interrupted ``assistant→tool`` tail. On
|
||||
resume the model re-issues the unanswered call → endless "thinking"/reboot loop.
|
||||
These pure helpers strip those tails before replay, for EVERY resume surface
|
||||
(messaging gateway and TUI/WebUI gateway alike).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -39,16 +32,36 @@ def is_interrupted_tool_result(content: Any) -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _call_name(call: Dict[str, Any]) -> str:
|
||||
return str((call.get("function") or {}).get("name") or "")
|
||||
|
||||
|
||||
def _call_id(call: Dict[str, Any]) -> str:
|
||||
return str(call.get("id") or call.get("call_id") or "")
|
||||
|
||||
|
||||
def _any_side_effecting(calls: List[Dict[str, Any]]) -> bool:
|
||||
return any(tool_may_have_side_effect(_call_name(call)) for call in calls)
|
||||
|
||||
|
||||
def _orphan_recovery(name: str, unknown_text: str, none_text: str) -> tuple:
|
||||
"""(effect_disposition, content) for an interrupted/dangling call named ``name``."""
|
||||
if tool_may_have_side_effect(name):
|
||||
return "unknown", unknown_text
|
||||
return "none", none_text
|
||||
|
||||
|
||||
def strip_interrupted_tool_tails(
|
||||
agent_history: List[Dict[str, Any]],
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Strip interrupted assistant→tool sequences from replay history.
|
||||
|
||||
Older interrupted gateway turns can be followed by a queued real user
|
||||
message, so the interrupted assistant/tool block is not necessarily the
|
||||
final tail by the time we rebuild replay history. Remove any contiguous
|
||||
assistant(tool_calls) + tool-result block that contains an interrupted tool
|
||||
result, while preserving successful tool-call sequences intact.
|
||||
The interrupted block is not necessarily the final tail (a queued real user
|
||||
message may follow it), so every contiguous assistant(tool_calls)+tool-result
|
||||
block containing an interrupted result is handled; successful sequences stay
|
||||
intact. Read-only blocks are dropped; blocks with a side-effecting call are
|
||||
KEPT with the interrupted results rewritten as orphan-recovery notices, since
|
||||
the effect may already have happened and erasing it would hide that.
|
||||
"""
|
||||
if not agent_history:
|
||||
return agent_history
|
||||
@@ -69,18 +82,8 @@ def strip_interrupted_tool_tails(
|
||||
for m in tool_results
|
||||
):
|
||||
calls = msg.get("tool_calls") or []
|
||||
if any(
|
||||
tool_may_have_side_effect(
|
||||
str((call.get("function") or {}).get("name") or "")
|
||||
)
|
||||
for call in calls
|
||||
):
|
||||
call_names = {
|
||||
str(call.get("id") or call.get("call_id") or ""): str(
|
||||
(call.get("function") or {}).get("name") or ""
|
||||
)
|
||||
for call in calls
|
||||
}
|
||||
if _any_side_effecting(calls):
|
||||
call_names = {_call_id(call): _call_name(call) for call in calls}
|
||||
cleaned.append(msg)
|
||||
for tool_result in tool_results:
|
||||
if not is_interrupted_tool_result(tool_result.get("content", "")):
|
||||
@@ -88,14 +91,11 @@ def strip_interrupted_tool_tails(
|
||||
continue
|
||||
recovered = dict(tool_result)
|
||||
name = call_names.get(str(tool_result.get("tool_call_id") or ""), "")
|
||||
recovered["effect_disposition"] = (
|
||||
"unknown" if tool_may_have_side_effect(name) else "none"
|
||||
)
|
||||
recovered["content"] = (
|
||||
recovered["effect_disposition"], recovered["content"] = _orphan_recovery(
|
||||
name,
|
||||
"[Orphan recovery: interrupted side-effecting tool may have "
|
||||
"executed; its effect is UNKNOWN. Inspect state before retrying.]"
|
||||
if recovered["effect_disposition"] == "unknown"
|
||||
else "[Orphan recovery: interrupted read-only tool did not complete.]"
|
||||
"executed; its effect is UNKNOWN. Inspect state before retrying.]",
|
||||
"[Orphan recovery: interrupted read-only tool did not complete.]",
|
||||
)
|
||||
cleaned.append(recovered)
|
||||
i = j
|
||||
@@ -122,24 +122,13 @@ def strip_dangling_tool_call_tail(
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Strip a trailing ``assistant(tool_calls)`` block left with NO answers.
|
||||
|
||||
When a tool call itself kills the gateway process (``docker restart``,
|
||||
``systemctl restart``, ``kill``, ``hermes gateway restart``), the process
|
||||
is terminated by SIGKILL *mid-call* — before the tool result is ever
|
||||
written and before the orderly shutdown rewind
|
||||
(``_drop_trailing_empty_response_scaffolding``) can run. The last thing
|
||||
persisted is the ``assistant`` message that issued the ``tool_calls``,
|
||||
with zero matching ``tool`` rows.
|
||||
|
||||
On resume the model sees an unanswered tool call at the tail and naturally
|
||||
re-issues it — which restarts the gateway again, producing the infinite
|
||||
reboot loop in #49201. ``strip_interrupted_tool_tails`` does not catch
|
||||
this because there is no tool result to inspect for an interrupt marker.
|
||||
|
||||
This strips that dangling tail at the source so there is nothing for the
|
||||
model to re-execute. It only acts when the tail is an
|
||||
``assistant(tool_calls)`` whose calls have NO corresponding ``tool``
|
||||
results — a completed assistant→tool pair (any tool answers present) is
|
||||
left untouched so genuine mid-progress tool loops still resume.
|
||||
A tool call that kills the gateway process itself (``docker restart``,
|
||||
``hermes gateway restart``) is SIGKILLed mid-call, before any tool result or
|
||||
the orderly shutdown rewind; the persisted tail is the assistant message with
|
||||
zero matching ``tool`` rows, which ``strip_interrupted_tool_tails`` cannot
|
||||
detect (no result to inspect). Only acts when the tail has NO tool answers —
|
||||
a partially answered block still resumes. Read-only tails are dropped;
|
||||
side-effecting ones get synthetic UNKNOWN-effect results instead of erasure.
|
||||
"""
|
||||
if not agent_history:
|
||||
return agent_history
|
||||
@@ -153,26 +142,18 @@ def strip_dangling_tool_call_tail(
|
||||
return agent_history
|
||||
|
||||
tool_calls = last.get("tool_calls") or []
|
||||
if any(
|
||||
tool_may_have_side_effect(
|
||||
str((call.get("function") or {}).get("name") or "")
|
||||
)
|
||||
for call in tool_calls
|
||||
):
|
||||
if _any_side_effecting(tool_calls):
|
||||
recovered = list(agent_history)
|
||||
for call in tool_calls:
|
||||
function = call.get("function") or {}
|
||||
name = str(function.get("name") or "unknown")
|
||||
call_id = str(call.get("id") or call.get("call_id") or "")
|
||||
disposition = "unknown" if tool_may_have_side_effect(name) else "none"
|
||||
content = (
|
||||
name = str((call.get("function") or {}).get("name") or "unknown")
|
||||
disposition, content = _orphan_recovery(
|
||||
name,
|
||||
"[Orphan recovery: this tool may have executed before Hermes stopped; "
|
||||
"its effect is UNKNOWN. Inspect current state before retrying.]"
|
||||
if disposition == "unknown"
|
||||
else "[Orphan recovery: this read-only tool did not complete and had no effect.]"
|
||||
"its effect is UNKNOWN. Inspect current state before retrying.]",
|
||||
"[Orphan recovery: this read-only tool did not complete and had no effect.]",
|
||||
)
|
||||
recovered.append(make_tool_result_message(
|
||||
name, content, call_id, effect_disposition=disposition,
|
||||
name, content, _call_id(call), effect_disposition=disposition,
|
||||
))
|
||||
logger.warning(
|
||||
"Recovered dangling side-effecting tool call(s) as UNKNOWN instead of erasing them"
|
||||
@@ -189,31 +170,23 @@ def strip_dangling_tool_call_tail(
|
||||
def sanitize_replay_history(
|
||||
agent_history: List[Dict[str, Any]],
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Apply both replay-tail strippers in the canonical order.
|
||||
|
||||
Convenience entry point for resume code paths: removes interrupted
|
||||
assistant→tool blocks anywhere in the history, then removes a dangling
|
||||
unanswered ``assistant(tool_calls)`` tail. Returns the same list object
|
||||
when there is nothing to strip.
|
||||
"""
|
||||
"""Both replay-tail strippers in canonical order (interrupted blocks, then
|
||||
dangling tail). Returns the same list object when nothing is stripped."""
|
||||
if not agent_history:
|
||||
return agent_history
|
||||
return strip_dangling_tool_call_tail(strip_interrupted_tool_tails(agent_history))
|
||||
|
||||
|
||||
# ──────────────────────────────────────────────────────────────────────
|
||||
# Stale dangerous-confirmation text expiry (#59607)
|
||||
# Stale dangerous-confirmation text expiry
|
||||
# ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
# How long a high-risk confirmation phrase remains valid.
|
||||
# Short on purpose: dangerous side effects should not survive any restart
|
||||
# or session resumption gap. The user can always re-confirm if needed.
|
||||
# Short on purpose: a dangerous confirmation must not survive any restart or
|
||||
# resume gap. The user can always re-confirm.
|
||||
_DANGEROUS_CONFIRMATION_EXPIRY_SECONDS = 60.0
|
||||
|
||||
# Confirmation phrases that unlock destructive host actions.
|
||||
# Substring match (case-insensitive) so that user variants (e.g. trailing
|
||||
# punctuation, additional context) still match. Add new patterns here when
|
||||
# new high-risk actions are introduced.
|
||||
# Confirmation phrases that unlock destructive host actions; case-insensitive
|
||||
# substring match so trailing punctuation / extra context still matches.
|
||||
_DANGEROUS_CONFIRMATION_PATTERNS: tuple = (
|
||||
"confirm forced restart",
|
||||
"confirm forced reboot",
|
||||
@@ -229,9 +202,8 @@ _DANGEROUS_CONFIRMATION_PATTERNS: tuple = (
|
||||
"確認重啟",
|
||||
)
|
||||
|
||||
# Replacement text for an expired confirmation. Redacting in place (rather
|
||||
# than deleting the message) preserves strict user/assistant role
|
||||
# alternation in the replayed history.
|
||||
# Redacting in place (rather than deleting the message) preserves strict
|
||||
# user/assistant role alternation in the replayed history.
|
||||
_EXPIRED_CONFIRMATION_SENTINEL = (
|
||||
"[A high-risk confirmation previously given here has EXPIRED and must "
|
||||
"not be acted on. Ask the user to re-confirm explicitly before "
|
||||
@@ -240,12 +212,7 @@ _EXPIRED_CONFIRMATION_SENTINEL = (
|
||||
|
||||
|
||||
def is_dangerous_confirmation(content: Any) -> bool:
|
||||
"""Return True if a user-message text matches a known dangerous confirmation.
|
||||
|
||||
Used by ``strip_stale_dangerous_confirmations`` to decide which
|
||||
transcript rows to expire. Substring + case-insensitive so that
|
||||
``"Please confirm forced restart, the host is critical"`` still matches.
|
||||
"""
|
||||
"""True if user-message text contains a known dangerous confirmation phrase."""
|
||||
if not isinstance(content, str):
|
||||
return False
|
||||
text = content.strip().lower()
|
||||
@@ -258,38 +225,15 @@ def strip_stale_dangerous_confirmations(
|
||||
now: float,
|
||||
expiry_seconds: float = _DANGEROUS_CONFIRMATION_EXPIRY_SECONDS,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Expire stale dangerous-confirmation text in user messages (#59607).
|
||||
"""Expire stale dangerous-confirmation text in user messages.
|
||||
|
||||
When a high-risk side effect (e.g. host restart via ``shutdown.exe``)
|
||||
runs, the user's plain-text confirmation phrase is persisted in the
|
||||
conversation transcript. If the host restart killed the gateway
|
||||
process before the assistant's tool result was written, the
|
||||
transcript tail ends on the assistant's text response — and the
|
||||
dangerous confirmation text remains in the user role.
|
||||
|
||||
On the next inbound message — possibly a casual "are you there?" from
|
||||
the user minutes later — the LLM sees the stale confirmation and may
|
||||
interpret the new turn as a fresh re-confirmation, re-executing the
|
||||
destructive action. This is the failure mode reported in #59607.
|
||||
|
||||
Expired confirmations are REDACTED IN PLACE, not removed: deleting a
|
||||
user message from the incident tail (``user(confirm) →
|
||||
assistant("OK, restarting")``) would leave two consecutive assistant
|
||||
messages, violating the strict role-alternation invariant providers
|
||||
enforce. The message survives with its role intact; only the trigger
|
||||
text is replaced by a sentinel that tells the model the confirmation
|
||||
has expired.
|
||||
|
||||
Messages without a timestamp are left untouched (backward
|
||||
compatibility: legacy transcripts and in-memory test scaffolding have
|
||||
no timestamps). User messages that contain dangerous confirmation
|
||||
text but are within the expiry window are also left untouched — they
|
||||
represent a fresh confirmation that has not yet been acted on.
|
||||
|
||||
Complements 75ed07ace (which strips the *assistant* side of the
|
||||
broken tail) by handling the *user* side: a stale plain-text
|
||||
confirmation that the assistant has not yet responded to in a way
|
||||
the resume logic recognises.
|
||||
If a host restart killed the gateway before the tool result was written, the
|
||||
user's confirmation phrase survives in the transcript; a casual "are you
|
||||
there?" minutes later can read to the model as a fresh re-confirmation and
|
||||
re-execute the destructive action. Expired confirmations are REDACTED IN
|
||||
PLACE (deleting the message would leave two consecutive assistant turns).
|
||||
Messages without a timestamp (legacy transcripts, test scaffolding) and
|
||||
confirmations still inside the expiry window are left untouched.
|
||||
"""
|
||||
if not agent_history:
|
||||
return agent_history
|
||||
@@ -312,10 +256,8 @@ def strip_stale_dangerous_confirmations(
|
||||
)
|
||||
redacted = dict(msg)
|
||||
redacted["content"] = _EXPIRED_CONFIRMATION_SENTINEL
|
||||
# Drop the api_content sidecar: it carries the exact bytes
|
||||
# previously sent — i.e. the dangerous confirmation this
|
||||
# redaction exists to expire. Replaying it verbatim would
|
||||
# undo the redaction on the wire.
|
||||
# The api_content sidecar carries the exact bytes previously sent
|
||||
# — the confirmation itself; replaying it would undo the redaction.
|
||||
drop_stale_api_content(redacted)
|
||||
cleaned.append(redacted)
|
||||
continue
|
||||
|
||||
+39
-47
@@ -1,13 +1,10 @@
|
||||
"""Single source of truth for the agent working directory.
|
||||
|
||||
`TERMINAL_CWD` is the runtime carrier for the configured working directory
|
||||
(design #19214/#19242: `terminal.cwd` is bridged once to `TERMINAL_CWD` at
|
||||
gateway/cron startup). The local-CLI backend deliberately leaves it unset and
|
||||
relies on the launch dir. Reading it in one place keeps the system prompt, the
|
||||
tool surfaces, and context-file discovery agreeing on where the agent lives.
|
||||
|
||||
Multi-session gateways can pin a logical cwd via the `_SESSION_CWD`
|
||||
contextvar; CLI/cron fall through to `TERMINAL_CWD`/launch cwd.
|
||||
(`terminal.cwd` is bridged to it once at gateway/cron startup; the local CLI
|
||||
leaves it unset and relies on the launch dir). Reading it in one place keeps the
|
||||
system prompt, tool surfaces, and context-file discovery agreeing on where the
|
||||
agent lives. Multi-session gateways can pin a logical cwd via `_SESSION_CWD`.
|
||||
"""
|
||||
|
||||
import logging
|
||||
@@ -22,18 +19,15 @@ _UNSET: Any = object()
|
||||
|
||||
_SESSION_CWD: ContextVar = ContextVar("HERMES_SESSION_CWD", default=_UNSET)
|
||||
|
||||
# The Python package/source root (this file lives at <root>/agent/runtime_cwd.py).
|
||||
# When a backend is launched from, or self-spawns into, this tree (the desktop
|
||||
# app default), an os.getcwd() fallback would inject this repo's contributor
|
||||
# AGENTS.md as authoritative project context. Context discovery must never
|
||||
# resolve here.
|
||||
# The package/source root (<root>/agent/runtime_cwd.py). A backend launched from
|
||||
# or self-spawned into this tree (desktop default) must never let an os.getcwd()
|
||||
# fallback inject this repo's contributor AGENTS.md as project context.
|
||||
_PACKAGE_ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
|
||||
def _is_install_tree(p: Path) -> bool:
|
||||
# True only when p IS the package root or sits inside it. Ancestors of the
|
||||
# package root (a user home that happens to contain the checkout, a --user
|
||||
# site-packages parent) are legitimate workspaces and must not be blocked.
|
||||
"""True only when ``p`` IS the package root or sits inside it — ancestors
|
||||
(a home dir containing the checkout) are legitimate workspaces."""
|
||||
try:
|
||||
p = p.resolve()
|
||||
except Exception:
|
||||
@@ -61,9 +55,9 @@ def _terminal_cwd_env() -> str:
|
||||
"""Scope-aware TERMINAL_CWD read (tools.terminal_scope.terminal_env).
|
||||
|
||||
Under gateway multiplexing the per-turn terminal scope carries the active
|
||||
profile's cwd; the process-global env var may hold another profile's
|
||||
value. Only an import failure falls back: an active refusal scope must
|
||||
raise, not silently resolve the launch profile's cwd.
|
||||
profile's cwd; the process-global env var may hold another profile's. Only
|
||||
an ImportError falls back: an active refusal scope must raise, not silently
|
||||
resolve the launch profile's cwd.
|
||||
"""
|
||||
try:
|
||||
from tools.terminal_scope import terminal_env
|
||||
@@ -75,51 +69,49 @@ def _terminal_cwd_env() -> str:
|
||||
def scope_terminal_cwd() -> str:
|
||||
"""Public wrapper — the scope-aware TERMINAL_CWD value (may be empty).
|
||||
|
||||
Shared by agent_init / skill_utils / code_execution_tool so every cwd
|
||||
consumer reads through the per-turn terminal scope under gateway
|
||||
multiplexing instead of the process-global env var.
|
||||
Shared by agent_init / skill_utils / code_execution_tool so every cwd consumer
|
||||
reads through the per-turn terminal scope under gateway multiplexing.
|
||||
"""
|
||||
return _terminal_cwd_env()
|
||||
|
||||
|
||||
def resolve_agent_cwd() -> Path:
|
||||
def _resolve_configured_cwd(*, override_is_final: bool) -> Path | None:
|
||||
"""Session override, then TERMINAL_CWD; each validated as a real directory.
|
||||
|
||||
``override_is_final``: a set-but-missing session override yields None
|
||||
instead of falling through to TERMINAL_CWD.
|
||||
"""
|
||||
override = _session_cwd_override()
|
||||
if override:
|
||||
p = Path(override).expanduser()
|
||||
if p.is_dir():
|
||||
return p
|
||||
logger.warning("configured working directory does not exist: %s", override)
|
||||
if override_is_final:
|
||||
return None
|
||||
raw = _terminal_cwd_env().strip()
|
||||
if raw:
|
||||
p = Path(raw).expanduser()
|
||||
if p.is_dir():
|
||||
return p
|
||||
logger.warning("TERMINAL_CWD does not exist: %s", raw)
|
||||
return Path(os.getcwd())
|
||||
return None
|
||||
|
||||
|
||||
def resolve_agent_cwd() -> Path:
|
||||
"""Configured cwd, else the launch dir (os.getcwd() — its OSError on a
|
||||
deleted cwd deliberately propagates; the caller owns that guard)."""
|
||||
p = _resolve_configured_cwd(override_is_final=False)
|
||||
return p if p is not None else Path(os.getcwd())
|
||||
|
||||
|
||||
def resolve_context_cwd() -> Path | None:
|
||||
# None means "no configured cwd": build_context_files_prompt then falls back
|
||||
# to the launch dir (os.getcwd()), correct for a local CLI launched inside a
|
||||
# real project. A configured path is validated here (previously it was passed
|
||||
# through unchecked, diverging from resolve_agent_cwd). An explicitly
|
||||
# configured path is otherwise honored verbatim — including the Hermes
|
||||
# source tree itself, which is a legitimate workspace when the user is
|
||||
# developing Hermes (per-surface policy for fallback-picked directories
|
||||
# lives in build_context_files_prompt; see #64590).
|
||||
override = _session_cwd_override()
|
||||
if override:
|
||||
p = Path(override).expanduser()
|
||||
if not p.is_dir():
|
||||
logger.warning("configured working directory does not exist: %s", override)
|
||||
else:
|
||||
return p
|
||||
return None
|
||||
raw = _terminal_cwd_env().strip()
|
||||
if raw:
|
||||
p = Path(raw).expanduser()
|
||||
if not p.is_dir():
|
||||
logger.warning("TERMINAL_CWD does not exist: %s", raw)
|
||||
else:
|
||||
return p
|
||||
return None
|
||||
"""Configured cwd for context-file discovery, or None for "no configured cwd".
|
||||
|
||||
None makes build_context_files_prompt fall back to the launch dir (correct
|
||||
for a local CLI launched inside a real project). A configured path is
|
||||
validated here; an existing one is honored verbatim — including the Hermes
|
||||
source tree itself, a legitimate workspace when developing Hermes
|
||||
(fallback-directory policy lives in build_context_files_prompt).
|
||||
"""
|
||||
return _resolve_configured_cwd(override_is_final=True)
|
||||
|
||||
+50
-196
@@ -1,87 +1,38 @@
|
||||
"""Skill bundles — aliases that load multiple skills under one slash command.
|
||||
|
||||
A skill bundle is a small YAML file that names a set of skills to load
|
||||
together. Invoking ``/<bundle-name>`` from the CLI or gateway loads every
|
||||
referenced skill's full content into a single user message, the same way
|
||||
``/<skill-name>`` does — but for N skills at once.
|
||||
|
||||
Storage
|
||||
-------
|
||||
Bundles live in ``~/.hermes/skill-bundles/*.yaml`` (and the equivalent
|
||||
profile-aware directory under ``HERMES_HOME``). Each file looks like::
|
||||
|
||||
name: backend-dev
|
||||
description: Backend feature work — code review, testing, PR workflow.
|
||||
skills:
|
||||
- github-code-review
|
||||
- test-driven-development
|
||||
- github-pr-workflow
|
||||
instruction: |
|
||||
Optional extra guidance to inject above the skill bodies.
|
||||
|
||||
The file's stem is treated as a fallback name when ``name:`` is absent, so
|
||||
dropping a YAML into the directory is enough to register a new bundle.
|
||||
|
||||
Conflict resolution
|
||||
-------------------
|
||||
If a bundle and a skill share the same slash name, the bundle wins. The
|
||||
slash command dispatch checks bundles first, then falls back to skills.
|
||||
This is the intended behavior — a user who names a bundle ``research``
|
||||
explicitly wants ``/research`` to mean their bundle, not whatever skill
|
||||
happens to share the slug.
|
||||
|
||||
Public API
|
||||
----------
|
||||
- :func:`get_skill_bundles` — return ``{"/slug": bundle_info}``
|
||||
- :func:`resolve_bundle_command_key` — map a user-typed command to its slug
|
||||
- :func:`build_bundle_invocation_message` — produce the full user message
|
||||
- :func:`reload_bundles` — re-scan disk and return a diff
|
||||
- :func:`list_bundles` — return rich info for display (``hermes bundles``)
|
||||
- :func:`save_bundle` / :func:`delete_bundle` — file-level operations
|
||||
Bundles are YAML files in ``<HERMES_HOME>/skill-bundles/`` (``name``,
|
||||
``description``, ``skills: [...]``, optional ``instruction``; the file stem is
|
||||
the fallback name). ``/<bundle>`` loads every member skill into one user
|
||||
message. If a bundle and a skill share a slug, the bundle wins — slash dispatch
|
||||
checks bundles first, on purpose.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Optional, Tuple
|
||||
|
||||
import yaml
|
||||
|
||||
from hermes_constants import get_hermes_home
|
||||
from agent.skill_commands import diff_command_snapshots, slugify_skill_name as _slugify
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Slug normalization — matches agent/skill_commands.py so a bundle and a
|
||||
# skill called "Foo Bar" both resolve to "/foo-bar".
|
||||
_BUNDLE_INVALID_CHARS = re.compile(r"[^a-z0-9-]")
|
||||
_BUNDLE_MULTI_HYPHEN = re.compile(r"-{2,}")
|
||||
|
||||
_bundles_cache: Dict[str, Dict[str, Any]] = {}
|
||||
_bundles_cache_mtime: Optional[float] = None
|
||||
|
||||
|
||||
def _bundles_dir() -> Path:
|
||||
"""Return the canonical bundles directory under HERMES_HOME.
|
||||
|
||||
Honors ``HERMES_BUNDLES_DIR`` for tests; falls back to
|
||||
``<HERMES_HOME>/skill-bundles``.
|
||||
"""
|
||||
"""Bundles directory: ``HERMES_BUNDLES_DIR`` override (tests) or ``<HERMES_HOME>/skill-bundles``."""
|
||||
override = os.environ.get("HERMES_BUNDLES_DIR")
|
||||
if override:
|
||||
return Path(override).expanduser()
|
||||
return get_hermes_home() / "skill-bundles"
|
||||
|
||||
|
||||
def _slugify(name: str) -> str:
|
||||
cmd = name.lower().replace(" ", "-").replace("_", "-")
|
||||
cmd = _BUNDLE_INVALID_CHARS.sub("", cmd)
|
||||
cmd = _BUNDLE_MULTI_HYPHEN.sub("-", cmd).strip("-")
|
||||
return cmd
|
||||
|
||||
|
||||
def _iter_bundle_files() -> List[Path]:
|
||||
base = _bundles_dir()
|
||||
if not base.exists():
|
||||
@@ -93,19 +44,9 @@ def _iter_bundle_files() -> List[Path]:
|
||||
|
||||
|
||||
def _max_mtime(files: List[Path]) -> float:
|
||||
"""Highest mtime across the bundle files plus the dir itself.
|
||||
|
||||
Watching the directory mtime catches deletions; watching individual
|
||||
files catches edits. Together they're a cheap freshness check.
|
||||
"""
|
||||
base = _bundles_dir()
|
||||
"""Highest mtime across the bundle files plus the dir itself (dir mtime catches deletions)."""
|
||||
mtimes = []
|
||||
if base.exists():
|
||||
try:
|
||||
mtimes.append(base.stat().st_mtime)
|
||||
except OSError:
|
||||
pass
|
||||
for f in files:
|
||||
for f in [_bundles_dir(), *files]:
|
||||
try:
|
||||
mtimes.append(f.stat().st_mtime)
|
||||
except OSError:
|
||||
@@ -114,11 +55,7 @@ def _max_mtime(files: List[Path]) -> float:
|
||||
|
||||
|
||||
def _load_bundle_file(path: Path) -> Optional[Dict[str, Any]]:
|
||||
"""Parse a single bundle YAML file. Returns ``None`` on any error.
|
||||
|
||||
Errors are logged at WARNING level. We don't raise — a broken bundle
|
||||
shouldn't take down slash command discovery.
|
||||
"""
|
||||
"""Parse one bundle YAML; ``None`` (logged) on any error so a broken bundle can't break discovery."""
|
||||
try:
|
||||
raw = path.read_text(encoding="utf-8")
|
||||
except OSError as exc:
|
||||
@@ -166,12 +103,7 @@ def _load_bundle_file(path: Path) -> Optional[Dict[str, Any]]:
|
||||
|
||||
|
||||
def scan_bundles() -> Dict[str, Dict[str, Any]]:
|
||||
"""Scan the bundles directory and rebuild the cache.
|
||||
|
||||
Returns the same mapping as :func:`get_skill_bundles` — ``"/slug"`` →
|
||||
bundle info dict. Later bundles with a duplicate slug are skipped with
|
||||
a warning (first wins, alphabetical order).
|
||||
"""
|
||||
"""Rebuild the ``"/slug"`` -> bundle info cache; duplicate slugs keep the first (alphabetical)."""
|
||||
global _bundles_cache, _bundles_cache_mtime
|
||||
files = _iter_bundle_files()
|
||||
out: Dict[str, Dict[str, Any]] = {}
|
||||
@@ -193,25 +125,15 @@ def scan_bundles() -> Dict[str, Dict[str, Any]]:
|
||||
|
||||
|
||||
def get_skill_bundles() -> Dict[str, Dict[str, Any]]:
|
||||
"""Return the current bundle mapping, rescanning when disk changed.
|
||||
|
||||
Cheap to call repeatedly: only rescans when the bundles directory or
|
||||
any bundle file's mtime is newer than the cached snapshot.
|
||||
"""
|
||||
files = _iter_bundle_files()
|
||||
current_mtime = _max_mtime(files)
|
||||
"""Current bundle mapping; rescans only when a bundle file or the dir mtime changed."""
|
||||
current_mtime = _max_mtime(_iter_bundle_files())
|
||||
if not _bundles_cache or _bundles_cache_mtime != current_mtime:
|
||||
scan_bundles()
|
||||
return _bundles_cache
|
||||
|
||||
|
||||
def resolve_bundle_command_key(command: str) -> Optional[str]:
|
||||
"""Resolve a user-typed command to its canonical bundle slash key.
|
||||
|
||||
Hyphens and underscores are treated interchangeably to mirror the
|
||||
skill-command behavior (Telegram converts hyphens to underscores in
|
||||
bot command names).
|
||||
"""
|
||||
"""Resolve a user-typed command to its ``/slug`` key (``_`` ≡ ``-``, as Telegram rewrites hyphens)."""
|
||||
if not command:
|
||||
return None
|
||||
cmd_key = f"/{command.replace('_', '-')}"
|
||||
@@ -219,35 +141,17 @@ def resolve_bundle_command_key(command: str) -> Optional[str]:
|
||||
|
||||
|
||||
def reload_bundles() -> Dict[str, Any]:
|
||||
"""Re-scan the bundles directory and return a diff.
|
||||
|
||||
Mirrors :func:`agent.skill_commands.reload_skills` so callers can use
|
||||
the same display logic. Returns a dict with ``added``, ``removed``,
|
||||
``unchanged``, and ``total`` keys.
|
||||
"""
|
||||
"""Re-scan and return an ``added``/``removed``/``unchanged``/``total`` diff (same shape as reload_skills)."""
|
||||
def _snapshot(cmds: Dict[str, Dict[str, Any]]) -> Dict[str, str]:
|
||||
return {k.lstrip("/"): (v or {}).get("description", "") for k, v in cmds.items()}
|
||||
|
||||
before = _snapshot(_bundles_cache)
|
||||
new = scan_bundles()
|
||||
after = _snapshot(new)
|
||||
|
||||
added_names = sorted(set(after) - set(before))
|
||||
removed_names = sorted(set(before) - set(after))
|
||||
unchanged = sorted(set(after) & set(before))
|
||||
|
||||
return {
|
||||
"added": [{"name": n, "description": after[n]} for n in added_names],
|
||||
"removed": [{"name": n, "description": before[n]} for n in removed_names],
|
||||
"unchanged": unchanged,
|
||||
"total": len(after),
|
||||
}
|
||||
return diff_command_snapshots(before, _snapshot(scan_bundles()))
|
||||
|
||||
|
||||
def list_bundles() -> List[Dict[str, Any]]:
|
||||
"""Return a sorted list of bundle info dicts for display."""
|
||||
bundles = get_skill_bundles()
|
||||
return sorted(bundles.values(), key=lambda b: b["slug"])
|
||||
return sorted(get_skill_bundles().values(), key=lambda b: b["slug"])
|
||||
|
||||
|
||||
def build_bundle_invocation_message(
|
||||
@@ -256,34 +160,20 @@ def build_bundle_invocation_message(
|
||||
task_id: str | None = None,
|
||||
platform: str | None = None,
|
||||
) -> Optional[Tuple[str, List[str], List[str]]]:
|
||||
"""Build the user message content for a bundle slash command invocation.
|
||||
"""Build the user message for a bundle invocation.
|
||||
|
||||
Returns ``(message, loaded_skill_names, missing_skill_names)`` or
|
||||
``None`` if the bundle wasn't found.
|
||||
|
||||
A bundle that references skills the user doesn't have installed still
|
||||
loads — the agent gets a note about which ones were skipped. This is
|
||||
the same forgiving stance ``build_preloaded_skills_prompt`` uses for
|
||||
``-s`` CLI preloading.
|
||||
|
||||
Disabled skills are also skipped: bundles load members via
|
||||
``_load_skill_payload`` directly, bypassing the scan-time disabled
|
||||
filter in ``get_skill_commands()``, so the disabled list must be
|
||||
re-applied here. ``platform`` scopes the check to a specific
|
||||
platform's ``skills.platform_disabled`` config (gateway dispatch
|
||||
passes it explicitly because the gateway handles multiple platforms
|
||||
in one process); when *None*, the platform resolves from session env
|
||||
vars and the global disabled list still applies. Mirrors the
|
||||
stacked-skill gate in gateway dispatch (#58888).
|
||||
Returns ``(message, loaded_skill_names, missing_skill_names)`` or ``None``
|
||||
if the bundle wasn't found. Uninstalled members are skipped with a note.
|
||||
Disabled members are skipped too: bundles load via ``_load_skill_payload``,
|
||||
bypassing the scan-time disabled filter, so the list is re-applied here.
|
||||
``platform`` scopes that check (gateway passes it; None resolves from env).
|
||||
"""
|
||||
bundles = get_skill_bundles()
|
||||
info = bundles.get(cmd_key)
|
||||
info = get_skill_bundles().get(cmd_key)
|
||||
if not info:
|
||||
return None
|
||||
|
||||
# Late import to avoid pulling tools/* at module import time and to
|
||||
# keep skill_bundles cheap to import in test environments.
|
||||
from agent.skill_commands import _load_skill_payload, _build_skill_message
|
||||
# Late import keeps skill_bundles cheap to import (no tools/* at import time).
|
||||
from agent.skill_commands import _load_skill_payload, _render_skill_block, _scaffold_header
|
||||
|
||||
try:
|
||||
from agent.skill_utils import get_disabled_skill_names
|
||||
@@ -298,10 +188,8 @@ def build_bundle_invocation_message(
|
||||
seen: set[str] = set()
|
||||
|
||||
bundle_name = info["name"]
|
||||
skills = info["skills"]
|
||||
extra_instruction = info.get("instruction") or ""
|
||||
|
||||
for skill_id in skills:
|
||||
for skill_id in info["skills"]:
|
||||
identifier = (skill_id or "").strip()
|
||||
if not identifier or identifier in seen:
|
||||
continue
|
||||
@@ -311,66 +199,36 @@ def build_bundle_invocation_message(
|
||||
if not loaded:
|
||||
missing.append(identifier)
|
||||
continue
|
||||
loaded_skill, skill_dir, skill_name = loaded
|
||||
skill_name = loaded[2]
|
||||
|
||||
# Per-platform / global disabled gate. Checked against the loaded
|
||||
# skill's canonical name (identifiers may be paths or aliases).
|
||||
# Gate on the loaded skill's canonical name (identifiers may be paths or aliases).
|
||||
if skill_name in disabled_names or identifier in disabled_names:
|
||||
disabled.append(skill_name or identifier)
|
||||
continue
|
||||
|
||||
try:
|
||||
from tools.skill_usage import bump_use
|
||||
bump_use(skill_name, task_id=task_id)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
activation_note = (
|
||||
f'[Loaded as part of the "{bundle_name}" skill bundle.]'
|
||||
)
|
||||
skill_blocks.append(
|
||||
_build_skill_message(
|
||||
loaded_skill,
|
||||
skill_dir,
|
||||
activation_note,
|
||||
session_id=task_id,
|
||||
)
|
||||
)
|
||||
skill_blocks.append(_render_skill_block(
|
||||
loaded,
|
||||
f'[Loaded as part of the "{bundle_name}" skill bundle.]',
|
||||
task_id,
|
||||
))
|
||||
loaded_names.append(skill_name)
|
||||
|
||||
if not skill_blocks:
|
||||
return None
|
||||
|
||||
# Header — tells the agent this is a bundle, lists the skills, and
|
||||
# provides any author-supplied instruction.
|
||||
header_lines = [
|
||||
f'[IMPORTANT: The user has invoked the "{bundle_name}" skill bundle, '
|
||||
f"loading {len(loaded_names)} skills together. Treat every skill below "
|
||||
"as active guidance for this turn.]",
|
||||
"",
|
||||
f"Bundle: {bundle_name}",
|
||||
f"Skills loaded: {', '.join(loaded_names)}",
|
||||
]
|
||||
if missing:
|
||||
header_lines.append(f"Skills missing (skipped): {', '.join(missing)}")
|
||||
if disabled:
|
||||
header_lines.append(
|
||||
f"Skills disabled for this platform (skipped): {', '.join(disabled)}"
|
||||
)
|
||||
if extra_instruction:
|
||||
header_lines.extend(["", f"Bundle instruction: {extra_instruction}"])
|
||||
if user_instruction:
|
||||
header_lines.extend(
|
||||
["", f"User instruction: {user_instruction}"]
|
||||
)
|
||||
|
||||
header = "\n".join(header_lines)
|
||||
header = _scaffold_header(
|
||||
f'"{bundle_name}" skill bundle',
|
||||
loaded_names,
|
||||
lead_lines=[f"Bundle: {bundle_name}"],
|
||||
missing=missing,
|
||||
disabled=disabled,
|
||||
extra_instruction=info.get("instruction") or "",
|
||||
user_instruction=user_instruction,
|
||||
)
|
||||
return ("\n\n".join([header, *skill_blocks]), loaded_names, missing)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# File-level CRUD helpers — used by `hermes bundles` CLI subcommand.
|
||||
# ---------------------------------------------------------------------------
|
||||
# ── File-level CRUD — used by `hermes bundles` ─────────────────────────────
|
||||
|
||||
|
||||
def bundle_path_for(name: str) -> Path:
|
||||
@@ -388,10 +246,10 @@ def save_bundle(
|
||||
instruction: str = "",
|
||||
overwrite: bool = False,
|
||||
) -> Path:
|
||||
"""Write a bundle to disk and invalidate the cache.
|
||||
"""Write a bundle to disk and refresh the cache.
|
||||
|
||||
Raises ``FileExistsError`` if the target exists and ``overwrite`` is
|
||||
False. Raises ``ValueError`` if the inputs are unusable.
|
||||
Raises ``FileExistsError`` if the target exists and not ``overwrite``;
|
||||
``ValueError`` for unusable inputs.
|
||||
"""
|
||||
name = (name or "").strip()
|
||||
if not name:
|
||||
@@ -415,15 +273,12 @@ def save_bundle(
|
||||
yaml.safe_dump(payload, sort_keys=False, allow_unicode=True),
|
||||
encoding="utf-8",
|
||||
)
|
||||
scan_bundles() # refresh cache
|
||||
scan_bundles()
|
||||
return path
|
||||
|
||||
|
||||
def delete_bundle(name: str) -> Path:
|
||||
"""Delete a bundle by name. Returns the deleted path.
|
||||
|
||||
Raises ``FileNotFoundError`` if the bundle doesn't exist.
|
||||
"""
|
||||
"""Delete a bundle by name and return its path; ``FileNotFoundError`` if absent."""
|
||||
path = bundle_path_for(name)
|
||||
if not path.exists():
|
||||
raise FileNotFoundError(f"No bundle at {path}")
|
||||
@@ -434,5 +289,4 @@ def delete_bundle(name: str) -> Path:
|
||||
|
||||
def get_bundle(name: str) -> Optional[Dict[str, Any]]:
|
||||
"""Look up a bundle by name (slug-normalized)."""
|
||||
slug = _slugify(name)
|
||||
return get_skill_bundles().get(f"/{slug}")
|
||||
return get_skill_bundles().get(f"/{_slugify(name)}")
|
||||
|
||||
+310
-452
File diff suppressed because it is too large
Load Diff
+244
-543
File diff suppressed because it is too large
Load Diff
+56
-137
@@ -1,16 +1,11 @@
|
||||
"""Progressive subdirectory hint discovery.
|
||||
|
||||
As the agent navigates into subdirectories via tool calls (read_file, terminal,
|
||||
search_files, etc.), this module discovers and loads project context files
|
||||
(AGENTS.md, CLAUDE.md, .cursorrules) from those directories. Discovered hints
|
||||
are appended to the tool result so the model gets relevant context at the moment
|
||||
it starts working in a new area of the codebase.
|
||||
|
||||
This complements the startup context loading in ``prompt_builder.py`` which only
|
||||
loads from the CWD. Subdirectory hints are discovered lazily and injected into
|
||||
the conversation without modifying the system prompt (preserving prompt caching).
|
||||
|
||||
Inspired by Block/goose's SubdirectoryHintTracker.
|
||||
As the agent navigates into subdirectories via tool calls, this module loads
|
||||
project context files (AGENTS.md, CLAUDE.md, .cursorrules) from those
|
||||
directories and appends them to the tool result — context arrives without
|
||||
touching the system prompt (preserving prompt caching). Complements the
|
||||
startup CWD-only loading in ``prompt_builder.py``. Inspired by goose's
|
||||
SubdirectoryHintTracker.
|
||||
"""
|
||||
|
||||
import hashlib
|
||||
@@ -24,32 +19,21 @@ from agent.prompt_builder import _scan_context_content
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Context files to look for in subdirectories, in priority order.
|
||||
# Same filenames as prompt_builder.py but we load ALL found (not first-wins)
|
||||
# since different subdirectories may use different conventions.
|
||||
# Same filenames as prompt_builder.py, in priority order (first match wins per dir).
|
||||
_HINT_FILENAMES = [
|
||||
"AGENTS.override.md",
|
||||
"AGENTS.md", "agents.md",
|
||||
"CLAUDE.md", "claude.md",
|
||||
".cursorrules",
|
||||
]
|
||||
|
||||
# Maximum chars per hint file to prevent context bloat
|
||||
_MAX_HINT_CHARS = 8_000
|
||||
|
||||
# Tool argument keys that typically contain file paths
|
||||
_PATH_ARG_KEYS = {"path", "file_path", "workdir"}
|
||||
|
||||
# Tools that take shell commands where we should extract paths
|
||||
_COMMAND_TOOLS = {"terminal"}
|
||||
|
||||
# How many parent directories to walk up when looking for hints.
|
||||
# Prevents scanning all the way to / for deeply nested paths.
|
||||
# Ancestor levels walked per path — bounds the scan for deeply nested paths.
|
||||
_MAX_ANCESTOR_WALK = 5
|
||||
|
||||
# Directory names that never contain authoritative project context.
|
||||
# Backups, vendored deps, VCS internals, and caches routinely hold *copies* of
|
||||
# AGENTS.md; loading those duplicates real context and inflates the prompt.
|
||||
# Directories that hold *copies* of context files (backups, vendored deps,
|
||||
# VCS internals, caches), never authoritative project context.
|
||||
_EXCLUDED_DIR_NAMES = frozenset({
|
||||
"node_modules", "venv", ".venv", "__pycache__",
|
||||
".git", ".hg", ".svn",
|
||||
@@ -61,7 +45,7 @@ _EXCLUDED_DIR_NAMES = frozenset({
|
||||
|
||||
|
||||
def _is_ancestor_or_same(a: Path, b: Path) -> bool:
|
||||
"""Check if *a* is the same as or an ancestor of *b* (parent directory check)."""
|
||||
"""True if *a* is *b* or one of its ancestors."""
|
||||
try:
|
||||
b.relative_to(a)
|
||||
return True
|
||||
@@ -72,34 +56,21 @@ def _is_ancestor_or_same(a: Path, b: Path) -> bool:
|
||||
class SubdirectoryHintTracker:
|
||||
"""Track which directories the agent visits and load hints on first access.
|
||||
|
||||
Usage::
|
||||
|
||||
tracker = SubdirectoryHintTracker(working_dir="/path/to/project")
|
||||
|
||||
# After each tool call:
|
||||
hints = tracker.check_tool_call("read_file", {"path": "backend/src/main.py"})
|
||||
if hints:
|
||||
tool_result += hints # append to the tool result string
|
||||
Usage: after each tool call, ``hints = tracker.check_tool_call(name, args)``
|
||||
and append the returned text to the tool result.
|
||||
"""
|
||||
|
||||
def __init__(self, working_dir: Optional[str] = None):
|
||||
self.working_dir = Path(working_dir or os.getcwd()).resolve()
|
||||
self._loaded_dirs: Set[Path] = set()
|
||||
# Content digests already injected — prevents re-sending the same file
|
||||
# reachable through symlinks, hardlinks, or duplicated copies.
|
||||
# The working dir is pre-marked loaded (startup context handles it).
|
||||
self._loaded_dirs: Set[Path] = {self.working_dir}
|
||||
# Content digests already injected: the same file reached through
|
||||
# symlinks/hardlinks/copies is never re-sent.
|
||||
self._loaded_digests: Set[str] = set()
|
||||
# Pre-mark the working dir as loaded (startup context handles it)
|
||||
self._loaded_dirs.add(self.working_dir)
|
||||
self._seed_working_dir_digest()
|
||||
|
||||
def _seed_working_dir_digest(self) -> None:
|
||||
"""Record the CWD context file's digest so it is never re-injected.
|
||||
|
||||
``prompt_builder`` already loads the working directory's context file at
|
||||
startup. Seeding its digest here means the same content reached through
|
||||
a different path (a symlink farm, a shared workspace) is recognised as a
|
||||
duplicate instead of being sent a second time.
|
||||
"""
|
||||
"""Record the CWD context file's digest (prompt_builder already loaded it)."""
|
||||
for filename in _HINT_FILENAMES:
|
||||
candidate = self.working_dir / filename
|
||||
try:
|
||||
@@ -119,23 +90,14 @@ class SubdirectoryHintTracker:
|
||||
tool_name: str,
|
||||
tool_args: Dict[str, Any],
|
||||
) -> Optional[str]:
|
||||
"""Check tool call arguments for new directories and load any hint files.
|
||||
|
||||
Returns formatted hint text to append to the tool result, or None.
|
||||
"""
|
||||
dirs = self._extract_directories(tool_name, tool_args)
|
||||
if not dirs:
|
||||
return None
|
||||
|
||||
"""Return formatted hint text for newly visited directories, or None."""
|
||||
all_hints = []
|
||||
for d in dirs:
|
||||
for d in self._extract_directories(tool_name, tool_args):
|
||||
hints = self._load_hints_for_directory(d)
|
||||
if hints:
|
||||
all_hints.append(hints)
|
||||
|
||||
if not all_hints:
|
||||
return None
|
||||
|
||||
return "\n\n" + "\n\n".join(all_hints)
|
||||
|
||||
def _extract_directories(
|
||||
@@ -143,39 +105,30 @@ class SubdirectoryHintTracker:
|
||||
) -> list:
|
||||
"""Extract directory paths from tool call arguments."""
|
||||
candidates: Set[Path] = set()
|
||||
|
||||
# Direct path arguments
|
||||
for key in _PATH_ARG_KEYS:
|
||||
val = args.get(key)
|
||||
if isinstance(val, str) and val.strip():
|
||||
self._add_path_candidate(val, candidates)
|
||||
|
||||
# Shell commands — extract path-like tokens
|
||||
if tool_name in _COMMAND_TOOLS:
|
||||
cmd = args.get("command", "")
|
||||
if isinstance(cmd, str):
|
||||
self._extract_paths_from_command(cmd, candidates)
|
||||
|
||||
return list(candidates)
|
||||
|
||||
def _add_path_candidate(self, raw_path: str, candidates: Set[Path]):
|
||||
"""Resolve a raw path and add its directory + ancestors to candidates.
|
||||
"""Add a raw path's directory and its ancestors to candidates.
|
||||
|
||||
Walks up from the resolved directory toward the filesystem root,
|
||||
stopping at the first directory already in ``_loaded_dirs`` (or after
|
||||
``_MAX_ANCESTOR_WALK`` levels). This ensures that reading
|
||||
``project/src/main.py`` discovers ``project/AGENTS.md`` even when
|
||||
``project/src/`` has no hint files of its own.
|
||||
Walks up toward the root, stopping at the first already-loaded
|
||||
directory or after ``_MAX_ANCESTOR_WALK`` levels, so reading
|
||||
``project/src/main.py`` still discovers ``project/AGENTS.md``.
|
||||
"""
|
||||
try:
|
||||
p = Path(raw_path).expanduser()
|
||||
if not p.is_absolute():
|
||||
p = self.working_dir / p
|
||||
p = p.resolve()
|
||||
# Use parent if it's a file path (has extension or doesn't exist as dir)
|
||||
if p.suffix or (p.exists() and p.is_file()):
|
||||
p = p.parent
|
||||
# Walk up ancestors — stop at already-loaded or root
|
||||
for _ in range(_MAX_ANCESTOR_WALK):
|
||||
if p in self._loaded_dirs:
|
||||
break
|
||||
@@ -189,32 +142,34 @@ class SubdirectoryHintTracker:
|
||||
pass
|
||||
|
||||
def _extract_paths_from_command(self, cmd: str, candidates: Set[Path]):
|
||||
"""Extract path-like tokens from a shell command string."""
|
||||
"""Extract path-like tokens (contain / or .; not flags or URLs) from a shell command."""
|
||||
try:
|
||||
tokens = shlex.split(cmd)
|
||||
except ValueError:
|
||||
tokens = cmd.split()
|
||||
|
||||
for token in tokens:
|
||||
# Skip flags
|
||||
if token.startswith("-"):
|
||||
continue
|
||||
# Must look like a path (contains / or .)
|
||||
if "/" not in token and "." not in token:
|
||||
continue
|
||||
# Skip URLs
|
||||
if token.startswith(("http://", "https://", "git@")):
|
||||
continue
|
||||
self._add_path_candidate(token, candidates)
|
||||
|
||||
def _is_valid_subdir(self, path: Path) -> bool:
|
||||
"""Check if path is a valid directory to scan for hints.
|
||||
def _within_working_dir(self, path: Path) -> bool:
|
||||
"""Reject paths outside the working-dir tree.
|
||||
|
||||
Only allow subdirectories within the working directory tree.
|
||||
This prevents loading AGENTS.md from outside the active workspace
|
||||
(e.g. ~/.codex/AGENTS.md, ~/.claude/CLAUDE.md), which causes
|
||||
cross-agent context contamination and instruction mixup.
|
||||
Loading ~/.codex/AGENTS.md or ~/.claude/CLAUDE.md would mix another
|
||||
agent's instructions into this session. ``is_relative_to`` handles
|
||||
symlinked paths; the ancestor check is a best-effort fallback.
|
||||
"""
|
||||
try:
|
||||
return path.is_relative_to(self.working_dir)
|
||||
except (OSError, ValueError):
|
||||
return _is_ancestor_or_same(self.working_dir, path)
|
||||
|
||||
def _is_valid_subdir(self, path: Path) -> bool:
|
||||
"""Directory inside the working-dir tree, not yet loaded, not an excluded copy dir."""
|
||||
try:
|
||||
if not path.is_dir():
|
||||
return False
|
||||
@@ -222,59 +177,31 @@ class SubdirectoryHintTracker:
|
||||
return False
|
||||
if path in self._loaded_dirs:
|
||||
return False
|
||||
# Reject paths outside the working directory tree.
|
||||
# path.resolve() may differ from working_dir.resolve() due to symlinks,
|
||||
# but path.is_relative_to(working_dir) handles both absolute and
|
||||
# symlinked paths correctly on Python 3.9+.
|
||||
try:
|
||||
if not path.is_relative_to(self.working_dir):
|
||||
return False
|
||||
except (OSError, ValueError):
|
||||
# Older Python or path resolution error — fall back to parent
|
||||
# check as a best-effort safeguard.
|
||||
if not _is_ancestor_or_same(self.working_dir, path):
|
||||
return False
|
||||
if self._is_excluded(path):
|
||||
if not self._within_working_dir(path):
|
||||
return False
|
||||
return True
|
||||
return not self._is_excluded(path)
|
||||
|
||||
def _is_excluded(self, path: Path) -> bool:
|
||||
"""True when the path sits inside a directory that holds copies, not context.
|
||||
"""True when a segment *below* the working dir is an excluded copy dir.
|
||||
|
||||
Directories the user is deliberately working inside are never excluded —
|
||||
if ``working_dir`` is itself under ``vendor/``, that segment is legitimate
|
||||
and only segments *below* the working dir are screened.
|
||||
Only segments under ``working_dir`` are screened: a user deliberately
|
||||
working inside ``vendor/`` keeps that segment legitimate.
|
||||
"""
|
||||
try:
|
||||
rel_parts = path.relative_to(self.working_dir).parts
|
||||
except ValueError:
|
||||
# Paths outside the working dir are already rejected by
|
||||
# _is_valid_subdir before this runs; treat as excluded defensively.
|
||||
return True
|
||||
return True # outside the tree — already rejected upstream
|
||||
return any(part in _EXCLUDED_DIR_NAMES for part in rel_parts)
|
||||
|
||||
def _load_hints_for_directory(self, directory: Path) -> Optional[str]:
|
||||
"""Load hint files from a directory. Returns formatted text or None.
|
||||
|
||||
Only loads hints from directories within the working directory tree.
|
||||
"""
|
||||
"""Load the first hint file in *directory*; formatted text or None."""
|
||||
self._loaded_dirs.add(directory)
|
||||
|
||||
# Reject paths outside the working directory tree.
|
||||
try:
|
||||
if not directory.is_relative_to(self.working_dir):
|
||||
logger.debug(
|
||||
"Skipping hint files in %s — outside working_dir %s",
|
||||
directory, self.working_dir,
|
||||
)
|
||||
return None
|
||||
except (OSError, ValueError):
|
||||
if not _is_ancestor_or_same(self.working_dir, directory):
|
||||
logger.debug(
|
||||
"Skipping hint files in %s — outside working_dir %s",
|
||||
directory, self.working_dir,
|
||||
)
|
||||
return None
|
||||
if not self._within_working_dir(directory):
|
||||
logger.debug(
|
||||
"Skipping hint files in %s — outside working_dir %s",
|
||||
directory, self.working_dir,
|
||||
)
|
||||
return None
|
||||
|
||||
found_hints = []
|
||||
for filename in _HINT_FILENAMES:
|
||||
@@ -288,10 +215,6 @@ class SubdirectoryHintTracker:
|
||||
content = hint_path.read_text(encoding="utf-8").strip()
|
||||
if not content:
|
||||
continue
|
||||
# Skip content we've already injected. The same AGENTS.md is
|
||||
# routinely reachable through several paths (symlinked shared
|
||||
# workspaces, hardlinks, copied backups); re-sending it burns
|
||||
# context for zero new information.
|
||||
digest = hashlib.sha256(content.encode("utf-8")).hexdigest()
|
||||
if digest in self._loaded_digests:
|
||||
logger.debug(
|
||||
@@ -301,14 +224,13 @@ class SubdirectoryHintTracker:
|
||||
)
|
||||
break
|
||||
self._loaded_digests.add(digest)
|
||||
# Same security scan as startup context loading
|
||||
# Same security scan as startup context loading.
|
||||
content = _scan_context_content(content, filename)
|
||||
if len(content) > _MAX_HINT_CHARS:
|
||||
content = (
|
||||
content[:_MAX_HINT_CHARS]
|
||||
+ f"\n\n[...truncated {filename}: {len(content):,} chars total]"
|
||||
)
|
||||
# Best-effort relative path for display
|
||||
rel_path = str(hint_path)
|
||||
try:
|
||||
rel_path = str(hint_path.relative_to(self.working_dir))
|
||||
@@ -320,20 +242,17 @@ class SubdirectoryHintTracker:
|
||||
except (ValueError, RuntimeError):
|
||||
pass # keep absolute
|
||||
found_hints.append((rel_path, content))
|
||||
# First match wins per directory (like startup loading)
|
||||
break
|
||||
break # first match wins per directory (like startup loading)
|
||||
except Exception as exc:
|
||||
logger.debug("Could not read %s: %s", hint_path, exc)
|
||||
|
||||
if not found_hints:
|
||||
return None
|
||||
|
||||
sections = []
|
||||
for rel_path, content in found_hints:
|
||||
sections.append(
|
||||
f"[Subdirectory context discovered: {rel_path}]\n{content}"
|
||||
)
|
||||
|
||||
sections = [
|
||||
f"[Subdirectory context discovered: {rel_path}]\n{content}"
|
||||
for rel_path, content in found_hints
|
||||
]
|
||||
logger.debug(
|
||||
"Loaded subdirectory hints from %s: %s",
|
||||
directory,
|
||||
|
||||
+92
-226
@@ -1,29 +1,10 @@
|
||||
"""Stateful scrubber for reasoning/thinking blocks in streamed assistant text.
|
||||
|
||||
``run_agent._strip_think_blocks`` is regex-based and correct for a complete
|
||||
string, but when it runs *per-delta* in ``_fire_stream_delta`` it destroys
|
||||
the state that downstream consumers (CLI ``_stream_delta``, gateway
|
||||
``GatewayStreamConsumer._filter_and_accumulate``) rely on.
|
||||
|
||||
Concretely, when MiniMax-M2.7 streams
|
||||
|
||||
delta1 = "<think>"
|
||||
delta2 = "Let me check their config"
|
||||
delta3 = "</think>"
|
||||
|
||||
the per-delta regex erases delta1 entirely (case 2: unterminated-open at
|
||||
boundary matches ``^<think>...``), so the downstream state machine never
|
||||
sees the open tag, treats delta2 as regular content, and leaks reasoning
|
||||
to the user. Consumers that don't run their own state machine (ACP,
|
||||
api_server, TTS) never had any defence at all — they just emitted
|
||||
whatever survived the upstream regex.
|
||||
|
||||
This module centralises the tag-suppression state machine at the
|
||||
upstream layer so every stream_delta_callback sees text that has
|
||||
already had reasoning blocks removed. Partial tags at delta
|
||||
boundaries are held back until the next delta resolves them, and
|
||||
end-of-stream flushing surfaces any held-back prose that turned out
|
||||
not to be a real tag.
|
||||
The regex ``run_agent._strip_think_blocks`` is correct for a complete string but,
|
||||
run per-delta, erases an opening ``<think>`` that arrives alone in one delta, so
|
||||
downstream state machines never see the open tag and leak reasoning. This class
|
||||
centralises tag suppression upstream: partial tags at delta boundaries are held
|
||||
back until resolved, and ``flush()`` releases held-back prose that was not a tag.
|
||||
|
||||
Usage::
|
||||
|
||||
@@ -33,25 +14,15 @@ Usage::
|
||||
if visible:
|
||||
emit(visible)
|
||||
tail = scrubber.flush() # at end of stream
|
||||
if tail:
|
||||
emit(tail)
|
||||
|
||||
The scrubber is re-entrant per agent instance. Call ``reset()`` at
|
||||
the top of each new turn so a hung block from an interrupted prior
|
||||
stream cannot taint the next turn's output.
|
||||
Call ``reset()`` at the top of each turn so an interrupted block cannot taint
|
||||
the next turn. Tags handled (case-insensitive): ``<think>``, ``<thinking>``,
|
||||
``<reasoning>``, ``<thought>``, ``<REASONING_SCRATCHPAD>``.
|
||||
|
||||
Tag variants handled (case-insensitive):
|
||||
``<think>``, ``<thinking>``, ``<reasoning>``, ``<thought>``,
|
||||
``<REASONING_SCRATCHPAD>``.
|
||||
|
||||
Block-boundary rule for opens: an opening tag is only treated as a
|
||||
reasoning-block opener when it appears at the start of the stream,
|
||||
after a newline (optionally followed by whitespace), or when only
|
||||
whitespace has been emitted on the current line. This prevents prose
|
||||
that *mentions* the tag name (e.g. ``"use <think> tags here"``) from
|
||||
being incorrectly suppressed. Closed pairs (``<think>X</think>``) are
|
||||
always suppressed regardless of boundary; a closed pair is an
|
||||
intentional, bounded construct.
|
||||
Boundary rule: an opening tag only starts a block at a block boundary (stream
|
||||
start, after a newline, or with only whitespace emitted on the current line), so
|
||||
prose that *mentions* ``<think>`` is not suppressed. Closed pairs
|
||||
(``<think>X</think>``) are always suppressed — a closed pair is intentional.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -64,16 +35,10 @@ __all__ = ["StreamingThinkScrubber"]
|
||||
class StreamingThinkScrubber:
|
||||
"""Stateful scrubber for streaming reasoning/thinking blocks.
|
||||
|
||||
State machine:
|
||||
- ``_in_block``: True while inside an opened block, waiting for
|
||||
a close tag. All text inside is discarded.
|
||||
- ``_buf``: held-back partial-tag tail. Emitted / discarded on
|
||||
the next ``feed()`` call or by ``flush()``.
|
||||
- ``_last_emitted_ended_newline``: True iff the most recent
|
||||
emission to the consumer ended with ``\\n``, or nothing has
|
||||
been emitted yet (start-of-stream counts as a boundary). Used
|
||||
to decide whether an open tag at buffer position 0 is at a
|
||||
block boundary.
|
||||
State: ``_in_block`` (inside an open block; text discarded), ``_buf``
|
||||
(held-back partial-tag tail), ``_last_emitted_ended_newline`` (True iff the
|
||||
last emission ended with ``\\n`` or nothing has been emitted yet — decides
|
||||
whether an open tag at buffer position 0 sits at a block boundary).
|
||||
"""
|
||||
|
||||
_OPEN_TAG_NAMES: Tuple[str, ...] = (
|
||||
@@ -84,31 +49,33 @@ class StreamingThinkScrubber:
|
||||
"REASONING_SCRATCHPAD",
|
||||
)
|
||||
|
||||
# Materialise literal tag strings so the hot path does string
|
||||
# operations, not regex compilation per feed().
|
||||
# Literal tag strings so the hot path does string ops, not regex per feed().
|
||||
_OPEN_TAGS: Tuple[str, ...] = tuple(f"<{name}>" for name in _OPEN_TAG_NAMES)
|
||||
_CLOSE_TAGS: Tuple[str, ...] = tuple(f"</{name}>" for name in _OPEN_TAG_NAMES)
|
||||
|
||||
# Pre-compute the longest tag (for partial-tag hold-back bound).
|
||||
_MAX_TAG_LEN: int = max(len(tag) for tag in _OPEN_TAGS + _CLOSE_TAGS)
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.reset()
|
||||
|
||||
def reset(self) -> None:
|
||||
"""Reset all state. Call at the top of every new turn."""
|
||||
self._in_block: bool = False
|
||||
self._buf: str = ""
|
||||
self._last_emitted_ended_newline: bool = True
|
||||
|
||||
def reset(self) -> None:
|
||||
"""Reset all state. Call at the top of every new turn."""
|
||||
self._in_block = False
|
||||
self._buf = ""
|
||||
self._last_emitted_ended_newline = True
|
||||
def _emit(self, out: list[str], text: str) -> None:
|
||||
"""Append visible prose to *out* (orphan close tags stripped) and track the newline flag."""
|
||||
if text:
|
||||
text = self._strip_orphan_close_tags(text)
|
||||
if text:
|
||||
out.append(text)
|
||||
self._last_emitted_ended_newline = text.endswith("\n")
|
||||
|
||||
def feed(self, text: str) -> str:
|
||||
"""Feed one delta; return the scrubbed visible portion.
|
||||
|
||||
May return an empty string when the entire delta is reasoning
|
||||
content or is being held back pending resolution of a partial
|
||||
tag at the boundary.
|
||||
Returns "" when the whole delta is reasoning content or is held back
|
||||
pending resolution of a partial tag at the boundary.
|
||||
"""
|
||||
if not text:
|
||||
return ""
|
||||
@@ -118,130 +85,67 @@ class StreamingThinkScrubber:
|
||||
|
||||
while buf:
|
||||
if self._in_block:
|
||||
# Hunt for the earliest close tag.
|
||||
close_idx, close_len = self._find_first_tag(
|
||||
buf, self._CLOSE_TAGS,
|
||||
)
|
||||
close_idx, close_len = self._find_first_tag(buf, self._CLOSE_TAGS)
|
||||
if close_idx == -1:
|
||||
# No close yet — hold back a potential partial
|
||||
# close-tag prefix; discard everything else.
|
||||
# No close yet: hold back a possible partial close-tag prefix, drop the rest.
|
||||
held = self._max_partial_suffix(buf, self._CLOSE_TAGS)
|
||||
self._buf = buf[-held:] if held else ""
|
||||
return "".join(out)
|
||||
# Found close: discard block content + tag, continue.
|
||||
buf = buf[close_idx + close_len:]
|
||||
self._in_block = False
|
||||
continue
|
||||
|
||||
# Priority 1: closed <tag>X</tag> pair anywhere (no boundary gating —
|
||||
# even inline pairs are almost certainly leaked reasoning).
|
||||
# Priority 2: unterminated open tag at a block boundary (gated so
|
||||
# prose that mentions '<think>' isn't over-stripped). Earliest wins.
|
||||
pair = self._find_earliest_closed_pair(buf)
|
||||
open_idx, open_len = self._find_open_at_boundary(buf, out)
|
||||
if pair is not None and (open_idx == -1 or pair[0] <= open_idx):
|
||||
self._emit(out, buf[:pair[0]])
|
||||
buf = buf[pair[1]:]
|
||||
continue
|
||||
if open_idx != -1:
|
||||
self._emit(out, buf[:open_idx])
|
||||
self._in_block = True
|
||||
buf = buf[open_idx + open_len:]
|
||||
continue
|
||||
|
||||
# No resolvable tag: hold back any partial-tag prefix at the tail
|
||||
# so a tag split across deltas isn't missed, then emit the rest.
|
||||
held = max(
|
||||
self._max_partial_suffix(buf, self._OPEN_TAGS),
|
||||
self._max_partial_suffix(buf, self._CLOSE_TAGS),
|
||||
)
|
||||
if held:
|
||||
self._emit(out, buf[:-held])
|
||||
self._buf = buf[-held:]
|
||||
else:
|
||||
# Priority 1 — closed <tag>X</tag> pair anywhere in
|
||||
# buf. Closed pairs are always an intentional,
|
||||
# bounded construct (even mid-line prose containing
|
||||
# an open/close pair is almost certainly a model
|
||||
# leaking reasoning inline), so no boundary gating.
|
||||
pair = self._find_earliest_closed_pair(buf)
|
||||
# Priority 2 — unterminated open tag at a block
|
||||
# boundary. Boundary-gated so prose that mentions
|
||||
# '<think>' isn't over-stripped.
|
||||
open_idx, open_len = self._find_open_at_boundary(
|
||||
buf, out,
|
||||
)
|
||||
|
||||
# Pick whichever match comes earliest in the buffer.
|
||||
if pair is not None and (
|
||||
open_idx == -1 or pair[0] <= open_idx
|
||||
):
|
||||
start_idx, end_idx = pair
|
||||
preceding = buf[:start_idx]
|
||||
if preceding:
|
||||
preceding = self._strip_orphan_close_tags(preceding)
|
||||
if preceding:
|
||||
out.append(preceding)
|
||||
self._last_emitted_ended_newline = (
|
||||
preceding.endswith("\n")
|
||||
)
|
||||
buf = buf[end_idx:]
|
||||
continue
|
||||
|
||||
if open_idx != -1:
|
||||
# Unterminated open at boundary — emit preceding,
|
||||
# enter block, continue loop with remainder.
|
||||
preceding = buf[:open_idx]
|
||||
if preceding:
|
||||
preceding = self._strip_orphan_close_tags(preceding)
|
||||
if preceding:
|
||||
out.append(preceding)
|
||||
self._last_emitted_ended_newline = (
|
||||
preceding.endswith("\n")
|
||||
)
|
||||
self._in_block = True
|
||||
buf = buf[open_idx + open_len:]
|
||||
continue
|
||||
|
||||
# No resolvable tag structure in buf. Hold back any
|
||||
# partial-tag prefix at the tail so a split tag
|
||||
# across deltas isn't missed, then emit the rest.
|
||||
held = self._max_partial_suffix(buf, self._OPEN_TAGS)
|
||||
held_close = self._max_partial_suffix(
|
||||
buf, self._CLOSE_TAGS,
|
||||
)
|
||||
held = max(held, held_close)
|
||||
if held:
|
||||
emit_text = buf[:-held]
|
||||
self._buf = buf[-held:]
|
||||
else:
|
||||
emit_text = buf
|
||||
self._buf = ""
|
||||
if emit_text:
|
||||
emit_text = self._strip_orphan_close_tags(emit_text)
|
||||
if emit_text:
|
||||
out.append(emit_text)
|
||||
self._last_emitted_ended_newline = (
|
||||
emit_text.endswith("\n")
|
||||
)
|
||||
return "".join(out)
|
||||
self._emit(out, buf)
|
||||
return "".join(out)
|
||||
|
||||
return "".join(out)
|
||||
|
||||
def flush(self) -> str:
|
||||
"""End-of-stream flush.
|
||||
|
||||
If still inside an unterminated block, held-back content is
|
||||
discarded — leaking partial reasoning is worse than a
|
||||
truncated answer. Otherwise the held-back partial-tag tail is
|
||||
emitted verbatim (it turned out not to be a real tag prefix).
|
||||
|
||||
Always treats the next ``feed()`` as a fresh stream boundary.
|
||||
Intra-turn retries (thinking-only prefill, empty-response
|
||||
retry) flush then stream again without calling ``reset()``;
|
||||
leaving ``_last_emitted_ended_newline`` False made a new
|
||||
stream's opening ``<think>`` look mid-line and leak into the
|
||||
visible reply.
|
||||
Inside an unterminated block the held-back content is discarded (leaking
|
||||
partial reasoning is worse than a truncated answer); otherwise the
|
||||
held-back tail is emitted verbatim. Always resets the boundary flag:
|
||||
intra-turn retries flush then stream again without ``reset()``, and a
|
||||
stale False flag made the new stream's opening ``<think>`` look mid-line.
|
||||
"""
|
||||
if self._in_block:
|
||||
self._buf = ""
|
||||
self._in_block = False
|
||||
# Next feed() is a new stream — start-of-stream is a boundary.
|
||||
self._last_emitted_ended_newline = True
|
||||
return ""
|
||||
tail = self._buf
|
||||
tail = "" if self._in_block else self._buf
|
||||
self._buf = ""
|
||||
# Same for the non-block path: do NOT derive the boundary flag
|
||||
# from the flushed tail (e.g. a held-back '<'). End-of-stream
|
||||
# means the next feed() starts a new model response.
|
||||
self._in_block = False
|
||||
self._last_emitted_ended_newline = True
|
||||
if not tail:
|
||||
return ""
|
||||
return self._strip_orphan_close_tags(tail)
|
||||
return self._strip_orphan_close_tags(tail) if tail else ""
|
||||
|
||||
# ── internal helpers ───────────────────────────────────────────────
|
||||
|
||||
@staticmethod
|
||||
def _find_first_tag(
|
||||
buf: str, tags: Tuple[str, ...],
|
||||
) -> Tuple[int, int]:
|
||||
"""Return (earliest_index, tag_length) over *tags*, or (-1, 0).
|
||||
|
||||
Case-insensitive match.
|
||||
"""
|
||||
def _find_first_tag(buf: str, tags: Tuple[str, ...]) -> Tuple[int, int]:
|
||||
"""Return (earliest_index, tag_length) over *tags* (case-insensitive), or (-1, 0)."""
|
||||
buf_lower = buf.lower()
|
||||
best_idx = -1
|
||||
best_len = 0
|
||||
@@ -253,14 +157,10 @@ class StreamingThinkScrubber:
|
||||
return best_idx, best_len
|
||||
|
||||
def _find_earliest_closed_pair(self, buf: str):
|
||||
"""Return (start_idx, end_idx) of the earliest closed pair, else None.
|
||||
"""Return (start_idx, end_idx) of the earliest ``<tag>...</tag>`` pair, else None.
|
||||
|
||||
A closed pair is ``<tag>...</tag>`` of any variant. Matches are
|
||||
case-insensitive and non-greedy (the closest close tag after
|
||||
an open tag wins), matching the regex ``<tag>.*?</tag>``
|
||||
semantics of ``_strip_think_blocks`` case 1. When two tag
|
||||
variants could both match, the one whose open tag appears
|
||||
earlier wins.
|
||||
Case-insensitive and non-greedy (closest close after the open wins),
|
||||
matching ``_strip_think_blocks`` case 1; the earliest open tag wins.
|
||||
"""
|
||||
buf_lower = buf.lower()
|
||||
best: "tuple[int, int] | None" = None
|
||||
@@ -270,23 +170,15 @@ class StreamingThinkScrubber:
|
||||
open_idx = buf_lower.find(open_lower)
|
||||
if open_idx == -1:
|
||||
continue
|
||||
close_idx = buf_lower.find(
|
||||
close_lower, open_idx + len(open_lower),
|
||||
)
|
||||
close_idx = buf_lower.find(close_lower, open_idx + len(open_lower))
|
||||
if close_idx == -1:
|
||||
continue
|
||||
end_idx = close_idx + len(close_lower)
|
||||
if best is None or open_idx < best[0]:
|
||||
best = (open_idx, end_idx)
|
||||
best = (open_idx, close_idx + len(close_lower))
|
||||
return best
|
||||
|
||||
def _find_open_at_boundary(
|
||||
self, buf: str, already_emitted: list[str],
|
||||
) -> Tuple[int, int]:
|
||||
"""Return the earliest block-boundary open-tag (idx, len).
|
||||
|
||||
Returns (-1, 0) if no boundary-legal opener is present.
|
||||
"""
|
||||
def _find_open_at_boundary(self, buf: str, already_emitted: list[str]) -> Tuple[int, int]:
|
||||
"""Return the earliest block-boundary open-tag (idx, len), or (-1, 0)."""
|
||||
buf_lower = buf.lower()
|
||||
best_idx = -1
|
||||
best_len = 0
|
||||
@@ -305,50 +197,30 @@ class StreamingThinkScrubber:
|
||||
search_start = idx + 1
|
||||
return best_idx, best_len
|
||||
|
||||
def _is_block_boundary(
|
||||
self, buf: str, idx: int, already_emitted: list[str],
|
||||
) -> bool:
|
||||
def _is_block_boundary(self, buf: str, idx: int, already_emitted: list[str]) -> bool:
|
||||
"""True iff position *idx* in *buf* is a block boundary.
|
||||
|
||||
A block boundary is:
|
||||
- buf position 0 AND the most recent emission ended with
|
||||
a newline (or nothing has been emitted yet)
|
||||
- any position whose preceding text on the current line
|
||||
(since the last newline in buf) is whitespace-only, AND
|
||||
if there is no newline in the preceding buf portion, the
|
||||
most recent prior emission ended with a newline
|
||||
Boundary = position 0 with the prior emission ending in a newline (or
|
||||
nothing emitted yet), or any position whose preceding text on the current
|
||||
line is whitespace-only (when no newline precedes it in *buf*, the prior
|
||||
emission must also have ended with a newline).
|
||||
"""
|
||||
prior_newline = (
|
||||
already_emitted[-1].endswith("\n") if already_emitted else self._last_emitted_ended_newline
|
||||
)
|
||||
if idx == 0:
|
||||
# Check whether the last already-emitted chunk in THIS
|
||||
# feed() call ended with a newline, otherwise fall back
|
||||
# to the cross-feed flag.
|
||||
if already_emitted:
|
||||
return already_emitted[-1].endswith("\n")
|
||||
return self._last_emitted_ended_newline
|
||||
return prior_newline
|
||||
preceding = buf[:idx]
|
||||
last_nl = preceding.rfind("\n")
|
||||
if last_nl == -1:
|
||||
# No newline in buf before the tag — boundary only if the
|
||||
# prior emission ended with a newline AND everything since
|
||||
# is whitespace.
|
||||
if already_emitted:
|
||||
prior_newline = already_emitted[-1].endswith("\n")
|
||||
else:
|
||||
prior_newline = self._last_emitted_ended_newline
|
||||
return prior_newline and preceding.strip() == ""
|
||||
# Newline present — text between it and the tag must be
|
||||
# whitespace-only.
|
||||
return preceding[last_nl + 1:].strip() == ""
|
||||
|
||||
@classmethod
|
||||
def _max_partial_suffix(
|
||||
cls, buf: str, tags: Tuple[str, ...],
|
||||
) -> int:
|
||||
"""Return the longest buf-suffix that is a prefix of any tag.
|
||||
def _max_partial_suffix(cls, buf: str, tags: Tuple[str, ...]) -> int:
|
||||
"""Longest buf-suffix that is a strict prefix of any tag (case-insensitive).
|
||||
|
||||
Only prefixes strictly shorter than the tag itself count
|
||||
(full-length suffixes are the tag and are handled as matches,
|
||||
not held-back partials). Case-insensitive.
|
||||
Full-length matches are real tags handled elsewhere, not held-back partials.
|
||||
"""
|
||||
if not buf:
|
||||
return 0
|
||||
@@ -364,12 +236,7 @@ class StreamingThinkScrubber:
|
||||
|
||||
@classmethod
|
||||
def _strip_orphan_close_tags(cls, text: str) -> str:
|
||||
"""Remove any close tags from *text* (orphan-close handling).
|
||||
|
||||
An orphan close tag has no matching open in the current
|
||||
scrubber state; it's always noise, stripped with any trailing
|
||||
whitespace so the surrounding prose flows naturally.
|
||||
"""
|
||||
"""Remove close tags with no matching open (always noise) plus trailing whitespace."""
|
||||
if "</" not in text:
|
||||
return text
|
||||
text_lower = text.lower()
|
||||
@@ -382,8 +249,7 @@ class StreamingThinkScrubber:
|
||||
tag_lower = tag.lower()
|
||||
tag_len = len(tag_lower)
|
||||
if text_lower[i:i + tag_len] == tag_lower:
|
||||
# Skip the tag and any trailing whitespace,
|
||||
# matching _strip_think_blocks case 3.
|
||||
# Skip the tag and trailing whitespace (matches _strip_think_blocks case 3).
|
||||
j = i + tag_len
|
||||
while j < len(text) and text[j] in " \t\n\r":
|
||||
j += 1
|
||||
|
||||
@@ -111,16 +111,6 @@ class TestTrustGate:
|
||||
|
||||
|
||||
class TestPrecedence:
|
||||
def test_scan_order_project_first(self, project_env):
|
||||
_trust(project_env["config"], project_env["repo"])
|
||||
order = su.get_scan_ordered_skills_dirs()
|
||||
proj_dirs = {
|
||||
(project_env["repo"] / ".hermes" / "skills").resolve(),
|
||||
(project_env["repo"] / ".agents" / "skills").resolve(),
|
||||
}
|
||||
assert set(order[:2]) == proj_dirs
|
||||
assert order[2] == su.get_skills_dir()
|
||||
|
||||
def test_project_paths_are_readonly_owned(self, project_env):
|
||||
_trust(project_env["config"], project_env["repo"])
|
||||
p = project_env["repo"] / ".hermes" / "skills" / "repo-skill" / "SKILL.md"
|
||||
@@ -177,9 +167,9 @@ class TestQuarantine:
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_quarantine_cache(self):
|
||||
su._project_quarantine_cache_clear()
|
||||
su._PROJECT_QUARANTINE_CACHE.clear()
|
||||
yield
|
||||
su._project_quarantine_cache_clear()
|
||||
su._PROJECT_QUARANTINE_CACHE.clear()
|
||||
|
||||
def _add_malicious_skill(self, repo: Path) -> Path:
|
||||
d = repo / ".hermes" / "skills" / "evil-skill"
|
||||
@@ -231,7 +221,7 @@ class TestQuarantine:
|
||||
(evil_dir / "SKILL.md").write_text(
|
||||
"---\nname: evil-skill\ndescription: now actually benign\n---\nbody\n"
|
||||
)
|
||||
su._project_quarantine_cache_clear()
|
||||
su._PROJECT_QUARANTINE_CACHE.clear()
|
||||
assert su.is_quarantined_project_skill(evil_dir / "SKILL.md") is False
|
||||
|
||||
def test_scan_cache_outside_repo(self, project_env):
|
||||
|
||||
@@ -17,8 +17,8 @@ from unittest.mock import patch
|
||||
import agent.skill_bundles as skill_bundles
|
||||
import agent.skill_commands as skill_commands
|
||||
import tools.skills_tool as skills_tool
|
||||
import agent.prompt_cache_boundary as prompt_cache_boundary
|
||||
from agent.prompt_cache_boundary import (
|
||||
clear_stable_prefixes,
|
||||
find_stable_prefix,
|
||||
register_stable_prefix,
|
||||
)
|
||||
@@ -36,9 +36,11 @@ SKILL_BODY = "Inspect the report carefully and preserve the stable instructions.
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _isolated_registry():
|
||||
clear_stable_prefixes()
|
||||
with prompt_cache_boundary._lock:
|
||||
prompt_cache_boundary._prefixes.clear()
|
||||
yield
|
||||
clear_stable_prefixes()
|
||||
with prompt_cache_boundary._lock:
|
||||
prompt_cache_boundary._prefixes.clear()
|
||||
|
||||
|
||||
def _write_skill(skills_dir, name, body=SKILL_BODY):
|
||||
|
||||
@@ -88,7 +88,7 @@ def test_t20880_tool_heavy_native_loop_reproduction():
|
||||
|
||||
assert final_tool_marked
|
||||
assert shared_transaction_endpoint
|
||||
assert after_exchange.marker_count <= 4
|
||||
assert _count_cache_markers(after_exchange.messages, after_exchange.tools) <= 4
|
||||
|
||||
|
||||
class TestPromptCachePlan:
|
||||
@@ -114,7 +114,7 @@ class TestPromptCachePlan:
|
||||
assert plan.tools is not tools
|
||||
assert "cache_control" not in tools[-1]
|
||||
assert plan.tools[-1]["cache_control"] == MARKER
|
||||
assert plan.marker_count == 4
|
||||
assert _count_cache_markers(plan.messages, plan.tools) == 4
|
||||
|
||||
def test_unmarkable_endpoint_does_not_consume_a_slot(self):
|
||||
messages = [
|
||||
@@ -129,7 +129,7 @@ class TestPromptCachePlan:
|
||||
direct_native_tool_cache=True,
|
||||
)
|
||||
|
||||
assert plan.marker_count == 2
|
||||
assert _count_cache_markers(plan.messages, plan.tools) == 2
|
||||
assert "cache_control" not in plan.messages[-1]
|
||||
|
||||
def test_static_prefix_equal_to_whole_prompt_emits_no_empty_block(self):
|
||||
@@ -184,7 +184,7 @@ class TestPromptCachePlan:
|
||||
native_anthropic=True,
|
||||
direct_native_tool_cache=True,
|
||||
)
|
||||
assert plan.marker_count == 3
|
||||
assert _count_cache_markers(plan.messages, plan.tools) == 3
|
||||
assert len(plan.tools) == 0
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user