refactor(agent/prompt): remove dead code, unify duplicated helpers, compact docstrings across prompt/skill/redaction modules

Dead (zero refs): coding_system_blocks, get_friendly_tool_labels, get_scan_ordered_skills_dirs,
_project_quarantine_cache_clear, clear_stable_prefixes, _redact_http_request_target_query_params,
_has_http_method_substring, PromptCachePlan.marker_count, display _diff_* colour thunks (-> _diff_ansi),
pass-through RedactingFormatter.__init__.
Unified: _slugify -> slugify_skill_name; reload diff -> diff_command_snapshots; _is_summary_item ->
is_compaction_summary_message alias; sanitizer walkers -> _sanitize_messages/_sanitize_structure;
assignment redaction passes -> _redact_assignments/_should_redact_assignment; quiet-mode tool lines -> _CUTE_LINES table.
This commit is contained in:
Teknium
2026-09-02 13:28:22 -07:00
parent a3d33fe22f
commit be5c6a2fd8
23 changed files with 2570 additions and 5225 deletions
+138 -312
View File
@@ -1,52 +1,29 @@
"""Coding-context awareness — base Hermes, every interactive surface.
When the user runs Hermes inside a code workspace (CLI, TUI, desktop app, or an
editor over ACP), Hermes shifts into a **coding posture**. This module is the
single place that decides whether we're in that posture and what it implies,
so the rest of the codebase never re-derives "are we coding?" on its own.
When Hermes runs inside a code workspace (CLI, TUI, desktop, ACP editor) it
shifts into a **coding posture**. This module is the single place that decides
whether we're in that posture and what it implies, so nothing else re-derives
"are we coding?". The posture is a frozen :class:`RuntimeMode` selected from a
small :class:`ContextProfile` registry (``coding`` / ``general``); a profile is
*data* (toolset, operating brief, skill-index hints) that every domain reads:
Architecture — one seam, many consumers
----------------------------------------
The posture is modelled as a frozen :class:`RuntimeMode` selected from a small
:class:`ContextProfile` registry (today: ``coding`` and ``general``). A profile
is *data* — it declares the toolset to collapse to, the operating brief to
inject, and hints for other domains (model routing, memory, subagents). Every
domain reads the same resolved object instead of probing git/config itself:
* System prompt — ``RuntimeMode.system_prompt_parts()`` → operating brief +
live git/workspace snapshot (``agent/system_prompt.py``).
* Toolset — ``RuntimeMode.toolset_selection()`` → ``coding`` toolset + enabled
MCP servers, ONLY under the opt-in ``focus`` mode. The default posture is
prompt-only and never strips a toolset the user explicitly enabled.
* Delegation — subagents inherit the toolset and prompt builder, so the
posture propagates for free.
* **System prompt** — ``RuntimeMode.system_blocks()`` → the operating brief +
a live git/workspace snapshot (``agent/system_prompt.py``).
* **Toolset** — ``RuntimeMode.toolset_selection()`` → the ``coding`` toolset
plus the user's enabled MCP servers (``cli.py`` / ``tui_gateway``). Only
under the opt-in ``focus`` mode: the default posture is prompt-only and
never touches the user's configured toolsets (toolsets like messaging /
smart-home / music are off-by-default anyway, and someone who explicitly
enabled image-gen or Spotify shouldn't lose it for being in a git repo).
* **Delegation** — subagents inherit the parent's toolset and run through the
same prompt builder, so the coding posture propagates to children for free.
* **Model / memory / compression** — declared on the profile
(``model_hint``, ``memory_policy``) as the extension seam; consumers read
``mode.profile`` rather than re-deciding.
Cache safety: the mode is resolved once and immutable; the workspace snapshot
is built once at prompt-build time and never re-probed per turn (the brief
tells the model to re-check with ``git``). A ``/coding`` flip takes effect next
session.
Cache safety
------------
The mode is resolved **once** and is immutable. The workspace snapshot is built
once at prompt-build time and baked into the *stable* system-prompt tier — never
re-probed per turn (that would shatter the prompt cache). Branch and dirty state
drift mid-session, so the brief tells the model to re-check with ``git`` before
acting on the snapshot. A ``/coding`` flip therefore only takes effect next
session (deferred), the same contract as ``/skills install`` vs ``--now``.
Activation (config ``agent.coding_context``):
* ``auto`` (default) — posture (brief + snapshot) on an interactive coding
surface sitting in a code workspace (git repo or recognised project root).
Prompt-only; toolsets and the skill index untouched.
* ``focus`` — like ``auto``, but additionally collapses the toolset to the
``coding`` set + enabled MCP servers and demotes non-coding skill
categories to names-only in the prompt's skill index (no skill is ever
hidden). Explicit opt-in for a lean schema.
* ``on`` — force the posture anywhere (incl. non-workspaces). Prompt-only.
* ``off`` — disable entirely.
Activation (config ``agent.coding_context``): ``auto`` (default) — posture on an
interactive surface in a code workspace, prompt-only; ``focus`` — also collapse
the toolset and demote non-coding skill categories to names-only (never
hidden); ``on`` — force the posture anywhere; ``off`` — disable.
"""
from __future__ import annotations
@@ -67,8 +44,7 @@ logger = logging.getLogger("hermes.coding_context")
CODING_TOOLSET = "coding"
# Surfaces where a coding posture makes sense under ``auto``. Messaging
# platforms (telegram, discord, slack, …) are intentionally absent — a chat bot
# in a group is not pair-programming.
# platforms are intentionally absent — a chat bot in a group is not pairing.
INTERACTIVE_CODING_PLATFORMS = {"cli", "tui", "acp", "desktop", ""}
# Project-root signals that mark a directory as a code workspace even when it
@@ -85,10 +61,8 @@ _PROJECT_MARKERS = (
# Agent-instruction files surfaced separately from manifests in the snapshot.
_CONTEXT_FILES = ("AGENTS.md", "CLAUDE.md", ".cursorrules")
# Source-file extensions that make a git repo a *code* workspace even with no
# manifest. Without this, `git init` on a notes/writing/research folder (a huge
# non-coding use case) would flip the whole session into the coding posture just
# for having a `.git`. A manifest still wins on its own (see `_PROJECT_MARKERS`).
# Source extensions that make a manifest-less git repo a *code* workspace, so
# `git init` on a notes/writing folder does not flip the session into coding.
_CODE_EXTENSIONS = frozenset({
".py", ".pyi", ".ipynb", ".js", ".jsx", ".ts", ".tsx", ".mjs", ".cjs",
".go", ".rs", ".java", ".kt", ".kts", ".scala", ".rb", ".php", ".c", ".h",
@@ -97,24 +71,16 @@ _CODE_EXTENSIONS = frozenset({
".hs", ".clj", ".erl", ".pl",
})
# Dirs never worth scanning for the code check (deps/build/vcs/venv noise).
_CODE_SCAN_SKIP_DIRS = frozenset({
".git", "node_modules", "venv", ".venv", "__pycache__", "dist", "build",
"target", ".next", ".turbo", "vendor",
})
# Bounded sweep: a code workspace reveals itself in the first handful of entries.
_CODE_SCAN_MAX_ENTRIES = 500
def _has_code_files(root: Path) -> bool:
"""Cheap, bounded check for source files in a repo's top two levels.
Lets a git repo of loose scripts (no manifest) still read as a code
workspace while a bare notes/writing repo does not. Scans the root and its
immediate subdirectories only, capped at ``_CODE_SCAN_MAX_ENTRIES`` stats —
a handful of readdirs at session start, not a full walk.
"""
"""Bounded check for source files in the root and its immediate subdirs."""
seen = 0
stack = [(root, True)]
while stack:
@@ -138,6 +104,7 @@ def _has_code_files(root: Path) -> bool:
continue
return False
# Lockfile → package manager, checked in priority order.
_PY_LOCKFILES = (("uv.lock", "uv"), ("poetry.lock", "poetry"), ("Pipfile.lock", "pipenv"))
_JS_LOCKFILES = (
@@ -153,21 +120,13 @@ _MAX_FACT_FILE_BYTES = 256 * 1024
_GIT_TIMEOUT = 2.5
# Per-model edit-format steering. Matching the edit tool format to how a model
# was trained reduces mistakes and wasted reasoning (OpenAI/Codex handle
# patch-style diffs best; Anthropic models — and most open-weight coding
# models, whose RL scaffolds use str_replace-style editors — do best with
# string-replacement). Our `patch` tool exposes both: mode="patch" (V4A
# multi-file) and mode="replace" (find-and-swap). We nudge each family toward
# its native format. Unknown families get nothing (the brief's neutral wording
# stands). Substrings match the model id; aligned with TOOL_USE_ENFORCEMENT_MODELS.
#
# GPT/Codex get V4A for ALL edits, single-file included: in codex-rs,
# apply_patch (V4A — apply_patch.lark) is the ONLY file editor, no
# str_replace-style tool exists, and the shipped model prompts say to use
# apply_patch even "for single file edits" — so a replace-mode nudge would
# steer those models toward a format their first-party harness never taught
# them.
# Per-model edit-format steering: nudge each family toward the `patch` mode it
# was trained on (unknown families get nothing). GPT/Codex get V4A for ALL
# edits incl. single-file — codex-rs ships apply_patch as its ONLY editor and
# its prompts say to use it even for single files, so a replace-mode nudge
# would steer them toward a format their first-party harness never taught.
# Anthropic and most open-weight coding models were RL'd on str_replace-style
# editors. Substrings match the model id; aligned with TOOL_USE_ENFORCEMENT_MODELS.
_EDIT_FORMAT_GUIDANCE: dict[str, tuple[tuple[str, ...], str]] = {
"patch": (
("gpt", "codex"),
@@ -188,12 +147,7 @@ _EDIT_FORMAT_GUIDANCE: dict[str, tuple[tuple[str, ...], str]] = {
def _model_family(model: Optional[str]) -> Optional[str]:
"""Classify a model id into an edit-format family key, or ``None``.
Used to steer the coding posture toward the edit tool format a model was
trained on. Family-agnostic by design: an unrecognised model gets ``None``
and the operating brief's neutral edit wording applies.
"""
"""Edit-format family key for a model id, or ``None`` (neutral wording applies)."""
if not model:
return None
lowered = model.lower()
@@ -206,14 +160,11 @@ def _model_family(model: Optional[str]) -> Optional[str]:
def _edit_format_line(model: Optional[str]) -> str:
"""The edit-format guidance line for this model's family (``""`` if none)."""
family = _model_family(model)
if family is None:
return ""
return _EDIT_FORMAT_GUIDANCE[family][1]
return "" if family is None else _EDIT_FORMAT_GUIDANCE[family][1]
# Operating brief for the coding posture. Tool names referenced here (read_file,
# search_files, patch, write_file, terminal, todo) are in the coding toolset and
# in _HERMES_CORE_TOOLS, so they're present on every surface this fires on.
# Operating brief for the coding posture. Tool names referenced here are in the
# coding toolset and in _HERMES_CORE_TOOLS, so they exist on every surface this fires on.
CODING_AGENT_GUIDANCE = (
"You are a coding agent pairing with the user inside their codebase. "
"Operate like a careful senior engineer.\n"
@@ -264,6 +215,12 @@ CODING_AGENT_GUIDANCE = (
"answer, not a preamble."
)
_TODO_SENTENCE = (
"- Track multi-step work with `todo_list`. Reference code as "
"`path:line` instead of pasting whole files."
)
_NO_TODO_SENTENCE = "- Reference code as `path:line` instead of pasting whole files."
# ── Context profiles (declarative posture definitions) ──────────────────────
@@ -272,35 +229,22 @@ CODING_AGENT_GUIDANCE = (
class ContextProfile:
"""A named operating posture. Pure data — consumers read these fields.
``toolset`` — collapse to this toolset (+ enabled MCP) when no explicit
selection is pinned; ``None`` keeps the platform default.
``guidance`` — operating brief injected into the stable system prompt;
``""`` injects nothing.
``model_hint`` — routing preference key for smart model routing
(extension seam; not yet consumed by the router).
``memory_policy``— memory namespace/weighting hint (extension seam).
``compact_skill_categories`` — skill categories DEMOTED to names-only in
the system-prompt skill index under the opt-in ``focus``
mode. Never hidden: every skill name stays visible
(so memory-anchored recall keeps working) — only the
descriptions are dropped to cut index noise. Deny-list
semantics so unknown/custom categories keep full
entries.
``toolset``: collapse to this toolset (+ enabled MCP) under ``focus``;
``None`` keeps the platform default. ``guidance``: operating brief for the
stable system prompt. ``model_hint``: routing preference (extension seam).
``compact_skill_categories``: categories DEMOTED to names-only in the skill
index under ``focus`` — deny-list, never hidden, so recall keeps working.
"""
name: str
toolset: Optional[str] = None
guidance: str = ""
model_hint: Optional[str] = None
memory_policy: str = "default"
compact_skill_categories: tuple[str, ...] = ()
# Skill categories that are clearly not part of a coding workflow. Demoted to
# names-only in the prompt's skill index under the opt-in ``focus`` mode only
# (deny-list — anything not listed here, incl. custom user categories, keeps
# full entries). Coding-adjacent categories (devops, github, mcp,
# data-science, diagramming, research, security, …) are intentionally absent.
# Clearly non-coding skill categories (deny-list: custom categories keep full
# entries). Coding-adjacent ones (devops, github, mcp, research, …) are absent.
_NON_CODING_SKILL_CATEGORIES = (
"apple", "communication", "cooking", "creative", "email", "finance",
"gaming", "gifs", "health", "media", "music", "note-taking",
@@ -315,7 +259,6 @@ CODING_PROFILE = ContextProfile(
toolset=CODING_TOOLSET,
guidance=CODING_AGENT_GUIDANCE,
model_hint="coding",
memory_policy="project",
compact_skill_categories=_NON_CODING_SKILL_CATEGORIES,
)
@@ -332,45 +275,38 @@ def get_profile(name: str) -> ContextProfile:
# ── Helpers ─────────────────────────────────────────────────────────────────
_MODE_ALIASES = {
**dict.fromkeys(("focus", "strict", "lean"), "focus"),
**dict.fromkeys(("on", "true", "yes", "1", "always"), "on"),
**dict.fromkeys(("off", "false", "no", "0", "never"), "off"),
}
def _coding_mode(config: Optional[dict[str, Any]]) -> str:
"""Return the normalized ``agent.coding_context`` mode (auto/focus/on/off)."""
def _agent_config_value(config: Optional[dict[str, Any]], key: str, default: Any, *, readonly: bool) -> Any:
"""``config["agent"][key]``, loading config when none was passed."""
if config is None:
try:
from hermes_cli.config import load_config_readonly
from hermes_cli.config import load_config, load_config_readonly
config = load_config_readonly()
config = load_config_readonly() if readonly else load_config()
except Exception:
config = {}
raw = ((config or {}).get("agent", {}) or {}).get("coding_context", "auto")
mode = str(raw).strip().lower()
if mode in {"focus", "strict", "lean"}:
return "focus"
if mode in {"on", "true", "yes", "1", "always"}:
return "on"
if mode in {"off", "false", "no", "0", "never"}:
return "off"
return "auto"
return ((config or {}).get("agent", {}) or {}).get(key, default)
def _coding_mode(config: Optional[dict[str, Any]]) -> str:
"""Normalized ``agent.coding_context`` mode (auto/focus/on/off)."""
raw = _agent_config_value(config, "coding_context", "auto", readonly=True)
return _MODE_ALIASES.get(str(raw).strip().lower(), "auto")
def _coding_instructions(config: Optional[dict[str, Any]]) -> str:
"""Standing operator instructions for the coding posture (config).
"""Standing operator instructions (``agent.coding_instructions``: str or list).
``agent.coding_instructions`` — a string or list of strings appended to the
coding brief as an extra stable system block, so a user can pin project-wide
coding-workflow rules (e.g. "for UI work don't run tsc/lint until I approve;
clean the diff before committing") without editing the shipped brief.
Cache-safe: resolved once per session into the stable system-prompt tier,
like the rest of the posture.
Appended to the brief as an extra stable block so a user can pin
project-wide workflow rules without editing the shipped brief.
"""
if config is None:
try:
from hermes_cli.config import load_config
config = load_config()
except Exception:
config = {}
raw = ((config or {}).get("agent", {}) or {}).get("coding_instructions", "")
raw = _agent_config_value(config, "coding_instructions", "", readonly=False)
if isinstance(raw, (list, tuple)):
return "\n".join(str(item).strip() for item in raw if str(item).strip())
return str(raw or "").strip()
@@ -403,19 +339,14 @@ def _home() -> Optional[Path]:
def _marker_root(cwd: Path) -> Optional[Path]:
"""Nearest ancestor that looks like a project root, or ``None``.
"""Nearest ancestor (≤6 levels) that looks like a project root, or ``None``.
Walks up at most a few levels so a manifest in the workspace root counts
even when the user is in a subdirectory. ``$HOME`` itself is skipped — a
Makefile or AGENTS.md sitting in the home directory is global user config,
not a project-root signal.
``$HOME`` and the shared temp root are skipped: a Makefile/AGENTS.md in the
home dir is global user config, and a stray manifest in /tmp must not flip
every session whose cwd lives under it into the coding posture.
"""
current = cwd.resolve()
home = _home()
# Shared world-writable temp roots are never project roots: a stray
# manifest in /tmp (left by any process) must not flip every session
# whose cwd lives under the temp dir into the coding posture. Same
# reasoning as the $HOME skip below.
try:
temp_root = Path(tempfile.gettempdir()).resolve()
except Exception:
@@ -435,17 +366,10 @@ def _detect_profile_name(mode: str, platform: str, cwd_str: str) -> str:
"""Resolve which profile applies.
``auto``/``focus``: coding when the surface is interactive AND the cwd is a
code workspace (a git repo or a recognised project root). ``on``: always
coding. ``off``: always general.
A git repo rooted at ``$HOME`` (the dotfiles pattern) is NOT a workspace
signal — without the guard, every session anywhere under a dotfiles-managed
home directory would silently flip to the coding posture.
Detection is intentionally not memoized: it's a handful of ``stat`` calls,
and callers resolve the mode once per session anyway. Caching here would
risk a stale posture if a long-lived process (gateway/TUI) serves sessions
from different working directories.
code workspace (project root, or a git repo that actually holds code).
``on``: always coding. ``off``: always general. A git repo rooted at
``$HOME`` (dotfiles) is NOT a workspace signal. Deliberately not memoized:
a long-lived gateway/TUI process serves sessions from different cwds.
"""
if mode == "off":
return GENERAL_PROFILE.name
@@ -454,16 +378,10 @@ def _detect_profile_name(mode: str, platform: str, cwd_str: str) -> str:
if platform and platform.strip().lower() not in INTERACTIVE_CODING_PLATFORMS:
return GENERAL_PROFILE.name
cwd = Path(cwd_str)
# A recognized project root (manifest / AGENTS.md / .cursorrules) is a code
# workspace on its own — cheap stat checks, no scan.
if _marker_root(cwd) is not None:
return CODING_PROFILE.name
git_root = _git_root(cwd)
if git_root is not None and git_root == _home():
git_root = None # dotfiles repo at $HOME — not a code workspace
# A bare git repo only counts when it actually holds code, so `git init` on a
# notes/writing/research folder stays in the general posture.
if git_root is not None and _has_code_files(git_root):
if git_root is not None and git_root != _home() and _has_code_files(git_root):
return CODING_PROFILE.name
return GENERAL_PROFILE.name
@@ -475,23 +393,18 @@ def _detect_profile_name(mode: str, platform: str, cwd_str: str) -> str:
class RuntimeMode:
"""The resolved operating posture for a session. Immutable by construction.
Built once via :func:`resolve_runtime_mode` and consumed by every domain
that cares about the coding/general distinction. Never mutate or re-resolve
mid-session — that would break the prompt cache.
Built once via :func:`resolve_runtime_mode`; never re-resolved mid-session
(that would break the prompt cache).
"""
profile: ContextProfile
surface: str
cwd: Path
# The normalized ``agent.coding_context`` mode this posture was resolved
# under (auto/focus/on/off). Toolset collapse is gated on ``focus``.
# Normalized ``agent.coding_context`` mode; toolset collapse is gated on ``focus``.
config_mode: str = "auto"
# The model id this session runs (e.g. "anthropic/claude-opus-4.8"). Used
# only to steer edit-format guidance toward the model's family — see
# ``_edit_format_line``. Fixed for the session, so cache-safe.
# Model id, used only to steer edit-format guidance (fixed per session).
model: Optional[str] = None
# Standing operator instructions (``agent.coding_instructions``), appended
# as an extra stable system block. Empty unless the user configures it.
# ``agent.coding_instructions``, appended as an extra stable block.
instructions: str = ""
@property
@@ -505,94 +418,58 @@ class RuntimeMode:
def toolset_selection(self, config: Optional[dict[str, Any]] = None) -> Optional[list[str]]:
"""Toolset list for this posture, or ``None`` to keep the platform default.
Non-``None`` only under the opt-in ``focus`` mode. The default posture
is prompt-only: most strippable toolsets are off-by-default anyway, and
a user who explicitly enabled one (image-gen for frontend/game assets,
messaging for build notifications, …) keeps it while coding.
Callers apply this only when the user hasn't pinned an explicit
selection (``--toolsets``, ``HERMES_TUI_TOOLSETS``, …); they never
override a pin. Returns the profile's toolset plus enabled MCP servers.
Non-``None`` only under ``focus``. Callers apply it only when the user
hasn't pinned an explicit selection (``--toolsets``, ``HERMES_TUI_TOOLSETS``).
"""
if self.config_mode != "focus":
return None
if self.profile.toolset is None:
if self.config_mode != "focus" or self.profile.toolset is None:
return None
return [self.profile.toolset, *_enabled_mcp_servers(config)]
def system_prompt_parts(
self, valid_tool_names=None
) -> tuple[list[str], list[str], list[str]]:
"""Return prefix, workspace, and trailing posture blocks separately.
"""Return (prefix, workspace, trailing) posture blocks.
The operating brief carries a model-family edit-format nudge appended
to it (one cached string, not a separate block) so the model is steered
toward the `patch` mode it handles best — see ``_edit_format_line``.
``valid_tool_names`` (when provided) tailors the brief to the session's
toolset: the ``todo`` tracking sentence is dropped when the todo tool
isn't loaded (e.g. Blank Slate), so the brief never references a tool
the model can't call. The toolset is fixed at session construction,
so the rendered brief is deterministic per session — cache-safe.
The three lists preserve the historical flat prompt order: the brief,
the live workspace snapshot, then configured operator instructions.
Prompt assembly can therefore put a cache boundary before the snapshot
without changing the persisted system-prompt bytes.
The brief carries the model-family edit-format nudge appended to it
(one cached string). ``valid_tool_names`` drops the ``todo_list``
sentence when that tool isn't loaded (e.g. Blank Slate). The three
lists preserve the historical flat order — brief, workspace snapshot,
operator instructions — so prompt assembly can put a cache boundary
before the snapshot without changing the persisted bytes.
"""
if not self.is_coding:
return [], [], []
prefix: list[str] = []
workspace_parts: list[str] = []
trailing: list[str] = []
if self.profile.guidance:
brief = self.profile.guidance
if valid_tool_names is not None and "todo_list" not in valid_tool_names:
brief = brief.replace(
"- Track multi-step work with `todo_list`. Reference code as "
"`path:line` instead of pasting whole files.",
"- Reference code as `path:line` instead of pasting "
"whole files.",
)
brief = brief.replace(_TODO_SENTENCE, _NO_TODO_SENTENCE)
edit_line = _edit_format_line(self.model)
if edit_line:
brief = f"{brief}\n{edit_line}"
prefix.append(brief)
workspace = build_coding_workspace_block(self.cwd)
if workspace:
workspace_parts.append(workspace)
# Operator instructions ride their own block so the brief (block 0) stays
# byte-stable and cache-keyed independently of user config.
if self.instructions:
trailing.append(f"Operator instructions (from config):\n{self.instructions}")
workspace_parts = [workspace] if workspace else []
# Operator instructions ride their own block so the brief stays
# byte-stable independently of user config.
trailing = (
[f"Operator instructions (from config):\n{self.instructions}"]
if self.instructions else []
)
return prefix, workspace_parts, trailing
def system_blocks(self) -> list[str]:
"""Return posture blocks in their historical display order.
``system_prompt_parts`` is the cache-aware API. This compatibility
helper retains the public flat list for callers outside prompt assembly.
"""
"""Posture blocks as one flat list in historical order (compat helper)."""
prefix, workspace, trailing = self.system_prompt_parts()
return [*prefix, *workspace, *trailing]
def compact_skill_categories(self) -> frozenset[str]:
"""Skill categories to demote to names-only in the prompt's skill index.
"""Skill categories to demote to names-only in the skill index.
Gated on the opt-in ``focus`` mode, like the toolset collapse: the
default posture leaves the skill index untouched. Users who didn't ask
for a lean prompt keep full entries for every category — index changes
under ``auto`` proved too surprising in practice, even names-only ones
(a demoted description is information the model no longer weighs when
deciding what to load).
Demoted — never hidden — even under ``focus``. An earlier revision
fully pruned these categories from the index, which caused silent
capability loss in a real workflow: agent-created skills are the
model's accumulated project memory (server-ops runbooks, learned
pitfalls, …), and models do not reliably reach for ``skills_list`` to
rediscover what the index stopped showing them. Names-only keeps every
skill loadable on recall while still cutting the description noise.
Gated on ``focus`` like the toolset collapse — index changes under
``auto`` proved too surprising. Demoted, never hidden: fully pruning
them caused silent capability loss (agent-created skills are the
model's project memory and models don't reliably re-run ``skills_list``).
"""
if not self.is_coding or self.config_mode != "focus":
return frozenset()
@@ -606,14 +483,10 @@ def resolve_runtime_mode(
config: Optional[dict[str, Any]] = None,
model: Optional[str] = None,
) -> RuntimeMode:
"""Resolve the operating posture once. Cheap — a handful of ``stat`` calls.
"""Resolve the operating posture once (a handful of ``stat`` calls).
This is the single entry point every domain should call. The returned
object is immutable and safe to cache for the session. Detection itself is
intentionally *not* memoized (see ``_detect_profile_name``) so a long-lived
process can't pin a stale posture; callers resolve once per session and
hold the result. ``model`` is recorded only to steer edit-format guidance;
it never affects detection.
The single entry point every domain should call; the result is immutable
and safe to hold for the session. ``model`` only steers edit-format guidance.
"""
resolved_cwd = _resolve_cwd(cwd)
mode = _coding_mode(config)
@@ -649,32 +522,12 @@ def coding_selection(
cwd: Optional[str | Path] = None,
config: Optional[dict[str, Any]] = None,
) -> Optional[list[str]]:
"""Toolset selection for the coding posture.
``None`` unless the user opted into ``focus`` mode AND the posture is
active — the default coding posture never overrides configured toolsets.
"""
"""Toolset selection for the coding posture (``None`` unless ``focus`` and active)."""
return resolve_runtime_mode(
platform=platform, cwd=cwd, config=config
).toolset_selection(config)
def coding_system_blocks(
*,
platform: Optional[str] = None,
cwd: Optional[str | Path] = None,
config: Optional[dict[str, Any]] = None,
model: Optional[str] = None,
) -> list[str]:
"""Stable system-prompt blocks for the current posture (empty when general).
``model`` steers the brief's edit-format nudge toward the model's family.
"""
return resolve_runtime_mode(
platform=platform, cwd=cwd, config=config, model=model
).system_blocks()
def coding_system_prompt_parts(
*,
platform: Optional[str] = None,
@@ -695,25 +548,14 @@ def coding_compact_skill_categories(
cwd: Optional[str | Path] = None,
config: Optional[dict[str, Any]] = None,
) -> frozenset[str]:
"""Skill categories the active posture demotes to names-only in the index.
Empty outside the coding posture and outside the opt-in ``focus`` mode —
the default posture never touches the skill index. Under ``focus``,
demoted — never hidden: every skill name stays in the index and remains
loadable via ``skill_view`` / ``skills_list``; only descriptions are
dropped.
"""
"""Skill categories the active posture demotes to names-only (empty outside ``focus``)."""
return resolve_runtime_mode(
platform=platform, cwd=cwd, config=config
).compact_skill_categories()
def _enabled_mcp_servers(config: Optional[dict[str, Any]]) -> list[str]:
"""Names of MCP servers the user has enabled — kept in the coding posture.
MCP servers (figma, browser, tophat, …) are explicitly configured and part
of the coding workflow, not noise to strip.
"""
"""Names of MCP servers the user has enabled — kept in the coding posture."""
try:
from hermes_cli.config import read_raw_config
from hermes_cli.tools_config import _parse_enabled_flag
@@ -735,10 +577,9 @@ def _enabled_mcp_servers(config: Optional[dict[str, Any]]) -> list[str]:
def _git(cwd: Path, *args: str) -> str:
"""``git -C <cwd> <args>`` → stripped stdout, or ``""`` on any failure.
Uses the shared :func:`bounded_git_probe` so the post-kill cleanup is bounded
on Windows — a plain ``subprocess.run(timeout=...)`` here deadlocked the agent
turn inside ``build_coding_workspace_block`` when a killed git left a suspended
descendant holding the pipe handles (issue #66037).
:func:`bounded_git_probe` bounds the post-kill cleanup on Windows — a plain
``subprocess.run(timeout=...)`` deadlocked when a killed git left a
suspended descendant holding the pipe handles.
"""
return bounded_git_probe(["git", "-C", str(cwd), *args], timeout=_GIT_TIMEOUT)
@@ -780,12 +621,7 @@ def _read_small(path: Path) -> str:
@dataclass(frozen=True)
class ProjectFacts:
"""Structured project facts — the model's verify loop, detected once.
The same data that feeds the workspace snapshot, exposed structurally so
non-prompt consumers (e.g. the desktop verify UI) read it instead of
re-detecting and drifting from the prompt.
"""
"""Structured project facts — exposed so non-prompt consumers (desktop verify UI) don't re-detect."""
manifests: list[str]
package_managers: list[str]
@@ -796,9 +632,8 @@ class ProjectFacts:
def detect_project_facts(root: Path) -> ProjectFacts:
"""Detect manifests, package manager(s), verify commands, and context files.
Cheap: stat calls plus reads of a couple of small files. The single source
of truth for both the prompt snapshot (:func:`_project_facts`) and the
gateway's ``project.facts`` — so the UI never re-sniffs verify commands.
Single source of truth for the prompt snapshot and the gateway's
``project.facts``. Cheap: stat calls plus a couple of small file reads.
"""
manifests = [m for m in _PROJECT_MARKERS if m not in _CONTEXT_FILES and (root / m).is_file()]
package_managers = list(
@@ -833,16 +668,9 @@ def detect_project_facts(root: Path) -> ProjectFacts:
def _project_facts(root: Path) -> list[str]:
"""Render :func:`detect_project_facts` as workspace-snapshot lines.
Hands the model its *verify loop* up front — which manifest, which package
manager, and the exact test/lint/build commands — instead of making it
rediscover them every session. Built once at prompt-build time; the string
output must stay byte-stable to preserve the prompt cache.
"""
"""Render :func:`detect_project_facts` as workspace-snapshot lines (byte-stable)."""
f = detect_project_facts(root)
facts: list[str] = []
if f.manifests:
line = f"- Project: {', '.join(f.manifests[:6])}"
if f.package_managers:
@@ -852,22 +680,25 @@ def _project_facts(root: Path) -> list[str]:
facts.append(f"- Verify: {'; '.join(f.verify_commands)}")
if f.context_files:
facts.append(f"- Context files: {', '.join(f.context_files)}")
return facts
def _workspace_roots(cwd: Optional[str | Path]) -> tuple[Optional[Path], Optional[Path]]:
"""(git_root, workspace_root) for *cwd*; workspace root is git root else marker root."""
resolved = _resolve_cwd(cwd)
git_root = _git_root(resolved)
return git_root, git_root or _marker_root(resolved)
def project_facts_for(cwd: Optional[str | Path] = None) -> Optional[dict[str, Any]]:
"""Structured project facts for ``cwd`` — ``None`` outside a workspace.
Same detection the system-prompt snapshot uses (git root, else marker root),
exposed for non-prompt consumers (the desktop verify UI) so they never
re-derive "are we coding?" or duplicate the verify-command sniffing.
Same detection the system-prompt snapshot uses, exposed for non-prompt
consumers (the desktop verify UI).
"""
resolved = _resolve_cwd(cwd)
root = _git_root(resolved) or _marker_root(resolved)
_, root = _workspace_roots(cwd)
if root is None:
return None
f = detect_project_facts(root)
return {
"root": str(root),
@@ -881,13 +712,10 @@ def project_facts_for(cwd: Optional[str | Path] = None) -> Optional[dict[str, An
def build_coding_workspace_block(cwd: Optional[str | Path] = None) -> str:
"""Workspace snapshot for the system prompt (empty outside a workspace).
Git state (branch/status/commits) when the cwd is in a repo, plus detected
project facts (manifest, package manager, verify commands, context files)
— so marker-only (non-git) projects still get a snapshot.
Git state when the cwd is in a repo, plus detected project facts — so
marker-only (non-git) projects still get a snapshot.
"""
resolved = _resolve_cwd(cwd)
git_root = _git_root(resolved)
root = git_root or _marker_root(resolved)
git_root, root = _workspace_roots(cwd)
if root is None:
return ""
@@ -908,11 +736,9 @@ def build_coding_workspace_block(cwd: Optional[str | Path] = None) -> str:
elif head == "(detached)":
lines.append("- Branch: (detached HEAD)")
# Linked worktree: the per-worktree git dir differs from the shared common dir.
# We surface the fact that it's a worktree (so the model knows branches/stashes
# are shared state) but deliberately do NOT expose the primary tree path —
# giving the model a second absolute path causes it to sometimes run commands
# in the wrong directory.
# Linked worktree: say so (branches/stashes are shared state) but do
# NOT expose the primary tree path — a second absolute path makes the
# model run commands in the wrong directory.
git_dir, common_dir = _git(root, "rev-parse", "--git-dir"), _git(root, "rev-parse", "--git-common-dir")
if git_dir and common_dir and Path(git_dir).resolve() != Path(common_dir).resolve():
lines.append("- Worktree: linked (git state shared with primary tree)")
+98 -161
View File
@@ -1,3 +1,5 @@
"""@-reference expansion (``@file:``, ``@folder:``, ``@diff``, ``@git:``, ``@url:`` + plugin prefixes)."""
from __future__ import annotations
import asyncio
@@ -7,6 +9,7 @@ import mimetypes
import os
import re
import subprocess
from abc import ABC, abstractmethod
from dataclasses import dataclass, field
from pathlib import Path
from typing import Awaitable, Callable
@@ -20,10 +23,8 @@ from hermes_cli._subprocess_compat import (
)
from hermes_cli.sizefmt import format_bytes
from abc import ABC, abstractmethod
# ---------------------------------------------------------------------------
# Plugin context-reference provider API (Issue #26193)
# Plugin context-reference provider API
# ---------------------------------------------------------------------------
BUILTIN_PREFIXES = frozenset({"diff", "staged", "file", "folder", "git", "url"})
@@ -43,11 +44,7 @@ class ContextCompletionItem:
class ContextReferenceProvider(ABC):
"""Base class for plugin-registered @-prefix context reference providers.
Plugins subclass this and register via
``PluginContext.register_context_reference()``.
"""
"""Base class for plugin @-prefix providers, registered via ``PluginContext.register_context_reference()``."""
prefix: str = "" # e.g. "issue", "channel", "doc"
description: str = "" # shown in autocomplete meta column
@@ -86,8 +83,7 @@ _QUOTED_REFERENCE_VALUE = r'(?:`[^`\n]+`|"[^"\n]+"|\'[^\'\n]+\')'
REFERENCE_PATTERN = re.compile(
rf"(?<![\w/])@(?:(?P<simple>diff|staged)\b|(?P<kind>file|folder|git|url):(?P<value>{_QUOTED_REFERENCE_VALUE}(?::\d+(?:-\d+)?)?|\S+))"
)
# Plugin fallback pattern – catches any @<word>:<value> not handled by the
# built-in regex so that plugin-registered prefixes can be resolved.
# Plugin fallback: any @<word>:<value> the built-in regex did not claim.
_PLUGIN_REFERENCE_PATTERN = re.compile(
rf"(?<![\w/])@(?P<kind>[a-zA-Z][a-zA-Z0-9_-]*):(?P<value>{_QUOTED_REFERENCE_VALUE}(?::\d+(?:-\d+)?)?|\S+)"
)
@@ -138,9 +134,8 @@ class ContextReferenceResult:
def format_reference_value(value: str) -> str:
"""Quote a reference value so ``REFERENCE_PATTERN`` reads it back whole.
The unquoted alternative in the pattern is ``\\S+``, so a path containing a
space parses as a truncated ref with the tail left behind as loose text.
Mirrors ``formatRefValue`` in the desktop's directive-text.tsx.
The unquoted alternative is ``\\S+``, so a path with a space would parse as a
truncated ref. Mirrors ``formatRefValue`` in the desktop's directive-text.tsx.
"""
if not _NEEDS_QUOTING.search(value):
return value
@@ -158,26 +153,14 @@ def parse_context_references(message: str) -> list[ContextReference]:
for match in REFERENCE_PATTERN.finditer(message):
simple = match.group("simple")
if simple:
refs.append(
ContextReference(
raw=match.group(0),
kind=simple,
target="",
start=match.start(),
end=match.end(),
)
)
refs.append(ContextReference(raw=match.group(0), kind=simple, target="", start=match.start(), end=match.end()))
continue
kind = match.group("kind")
value = _strip_trailing_punctuation(match.group("value") or "")
line_start = None
line_end = None
target = _strip_reference_wrappers(value)
if kind == "file":
target, line_start, line_end = _parse_file_reference_value(value)
else:
target, line_start, line_end = _strip_reference_wrappers(value), None, None
refs.append(
ContextReference(
raw=match.group(0),
@@ -190,26 +173,24 @@ def parse_context_references(message: str) -> list[ContextReference]:
)
)
# Second pass: resolve plugin-registered prefixes the built-in pattern missed
# Second pass: plugin-registered prefixes the built-in pattern missed.
if _context_reference_providers:
for match in _PLUGIN_REFERENCE_PATTERN.finditer(message):
kind = match.group("kind")
if kind in BUILTIN_PREFIXES:
if kind in BUILTIN_PREFIXES or kind not in _context_reference_providers:
continue
# Skip if already captured by the built-in pattern
if any(r.kind == kind and r.start == match.start() for r in refs):
continue
if kind in _context_reference_providers:
value = _strip_trailing_punctuation(match.group("value") or "")
refs.append(
ContextReference(
raw=match.group(0),
kind=kind,
target=_strip_reference_wrappers(value),
start=match.start(),
end=match.end(),
)
value = _strip_trailing_punctuation(match.group("value") or "")
refs.append(
ContextReference(
raw=match.group(0),
kind=kind,
target=_strip_reference_wrappers(value),
start=match.start(),
end=match.end(),
)
)
return refs
@@ -222,6 +203,7 @@ def preprocess_context_references(
url_fetcher: Callable[[str], str | Awaitable[str]] | None = None,
allowed_root: str | Path | None = None,
) -> ContextReferenceResult:
"""Sync wrapper; safe both without a loop (CLI) and inside a running loop (gateway)."""
coro = preprocess_context_references_async(
message,
cwd=cwd,
@@ -229,7 +211,6 @@ def preprocess_context_references(
url_fetcher=url_fetcher,
allowed_root=allowed_root,
)
# Safe for both CLI (no loop) and gateway (loop already running).
try:
loop = asyncio.get_running_loop()
except RuntimeError:
@@ -254,31 +235,17 @@ async def preprocess_context_references_async(
return ContextReferenceResult(message=message, original_message=message)
cwd_path = Path(cwd).expanduser().resolve()
# Default to the current working directory so @ references cannot escape
# the active workspace unless a caller explicitly widens the root.
allowed_root_path = (
Path(allowed_root).expanduser().resolve() if allowed_root is not None else cwd_path
)
# Default root = cwd so @ references cannot escape the workspace unless a caller widens it.
allowed_root_path = Path(allowed_root).expanduser().resolve() if allowed_root is not None else cwd_path
warnings: list[str] = []
blocks: list[str] = []
injected_tokens = 0
# Expand all references concurrently. Each _expand_reference is independent
# (no shared state during expansion) — a message with several @url: refs
# would otherwise pay one full web_extract round-trip per ref in series.
# gather preserves positional order, so we reassemble warnings/blocks in the
# original ref order exactly as the prior serial loop did; the token-budget
# check below is unchanged (it runs once, after all refs are expanded).
# Expand concurrently (each ref is independent; several @url: refs would otherwise
# serialize web_extract round-trips). gather preserves order, so warnings/blocks
# are assembled in ref order; the token-budget check runs once afterwards.
expanded = await asyncio.gather(
*(
_expand_reference(
ref,
cwd_path,
url_fetcher=url_fetcher,
allowed_root=allowed_root_path,
)
for ref in refs
)
*(_expand_reference(ref, cwd_path, url_fetcher=url_fetcher, allowed_root=allowed_root_path) for ref in refs)
)
for warning, block in expanded:
if warning:
@@ -302,17 +269,14 @@ async def preprocess_context_references_async(
expanded=False,
blocked=True,
)
if injected_tokens > soft_limit:
warnings.append(
f"@ context injection warning: {injected_tokens} tokens exceeds the 25% soft limit ({soft_limit})."
)
# Leave the `@file:`/`@folder:` tokens where the user typed them. The token
# IS the reference, not scaffolding around it: clients render each one as an
# inline chip, so stripping them left a sentence with a hole in it ("review
# and ship") and made the desktop re-derive the refs from the attached block
# to show them as a detached list above the prose.
# The `@file:`/`@folder:` tokens stay where the user typed them: the token IS the
# reference (clients render it as an inline chip); stripping it left a hole in the
# sentence and forced the desktop to re-derive refs from the attached block.
final = message
if warnings:
final = f"{final}\n\n--- Context Warnings ---\n" + "\n".join(f"- {warning}" for warning in warnings)
@@ -337,6 +301,7 @@ async def _expand_reference(
url_fetcher: Callable[[str], str | Awaitable[str]] | None = None,
allowed_root: Path | None = None,
) -> tuple[str | None, str | None]:
"""Return ``(warning, block)`` for one reference; exactly one side is set."""
try:
if ref.kind == "file":
return _expand_file_reference(ref, cwd, allowed_root=allowed_root)
@@ -357,7 +322,6 @@ async def _expand_reference(
except Exception as exc:
return f"{ref.raw}: {exc}", None
# Plugin-provided context references
provider = _context_reference_providers.get(ref.kind)
if provider is not None:
try:
@@ -383,13 +347,8 @@ def _expand_file_reference(
if not path.is_file():
return f"{ref.raw}: path is not a file", None
if _is_binary_file(path):
# A binary file can't be inlined as text, but it IS on disk (the agent's
# tools run where this resolves — the local cwd, or the staged copy in a
# remote session workspace). Returning a bare "not supported" warning
# with no content was a dead end: the model saw a failure and gave up
# (told the user the file type wasn't supported). Instead, hand it an
# actionable block — the path, type, size, and a nudge to use its tools —
# so it can read/convert/view the file itself.
# A bare "not supported" warning was a dead end (the model gave up); the file IS
# on disk where the agent's tools run, so hand it an actionable block instead.
return None, _binary_reference_block(ref, path)
text = path.read_text(encoding="utf-8")
@@ -400,8 +359,7 @@ def _expand_file_reference(
text = "\n".join(lines[start_idx:end_idx])
lang = _code_fence_language(path)
label = ref.raw
return None, f"📄 {label} ({estimate_tokens_rough(text)} tokens)\n```{lang}\n{text}\n```"
return None, f"📄 {ref.raw} ({estimate_tokens_rough(text)} tokens)\n```{lang}\n{text}\n```"
def _expand_folder_reference(
@@ -416,37 +374,45 @@ def _expand_folder_reference(
return f"{ref.raw}: folder not found", None
if not path.is_dir():
return f"{ref.raw}: path is not a folder", None
listing = _build_folder_listing(path, cwd)
return None, f"📁 {ref.raw} ({estimate_tokens_rough(listing)} tokens)\n{listing}"
def _run_quiet(
cmd: list[str], cwd: Path, timeout: int, env: dict | None = None
) -> subprocess.CompletedProcess:
"""subprocess.run with captured text output, no stdin, and no console flash on Windows."""
popen_kwargs: dict = {"creationflags": windows_hide_flags()} if IS_WINDOWS else {}
if env is not None:
popen_kwargs["env"] = env
return subprocess.run(
cmd,
cwd=cwd,
capture_output=True,
text=True, encoding='utf-8', errors='replace',
timeout=timeout,
stdin=subprocess.DEVNULL,
**popen_kwargs,
)
def _expand_git_reference(
ref: ContextReference,
cwd: Path,
args: list[str],
label: str,
) -> tuple[str | None, str | None]:
_popen_kwargs = {"creationflags": windows_hide_flags()} if IS_WINDOWS else {}
try:
result = subprocess.run(
["git", *harden_git_argv(args)],
cwd=cwd,
capture_output=True,
text=True, encoding='utf-8', errors='replace',
timeout=30,
stdin=subprocess.DEVNULL,
env=noninteractive_git_env(),
**_popen_kwargs,
# Repo-supplied config/attributes must never execute code (GHSA-7x36-8jrh-v4pw).
result = _run_quiet(
["git", *harden_git_argv(args)], cwd, 30, env=noninteractive_git_env()
)
except subprocess.TimeoutExpired:
return f"{ref.raw}: git command timed out (30s)", None
if result.returncode != 0:
stderr = (result.stderr or "").strip() or "git command failed"
return f"{ref.raw}: {stderr}", None
content = result.stdout.strip()
if not content:
content = "(no output)"
content = result.stdout.strip() or "(no output)"
return None, f"🧾 {label} ({estimate_tokens_rough(content)} tokens)\n```diff\n{content}\n```"
@@ -466,8 +432,7 @@ async def _default_url_fetcher(url: str) -> str:
from tools.web_tools import web_extract_tool
raw = await web_extract_tool([url], format="markdown")
payload = json.loads(raw)
docs = payload.get("results", [])
docs = json.loads(raw).get("results", [])
if not docs:
return ""
doc = docs[0]
@@ -488,6 +453,7 @@ def _resolve_path(cwd: Path, target: str, *, allowed_root: Path | None = None) -
def _ensure_reference_path_allowed(path: Path) -> None:
"""Refuse credential/internal paths. Fails CLOSED: the gateway feeds untrusted remote text here."""
from hermes_constants import get_hermes_home
home = Path(os.path.expanduser("~")).resolve()
hermes_home = get_hermes_home().resolve()
@@ -499,7 +465,6 @@ def _ensure_reference_path_allowed(path: Path) -> None:
if path in blocked_exact:
raise ValueError("path is a sensitive credential file and cannot be attached")
for blocked_dir in blocked_dirs:
try:
path.relative_to(blocked_dir)
@@ -507,16 +472,9 @@ def _ensure_reference_path_allowed(path: Path) -> None:
continue
raise ValueError("path is a sensitive credential or internal Hermes path and cannot be attached")
# Anchor to the canonical read deny-list (agent/file_safety.get_read_block_error),
# the single source of truth used by the file/terminal read path. The narrow
# list above predates that guard and never caught the real credential stores:
# provider keys (auth.json), Anthropic OAuth tokens (.anthropic_oauth.json),
# MCP OAuth material (mcp-tokens/), webhook HMAC secrets, and project-local
# .env files. That gap matters because the gateway feeds UNTRUSTED remote
# message text into reference expansion, so `@file:~/.hermes/auth.json` from a
# chat peer would otherwise read the operator's keys straight into context.
# Routing through the canonical guard closes the gap today and keeps this path
# protected automatically whenever that deny-list grows.
# Anchor to the canonical read deny-list (agent/file_safety.get_read_block_error): the
# narrow list above never caught auth.json, .anthropic_oauth.json, mcp-tokens/, webhook
# secrets or project .env files, and it grows automatically with that deny-list.
try:
from agent.file_safety import get_read_block_error
@@ -527,13 +485,8 @@ def _ensure_reference_path_allowed(path: Path) -> None:
except ValueError:
raise
except Exception:
# Fail CLOSED on the security path. This guard exists specifically to
# cover credential stores the narrow list above misses (auth.json,
# .anthropic_oauth.json, mcp-tokens/, ...). If the canonical lookup
# ever fails, silently falling through would re-open that exact hole —
# the gateway feeds untrusted remote text here, so a probe could then
# attach the operator's keys. Refuse instead: a spurious block on a
# legitimate file is a recoverable annoyance; a leaked credential is not.
# If the canonical lookup fails, falling through would re-open the exact hole this
# guard closes; a spurious block is recoverable, a leaked credential is not.
raise ValueError(
"path could not be verified against the credential deny-list and cannot be attached"
)
@@ -583,27 +536,26 @@ def _parse_file_reference_value(value: str) -> tuple[str, int | None, int | None
return _strip_reference_wrappers(value), None, None
_TEXT_EXTENSIONS = (".py", ".md", ".txt", ".json", ".yaml", ".yml", ".toml", ".js", ".ts")
def _is_binary_file(path: Path) -> bool:
mime, _ = mimetypes.guess_type(path.name)
if mime and not mime.startswith("text/") and not any(
path.name.endswith(ext) for ext in (".py", ".md", ".txt", ".json", ".yaml", ".yml", ".toml", ".js", ".ts")
):
if mime and not mime.startswith("text/") and not path.name.endswith(_TEXT_EXTENSIONS):
return True
chunk = path.read_bytes()[:4096]
return b"\x00" in chunk
return b"\x00" in path.read_bytes()[:4096]
def _build_folder_listing(path: Path, cwd: Path, limit: int = 200) -> str:
lines = [f"{path.relative_to(cwd)}/"]
entries = _iter_visible_entries(path, cwd, limit=limit)
base_depth = len(path.relative_to(cwd).parts)
for entry in entries:
rel = entry.relative_to(cwd)
indent = " " * max(len(rel.parts) - len(path.relative_to(cwd).parts) - 1, 0)
indent = " " * max(len(entry.relative_to(cwd).parts) - base_depth - 1, 0)
if entry.is_dir():
lines.append(f"{indent}- {entry.name}/")
else:
meta = _file_metadata(entry)
lines.append(f"{indent}- {entry.name} ({meta})")
lines.append(f"{indent}- {entry.name} ({_file_metadata(entry)})")
if len(entries) >= limit:
lines.append("- ...")
return "\n".join(lines)
@@ -629,29 +581,16 @@ def _iter_visible_entries(path: Path, cwd: Path, limit: int) -> list[Path]:
dirs[:] = sorted(d for d in dirs if not d.startswith(".") and d != "__pycache__")
files = sorted(f for f in files if not f.startswith("."))
root_path = Path(root)
for d in dirs:
output.append(root_path / d)
if len(output) >= limit:
return output
for f in files:
output.append(root_path / f)
for name in dirs + files:
output.append(root_path / name)
if len(output) >= limit:
return output
return output
def _rg_files(path: Path, cwd: Path, limit: int) -> list[Path] | None:
_popen_kwargs = {"creationflags": windows_hide_flags()} if IS_WINDOWS else {}
try:
result = subprocess.run(
["rg", "--files", str(path.relative_to(cwd))],
cwd=cwd,
capture_output=True,
text=True, encoding='utf-8', errors='replace',
timeout=10,
stdin=subprocess.DEVNULL,
**_popen_kwargs,
)
result = _run_quiet(["rg", "--files", str(path.relative_to(cwd))], cwd, 10)
except (FileNotFoundError, OSError, subprocess.TimeoutExpired):
return None
if result.returncode != 0:
@@ -661,19 +600,15 @@ def _rg_files(path: Path, cwd: Path, limit: int) -> list[Path] | None:
def _agent_visible_path(path: Path) -> str:
"""Map a host path to the path the agent's tools can read in the active backend.
"""Map a host path to what the agent's tools can read in the active backend.
Under a container backend (docker) the gateway host path dangles inside the
sandbox — the container has its own filesystem and the host path is not
mounted. Files staged into an auto-mounted cache dir (``images/``,
``attachments/``, ...) are translated to their in-container path via the
existing ``tools.credential_files`` machinery (#76577). Falls back to the
host path when the backend is local or translation is unavailable.
Under a container backend the host path dangles inside the sandbox; files staged
into an auto-mounted cache dir are translated via ``tools.credential_files``.
Falls back to the host path when the backend is local or translation fails.
"""
try:
# Desktop/in-process gateways may not have bridged ``terminal.*``
# config into ``TERMINAL_ENV`` at startup; run the idempotent bridge so
# the credential_files translation gate sees the active backend.
# In-process gateways may not have bridged terminal.* config into TERMINAL_ENV
# yet; run the idempotent bridge so the translation gate sees the active backend.
from tools.terminal_tool import _ensure_terminal_env_bridged
_ensure_terminal_env_bridged()
@@ -709,18 +644,20 @@ def _file_metadata(path: Path) -> str:
return f"{line_count} lines"
_FENCE_LANGUAGES = {
".py": "python",
".js": "javascript",
".ts": "typescript",
".tsx": "tsx",
".jsx": "jsx",
".json": "json",
".md": "markdown",
".sh": "bash",
".yml": "yaml",
".yaml": "yaml",
".toml": "toml",
}
def _code_fence_language(path: Path) -> str:
mapping = {
".py": "python",
".js": "javascript",
".ts": "typescript",
".tsx": "tsx",
".jsx": "jsx",
".json": "json",
".md": "markdown",
".sh": "bash",
".yml": "yaml",
".yaml": "yaml",
".toml": "toml",
}
return mapping.get(path.suffix.lower(), "")
return _FENCE_LANGUAGES.get(path.suffix.lower(), "")
+342 -608
View File
File diff suppressed because it is too large Load Diff
+33 -119
View File
@@ -1,32 +1,8 @@
"""Lightweight internationalization (i18n) for Hermes static user-facing messages.
"""Lightweight i18n for Hermes' static user-facing strings (approval prompts, a few gateway replies).
Scope (thin slice, by design): only the highest-impact static strings shown
to the user by Hermes itself -- approval prompts, a handful of gateway slash
command replies, restart-drain notices. Agent-generated output, log lines,
error tracebacks, tool outputs, and slash-command descriptions all stay in
English.
Catalog files live under ``locales/<lang>.yaml`` at the repo root. Each
catalog is a flat dict keyed by dotted paths (e.g. ``approval.choose`` or
``gateway.approval_expired``). Missing keys fall back to English; if English
is missing too, the key path itself is returned so a broken catalog never
crashes the agent.
Usage::
from agent.i18n import t
print(t("approval.choose_long")) # current lang
print(t("gateway.draining", count=3)) # {count} formatted
print(t("approval.choose_long", lang="zh")) # explicit override
Language resolution order:
1. Explicit ``lang=`` argument passed to :func:`t`
2. ``HERMES_LANGUAGE`` environment variable (for tests / quick override)
3. ``display.language`` from config.yaml
4. ``"en"`` (baseline)
Supported languages: en, zh, zh-hant, ja, de, es, fr, tr, uk, af, ko, it, ga,
pt, ru, hu, ar. Unknown values fall back to en.
Catalogs are ``locales/<lang>.yaml`` flattened to dotted keys. Missing keys
fall back to English, then to the key itself, so a broken catalog never crashes.
Language resolution: explicit ``lang=`` > ``HERMES_LANGUAGE`` > ``display.language`` > ``en``.
"""
from __future__ import annotations
@@ -46,15 +22,13 @@ SUPPORTED_LANGUAGES: tuple[str, ...] = (
)
DEFAULT_LANGUAGE = "en"
# Accept a few natural aliases so users who type "chinese" / "zh-CN" / "jp"
# get the right catalog instead of silently falling back to English.
# Natural aliases so "chinese" / "zh-CN" / "jp" hit the right catalog instead of
# silently falling back to English. Bare "chinese" defaults to Simplified;
# Taiwan/HK/Macau tags route to the distinct Traditional catalog. pt-br shares
# the pt catalog (no separate br one).
_LANGUAGE_ALIASES: dict[str, str] = {
"english": "en", "en-us": "en", "en-gb": "en",
# Simplified Chinese — explicit codes route here; bare "chinese" / "mandarin"
# also default to Simplified since that's the larger user base.
"chinese": "zh", "mandarin": "zh", "zh-cn": "zh", "zh-hans": "zh", "zh-sg": "zh",
# Traditional Chinese — distinct catalog. Cover Taiwan / Hong Kong / Macau
# locale tags plus the common "traditional" alias.
"traditional-chinese": "zh-hant", "traditional_chinese": "zh-hant",
"zh-tw": "zh-hant", "zh-hk": "zh-hant", "zh-mo": "zh-hant",
"japanese": "ja", "jp": "ja", "ja-jp": "ja",
@@ -63,23 +37,14 @@ _LANGUAGE_ALIASES: dict[str, str] = {
"french": "fr", "français": "fr", "france": "fr", "fr-fr": "fr", "fr-be": "fr", "fr-ca": "fr", "fr-ch": "fr",
"ukrainian": "uk", "ukrainisch": "uk", "українська": "uk", "uk-ua": "uk", "ua": "uk",
"turkish": "tr", "türkçe": "tr", "tr-tr": "tr",
# Afrikaans — South African Dutch-derived language; "af-ZA" is the common BCP-47 tag.
"afrikaans": "af", "af-za": "af",
# Korean
"korean": "ko", "한국어": "ko", "ko-kr": "ko",
# Italian
"italian": "it", "italiano": "it", "it-it": "it", "it-ch": "it",
# Irish (Gaeilge) — ga is the BCP-47 code
"irish": "ga", "gaeilge": "ga", "ga-ie": "ga",
# Portuguese — bare "portuguese" routes to European Portuguese; pt-br
# is in the same family but rendered identically here (no separate br catalog).
"portuguese": "pt", "português": "pt", "portugues": "pt",
"pt-pt": "pt", "pt-br": "pt", "brazilian": "pt", "brasileiro": "pt",
# Russian
"russian": "ru", "русский": "ru", "ru-ru": "ru",
# Hungarian
"hungarian": "hu", "magyar": "hu", "hu-hu": "hu",
# Arabic — bare "arabic"/endonym plus the common regional BCP-47 tags.
"arabic": "ar", "العربية": "ar",
"ar-sa": "ar", "ar-eg": "ar", "ar-ae": "ar", "ar-ma": "ar", "ar-dz": "ar",
}
@@ -89,18 +54,10 @@ _catalog_lock = threading.Lock()
def _locales_dir() -> Path:
"""Return the directory containing locale YAML files.
"""Locale dir: ``HERMES_BUNDLED_LOCALES`` (sealed packaging, e.g. Nix) if it exists, else ``<repo-root>/locales``.
Resolution order, first existing wins:
1. ``HERMES_BUNDLED_LOCALES`` env var -- set by the Nix wrapper (or any
sealed-packaging system) to point at the installed catalog directory.
2. ``<repo-root>/locales`` -- source checkouts and editable installs,
where the working tree sits next to ``agent/``.
Falling through to the source-style path (even when missing) keeps
``_load_catalog`` error messages informative -- it logs the path it
looked at -- rather than raising.
The source path is returned even when missing so ``_load_catalog`` can log
the path it looked at rather than raise.
"""
override = os.getenv("HERMES_BUNDLED_LOCALES", "").strip()
if override:
@@ -112,19 +69,11 @@ def _locales_dir() -> Path:
"falling back to bundled/source locale resolution",
override,
)
# agent/i18n.py -> agent/ -> repo root (source checkout, editable install)
source_dir = Path(__file__).resolve().parent.parent / "locales"
return source_dir
return Path(__file__).resolve().parent.parent / "locales"
def _normalize_lang(value: Any) -> str:
"""Normalize a user-supplied language value to a supported code.
Accepts supported codes directly, common aliases (``chinese`` -> ``zh``),
and case-insensitive regional tags (``zh-CN`` -> ``zh``). Returns the
default language for unknown values.
"""
"""Map a user-supplied value (code, alias, or regional tag like ``zh-CN``) to a supported code, else default."""
if not isinstance(value, str):
return DEFAULT_LANGUAGE
key = value.strip().lower()
@@ -134,20 +83,20 @@ def _normalize_lang(value: Any) -> str:
return key
if key in _LANGUAGE_ALIASES:
return _LANGUAGE_ALIASES[key]
# Try stripping a region suffix (e.g. "pt-br" -> "pt" won't be supported,
# but "zh-CN" -> "zh" will).
base = key.split("-", 1)[0]
base = key.split("-", 1)[0] # strip region suffix
if base in SUPPORTED_LANGUAGES:
return base
return DEFAULT_LANGUAGE
def _load_catalog(lang: str) -> dict[str, str]:
"""Load and flatten one locale YAML file into a dotted-key dict.
def _cache_catalog(lang: str, flat: dict[str, str]) -> dict[str, str]:
with _catalog_lock:
_catalog_cache[lang] = flat
return flat
YAML files can be nested for human readability; this produces the flat
key space :func:`t` expects. Cached per-language for the process.
"""
def _load_catalog(lang: str) -> dict[str, str]:
"""Load one locale YAML flattened to dotted keys; cached per language (empty dict on any failure)."""
with _catalog_lock:
cached = _catalog_cache.get(lang)
if cached is not None:
@@ -156,46 +105,34 @@ def _load_catalog(lang: str) -> dict[str, str]:
path = _locales_dir() / f"{lang}.yaml"
if not path.is_file():
logger.debug("i18n catalog missing for %s at %s", lang, path)
with _catalog_lock:
_catalog_cache[lang] = {}
return {}
return _cache_catalog(lang, {})
try:
import yaml # PyYAML is already a hermes dependency
import yaml
with path.open("r", encoding="utf-8") as f:
raw = yaml.safe_load(f) or {}
except Exception as exc:
logger.warning("Failed to load i18n catalog %s: %s", path, exc)
with _catalog_lock:
_catalog_cache[lang] = {}
return {}
return _cache_catalog(lang, {})
flat: dict[str, str] = {}
_flatten_into(raw, "", flat)
with _catalog_lock:
_catalog_cache[lang] = flat
return flat
return _cache_catalog(lang, flat)
def _flatten_into(node: Any, prefix: str, out: dict[str, str]) -> None:
# Non-string, non-dict leaves are ignored -- catalogs are text-only.
if isinstance(node, dict):
for key, value in node.items():
child_key = f"{prefix}.{key}" if prefix else str(key)
_flatten_into(value, child_key, out)
elif isinstance(node, str):
out[prefix] = node
# Non-string, non-dict leaves are ignored -- catalogs are text-only.
@lru_cache(maxsize=1)
def _config_language_cached() -> str | None:
"""Read ``display.language`` from config.yaml once per process.
Cached because ``t()`` is called in hot paths (every approval prompt,
every gateway reply) and re-reading YAML each call would be wasteful.
``reset_language_cache()`` clears this when config changes at runtime
(e.g. after the setup wizard).
"""
"""``display.language`` from config.yaml, read once per process (``t()`` is a hot path)."""
try:
from hermes_cli.config import load_config_readonly
cfg = load_config_readonly()
@@ -208,11 +145,7 @@ def _config_language_cached() -> str | None:
def reset_language_cache() -> None:
"""Invalidate cached language resolution and catalogs.
Call after :func:`hermes_cli.config.save_config` if a running process
needs to pick up a changed ``display.language`` without restart.
"""
"""Invalidate cached language resolution and catalogs (call after ``save_config`` changes ``display.language``)."""
_config_language_cached.cache_clear()
with _catalog_lock:
_catalog_cache.clear()
@@ -223,41 +156,22 @@ def get_language() -> str:
env_lang = os.environ.get("HERMES_LANGUAGE")
if env_lang:
return _normalize_lang(env_lang)
cfg_lang = _config_language_cached()
if cfg_lang:
return cfg_lang
return DEFAULT_LANGUAGE
return _config_language_cached() or DEFAULT_LANGUAGE
def t(key: str, lang: str | None = None, **format_kwargs: Any) -> str:
"""Translate a dotted key to the active language.
"""Translate a dotted catalog key to the active (or explicit ``lang``) language.
Parameters
----------
key
Dotted path into the catalog, e.g. ``"approval.choose_long"``.
lang
Explicit language override. Takes precedence over env + config.
**format_kwargs
``str.format`` substitution arguments (``t("gateway.drain", count=3)``
expects a catalog entry with a ``{count}`` placeholder).
Returns
-------
The translated string, or the English fallback if the key is missing in
the target language, or the bare key if English is also missing.
``format_kwargs`` are applied with ``str.format``. Falls back to English,
then to the bare key; a format failure returns the unformatted string.
"""
target = _normalize_lang(lang) if lang else get_language()
catalog = _load_catalog(target)
value = catalog.get(key)
value = _load_catalog(target).get(key)
if value is None and target != DEFAULT_LANGUAGE:
# Fall through to English rather than showing a key path to the user.
value = _load_catalog(DEFAULT_LANGUAGE).get(key)
if value is None:
# Last-ditch: return the key itself. A broken catalog should not
# crash anything; it just looks ugly until someone fixes it.
logger.debug("i18n miss: key=%r lang=%r", key, target)
value = key
+72 -155
View File
@@ -1,30 +1,15 @@
"""CJK/wide-character-aware re-alignment of model-emitted markdown tables.
Models pad markdown tables assuming each character occupies one terminal
cell. CJK glyphs and most emoji render as two cells, so the model's
spacing collapses into drift the moment a table reaches a real terminal —
header pipes line up, every body row drifts right by N cells per CJK
char.
Models pad tables assuming one cell per character; CJK glyphs and most emoji
take two, so body rows drift right on real terminals. This rebuilds padding
with ``wcwidth.wcswidth`` while preserving pipes/dashes so the table still reads
as plain text in ``strip``/unrendered modes (Rich already aligns CJK itself).
This module rebuilds row padding using ``wcwidth.wcswidth`` (display
columns), preserving the table's pipes and dashes so it still reads as a
plain-text table in ``strip`` / unrendered display modes. Standard Rich
markdown rendering already aligns CJK correctly inside a wide enough
panel; this helper is for the paths that print the model's text more or
less verbatim.
The helper is deliberately conservative:
* Only contiguous ``| ... |`` blocks with a divider line are rewritten.
* Anything that does not look like a table is passed through unchanged.
* Single-line / mid-stream fragments are left alone — callers buffer
table rows and flush them once the block is complete.
There is a small, intentional caveat: ``wcwidth`` returns ``-1`` for some
emoji-with-variation-selector sequences (e.g. ``⚠️``); we clamp those to
0 so they do not corrupt the column width math. The 1-cell drift on
those specific glyphs is preferable to silently widening every table
that contains one.
Deliberately conservative: only contiguous ``| ... |`` blocks with a divider are
rewritten; everything else passes through; single-line/mid-stream fragments are
left alone (callers buffer rows and flush complete blocks). ``wcwidth`` returns
``-1`` for some emoji+variation-selector sequences (``⚠️``); those clamp to 0 —
a 1-cell drift on that glyph beats widening every table that contains one.
"""
from __future__ import annotations
@@ -47,13 +32,7 @@ _MIN_COL_WIDTH = 3 # matches the divider's minimum dash run.
def _disp_width(s: str) -> int:
"""``wcswidth`` clamped to a non-negative integer.
``wcswidth`` returns ``-1`` when it encounters a control char or an
unknown sequence; treat those as zero-width rather than letting a
negative number flow into ``max`` and break the column-width math.
"""
"""``wcswidth`` clamped to >= 0 (it returns -1 for control/unknown sequences)."""
w = wcswidth(s)
return w if w > 0 else 0
@@ -64,7 +43,6 @@ def _pad_to_width(s: str, target: int) -> str:
def split_table_row(row: str) -> List[str]:
"""Split ``| a | b | c |`` into ``["a", "b", "c"]`` with trims."""
s = row.strip()
if s.startswith("|"):
s = s[1:]
@@ -75,7 +53,6 @@ def split_table_row(row: str) -> List[str]:
def is_table_divider(row: str) -> bool:
"""True when ``row`` is a markdown table separator line."""
cells = split_table_row(row)
return len(cells) > 1 and all(_DIVIDER_CELL_RE.match(c) for c in cells)
@@ -83,76 +60,70 @@ def is_table_divider(row: str) -> bool:
def looks_like_table_row(row: str) -> bool:
"""True when ``row`` could plausibly be a markdown table row.
Used by streaming callers to decide whether to buffer an in-flight
line. We are intentionally permissive here — the realigner itself
only rewrites blocks that are accompanied by a divider, so a false
positive here at most delays the print of one line.
Intentionally permissive for streaming callers deciding whether to buffer a
line: the realigner only rewrites divider-backed blocks, so a false positive
at most delays printing one line. A leading pipe is the strongest signal;
without it we accept >= 2 pipes so models that omit the leading pipe still match.
"""
if "|" not in row:
return False
stripped = row.strip()
if not stripped:
return False
# A leading pipe is the strongest signal; without it we still allow
# rows with at least two pipes so models that omit the leading pipe
# don't slip past us.
if stripped.startswith("|"):
return True
return stripped.count("|") >= 2
return stripped.startswith("|") or stripped.count("|") >= 2
def _render_block(rows: List[List[str]], available_width: int | None = None) -> List[str]:
"""Render ``rows`` (header + body, divider implied) at uniform widths.
If ``available_width`` is given and the rebuilt horizontal table
would exceed it, fall back to a vertical key-value rendering so
rows do not soft-wrap mid-cell — terminal soft-wrap destroys
column alignment visually even when the underlying bytes are
perfectly padded, which is exactly the "tables look broken"
user report this code path is meant to address.
When the horizontal table would exceed ``available_width`` fall back to a
vertical key-value rendering: terminal soft-wrap mid-cell destroys alignment
visually even when the bytes are perfectly padded.
"""
ncols = max(len(r) for r in rows)
rows = [r + [""] * (ncols - len(r)) for r in rows]
widths = [max(_MIN_COL_WIDTH, *(_disp_width(r[c]) for r in rows)) for c in range(ncols)]
widths = [
max(_MIN_COL_WIDTH, *(_disp_width(r[c]) for r in rows))
for c in range(ncols)
]
# Total horizontal width for the rendered row:
# `| ` + cell + ` ` for each column, plus the final closing `|`.
# `| ` + cell + ` ` per column, plus the closing `|`.
horizontal_width = sum(widths) + 3 * ncols + 1
if available_width is not None and horizontal_width > max(available_width, 20):
return _render_vertical(rows, ncols, available_width)
def _row(cells: List[str]) -> str:
return (
"| "
+ " | ".join(_pad_to_width(c, widths[k]) for k, c in enumerate(cells))
+ " |"
)
return "| " + " | ".join(_pad_to_width(c, widths[k]) for k, c in enumerate(cells)) + " |"
out = [_row(rows[0])]
out.append("|" + "|".join("-" * (w + 2) for w in widths) + "|")
for r in rows[1:]:
out.append(_row(r))
out = [_row(rows[0]), "|" + "|".join("-" * (w + 2) for w in widths) + "|"]
out.extend(_row(r) for r in rows[1:])
return out
def _hard_break(word: str, w: int) -> List[str]:
"""Split a single over-wide word into display-width-``w`` chunks."""
out: List[str] = []
buf = ""
bw = 0
for ch in word:
cw = _disp_width(ch) or 1
if bw + cw > w and buf:
out.append(buf)
buf = ch
bw = cw
else:
buf += ch
bw += cw
if buf:
out.append(buf)
return out
def _wrap_to_width(text: str, width: int) -> List[str]:
"""Soft-wrap ``text`` at word boundaries to fit ``width`` display cells.
"""Soft-wrap ``text`` at word boundaries to ``width`` display cells.
Falls back to hard-breaking the longest word if a single token is
wider than ``width``. Empty input yields a single empty string so
the caller's row count stays predictable.
Words wider than ``width`` are hard-broken. Empty input yields a single
empty string so the caller's row count stays predictable.
"""
if width <= 0 or not text:
return [text]
words = text.split()
if not words:
return [""]
@@ -161,100 +132,61 @@ def _wrap_to_width(text: str, width: int) -> List[str]:
current = ""
current_w = 0
def _hard_break(word: str, w: int) -> List[str]:
out: List[str] = []
buf = ""
bw = 0
for ch in word:
cw = _disp_width(ch) or 1
if bw + cw > w and buf:
out.append(buf)
buf = ch
bw = cw
else:
buf += ch
bw += cw
if buf:
out.append(buf)
return out
def _start(word: str, ww: int) -> None:
nonlocal current, current_w
if ww <= width:
current, current_w = word, ww
else:
pieces = _hard_break(word, width)
lines.extend(pieces[:-1])
current = pieces[-1] if pieces else ""
current_w = _disp_width(current)
for word in words:
ww = _disp_width(word)
if not current:
if ww <= width:
current = word
current_w = ww
else:
pieces = _hard_break(word, width)
lines.extend(pieces[:-1])
current = pieces[-1] if pieces else ""
current_w = _disp_width(current)
continue
if current_w + 1 + ww <= width:
_start(word, ww)
elif current_w + 1 + ww <= width:
current += " " + word
current_w += 1 + ww
else:
lines.append(current)
if ww <= width:
current = word
current_w = ww
else:
pieces = _hard_break(word, width)
lines.extend(pieces[:-1])
current = pieces[-1] if pieces else ""
current_w = _disp_width(current)
_start(word, ww)
if current:
lines.append(current)
return lines or [""]
def _render_vertical(
rows: List[List[str]], ncols: int, available_width: int
) -> List[str]:
"""Render a too-wide table as vertical ``Header: value`` rows.
def _render_vertical(rows: List[List[str]], ncols: int, available_width: int) -> List[str]:
"""Render a too-wide table as ``Header: value`` blocks (Claude Code's narrow fallback).
Mirrors Claude Code's narrow-terminal fallback in
``MarkdownTable.tsx``: each body row becomes a small block of
``Header: cell-value`` lines (continuation lines indented two
spaces) separated by a thin ``─`` divider between rows. Keeps
every line narrower than ``available_width`` so the terminal does
not soft-wrap mid-cell.
Each body row becomes one block with continuation lines indented two spaces,
blocks separated by a thin ``─`` rule; every line stays under ``available_width``.
"""
if not rows:
return []
headers = rows[0] + [""] * (ncols - len(rows[0]))
body = rows[1:]
labels = [h or f"Column {i + 1}" for i, h in enumerate(headers)]
sep_width = max(20, min(40, available_width - 2)) if available_width else 30
separator = "─" * sep_width
indent = " "
indent_w = _disp_width(indent)
cont_budget = max(10, available_width - _disp_width(indent))
out: List[str] = []
for ri, row in enumerate(body):
for ri, row in enumerate(rows[1:]):
if ri > 0:
out.append(separator)
for ci in range(ncols):
label = labels[ci]
value = row[ci] if ci < len(row) else ""
label_w = _disp_width(label)
first_budget = max(10, available_width - label_w - 2)
cont_budget = max(10, available_width - indent_w)
if not value:
out.append(f"{label}:")
continue
wrapped = _wrap_to_width(value, first_budget)
wrapped = _wrap_to_width(value, max(10, available_width - _disp_width(label) - 2))
out.append(f"{label}: {wrapped[0]}")
if len(wrapped) > 1:
# Re-flow continuation text at the wider continuation
# budget — words split across the narrower first-line
# budget should re-pack greedily for the rest.
cont_text = " ".join(wrapped[1:])
for cl in _wrap_to_width(cont_text, cont_budget):
# Re-flow continuation text at the wider continuation budget.
for cl in _wrap_to_width(" ".join(wrapped[1:]), cont_budget):
if cl.strip():
out.append(f"{indent}{cl}")
return out
@@ -263,16 +195,10 @@ def _render_vertical(
def realign_markdown_tables(text: str, available_width: int | None = None) -> str:
"""Rewrite every ``| ... |`` + divider block with wcwidth-aware padding.
Lines that are not part of a recognised table are returned verbatim,
so this is safe to apply to arbitrary assistant prose.
If ``available_width`` is given (terminal cells available for the
rendered table), tables wider than that are rendered as vertical
key-value pairs instead of a horizontal pipe-bordered grid. This
avoids the terminal soft-wrapping mid-cell, which destroys column
alignment visually even when the bytes are perfectly padded.
Non-table lines are returned verbatim, so this is safe on arbitrary prose.
With ``available_width`` (terminal cells), tables wider than that render as
vertical key-value pairs instead of soft-wrapping mid-cell.
"""
if "|" not in text:
return text
@@ -280,30 +206,21 @@ def realign_markdown_tables(text: str, available_width: int | None = None) -> st
out: List[str] = []
i = 0
n = len(lines)
while i < n:
line = lines[i]
# A table starts with a header row whose next line is a divider.
if (
"|" in line
and i + 1 < n
and is_table_divider(lines[i + 1])
):
if "|" in line and i + 1 < n and is_table_divider(lines[i + 1]):
header = split_table_row(line)
body: List[List[str]] = []
j = i + 2
while j < n and "|" in lines[j] and lines[j].strip():
if is_table_divider(lines[j]):
j += 1
continue
body.append(split_table_row(lines[j]))
if not is_table_divider(lines[j]):
body.append(split_table_row(lines[j]))
j += 1
if any(c for c in header) or body:
out.extend(_render_block([header] + body, available_width))
i = j
continue
out.append(line)
i += 1
return "\n".join(out)
File diff suppressed because it is too large Load Diff
+99 -203
View File
@@ -1,43 +1,24 @@
"""Native OpenAI Responses server-side compaction — gpt-5.6 on direct OpenAI routes only.
OpenAI's Responses API supports server-side compaction: include
``context_management=[{"type": "compaction", "compact_threshold": N}]`` in a
``/v1/responses`` request and, when the rendered input crosses N tokens, the
server summarizes older context into an opaque ``compaction`` output item
(``encrypted_content``, sealed to the issuing endpoint). Replaying that item
as an input item on later requests stands in for the pruned history, so the
model keeps long-horizon recall without the client ever seeing a summary.
Docs: https://developers.openai.com/api/docs/guides/compaction
Including ``context_management=[{"type": "compaction", "compact_threshold": N}]``
in a ``/v1/responses`` request makes the server summarize older context into an
opaque ``compaction`` item (``encrypted_content``, sealed to the issuing
endpoint) once the input crosses N tokens; replaying that item stands in for
the pruned history. Docs: https://developers.openai.com/api/docs/guides/compaction
Hermes' support is deliberately narrow (live verification, Aug 2026):
Support is deliberately narrow (live-verified):
* gpt-5.6 family only — gpt-5.1/5.2 fail server-side (HTTP 500 blocking, a
permanent stall streaming) with no structured "unsupported" rejection, so an
explicit model-family check is the only safe gate.
* Direct OpenAI routes only (api.openai.com or the ChatGPT Codex backend) —
other Responses surfaces would 400 on the field and cannot mint/decrypt the blob.
* **gpt-5.6 family only.** gpt-5.6 and its variants compact correctly.
Sending the field to gpt-5.1 / gpt-5.2 reliably fails server-side —
HTTP 500 on the blocking path and a permanent stall on the streaming
path (90s watchdog x 3 retries = a dead turn). There is no structured
"unsupported" rejection to downgrade on, so the only safe gate is an
explicit model-family check.
* **Direct OpenAI routes only:** api.openai.com (API key) or the ChatGPT
Codex backend (subscription OAuth). Every other Responses surface
(xAI, GitHub/Copilot, relays, local servers) never sees the field —
most would 400 on the unknown parameter, and none can mint or decrypt
the compaction blob.
Ownership model: Hermes' local compression stays fully armed as the
fallback owner. The native threshold is clamped safely below the local
compressor's trigger so the server compacts first; if it doesn't (native
disabled mid-session, provider hiccup, non-eligible route), the local
summarizer fires exactly as before. There is no new custody state — the
captured compaction items ride the existing ``codex_reasoning_items``
sidecar, which already handles persistence (state.db), gateway session
replay, cross-issuer stamping, and the encrypted-replay kill switch.
This module stays free of transport/adapter dependencies so the transport,
adapter, and conversation loop can share the gate without import cycles. The
two exceptions — ``agent.context_compressor`` and ``agent.message_content`` —
sit below this module in the dependency graph (neither imports
``native_compaction``), so importing their provenance/text primitives here
introduces no cycle.
Hermes' local compressor stays armed as fallback owner: the native threshold is
clamped below the local trigger so the server compacts first, and captured
compaction items ride the existing ``codex_reasoning_items`` sidecar (persistence,
replay, cross-issuer stamping, kill switch). This module stays free of
transport/adapter imports so transport, adapter, and loop share the gate
without cycles; ``context_compressor`` and ``message_content`` sit below it.
"""
from __future__ import annotations
@@ -52,14 +33,11 @@ from agent.message_content import flatten_message_text
logger = logging.getLogger(__name__)
# Native compaction fires this many tokens below the local compressor's
# trigger so the server always gets the first shot at compaction.
# trigger so the server always gets the first shot.
LOCAL_TRIGGER_SAFETY_MARGIN = 8_192
# Deterministic fallback when automatic mode cannot inspect a local trigger.
# Fallback when automatic mode has no local trigger to follow.
DEFAULT_COMPACT_THRESHOLD = 200_000
# Model-family gate. Substring match on the lowercased model id so dated
# snapshots (gpt-5.6-2026-07-xx) and variants (gpt-5.6-mini) stay eligible.
# Substring match so dated snapshots and variants (gpt-5.6-mini) stay eligible.
_ELIGIBLE_MODEL_MARKER = "gpt-5.6"
@@ -77,11 +55,10 @@ def resolve_native_compaction_capabilities(
) -> Dict[str, bool]:
"""Resolve the native-compaction capability for a runtime destination.
The result is deliberately explicit: a resolved ``False`` is different
from an unresolved capability and must survive model switches unchanged.
A resolved ``False`` is distinct from "unresolved" and must survive model
switches unchanged.
"""
normalized_provider = (provider or "").strip().lower()
direct_default = normalized_provider == "openai" and not base_url
direct_default = (provider or "").strip().lower() == "openai" and not base_url
eligible = is_native_compaction_model(model) and (
direct_default
or is_direct_openai_route(base_url, is_codex_backend=is_codex_backend)
@@ -110,10 +87,10 @@ def resolve_compact_threshold(
) -> int:
"""Resolve automatic mode or clamp an explicit native threshold.
An omitted or invalid setting follows the resolved local compressor trigger.
An explicit positive integer remains absolute unless it must be clamped so
native compaction fires first. ``local_trigger_tokens`` is
``ContextCompressor.threshold_tokens`` when a compressor is attached.
An omitted/invalid setting follows the local compressor trigger
(``ContextCompressor.threshold_tokens``) minus the safety margin. An
explicit positive integer is absolute unless it must be clamped so native
compaction fires first. Booleans are never thresholds.
"""
local = None
try:
@@ -139,7 +116,7 @@ def resolve_compact_threshold(
)
except (TypeError, ValueError):
configured = None
if isinstance(configured_threshold, bool) or configured is None or configured <= 0:
if configured is None or configured <= 0:
return upper if upper is not None else DEFAULT_COMPACT_THRESHOLD
if upper is None:
return configured
@@ -150,11 +127,7 @@ _checkpoint_suppression_logged = False
def _warn_native_compaction_suppressed_by_checkpoint_gate() -> None:
"""Log once per process that the checkpoint gate suppresses native compaction.
The suppression itself is re-evaluated per request; only the log line is
deduplicated so a long session does not repeat it on every API call.
"""
"""Log once per process; the suppression itself is re-evaluated per request."""
global _checkpoint_suppression_logged
if _checkpoint_suppression_logged:
return
@@ -175,27 +148,22 @@ def native_compaction_context_management(
) -> Optional[List[Dict[str, Any]]]:
"""Return the ``context_management`` payload for this request, or None.
None means "do not send the field" — the request is byte-identical to
pre-feature behavior. All gates are re-checked per request so a
mid-session model switch or the in-session kill switch
(``agent.codex_responses_native_compaction = False``, set by the
conversation loop's rejection recovery) takes effect on the next call.
None means "do not send the field" (request byte-identical to pre-feature).
Every gate is re-checked per request so a mid-session model switch or the
in-session kill switch (``agent.codex_responses_native_compaction = False``,
set by rejection recovery) takes effect on the next call.
"""
capabilities = getattr(agent, "runtime_capabilities", None)
if isinstance(capabilities, dict):
if not bool(capabilities.get("native_compaction", False)):
return None
if not bool(getattr(agent, "codex_responses_native_compaction", False)):
if isinstance(capabilities, dict) and not capabilities.get("native_compaction", False):
return None
# compression.enabled: false disables ALL automatic compaction, native
# included — mirrors the codex_app_server_auto contract.
if not bool(getattr(agent, "compression_enabled", True)):
if not getattr(agent, "codex_responses_native_compaction", False):
return None
# compression.checkpoint_required: server-side compaction is a lossy
# boundary the provider owns — no pre-compress checkpoint can run before
# the server replaces older context. Keep the checkpoint-aware Hermes
# compressor authoritative instead of silently letting the server
# compact. Explicit-True check matches the compress_context() gate.
# compression.enabled: false disables ALL automatic compaction, native included.
if not getattr(agent, "compression_enabled", True):
return None
# Server-side compaction is a lossy boundary the provider owns — no
# pre-compress checkpoint can run first — so the checkpoint-aware Hermes
# compressor stays authoritative. Explicit-True matches compress_context().
if getattr(agent, "compression_checkpoint_required", False) is True:
_warn_native_compaction_suppressed_by_checkpoint_gate()
return None
@@ -219,13 +187,10 @@ def native_compaction_context_management(
return [{"type": "compaction", "compact_threshold": threshold}]
# Retention budget for plaintext user messages carried across a native
# compaction boundary (mirrors Codex CLI's RETAINED_MESSAGE_TOKEN_BUDGET).
# Live verification (Aug 2026, gpt-5.6 @ api.openai.com): the server renders
# Retention budgets for plaintext user messages / local compression summaries
# carried across a native compaction boundary (mirrors Codex CLI's
# RETAINED_MESSAGE_TOKEN_BUDGET; the summary budget prevents summary inflation).
RETAINED_USER_MESSAGE_TOKEN_BUDGET = 64_000
# Retention budget for local compression summary messages carried across a native
# compaction boundary to prevent summary token inflation.
RETAINED_SUMMARY_TOKEN_BUDGET = 32_000
@@ -235,12 +200,7 @@ def _approx_tokens(text: str) -> int:
def _extract_item_text(item: Any) -> Optional[str]:
"""Extract measurable text from message content and fallback fields.
Returns None when the item carries no measurable text. Handles string
content, multipart lists (input_text/text/output_text), and nested
metadata text.
"""
"""Measurable text from a Responses item (string/multipart/metadata), or None."""
if not isinstance(item, dict):
return None
@@ -272,12 +232,10 @@ def _extract_item_text(item: Any) -> Optional[str]:
def _has_retainable_image_content(item: Any) -> bool:
"""Return True for a converted Responses message with a valid image part.
"""True for a converted Responses message with a valid ``input_image`` part.
The pruning boundary receives normalized Responses items, so only the
adapter-owned ``input_image`` shape is authority here. Unknown, malformed,
or empty multipart placeholders must not become durable history merely
because their list is non-empty.
Only the adapter-owned ``input_image`` shape counts: unknown or empty
multipart placeholders must not become durable history for being non-empty.
"""
if not isinstance(item, dict):
return False
@@ -295,26 +253,11 @@ def _has_retainable_image_content(item: Any) -> bool:
return False
def _is_summary_item(item: Any) -> bool:
"""True when *item* is a canonical Hermes compression-summary message.
Delegates entirely to
``agent.context_compressor.is_compaction_summary_message`` — the single
authoritative provenance check already used by every other summary
consumer (memory providers, frontends, the compactor itself). It prefers
the exact, truthy ``COMPRESSED_SUMMARY_METADATA_KEY`` marker and falls
back to the canonical prefix classifier (``SUMMARY_PREFIX`` /
``LEGACY_SUMMARY_PREFIX`` / historical prefixes, including the
merge-into-tail shape) for the case where the underscore-prefixed key
was already stripped by a wire sanitizer.
Deliberately NOT a second heuristic: no arbitrary underscore-key scan, no
inference from a falsy or unrelated metadata key, and no matching on
ad-hoc content headings like ``"## Summary"`` in ordinary text — any of
those can promote a normal user/assistant message (or adversarial
content) to durable retained history (#90975 review).
"""
return is_compaction_summary_message(item)
# Canonical provenance check (metadata marker, then canonical prefix classifier).
# Deliberately NOT a second heuristic: no underscore-key scan, no matching on
# ad-hoc headings — either could promote ordinary or adversarial content to
# durable retained history.
_is_summary_item = is_compaction_summary_message
def prune_pre_checkpoint_items(
@@ -326,48 +269,29 @@ def prune_pre_checkpoint_items(
) -> List[Dict[str, Any]]:
"""Restructure Responses input around the newest compaction checkpoint.
The server drops every input item that precedes a replayed ``compaction``
item (live-verified Aug 2026), so sending pre-checkpoint history is dead
weight AND silently erases the user's plaintext asks — including any
local-compression summary the agent already produced, which previously
vanished here because it carries ``role="assistant"``, not ``"user"``
(#90975). When a checkpoint is present, rebuild the wire as::
The server drops every input item preceding a replayed ``compaction`` item,
which silently erases the user's plaintext asks and any local-compression
summary (``role="assistant"``). With a checkpoint present, rebuild as::
[checkpoint run] + [retained user & summary messages (newest-first budget)] + [post]
- The NEWEST contiguous run of checkpoints wins.
- Retained user messages are kept verbatim within
``retained_user_token_budget``; the boundary message is head-truncated
when it only partially fits (string content only) — goals are usually
stated up front, so the head is the valuable end. A recognized
- User messages are kept verbatim within ``retained_user_token_budget``;
the boundary message is head-truncated when it only partially fits
(string content only — goals are stated up front). A recognized
image-only user message is retained whole at one-token cost.
- Compression summary messages (``_is_summary_item``, the canonical
``agent.context_compressor`` provenance check) are retained whole
within ``retained_summary_token_budget``. A summary is never
byte/character-sliced: Hermes summaries carry structural framing
(handoff prefix, end marker, merge-into-tail delimiters) that a blind
slice can corrupt, so one that doesn't fit whole is dropped instead.
A summary already retained once (identical text) is never duplicated,
so repeated checkpoints stay idempotent.
- ``enable_summary_retention`` is a function-level override (used by
tests and callers that need the pre-#90975 behavior back); it is not
wired to a user-facing config surface.
- Original relative chronological order between user messages and
summaries is preserved.
- ``item_sources`` (optional, parallel to ``items``) is the raw chat
message each Responses item was converted from. By the time a summary
reaches this function as a converted ``item`` it can already be lossy:
a merge-into-tail tool-result carrier becomes a typed
``function_call_output`` (no ``content``/``role`` survives the
conversion at all), and a merge-into-tail assistant carrier can be
shadowed by a stale exact ``codex_message_items`` replay captured
before the merge rewrote its content. When a source is provided and is
itself a canonical summary carrier (``is_compaction_summary_message``),
its content is read directly from the source — never from the
converted item — and it is retained as a synthesized
``role="assistant"`` message regardless of what shape the original
item took. Without ``item_sources`` (default), retention only sees
what survived conversion, matching pre-#90976 behavior (#90976).
- Summaries are retained whole within ``retained_summary_token_budget`` and
never sliced (their structural framing would corrupt); one that doesn't
fit is dropped. Identical summary text is never retained twice.
- Relative order between user messages and summaries is preserved.
- ``item_sources`` (parallel to ``items``) is the raw chat message each item
was converted from. Conversion can be lossy for summaries (a
merge-into-tail carrier becomes a typed ``function_call_output``, or an
assistant carrier is shadowed by a stale exact replay), so when a source
is itself a canonical summary carrier its content is read from the
SOURCE and retained as a synthesized ``role="assistant"`` message.
- ``enable_summary_retention`` is a function-level override for tests, not
a config surface.
"""
if not isinstance(items, list) or not items:
return items
@@ -379,7 +303,6 @@ def prune_pre_checkpoint_items(
if last_cp is None:
return items
# Extend backwards over the contiguous run ending at last_cp.
first_cp = last_cp
while (
first_cp > 0
@@ -403,14 +326,12 @@ def prune_pre_checkpoint_items(
seen_summary_texts: set = set()
def _try_retain_summary(text: Optional[str]) -> Optional[Dict[str, Any]]:
"""Check budget/dedup/cost for a summary; return cost info or None."""
"""Budget/dedup check for a summary; return its cost or None."""
if not text or summary_remaining <= 0 or text in seen_summary_texts:
return None
cost = _approx_tokens(text)
if cost > summary_remaining:
# Never byte-slice a summary's structural framing — drop it
# whole rather than corrupt the handoff prefix / end marker.
return None
return None # never slice a summary's structural framing
seen_summary_texts.add(text)
return {"cost": cost}
@@ -418,15 +339,10 @@ def prune_pre_checkpoint_items(
if not isinstance(item, dict):
continue
# Canonical source-based summary detection: reads the ORIGINAL chat
# message's own content, so it sees past a lossy conversion (a
# typed `function_call_output` wrapper, or a stale exact-replay
# message) that erased the summary from `item` itself (#90976).
# This is never a heuristic promotion of arbitrary item content —
# it only fires when the source message itself is a canonical,
# provenance-tagged summary carrier.
# Source-based detection sees past a lossy conversion; it only fires
# when the source itself is a provenance-tagged summary carrier.
if enable_summary_retention and isinstance(source, dict) and _is_summary_item(source):
text = flatten_message_text(source.get("content")) if isinstance(source, dict) else ""
text = flatten_message_text(source.get("content"))
text = text if text.strip() else None
result = _try_retain_summary(text)
if result:
@@ -438,9 +354,7 @@ def prune_pre_checkpoint_items(
summary_remaining -= result["cost"]
continue
# Skip typed non-message items (function_call_output etc. never
# carry role=user or a summary flag, but stay defensive about
# future shapes).
# Typed non-message items never carry role=user or a summary flag.
if "type" in item and item.get("type") != "message":
continue
@@ -476,8 +390,7 @@ def prune_pre_checkpoint_items(
retained_reversed.append(truncated)
user_remaining = 0
retained_ordered = list(reversed(retained_reversed))
result = checkpoint_run + retained_ordered + post
result = checkpoint_run + list(reversed(retained_reversed)) + post
logger.debug(
"Pruned pre-checkpoint items: %d input -> %d retained (user_rem=%d, summary_rem=%d)",
@@ -490,25 +403,21 @@ def prune_pre_checkpoint_items(
return result
_REJECTION_MARKERS = (
"unknown", "unsupported", "invalid", "unexpected", "not permitted",
"not allowed", "unrecognized", "extra field", "no such", "bad request",
"not supported",
)
def is_native_compaction_rejection(error: Any, status_code: Any = None) -> bool:
"""True when a provider error is a STRUCTURED rejection of the
context_management field.
"""True when a provider error is a STRUCTURED rejection of ``context_management``.
Used by the conversation loop's one-shot recovery: strip the field,
disable native compaction for the rest of the session, retry. Matching
is deliberately narrow — a transient 5xx/timeout whose body merely
ECHOES the request (and therefore contains the field name) must NOT
permanently downgrade native compaction for the session (#82777).
Two conditions, both required when a status is known:
* ``status_code`` is 400 (or unknown/None — some transports surface
only a message string; field-name matching alone is then the best
available signal, preserving pre-#82777 behavior for them), and
* the error text names ``context_management`` / ``compact_threshold``
alongside rejection language ("unknown", "unsupported", "invalid",
"unexpected", "not permitted"...). A bare field-name echo without
rejection language does not match.
Drives the loop's one-shot recovery (strip the field, disable for the
session, retry), so matching is narrow: a transient 5xx whose body merely
ECHOES the request must not permanently downgrade native compaction. Requires
``status_code`` 400 (or unknown — some transports surface only a message)
AND the field name alongside rejection language.
"""
text = str(error or "").lower()
if "context_management" not in text and "compact_threshold" not in text:
@@ -519,23 +428,15 @@ def is_native_compaction_rejection(error: Any, status_code: Any = None) -> bool:
return False
except (TypeError, ValueError):
pass
rejection_markers = (
"unknown", "unsupported", "invalid", "unexpected", "not permitted",
"not allowed", "unrecognized", "extra field", "no such", "bad request",
"not supported",
)
return any(marker in text for marker in rejection_markers)
return any(marker in text for marker in _REJECTION_MARKERS)
def has_compaction_checkpoint(items: Any) -> bool:
"""Does this ``codex_reasoning_items`` sidecar carry a compaction checkpoint?
A ``type: "compaction"`` item is the server-side stand-in for history that
has already been pruned — cumulative context, not per-turn reasoning. It
rides the same sidecar as ordinary reasoning items, so anything that
rewrites or discards that sidecar (or the message carrying it) has to ask
this question first: the checkpoint exists in exactly one place, and the
request that loses it loses the compacted history with it.
A ``type: "compaction"`` item is cumulative context, not per-turn
reasoning, and exists in exactly one place: anything that rewrites or
discards the sidecar must ask this first or lose the compacted history.
"""
return any(
isinstance(item, dict) and item.get("type") == "compaction"
@@ -547,16 +448,11 @@ def merge_interim_reasoning_items(
prior_items: Any,
new_items: Any,
) -> List[Dict[str, Any]]:
"""Merge ``codex_reasoning_items`` across Codex incomplete-continuation
dedup, preserving native compaction checkpoints.
"""Merge ``codex_reasoning_items`` across Codex incomplete-continuation dedup.
The incomplete-retry path updates a visually-duplicate interim assistant
message in place with the newer response's replay payload. A checkpoint
captured on the EARLIER response is a cumulative context carrier the
continuation won't re-emit (the replayed checkpoint keeps the server
render under threshold), so a blind overwrite drops the only copy and the
next request balloons back to full history. Rule: newer items win, but
prior checkpoints are prepended unless the newer payload carries its own.
A checkpoint captured on the EARLIER response is not re-emitted by the
continuation, so a blind overwrite drops the only copy. Rule: newer items
win, but prior checkpoints are prepended unless the newer payload has its own.
"""
kept_checkpoints = [
item
+74 -110
View File
@@ -1,13 +1,9 @@
"""
Contextual first-touch onboarding hints.
"""Contextual first-touch onboarding hints.
Instead of blocking first-run questionnaires, show a one-time hint the *first*
time a user hits a behavior fork — message-while-running, first long-running
tool, etc. Each hint is shown once per install (tracked in ``config.yaml`` under
``onboarding.seen.<flag>``) and then never again.
Keep this module tiny and dependency-free so both the CLI and gateway can import
it without pulling in heavy modules.
Each hint is shown once per install the *first* time a user hits a behavior
fork (message-while-running, first long tool, ...), tracked in ``config.yaml``
under ``onboarding.seen.<flag>``. Kept tiny and dependency-free so both the CLI
and gateway can import it.
"""
from __future__ import annotations
@@ -19,79 +15,75 @@ from typing import Any, Mapping, Optional
logger = logging.getLogger(__name__)
# -------------------------------------------------------------------------
# Flag names (stable — used as config.yaml keys under onboarding.seen)
# -------------------------------------------------------------------------
BUSY_INPUT_FLAG = "busy_input_prompt"
TOOL_PROGRESS_FLAG = "tool_progress_prompt"
OPENCLAW_RESIDUE_FLAG = "openclaw_residue_cleanup"
PROFILE_BUILD_FLAG = "profile_build_offered"
# -------------------------------------------------------------------------
# Hint content
# -------------------------------------------------------------------------
# ── Hint content ──────────────────────────────────────────────────────────
# Busy-input hints are keyed by the effective busy_input_mode that was just
# applied so the message matches reality; "interrupt" is the default branch.
_BUSY_INPUT_HINTS_GATEWAY = {
"queue": (
"💡 First-time tip — I queued your message instead of interrupting. "
"Send `/busy interrupt` to make new messages stop the current task "
"immediately, or `/busy status` to check. This notice won't appear again."
),
"steer": (
"💡 First-time tip — I steered your message into the current run; "
"it will arrive after the next tool call instead of interrupting. "
"Send `/busy interrupt` or `/busy queue` to change this, or "
"`/busy status` to check. This notice won't appear again."
),
"redirect": (
"💡 First-time tip — I redirected the current run using your message. "
"Completed work stays in context, and `/stop` still cancels the task. "
"Send `/busy queue` to wait for a separate turn, or `/busy status` "
"to check. This notice won't appear again."
),
}
_BUSY_INPUT_HINT_GATEWAY_DEFAULT = (
"💡 First-time tip — I just interrupted my current task to answer you. "
"Send `/busy queue` to queue follow-ups for after the current task instead, "
"`/busy steer` to inject them mid-run without interrupting, or "
"`/busy status` to check. This notice won't appear again."
)
_BUSY_INPUT_HINTS_CLI = {
"queue": (
"(tip) Your message was queued for the next turn. "
"Use /busy interrupt to make Enter stop the current run instead, "
"or /busy steer to inject mid-run. This tip only shows once."
),
"steer": (
"(tip) Your message was steered into the current run; it arrives "
"after the next tool call. Use /busy interrupt or /busy queue to "
"change this. This tip only shows once."
),
"redirect": (
"(tip) Your correction redirected the current run without discarding "
"completed work. Use /stop to cancel or /busy queue to wait for a "
"separate turn. This tip only shows once."
),
}
_BUSY_INPUT_HINT_CLI_DEFAULT = (
"(tip) Your message interrupted the current run. "
"Use /busy queue to queue messages for the next turn instead, "
"or /busy steer to inject mid-run. This tip only shows once."
)
def busy_input_hint_gateway(mode: str) -> str:
"""Hint shown the first time a user messages while the agent is busy.
``mode`` is the effective busy_input_mode that was just applied, so the
message matches reality ("I just interrupted…" vs "I just queued…").
"""
if mode == "queue":
return (
"💡 First-time tip — I queued your message instead of interrupting. "
"Send `/busy interrupt` to make new messages stop the current task "
"immediately, or `/busy status` to check. This notice won't appear again."
)
if mode == "steer":
return (
"💡 First-time tip — I steered your message into the current run; "
"it will arrive after the next tool call instead of interrupting. "
"Send `/busy interrupt` or `/busy queue` to change this, or "
"`/busy status` to check. This notice won't appear again."
)
if mode == "redirect":
return (
"💡 First-time tip — I redirected the current run using your message. "
"Completed work stays in context, and `/stop` still cancels the task. "
"Send `/busy queue` to wait for a separate turn, or `/busy status` "
"to check. This notice won't appear again."
)
return (
"💡 First-time tip — I just interrupted my current task to answer you. "
"Send `/busy queue` to queue follow-ups for after the current task instead, "
"`/busy steer` to inject them mid-run without interrupting, or "
"`/busy status` to check. This notice won't appear again."
)
"""Hint shown the first time a user messages while the agent is busy (markdown)."""
return _BUSY_INPUT_HINTS_GATEWAY.get(mode, _BUSY_INPUT_HINT_GATEWAY_DEFAULT)
def busy_input_hint_cli(mode: str) -> str:
"""CLI version of the busy-input hint (plain text, no markdown)."""
if mode == "queue":
return (
"(tip) Your message was queued for the next turn. "
"Use /busy interrupt to make Enter stop the current run instead, "
"or /busy steer to inject mid-run. This tip only shows once."
)
if mode == "steer":
return (
"(tip) Your message was steered into the current run; it arrives "
"after the next tool call. Use /busy interrupt or /busy queue to "
"change this. This tip only shows once."
)
if mode == "redirect":
return (
"(tip) Your correction redirected the current run without discarding "
"completed work. Use /stop to cancel or /busy queue to wait for a "
"separate turn. This tip only shows once."
)
return (
"(tip) Your message interrupted the current run. "
"Use /busy queue to queue messages for the next turn instead, "
"or /busy steer to inject mid-run. This tip only shows once."
)
return _BUSY_INPUT_HINTS_CLI.get(mode, _BUSY_INPUT_HINT_CLI_DEFAULT)
def tool_progress_hint_gateway() -> str:
@@ -110,13 +102,7 @@ def tool_progress_hint_cli() -> str:
def openclaw_residue_hint_cli() -> str:
"""Banner shown the first time Hermes starts and finds ``~/.openclaw/``.
Points users at ``hermes claw migrate`` (non-destructive port of config,
memory, and skills) first. ``hermes claw cleanup`` is mentioned as the
follow-up step for users who have already migrated and want to archive
the old directory — with a warning that archiving breaks OpenClaw.
"""
"""Banner shown the first time Hermes finds ``~/.openclaw/``: migrate first, cleanup (which breaks OpenClaw) after."""
return (
"A legacy OpenClaw directory was detected at ~/.openclaw/.\n"
"To port your config, memory, and skills over to Hermes, run "
@@ -129,10 +115,7 @@ def openclaw_residue_hint_cli() -> str:
def detect_openclaw_residue(home: Optional[Path] = None) -> bool:
"""Return True if an OpenClaw workspace directory is present in ``$HOME``.
Pure filesystem check — no side effects. ``home`` override exists for tests.
"""
"""True if ``$HOME/.openclaw`` is a directory (pure check; ``home`` override for tests)."""
base = home or Path.home()
try:
return (base / ".openclaw").is_dir()
@@ -140,25 +123,15 @@ def detect_openclaw_residue(home: Optional[Path] = None) -> bool:
return False
# -------------------------------------------------------------------------
# Onboarding profile-build path (opt-in, consent-gated)
# -------------------------------------------------------------------------
# ── Onboarding profile-build path (opt-in, consent-gated) ─────────────────
def profile_build_mode(config: Mapping[str, Any]) -> str:
"""Resolve the onboarding profile-build mode from config.
"""``config.onboarding.profile_build``: ``"off"`` never offers; anything else -> ``"ask"`` (offer on first contact).
Returns one of:
``"ask"`` — on first contact, OFFER to build a profile (default).
``"off"`` — never offer; the first-message note stays a plain intro.
Read from ``config.onboarding.profile_build``. Unknown / missing values
fall back to ``"ask"`` so the default experience offers the flow. Any
network/account lookups inside the flow are separately consented to in
conversation — this setting only governs whether the offer is made.
This only governs whether the offer is made; lookups inside the flow are
consented to separately in conversation.
"""
if not isinstance(config, Mapping):
return "ask"
onboarding = config.get("onboarding")
onboarding = config.get("onboarding") if isinstance(config, Mapping) else None
if not isinstance(onboarding, Mapping):
return "ask"
mode = onboarding.get("profile_build")
@@ -170,11 +143,9 @@ def profile_build_mode(config: Mapping[str, Any]) -> str:
def profile_build_directive() -> str:
"""System-note directive appended to the very first message ever.
Instructs the agent to run a short, opt-in, consent-gated profile-build
flow and persist confirmed facts to the user-profile memory store
(``memory`` tool, ``target="user"``). Phrased so the agent ASKS before any
lookup and never silently reads connected accounts — directly addressing
the privacy concern that reading email/accounts unprompted feels invasive.
Runs a short opt-in profile-build flow persisting to the user-profile memory
store; phrased so the agent ASKS before any lookup and never silently reads
connected accounts.
"""
return (
"\n\n[System note: This is the user's very first message ever. "
@@ -196,9 +167,7 @@ def profile_build_directive() -> str:
)
# -------------------------------------------------------------------------
# State read / write
# -------------------------------------------------------------------------
# ── State read / write ────────────────────────────────────────────────────
def _get_seen_dict(config: Mapping[str, Any]) -> Mapping[str, Any]:
onboarding = config.get("onboarding") if isinstance(config, Mapping) else None
@@ -214,12 +183,7 @@ def is_seen(config: Mapping[str, Any], flag: str) -> bool:
def mark_seen(config_path: Path, flag: str) -> bool:
"""Persist ``onboarding.seen.<flag> = True`` to ``config_path``.
Uses the atomic YAML writer so a concurrent process can't observe a
partially-written file. Returns True on success, False on any error
(including the config file being absent — onboarding is best-effort).
"""
"""Persist ``onboarding.seen.<flag> = True`` atomically; False on any error (best-effort)."""
try:
import yaml
from hermes_cli.config import atomic_config_write
@@ -239,7 +203,7 @@ def mark_seen(config_path: Path, flag: str) -> bool:
seen = {}
cfg["onboarding"]["seen"] = seen
if seen.get(flag) is True:
return True # already marked — nothing to do
return True
seen[flag] = True
atomic_config_write(config_path, cfg)
return True
+8 -29
View File
@@ -1,31 +1,16 @@
#!/usr/bin/env python3
"""``/plan`` — build the plan-mode prompt that turns the user's request into a
saved markdown implementation plan, with no execution.
"""``/plan`` — build the plan-mode prompt: a saved markdown implementation plan, no execution.
``/plan`` used to be a bundled skill (``skills/software-development/plan``)
whose auto-generated slash command fell off the capped Telegram/Discord command
menus for most installs (skills are the only tier trimmed at the platform
caps, alphabetically — ``plan`` sat past the cutoff). It is now a first-class
built-in: this module builds ONE prompt that instructs the live agent to
1. Stay in planning mode for the turn — read-only inspection is allowed,
but no implementation, no mutating commands, no side effects.
2. Write a concrete, bite-sized, TDD-shaped markdown plan under
``.hermes/plans/`` in the active workspace via ``write_file``.
There is no engine and no model-tool footprint: the agent does the work with
its existing toolset, so this works identically on local, Docker, and remote
terminal backends. Every surface (CLI ``/plan``, gateway ``/plan``, TUI
``/plan``) calls :func:`build_plan_prompt` and feeds the result to the agent
as a normal turn — same pattern as ``/learn`` and ``/init``, preserving
prompt-cache invariants (no system-prompt or history mutation).
A first-class built-in (the former bundled skill fell off capped Telegram/Discord
command menus). No engine, no model-tool footprint: every surface feeds
:func:`build_plan_prompt` to the agent as a normal turn, like ``/learn`` and
``/init``, so system prompt and history stay untouched (prompt-cache safe).
"""
from __future__ import annotations
# The plan-mode ground rules + authoring craft, distilled from the retired
# bundled skill (v2.0.0, writing-craft adapted from obra/superpowers).
# Embedded in the prompt so the agent plans the way a maintainer would.
# Plan-mode ground rules + authoring craft, distilled from the retired bundled
# skill (writing-craft adapted from obra/superpowers).
_PLAN_MODE_RULES = """\
For this turn, you are in PLAN MODE — planning only.
@@ -76,13 +61,7 @@ Interaction style:
def build_plan_prompt(task: str = "") -> str:
"""Build the plan-mode prompt for the live agent.
Args:
task: What to plan. Empty → infer the task from the current
conversation context (mirrors the retired skill's behavior and
issue #36821's "plan from context" expectation).
"""
"""Build the plan-mode prompt; empty *task* asks the agent to infer it from conversation context."""
task = (task or "").strip()
if task:
task_block = f"Task to plan:\n{task}\n"
+25 -57
View File
@@ -1,52 +1,30 @@
"""Builder-declared stable prefixes for Anthropic prompt caching (#81867).
"""Builder-declared stable prefixes for Anthropic prompt caching.
Skill, webhook, and cron builders concatenate a large static scaffold
(activation note + expanded skill body) with a small volatile invocation
tail (ticket payload, timestamps, run context) into one user-message
string. Only the builder knows the exact byte where the volatile tail
begins, so it registers the stable prefix here at construction time; the
cache planner consults the registry to place a cache breakpoint at that
boundary instead of caching the whole message as one atomic block.
Skill/webhook/cron builders concatenate a large static scaffold with a small
volatile invocation tail into one user-message string. Only the builder knows
where the tail begins, so it registers the stable prefix here and the cache
planner places a breakpoint at that boundary instead of caching the whole
message. Re-parsing marker strings out of the message at request time is
deliberately avoided: markers can legitimately appear inside skill bodies or
event payloads, and any delimiter heuristic then shrinks the cached prefix or
silently absorbs volatile bytes into it.
This deliberately avoids re-parsing scaffold marker strings out of the
message at request time: markers can legitimately appear inside skill
bodies or inside event payloads (e.g. a helpdesk ticket quoting an agent
transcript), and any delimiter-search heuristic then either shrinks the
cached prefix or — worse — silently absorbs volatile bytes into it,
reintroducing the per-invocation cache miss this exists to fix.
The registry is process-local by design. A freshly fired webhook/cron
invocation is always built and sent by the same process, which is the
only window where the split pays off. Any miss (restart, eviction,
historic message) falls back to the pre-existing whole-message policy.
Split-shape lifetime: the split is applied only while the skill message is
one of the plan's marked endpoints (the last few cacheable messages). Once
later turns rotate it out of that window it ships as a single string block
again, which changes the block boundary once and re-ingests the prefix from
that message onward exactly one time in a long-lived session. Webhook/cron
invocations — the workload this exists for — send the skill turn as the
newest message every time, so they always hit the split shape; the one-time
re-ingest only affects long interactive sessions and nets out far below the
per-invocation full rewrite this removes.
Process-local by design: a webhook/cron fire is built and sent by the same
process, and any miss (restart, eviction, historic message) falls back to the
whole-message policy. The split only applies while the message is one of the
plan's marked endpoints; once it rotates out it ships as one block again
(one-time re-ingest in long interactive sessions, never for webhook/cron).
"""
import threading
from collections import OrderedDict
from typing import Optional
# A couple dozen distinct active scaffolds (webhook routes x skills x cron
# jobs) is generous for one gateway process; beyond that, oldest entries
# fall back to whole-message caching rather than growing unboundedly.
# A couple dozen active scaffolds is generous for one gateway process.
_MAX_ENTRIES = 32
# Entries hold whole expanded skill bodies, so an entry count alone does not
# bound memory — a handful of large skills can retain tens of MB in a
# long-lived gateway process. Evict by total retained characters too (a
# conservative proxy for bytes: actual memory is 1–4x depending on the
# string's widest code point), always keeping the newest entry so a single
# oversized scaffold still gets a boundary instead of silently disabling
# the split.
# Entries hold whole expanded skill bodies, so also bound total retained chars
# (1-4x bytes). The newest entry is always kept so one oversized scaffold still
# gets a boundary instead of silently disabling the split.
_MAX_CHARS = 4 * 1024 * 1024
_lock = threading.Lock()
@@ -67,29 +45,19 @@ def register_stable_prefix(prefix: str) -> None:
def find_stable_prefix(content: str) -> Optional[str]:
"""Longest registered prefix that is a *proper* prefix of ``content`` with non-whitespace tail.
"""Longest registered *proper* prefix of ``content`` with a non-whitespace tail.
Proper with non-whitespace tail (``bool(content[len(prefix):].strip())``) so the
split never produces an empty or whitespace-only volatile text block, which
Anthropic rejects on the wire (HTTP 400).
A hit refreshes the entry's LRU position: a scaffold fired every minute
by cron must not be evicted by a burst of one-off skill invocations,
which would silently drop it back to whole-message caching.
The tail must be non-whitespace so the split never yields an empty text
block (Anthropic rejects it with HTTP 400). A hit refreshes the entry's LRU
position so a scaffold fired every minute by cron is not evicted by a
burst of one-off skill invocations.
"""
with _lock:
best: Optional[str] = None
for prefix in _prefixes:
if content.startswith(prefix) and bool(content[len(prefix):].strip()):
if content.startswith(prefix) and content[len(prefix):].strip():
if best is None or len(prefix) > len(best):
best = prefix
if best is not None:
# After the scan so the OrderedDict is never mutated mid-iteration.
_prefixes.move_to_end(best)
_prefixes.move_to_end(best) # after the scan: never mutate mid-iteration
return best
def clear_stable_prefixes() -> None:
"""Test isolation helper."""
with _lock:
_prefixes.clear()
+68 -174
View File
@@ -1,68 +1,30 @@
"""Rotation-stable logical cache scope for prompt_cache_key derivation.
Context-compression rotation (legacy ``compression.in_place: false`` mode)
mints a new physical ``session_id`` mid-conversation to segment the
transcript. The prompt-cache scope introduced by #79161 was derived from that
physical id, so every rotation moved the conversation into a fresh cache
bucket even though it is logically the same conversation continuing
(issue #79017).
Legacy compression rotation (``compression.in_place: false``) mints a new
physical ``session_id`` mid-conversation, which moved the conversation into a
fresh cache bucket each time. ``resolve_prompt_cache_scope()`` instead maps
the physical id to the ROOT of its compression lineage via
``SessionDB.get_compression_lineage()`` — NOT ``get_conversation_root`` /
``_conversation_root_id`` (the Portal-attribution walk), which follows
``parent_session_id`` blindly and would collapse /branch children and delegate
trees into one id. The two resolvers are intentionally different.
``resolve_prompt_cache_scope()`` maps the physical session id to the ROOT of
its *compression lineage* — the pre-rotation session id — using
``SessionDB.get_compression_lineage()``, whose fork-aware semantics
(hardened in #79193) give exactly the scope boundaries the cache key needs.
NOT ``SessionDB.get_conversation_root`` / ``run_agent._conversation_root_id``
(the Portal-attribution walk): that one follows ``parent_session_id`` blindly,
collapsing /branch children and whole delegate trees into one id, which would
violate the #79161 isolation this scope must preserve. The two resolvers are
intentionally different — do not "deduplicate" them.
Scope boundaries: rotation children walk back to the original segment; ``/new``
starts a fresh scope; ``/branch`` children, delegate subagents, and tool-tagged
children are explicit fork children with their own isolated scope; cron fires
keep their physical id (the per-fire timestamp is stripped later).
- compression-rotation children walk back to the original segment
(rotation-stable scope — the fix);
- ``/new`` starts a lineage-less session (fresh scope);
- ``/branch`` children (``_branched_from``), delegate subagents
(``_delegate_from``), and tool-tagged children (``source="tool"``) are
explicit fork children and keep their own isolated scope, preserving the
sibling/subagent isolation #79161 established;
- cron fires keep their physical ``cron_<job>_<ts>`` id here — the per-fire
timestamp is stripped later by ``_cache_scope_from_session_id`` exactly as
before.
Hosts that mint one physical id per RESPONSE (Studio group chat, ``/v1/responses``
with client-managed history) carry no lineage, so the walk returns the physical
id and the scope moves every reply. Hermes must not infer the conversation from
id SYNTAX (that collides client-supplied ids); the host declares it via
``gateway_session_key`` (``X-Hermes-Session-Key`` / ``build_session_key``),
consumed by ``declared_conversation_scope()``, which wins over the lineage walk.
The declared key is hashed to ``gwk_<sha256[:24]>`` because it embeds
platform/chat/user identifiers and leaves the process as a provider routing key.
A host that mints one physical ``session_id`` per RESPONSE (Hermes Studio's
group chat, and ``POST /v1/responses`` with client-managed history, which
mints ``str(uuid4())`` per request) re-keys every conversation-affinity hint
Hermes sends — ``prompt_cache_key`` on both OpenAI-wire transports, plus the
OpenRouter/Nous sticky ``session_id`` and xAI's ``x-grok-conv-id`` through
``portal_tags`` (issue #96811). Those rows carry no lineage, so the walk
above correctly returns the physical id and the scope moves every reply.
Hermes must not infer the logical conversation from the id's SYNTAX (that
rule collides independent client-supplied ids and merges Studio members
truncated past its 96-character boundary — the #79017 failure class). The
host has to declare it, and one carrier already means exactly that:
``gateway_session_key`` — the "stable per-chat key" (``agent:main:telegram:
dm:123``) built by ``gateway.session.build_session_key`` from the
``X-Hermes-Session-Key`` header, which branching deliberately does NOT key
off. ``declared_conversation_scope()`` consumes it, and it wins over the
lineage walk because it is stable across rotation AND across per-response
ids. Two boundaries it must not cross:
- explicit fork children (``/branch``, delegate subagents, tool children)
share their parent's chat key but are separate conversations — the row's
fork markers keep them on their own scope (#79161);
- background-review forks run on a clone of the live runtime, so they are
excluded by ``_persist_disabled`` for the same reason.
The declared key is hashed into ``gwk_<sha256[:24]>`` before it becomes a
scope: unlike a session id it embeds platform/chat/user identifiers, and
this value leaves the process verbatim as OpenRouter's sticky ``session_id``
and xAI's ``x-grok-conv-id``.
The resolution is memoized per (agent, session_id): the lineage walk runs
once per transcript segment — NOT per API call — and re-runs only when
rotation actually changes ``agent.session_id`` (per the no-DB-on-the-hot-path
constraint recorded on #79017). Default installs compact in place and never
rotate, so they hit the memo forever and behave byte-identically to before.
Resolution is memoized per (agent, session_id, db-present): the lineage walk
runs once per transcript segment, never per API call.
"""
import hashlib
@@ -72,16 +34,13 @@ from typing import Any, Optional
logger = logging.getLogger(__name__)
_MEMO_ATTR = "_prompt_cache_scope_memo"
# Namespace for a scope resolved from a host-declared conversation key.
_DECLARED_SCOPE_PREFIX = "gwk_"
def _lineage_root(session_id: str, session_db: Any) -> Optional[str]:
"""Return the compression-lineage root of *session_id*, or None.
"""Compression-lineage root of *session_id*, or None.
Defensive about the DB handle: test doubles and partially constructed
agents can hand back non-list results — anything that is not a non-empty
list/tuple whose first element is a non-empty string is ignored.
Tolerates non-list results from test doubles / partially built agents.
"""
if session_db is None:
return None
@@ -102,24 +61,13 @@ def _agent_source(
) -> str:
"""The ``sessions.source`` this agent's conversation is recorded under.
Read from the agent's own row when it exists, because that is the value
the peer queries below match on.
``row_source`` is that value when the caller already has it — the single
identity read in :func:`declared_conversation_scope` — where ``""`` means
"the row was read and carries no source". ``None`` means "not read yet"
and keeps the original lookup, which is the path a ``SessionDB`` without
:meth:`~hermes_state.SessionDB.declared_scope_identity` still takes.
Before the row lands — this module resolves the first scope ahead of
``_ensure_db_session`` — it uses the SAME resolver persistence will use,
``run_agent._session_source_for_agent``, not ``agent.platform``. The two
diverge whenever ``HERMES_SESSION_SOURCE`` overrides the platform, and the
divergence is not a cosmetic one: the declared scope is non-``None``
immediately, so ``resolve_prompt_cache_scope`` memoizes it for this session
id and never re-resolves once the authoritative row appears. Both sides of
a ``/new`` would then read the platform domain, miss the boundary recorded
under the override, and hash the same scope.
``row_source`` is the row's value when the caller already read it (``""``
= read, no source; ``None`` = not read yet, do the lookup). Before the row
lands, use the SAME resolver persistence uses
(``run_agent._session_source_for_agent``), not ``agent.platform``: they
diverge under ``HERMES_SESSION_SOURCE``, and the declared scope is memoized
immediately, so both sides of a ``/new`` would otherwise miss the boundary
recorded under the override and hash the same scope.
"""
if row_source is None and session_id and session_db is not None:
try:
@@ -132,8 +80,7 @@ def _agent_source(
return row_source
platform = getattr(agent, "platform", None)
try:
# Imported lazily: run_agent imports this module, and this is the
# single owner of the source a session row is created with.
# Lazy: run_agent imports this module.
from run_agent import _session_source_for_agent
source = str(_session_source_for_agent(platform) or "").strip()
@@ -145,23 +92,14 @@ def _agent_source(
def _conversation_generation(session_key: str, source: str, session_db: Any) -> str:
"""Return the durable generation for *session_key*'s current conversation.
"""Durable generation for *session_key*'s current conversation (``""`` if none).
The declared key names a chat and deliberately survives `/new` and policy
resets. Hashing it alone would therefore reuse one affinity scope across
distinct conversations, violating the #79017/#86733 contract: warm across
compression, cold across a conversation boundary.
``SessionDB.latest_conversation_boundary`` reads the monotonic
``conversation_generations`` counter for ``(source, session_key)``. The
counter advances in the same transaction that records an
``_RESET_END_REASONS`` boundary. It is independent of prunable session rows
and wall-clock time, so deletion, bulk pruning, and clock rollback cannot
reissue an old generation. Compression continues the current conversation
and does not advance it.
This lookup runs on the memoized resolution path, not once per API call.
Return ``""`` when the key has never reset or the DB exposes no generation.
The declared key names a chat and survives ``/new`` and policy resets, so
hashing it alone would reuse one scope across distinct conversations. The
``conversation_generations`` counter advances in the same transaction that
records a reset boundary and is independent of prunable rows and
wall-clock, so pruning or clock rollback cannot reissue a generation.
Compression does not advance it.
"""
reader = getattr(session_db, "latest_conversation_boundary", None)
if not callable(reader):
@@ -173,30 +111,17 @@ def _conversation_generation(session_key: str, source: str, session_db: Any) ->
def declared_conversation_scope(agent: Any) -> Optional[str]:
"""Return the host-declared logical conversation scope, or None.
"""Host-declared logical conversation scope (``gwk_<sha256[:24]>``), or None.
Resolved from ``agent._gateway_session_key`` (the ``X-Hermes-Session-Key``
/``build_session_key`` per-chat key) qualified by the conversation
generation currently live on it (:func:`_conversation_generation`), hashed
together into ``gwk_<sha256[:24]>`` so no platform/chat/user identifier
reaches a provider on the wire and the value stays inside every caller's
length/charset budget.
The key alone would outlive the conversation — it survives ``/new`` and the
idle/daily policy resets by design — so the generation is what makes this
carrier legal: stable across a host's per-response physical ids, and cold
on every conversation replacement.
None — meaning "fall back to the physical-id scope" — when no key was
declared, when this agent is a background-review fork (``_persist_disabled``:
it clones the live runtime, including the key), when the session row is an
explicit fork child (``/branch``, delegate, tool), and on any DB error
during either lookup.
Hashes ``(source, gateway_session_key, generation)`` so no platform/chat/
user identifier reaches a provider. None — fall back to the physical-id
scope — when no key is declared, when the agent is a background-review
fork (``_persist_disabled`` clones the live runtime incl. the key), when
the row is an explicit fork child, and on any DB error (fail closed rather
than merge a fork onto its parent's key).
"""
key = str(getattr(agent, "_gateway_session_key", "") or "").strip()
if not key:
return None
if getattr(agent, "_persist_disabled", False):
if not key or getattr(agent, "_persist_disabled", False):
return None
sid = str(getattr(agent, "session_id", None) or "")
db = getattr(agent, "_session_db", None)
@@ -204,13 +129,9 @@ def declared_conversation_scope(agent: Any) -> Optional[str]:
row_source: Optional[str] = None
if sid and db is not None:
try:
# One read for both halves of the row's identity: the fork verdict
# and the source the peer queries match on live on the same
# ``sessions`` row, and asking for them separately read it twice
# per resolution (@teknium1 on #98811). A SessionDB without the
# combined view keeps the original call, so nothing that predates
# it — including the doubles that certify the fail-closed contract
# below — changes behaviour.
# One read for both halves of the row identity (fork verdict +
# source). A SessionDB without the combined view keeps the
# original call.
identity = getattr(db, "declared_scope_identity", None)
if callable(identity):
is_fork, row_source = identity(sid)
@@ -219,11 +140,6 @@ def declared_conversation_scope(agent: Any) -> Optional[str]:
if is_fork:
return None
except Exception:
# Degrade to the physical-id scope rather than risk merging a
# fork onto its parent's key on a transient DB failure. The
# source read is inside this same guard for the same reason: it
# was always the second half of a read that had already failed
# closed here.
logger.debug("declared-scope fork check failed", exc_info=True)
return None
source = _agent_source(agent, sid, db, row_source)
@@ -231,65 +147,46 @@ def declared_conversation_scope(agent: Any) -> Optional[str]:
try:
generation = _conversation_generation(key, source, db)
except Exception:
# Same fail-closed rule as the fork check: an unqualified key
# spans /new, so degrade to the physical-id scope instead.
logger.debug("declared-scope generation read failed", exc_info=True)
return None
# The carrier is the SAME identity tuple the peer queries use: two hosts
# may legally declare the same key string under different sources, and the
# scope leaves this process as a routing key, so it must not collapse them.
# Same identity tuple the peer queries use: two hosts may declare the
# same key under different sources and must not collapse.
carrier = f"{source}|{key}|{generation}"
digest = hashlib.sha256(carrier.encode("utf-8", errors="replace")).hexdigest()[:24]
return f"{_DECLARED_SCOPE_PREFIX}{digest}"
def resolve_prompt_cache_scope(agent: Any) -> str:
"""Resolve the rotation-stable cache-scope id for *agent*'s conversation.
"""Rotation-stable cache-scope id for *agent*'s conversation.
Returns the host-declared conversation scope when one applies
(``declared_conversation_scope``), else the compression-lineage ROOT of
``agent.session_id`` (the physical id itself when the session has no
compression ancestry, no DB is attached, or the walk fails). The result is memoized on the agent
keyed by the current session id, so the DB walk happens once per
transcript segment rather than once per API call.
Declared scope when one applies, else the compression-lineage root of
``agent.session_id`` (the physical id when there is no ancestry, no DB, or
the walk fails). Memoized on the agent keyed by session id.
"""
sid = str(getattr(agent, "session_id", None) or "")
if not sid:
return ""
db = getattr(agent, "_session_db", None)
# Memo key includes DB presence: an agent that starts DB-less and gains a
# handle later (run_agent._get_session_db_for_recall lazily attaches one)
# DB presence is part of the key: an agent that gains a DB handle later
# must re-resolve instead of staying pinned to the physical id.
key = (sid, db is not None)
memo = getattr(agent, _MEMO_ATTR, None)
if isinstance(memo, tuple) and len(memo) == 2 and memo[0] == key:
return memo[1]
# A declared conversation key outranks the lineage walk: it is stable
# across compression rotation AND across a host's per-response ids, which
# the walk cannot see (#96811).
root = declared_conversation_scope(agent) or (
_lineage_root(sid, db) if db is not None else None
)
scope = root or sid
# Memoize on a successful walk, or when there is no DB to consult at all,
# or when the agent will never persist a row (background-review forks set
# _persist_disabled but still hold a DB handle — without this, every API
# call would re-run the lineage query forever).
# A failed/empty walk on a persisting agent is NOT memoized: falling back
# to the physical id is the correct degraded answer right now (row not
# persisted yet, transient DB error), but pinning it for the whole segment
# would keep the scope wrong after the session row lands.
if (
root is not None
or db is None
or getattr(agent, "_persist_disabled", False)
):
# Memoize on success, with no DB, or when the agent never persists a row
# (background-review forks hold a DB handle but set _persist_disabled).
# A failed/empty walk on a persisting agent is NOT memoized: the physical
# id is right for now (row not yet persisted, transient error) but would
# stay wrong for the whole segment once the row lands.
if root is not None or db is None or getattr(agent, "_persist_disabled", False):
try:
setattr(agent, _MEMO_ATTR, (key, scope))
except Exception:
# Frozen/slotted test doubles — resolution still works, just
# unmemoized.
pass
pass # frozen/slotted doubles: resolution works, just unmemoized
return scope
@@ -303,14 +200,11 @@ def declared_conversation_scope_safe(agent: Any) -> Optional[str]:
def resolve_prompt_cache_scope_safe(agent: Any) -> Optional[str]:
"""Never-raising variant of :func:`resolve_prompt_cache_scope`.
"""Never-raising variant of :func:`resolve_prompt_cache_scope` (None on failure/empty).
Returns None on any failure (or when there is no scope). Consumers treat
None/empty as "fall back to the physical session_id", so a resolution
failure degrades to pre-#79017 behavior instead of blocking the caller —
important at turn_context's call site, where an exception raised inside
the ``set_runtime_main(...)`` argument list would otherwise skip the whole
runtime binding, not just the cache scope.
Consumers treat None as "use the physical session_id"; at turn_context's
call site an exception inside the ``set_runtime_main(...)`` argument list
would skip the whole runtime binding, not just the cache scope.
"""
try:
return resolve_prompt_cache_scope(agent) or None
+88 -212
View File
@@ -1,13 +1,10 @@
"""Anthropic prompt caching strategy.
"""Anthropic prompt caching strategy — pure functions, no AIAgent dependency.
The default layout uses 4 cache_control breakpoints: the static system
prefix, the end of the system prompt, and the last 2 non-system messages.
When a static system prefix is unavailable, it falls back to one system
breakpoint plus the last 3 messages. All markers use the same TTL (5m or 1h).
This preserves intra-session caching while allowing new sessions to reuse the
stable system-prompt prefix.
Pure functions -- no class state, no AIAgent dependency.
Default layout: 4 cache_control breakpoints — the static system prefix, the end
of the system prompt, and the last 2 non-system messages. Without a static
prefix: one system breakpoint plus the last 3 messages. All markers share one
TTL (5m or 1h). This keeps intra-session caching while letting new sessions
reuse the stable system-prompt prefix.
"""
import copy
@@ -24,30 +21,18 @@ class PromptCachePlan:
messages: List[Dict[str, Any]]
tools: List[Dict[str, Any]]
@property
def marker_count(self) -> int:
"""Wire-visible cache markers in this plan (computed on demand).
Only tests consume this; keeping it lazy avoids walking every
message part and tool schema on the per-request hot path.
"""
return _count_cache_markers(self.messages, self.tools)
def envelope_tool_part_cache_markers_supported(
provider: str | None, base_url: str | None
) -> bool:
"""Whether the envelope-layout route honors part-level markers on role:tool.
OpenRouter (and Nous Portal, which proxies to it) relocate a
``cache_control`` sitting on a tool message's content part onto the
``tool_result`` block during their OpenAI→Anthropic translation, so the
marker is honored there. LiteLLM-style OpenAI-wire proxies instead map
content parts verbatim: the part-level marker lands at
``tool_result.content[0]``, which the Anthropic Messages schema forbids —
a non-retryable HTTP 400 that kills the whole turn (#89886). On those
routes tool messages must not carry part-level markers at all; the
breakpoint budget reallocates to the nearest eligible message instead.
OpenRouter (and Nous Portal, which proxies to it) relocate a part-level
``cache_control`` onto the ``tool_result`` block during OpenAI→Anthropic
translation. LiteLLM-style proxies copy parts verbatim, so the marker lands
at ``tool_result.content[0]`` — forbidden by the Anthropic schema, a
non-retryable 400. On those routes tool messages carry no part markers and
the breakpoint budget reallocates to the nearest eligible message.
"""
from agent.agent_runtime_helpers import _is_litellm_route
@@ -65,27 +50,19 @@ def _apply_cache_marker(
content = msg.get("content")
if role == "tool" and native_anthropic:
# Native Anthropic layout: top-level marker; the adapter moves it
# inside the tool_result block.
# Top-level marker; the native adapter moves it inside tool_result.
msg["cache_control"] = cache_marker
return
if role == "tool" and not tool_part_markers:
# Envelope route whose OpenAI→Anthropic translation copies content
# parts verbatim (LiteLLM et al.): a part-level marker becomes
# tool_result.content[0].cache_control → non-retryable 400 (#89886).
# LiteLLM-style envelope: a part marker becomes
# tool_result.content[0].cache_control → non-retryable 400.
return
if content is None or content == "":
if role == "tool" and not native_anthropic:
# OpenRouter rejects top-level cache_control on role:tool (silent
# hang) and an empty message has no content part to carry the
# marker — skip. Non-empty tool content falls through below and
# gets the marker on a content part, which OpenRouter honors.
return
if role == "assistant" and not native_anthropic:
# Empty assistant turns are pure tool_calls. A top-level marker
# here is ignored on the envelope layout, so skip.
# Envelope layout: OpenRouter rejects top-level cache_control on
# role:tool (silent hang), and ignores it on empty assistant turns
# (pure tool_calls) — neither has a content part to carry it.
if role in ("tool", "assistant") and not native_anthropic:
return
msg["cache_control"] = cache_marker
return
@@ -96,17 +73,13 @@ def _apply_cache_marker(
if stable_prefix is not None:
suffix = content[len(stable_prefix):]
if suffix.strip():
# Builder-declared boundary (#81867): the scaffold carries the
# breakpoint, the volatile invocation tail rides unmarked so a
# changed ticket ID or timestamp no longer invalidates the
# whole skill body. Request-local only — the canonical session
# message stays a plain string.
# Builder-declared boundary: the scaffold carries the
# breakpoint and the volatile tail rides unmarked, so a
# changed ticket ID/timestamp no longer invalidates the
# skill body. Request-local only — the stored message
# stays a plain string.
msg["content"] = [
{
"type": "text",
"text": stable_prefix,
"cache_control": cache_marker,
},
{"type": "text", "text": stable_prefix, "cache_control": cache_marker},
{"type": "text", "text": suffix},
]
return
@@ -126,17 +99,12 @@ def _can_carry_marker(
) -> bool:
"""True if a marker on this message is actually honored by the provider.
On the native Anthropic layout every message works (top-level markers are
relocated by the adapter). On the envelope layout (OpenRouter et al.) only
markers inside content parts are honored: empty-content messages (e.g.
assistant turns that are pure tool_calls) and empty tool messages would
receive a top-level marker the provider ignores — wasting one of the four
breakpoints. Skip those so the breakpoints land on messages that count.
``tool_part_markers=False`` (LiteLLM-style envelope routes, #89886)
additionally excludes ALL role:tool messages: their part-level marker
would be forwarded verbatim into ``tool_result.content[]`` and rejected
with a non-retryable 400, so the breakpoint must reallocate instead.
Native Anthropic honors every message (the adapter relocates top-level
markers). The envelope layout only honors markers inside content parts, so
empty-content messages would waste one of the four breakpoints; with
``tool_part_markers=False`` (LiteLLM-style routes) every role:tool message
is excluded too, since its part marker would be rejected with a 400.
Must agree with :func:`_apply_cache_marker`, which marks only the LAST part.
"""
if native_anthropic:
return True
@@ -146,10 +114,6 @@ def _can_carry_marker(
if content is None or content == "":
return False
if isinstance(content, list):
# _apply_cache_marker only marks the LAST content part, so the carrier
# predicate must agree: a list whose last element isn't a dict cannot
# actually receive a marker and would waste a breakpoint. Mirror the
# `content` truthiness + last-element-dict check in _apply_cache_marker.
return bool(content) and isinstance(content[-1], dict)
return isinstance(content, str)
@@ -162,66 +126,31 @@ def _build_marker(ttl: str) -> Dict[str, str]:
return marker
# Alibaba-family providers (Qwen routes). Their context cache documents a
# five-minute window (renewed on hit) and rejects the Anthropic 1h tier.
# Shared with agent_runtime_helpers.anthropic_prompt_cache_policy so the
# cache-policy opt-in and the TTL clamp can never desync (#84733).
# Alibaba-family providers (Qwen routes): documented five-minute context cache,
# Anthropic 1h tier rejected. Shared with
# agent_runtime_helpers.anthropic_prompt_cache_policy so the cache-policy
# opt-in and the TTL clamp never desync. Do NOT narrow this set to extend a
# TTL — it also drives the marker-layout opt-in, so narrowing DISABLES caching.
ALIBABA_FAMILY_PROVIDERS = frozenset({
"opencode",
"opencode-zen",
"opencode-go",
"opencode-zen",
"alibaba",
})
# --- 1h-tier membership: an ALLOW-list, deliberately minimal ----------------
#
# #84733 clamped 1h -> 5m for the whole alibaba/opencode family, reasoning from
# Alibaba's PUBLISHED Qwen docs. Wire measurement on the opencode-go route
# contradicts the docs. Controlled run: identical request, only the ttl flag
# varying, read back after 11 minutes with no intervening call (a read renews
# the window and would mask expiry):
#
# qwen3.8-max ttl=1h -> cache_read 2122 SURVIVED
# qwen3.8-max ttl=- -> cache_read 0 EXPIRED <- control
# glm-5.2 ttl=1h -> cache_read 2092 SURVIVED
# minimax-m2.5 ttl=1h -> cache_read 0 EXPIRED
#
# Read the two non-qwen rows for what they are: evidence about the ROUTE, not
# about traffic Hermes sends today. anthropic_prompt_cache_policy currently
# opts opencode-go in only for qwen models, so glm-5.2 and minimax-m2.5 on
# that route receive no cache_control marker at all and never reach this
# clamp in production. They constrain the route-level rule; they are not
# live paths.
#
# Only opencode-go is listed: it is the only route measured. Other opencode
# routes stay clamped because they were NOT measured, not because they are
# known bad. opencode-zen returns cache_creation.ephemeral_1h_input_tokens for
# Claude models, so it is a candidate -- but qwen on zen is unmeasured, so
# adding the provider wholesale would outrun the evidence.
#
# WARNING: opencode-go labels EVERY write `ephemeral_5m_input_tokens` whatever
# ttl was requested. That label is NOT evidence of the retention window -- it
# is what made the original docs-based reasoning look confirmed. Verify only
# with a delayed read past 5 minutes and no intervening call.
#
# NOTE: kept separate from ALIBABA_FAMILY_PROVIDERS on purpose. That set also
# drives the cache-marker-layout OPT-IN in
# agent_runtime_helpers.anthropic_prompt_cache_policy; narrowing it would
# silently DISABLE caching for qwen on opencode-go rather than extend its TTL.
# 1h-tier ALLOW-list: only routes wire-measured to retain a 1h marker (delayed
# read past 5 minutes with no intervening call — an intervening read renews the
# window and masks expiry). Other opencode routes stay clamped because they are
# UNMEASURED, not known-bad. Note opencode-go labels every write
# `ephemeral_5m_input_tokens` regardless of requested ttl; that label is not
# evidence of the retention window.
MEASURED_1H_PROVIDERS = frozenset({
"opencode-go",
})
# Models measured to ignore the 1h tier even on a 1h-capable route.
#
# SCOPE: consulted only for providers already in MEASURED_1H_PROVIDERS. The
# measurement was taken on the opencode-go route, so it says nothing about the
# same model reached some other way -- and MiniMax on its own
# Anthropic-compatible endpoint IS a separate, cache-eligible route
# (anthropic_prompt_cache_policy opts it in by provider id / host match).
# Checking this set globally would have silently regressed that unrelated
# route's configured 1h to 5m off the back of an opencode-go observation.
# Models measured to ignore the 1h tier on a MEASURED_1H_PROVIDERS route.
# Consulted only there: the same model on its own Anthropic-compatible endpoint
# is a separate cache-eligible route and must not inherit this clamp.
NO_1H_TIER_MODELS = frozenset({
"minimax-m2.5",
})
@@ -235,9 +164,8 @@ def _flat_model(model: str) -> str:
def is_qwen_model(model: str) -> bool:
"""True when ``model`` names a Qwen-family model (case-insensitive).
Shared by the TTL clamp below and
``agent_runtime_helpers.anthropic_prompt_cache_policy`` so the
cache-policy opt-in and the clamp can never desync (#84733).
Shared with ``agent_runtime_helpers.anthropic_prompt_cache_policy`` so the
cache-policy opt-in and the TTL clamp never desync.
"""
return "qwen" in (model or "").lower()
@@ -250,30 +178,18 @@ def effective_cache_ttl(
) -> str:
"""Clamp a requested cache TTL to what the destination route supports.
Qwen/Alibaba context caching documents an explicit five-minute window
(renewed on hit); the Anthropic ``1h`` tier is ignored/rejected there,
so a configured ``1h`` regresses to ``5m`` instead of shipping a marker
the provider drops and creating a false 1h-cache expectation (#84733).
Exception: routes in ``MEASURED_1H_PROVIDERS`` were wire-measured to
honour the tier (delayed read past 5 minutes) and keep ``1h`` — minus
any model in ``NO_1H_TIER_MODELS`` measured to ignore it on that route.
All other caching routes keep the requested TTL.
``None`` (caching active with no explicit tier) resolves to ``5m``.
Qwen/Alibaba routes document a five-minute window and drop the ``1h``
tier, so a configured ``1h`` regresses to ``5m`` there instead of creating
a false 1h-cache expectation — except on ``MEASURED_1H_PROVIDERS``, which
keep ``1h`` minus any ``NO_1H_TIER_MODELS`` model. The measured-route check
runs BEFORE the generic Qwen clamp, which would otherwise swallow every
Qwen model on it. ``None`` resolves to ``5m``.
"""
if ttl != "1h":
return ttl or "5m"
if (provider or "").lower() in MEASURED_1H_PROVIDERS:
# Route measured to honour the tier -- checked BEFORE the generic
# is_qwen_model clamp below, which would otherwise swallow every Qwen
# model on it. Within the route, a model measured to ignore the tier
# still wins; the denial stays nested here so an opencode-go
# observation cannot leak out and reclamp the same model on an
# unrelated route.
return "5m" if _flat_model(model) in NO_1H_TIER_MODELS else "1h"
if is_qwen_model(model):
return "5m"
if (provider or "").lower() in ALIBABA_FAMILY_PROVIDERS:
if is_qwen_model(model) or (provider or "").lower() in ALIBABA_FAMILY_PROVIDERS:
return "5m"
return "1h"
@@ -289,21 +205,13 @@ def _apply_system_cache_markers(
) -> int:
"""Mark the static system prefix (and optionally the full prompt).
The system prompt remains one stored string. Splitting it only in the
outgoing request keeps session persistence and non-Anthropic transports
unchanged while making the stable prefix independently cacheable.
``mark_suffix=False`` is the tool-cache-plan layout: only the static
prefix carries a marker, the volatile suffix rides unmarked (its
breakpoint budget is spent on the tools array instead).
``fallback_to_whole=False`` skips marking entirely when the prefix
split is not possible (no prefix, mismatched prefix, non-string
content) instead of marking the whole message.
When the prompt IS exactly the static prefix (empty suffix), the whole
message is marked as a single block — never a two-part split with an
empty text block, which Anthropic rejects.
The system prompt stays one stored string; it is split only in the
outgoing request so persistence and non-Anthropic transports are
unchanged. ``mark_suffix=False`` is the tool-cache-plan layout (suffix
unmarked, its budget spent on the tools array). ``fallback_to_whole=False``
marks nothing when the prefix split is impossible. When the prompt IS the
prefix (empty/whitespace suffix) the whole message is marked as one block —
never a split with an empty text block, which Anthropic rejects.
Returns the number of markers applied (0, 1, or 2).
"""
@@ -320,17 +228,10 @@ def _apply_system_cache_markers(
if mark_suffix:
suffix_part["cache_control"] = cache_marker
message["content"] = [
{
"type": "text",
"text": static_system_prefix,
"cache_control": cache_marker,
},
{"type": "text", "text": static_system_prefix, "cache_control": cache_marker},
suffix_part,
]
return 2 if mark_suffix else 1
# Empty/whitespace-only suffix: the stored prompt IS the static prefix. Mark it as
# one whole block — a [marked-prefix, ""] split would put an empty
# text block on the wire (HTTP 400 on native Anthropic).
_apply_cache_marker(message, cache_marker, native_anthropic=native_anthropic)
return 1
@@ -345,27 +246,17 @@ def strip_anthropic_cache_control(
) -> List[Dict[str, Any]]:
"""Remove ``cache_control`` markers and undo decoration-produced list shapes.
Used before re-applying decoration after a mid-turn provider failover so
the mutated, undecorated shape (image shrink / ASCII cleanup / etc.) is
preserved while markers match the *new* provider's cache policy (#72626).
Used before re-decorating after a mid-turn provider failover, so the
mutated undecorated shape is preserved while markers match the new
provider's policy. Flattening back to a plain string is restricted to the
exact shapes :func:`apply_anthropic_cache_control` produces from string
content — a single text part, the two-part ``[static, volatile]`` system
split, or the two-part skill split — so the ``""``-join is provably
byte-exact; organic multi-part text and parts with extra keys keep their
structure. Marker removal is copy-on-write on part dicts: parts can alias
caller-held lists and stripping must never rewrite the stored transcript.
Flattening back to a plain string is restricted to the exact shapes
:func:`apply_anthropic_cache_control` produces from string content —
a single ``{"type": "text"}`` part, the two-part ``[static, volatile]``
system split, or the two-part builder-declared skill split (recognised
by its marker-on-the-first-part shape, so flattening never depends on
the prefix registry still holding the entry) — so the ``""``-join is
provably byte-exact. Organic
multi-part text (merged user turns, imported transcripts) and parts
carrying extra keys (``citations`` etc.) keep their structure; only
per-part markers are removed. Marker removal is copy-on-write on the
part dicts: content parts can alias caller-held message lists (the main
send path now hands structurally-cloned copies via
_clone_message_for_send, but other callers may pass shallow copies),
and stripping must never rewrite the stored transcript.
Mutates the top-level message dicts of ``api_messages`` in place and
returns the same list.
Mutates the top-level message dicts in place and returns the same list.
"""
for msg in api_messages:
if not isinstance(msg, dict):
@@ -374,13 +265,11 @@ def strip_anthropic_cache_control(
content = msg.get("content")
if not isinstance(content, list):
continue
# Two-part skill-invocation split (#81867). The builder-declared
# boundary is the only decoration that marks the *first* part of a
# user message: list content otherwise receives its marker on the
# last part, and the two-part [static, volatile] split is role-gated
# to system. So the shape alone identifies it, and flattening stays
# correct even when the prefix registry has since evicted the entry
# (failover re-decorates a request built many messages ago, #72626).
# The builder-declared skill split is the only decoration that marks
# the FIRST part of a user message (list content is otherwise marked
# on the last part; the [static, volatile] split is system-only), so
# the shape alone identifies it even after the prefix registry has
# evicted the entry.
skill_split_shape = (
msg.get("role") == "user"
and len(content) == 2
@@ -505,9 +394,9 @@ def build_prompt_cache_plan(
) -> PromptCachePlan:
"""Build isolated cache sections for one resolved request destination.
``tool_part_markers=False`` (LiteLLM-style envelope routes, #89886)
keeps ``cache_control`` off role:tool content parts; breakpoints
reallocate to the nearest eligible non-tool message.
``tool_part_markers=False`` (LiteLLM-style envelope routes) keeps
``cache_control`` off role:tool content parts; breakpoints reallocate to
the nearest eligible non-tool message.
"""
messages = copy.deepcopy(api_messages or [])
strip_anthropic_cache_control(messages)
@@ -558,23 +447,13 @@ def apply_anthropic_cache_control(
) -> List[Dict[str, Any]]:
"""Apply Anthropic cache-control markers to API messages.
When ``static_system_prefix`` exactly matches the beginning of a string
system prompt, it receives an early marker and the full system prompt gets
a trailing marker. The remaining two markers target the latest cacheable
non-system messages. Without that prefix, the legacy system-and-3 layout
is retained.
Idempotent: pre-existing ``cache_control`` markers are stripped from a
per-message copy before new ones are placed, so calling this twice (or
handing it messages a prior call already marked) can never accumulate
past 4 markers. Only messages that already carry a marker pay the copy
cost — a shallow top-level copy suffices because
:func:`strip_anthropic_cache_control` is copy-on-write on content parts —
and the rest of the copy-on-write contract is unchanged (#90971).
``tool_part_markers=False`` (LiteLLM-style envelope routes, #89886)
keeps markers off role:tool messages entirely; the breakpoint budget
reallocates to the nearest eligible non-tool message.
With a matching ``static_system_prefix`` the prefix gets an early marker
and the full system prompt a trailing one; the remaining two markers go to
the latest cacheable non-system messages. Without it, the legacy
system-and-3 layout applies. Idempotent: pre-existing markers are stripped
from a per-message copy first, so repeated calls never accumulate past 4
markers; a shallow top-level copy suffices because
:func:`strip_anthropic_cache_control` is copy-on-write on content parts.
Returns:
Shallow copy of message list with selective deep copies of modified messages.
@@ -594,9 +473,6 @@ def apply_anthropic_cache_control(
and any(isinstance(part, dict) and "cache_control" in part for part in content)
)
if has_marker:
# Shallow top-level copy is enough: strip pops the top-level key
# and rebuilds content lists/part dicts copy-on-write, so the
# caller's message (and any aliased parts) are never mutated.
messages[i] = strip_anthropic_cache_control([dict(msg)])[0]
breakpoints_used = 0
+415 -766
View File
File diff suppressed because it is too large Load Diff
+71 -129
View File
@@ -1,18 +1,11 @@
"""Replay-history sanitization shared across resume code paths.
When a session's last turn dies mid-tool-loop — the process is killed by a
restart/shutdown command, a stale-timeout fires, or an interrupt lands before
the tool result is written — the persisted transcript can end with a dangling
``assistant(tool_calls)`` (no matching ``tool`` answer) or an interrupted
``assistant→tool`` block. On resume the model sees that broken tail and
re-issues the unanswered call, producing an endless "thinking"/reboot loop
(#49201, #29086).
These pure helpers strip those tails before the history is replayed to the
model. They were originally local to ``gateway/run.py`` (which fixed the
messaging-gateway path) and are extracted here so every resume surface — the
messaging gateway AND the TUI/WebUI gateway — shares the same cleanup instead
of the WebUI path silently skipping it.
A session whose last turn died mid-tool-loop (process killed by a restart
command, stale timeout, interrupt before the tool result was written) persists
a dangling ``assistant(tool_calls)`` or interrupted ``assistant→tool`` tail. On
resume the model re-issues the unanswered call → endless "thinking"/reboot loop.
These pure helpers strip those tails before replay, for EVERY resume surface
(messaging gateway and TUI/WebUI gateway alike).
"""
from __future__ import annotations
@@ -39,16 +32,36 @@ def is_interrupted_tool_result(content: Any) -> bool:
return False
def _call_name(call: Dict[str, Any]) -> str:
return str((call.get("function") or {}).get("name") or "")
def _call_id(call: Dict[str, Any]) -> str:
return str(call.get("id") or call.get("call_id") or "")
def _any_side_effecting(calls: List[Dict[str, Any]]) -> bool:
return any(tool_may_have_side_effect(_call_name(call)) for call in calls)
def _orphan_recovery(name: str, unknown_text: str, none_text: str) -> tuple:
"""(effect_disposition, content) for an interrupted/dangling call named ``name``."""
if tool_may_have_side_effect(name):
return "unknown", unknown_text
return "none", none_text
def strip_interrupted_tool_tails(
agent_history: List[Dict[str, Any]],
) -> List[Dict[str, Any]]:
"""Strip interrupted assistant→tool sequences from replay history.
Older interrupted gateway turns can be followed by a queued real user
message, so the interrupted assistant/tool block is not necessarily the
final tail by the time we rebuild replay history. Remove any contiguous
assistant(tool_calls) + tool-result block that contains an interrupted tool
result, while preserving successful tool-call sequences intact.
The interrupted block is not necessarily the final tail (a queued real user
message may follow it), so every contiguous assistant(tool_calls)+tool-result
block containing an interrupted result is handled; successful sequences stay
intact. Read-only blocks are dropped; blocks with a side-effecting call are
KEPT with the interrupted results rewritten as orphan-recovery notices, since
the effect may already have happened and erasing it would hide that.
"""
if not agent_history:
return agent_history
@@ -69,18 +82,8 @@ def strip_interrupted_tool_tails(
for m in tool_results
):
calls = msg.get("tool_calls") or []
if any(
tool_may_have_side_effect(
str((call.get("function") or {}).get("name") or "")
)
for call in calls
):
call_names = {
str(call.get("id") or call.get("call_id") or ""): str(
(call.get("function") or {}).get("name") or ""
)
for call in calls
}
if _any_side_effecting(calls):
call_names = {_call_id(call): _call_name(call) for call in calls}
cleaned.append(msg)
for tool_result in tool_results:
if not is_interrupted_tool_result(tool_result.get("content", "")):
@@ -88,14 +91,11 @@ def strip_interrupted_tool_tails(
continue
recovered = dict(tool_result)
name = call_names.get(str(tool_result.get("tool_call_id") or ""), "")
recovered["effect_disposition"] = (
"unknown" if tool_may_have_side_effect(name) else "none"
)
recovered["content"] = (
recovered["effect_disposition"], recovered["content"] = _orphan_recovery(
name,
"[Orphan recovery: interrupted side-effecting tool may have "
"executed; its effect is UNKNOWN. Inspect state before retrying.]"
if recovered["effect_disposition"] == "unknown"
else "[Orphan recovery: interrupted read-only tool did not complete.]"
"executed; its effect is UNKNOWN. Inspect state before retrying.]",
"[Orphan recovery: interrupted read-only tool did not complete.]",
)
cleaned.append(recovered)
i = j
@@ -122,24 +122,13 @@ def strip_dangling_tool_call_tail(
) -> List[Dict[str, Any]]:
"""Strip a trailing ``assistant(tool_calls)`` block left with NO answers.
When a tool call itself kills the gateway process (``docker restart``,
``systemctl restart``, ``kill``, ``hermes gateway restart``), the process
is terminated by SIGKILL *mid-call* — before the tool result is ever
written and before the orderly shutdown rewind
(``_drop_trailing_empty_response_scaffolding``) can run. The last thing
persisted is the ``assistant`` message that issued the ``tool_calls``,
with zero matching ``tool`` rows.
On resume the model sees an unanswered tool call at the tail and naturally
re-issues it — which restarts the gateway again, producing the infinite
reboot loop in #49201. ``strip_interrupted_tool_tails`` does not catch
this because there is no tool result to inspect for an interrupt marker.
This strips that dangling tail at the source so there is nothing for the
model to re-execute. It only acts when the tail is an
``assistant(tool_calls)`` whose calls have NO corresponding ``tool``
results — a completed assistant→tool pair (any tool answers present) is
left untouched so genuine mid-progress tool loops still resume.
A tool call that kills the gateway process itself (``docker restart``,
``hermes gateway restart``) is SIGKILLed mid-call, before any tool result or
the orderly shutdown rewind; the persisted tail is the assistant message with
zero matching ``tool`` rows, which ``strip_interrupted_tool_tails`` cannot
detect (no result to inspect). Only acts when the tail has NO tool answers —
a partially answered block still resumes. Read-only tails are dropped;
side-effecting ones get synthetic UNKNOWN-effect results instead of erasure.
"""
if not agent_history:
return agent_history
@@ -153,26 +142,18 @@ def strip_dangling_tool_call_tail(
return agent_history
tool_calls = last.get("tool_calls") or []
if any(
tool_may_have_side_effect(
str((call.get("function") or {}).get("name") or "")
)
for call in tool_calls
):
if _any_side_effecting(tool_calls):
recovered = list(agent_history)
for call in tool_calls:
function = call.get("function") or {}
name = str(function.get("name") or "unknown")
call_id = str(call.get("id") or call.get("call_id") or "")
disposition = "unknown" if tool_may_have_side_effect(name) else "none"
content = (
name = str((call.get("function") or {}).get("name") or "unknown")
disposition, content = _orphan_recovery(
name,
"[Orphan recovery: this tool may have executed before Hermes stopped; "
"its effect is UNKNOWN. Inspect current state before retrying.]"
if disposition == "unknown"
else "[Orphan recovery: this read-only tool did not complete and had no effect.]"
"its effect is UNKNOWN. Inspect current state before retrying.]",
"[Orphan recovery: this read-only tool did not complete and had no effect.]",
)
recovered.append(make_tool_result_message(
name, content, call_id, effect_disposition=disposition,
name, content, _call_id(call), effect_disposition=disposition,
))
logger.warning(
"Recovered dangling side-effecting tool call(s) as UNKNOWN instead of erasing them"
@@ -189,31 +170,23 @@ def strip_dangling_tool_call_tail(
def sanitize_replay_history(
agent_history: List[Dict[str, Any]],
) -> List[Dict[str, Any]]:
"""Apply both replay-tail strippers in the canonical order.
Convenience entry point for resume code paths: removes interrupted
assistant→tool blocks anywhere in the history, then removes a dangling
unanswered ``assistant(tool_calls)`` tail. Returns the same list object
when there is nothing to strip.
"""
"""Both replay-tail strippers in canonical order (interrupted blocks, then
dangling tail). Returns the same list object when nothing is stripped."""
if not agent_history:
return agent_history
return strip_dangling_tool_call_tail(strip_interrupted_tool_tails(agent_history))
# ──────────────────────────────────────────────────────────────────────
# Stale dangerous-confirmation text expiry (#59607)
# Stale dangerous-confirmation text expiry
# ──────────────────────────────────────────────────────────────────────
# How long a high-risk confirmation phrase remains valid.
# Short on purpose: dangerous side effects should not survive any restart
# or session resumption gap. The user can always re-confirm if needed.
# Short on purpose: a dangerous confirmation must not survive any restart or
# resume gap. The user can always re-confirm.
_DANGEROUS_CONFIRMATION_EXPIRY_SECONDS = 60.0
# Confirmation phrases that unlock destructive host actions.
# Substring match (case-insensitive) so that user variants (e.g. trailing
# punctuation, additional context) still match. Add new patterns here when
# new high-risk actions are introduced.
# Confirmation phrases that unlock destructive host actions; case-insensitive
# substring match so trailing punctuation / extra context still matches.
_DANGEROUS_CONFIRMATION_PATTERNS: tuple = (
"confirm forced restart",
"confirm forced reboot",
@@ -229,9 +202,8 @@ _DANGEROUS_CONFIRMATION_PATTERNS: tuple = (
"確認重啟",
)
# Replacement text for an expired confirmation. Redacting in place (rather
# than deleting the message) preserves strict user/assistant role
# alternation in the replayed history.
# Redacting in place (rather than deleting the message) preserves strict
# user/assistant role alternation in the replayed history.
_EXPIRED_CONFIRMATION_SENTINEL = (
"[A high-risk confirmation previously given here has EXPIRED and must "
"not be acted on. Ask the user to re-confirm explicitly before "
@@ -240,12 +212,7 @@ _EXPIRED_CONFIRMATION_SENTINEL = (
def is_dangerous_confirmation(content: Any) -> bool:
"""Return True if a user-message text matches a known dangerous confirmation.
Used by ``strip_stale_dangerous_confirmations`` to decide which
transcript rows to expire. Substring + case-insensitive so that
``"Please confirm forced restart, the host is critical"`` still matches.
"""
"""True if user-message text contains a known dangerous confirmation phrase."""
if not isinstance(content, str):
return False
text = content.strip().lower()
@@ -258,38 +225,15 @@ def strip_stale_dangerous_confirmations(
now: float,
expiry_seconds: float = _DANGEROUS_CONFIRMATION_EXPIRY_SECONDS,
) -> List[Dict[str, Any]]:
"""Expire stale dangerous-confirmation text in user messages (#59607).
"""Expire stale dangerous-confirmation text in user messages.
When a high-risk side effect (e.g. host restart via ``shutdown.exe``)
runs, the user's plain-text confirmation phrase is persisted in the
conversation transcript. If the host restart killed the gateway
process before the assistant's tool result was written, the
transcript tail ends on the assistant's text response — and the
dangerous confirmation text remains in the user role.
On the next inbound message — possibly a casual "are you there?" from
the user minutes later — the LLM sees the stale confirmation and may
interpret the new turn as a fresh re-confirmation, re-executing the
destructive action. This is the failure mode reported in #59607.
Expired confirmations are REDACTED IN PLACE, not removed: deleting a
user message from the incident tail (``user(confirm) →
assistant("OK, restarting")``) would leave two consecutive assistant
messages, violating the strict role-alternation invariant providers
enforce. The message survives with its role intact; only the trigger
text is replaced by a sentinel that tells the model the confirmation
has expired.
Messages without a timestamp are left untouched (backward
compatibility: legacy transcripts and in-memory test scaffolding have
no timestamps). User messages that contain dangerous confirmation
text but are within the expiry window are also left untouched — they
represent a fresh confirmation that has not yet been acted on.
Complements 75ed07ace (which strips the *assistant* side of the
broken tail) by handling the *user* side: a stale plain-text
confirmation that the assistant has not yet responded to in a way
the resume logic recognises.
If a host restart killed the gateway before the tool result was written, the
user's confirmation phrase survives in the transcript; a casual "are you
there?" minutes later can read to the model as a fresh re-confirmation and
re-execute the destructive action. Expired confirmations are REDACTED IN
PLACE (deleting the message would leave two consecutive assistant turns).
Messages without a timestamp (legacy transcripts, test scaffolding) and
confirmations still inside the expiry window are left untouched.
"""
if not agent_history:
return agent_history
@@ -312,10 +256,8 @@ def strip_stale_dangerous_confirmations(
)
redacted = dict(msg)
redacted["content"] = _EXPIRED_CONFIRMATION_SENTINEL
# Drop the api_content sidecar: it carries the exact bytes
# previously sent — i.e. the dangerous confirmation this
# redaction exists to expire. Replaying it verbatim would
# undo the redaction on the wire.
# The api_content sidecar carries the exact bytes previously sent
# — the confirmation itself; replaying it would undo the redaction.
drop_stale_api_content(redacted)
cleaned.append(redacted)
continue
+39 -47
View File
@@ -1,13 +1,10 @@
"""Single source of truth for the agent working directory.
`TERMINAL_CWD` is the runtime carrier for the configured working directory
(design #19214/#19242: `terminal.cwd` is bridged once to `TERMINAL_CWD` at
gateway/cron startup). The local-CLI backend deliberately leaves it unset and
relies on the launch dir. Reading it in one place keeps the system prompt, the
tool surfaces, and context-file discovery agreeing on where the agent lives.
Multi-session gateways can pin a logical cwd via the `_SESSION_CWD`
contextvar; CLI/cron fall through to `TERMINAL_CWD`/launch cwd.
(`terminal.cwd` is bridged to it once at gateway/cron startup; the local CLI
leaves it unset and relies on the launch dir). Reading it in one place keeps the
system prompt, tool surfaces, and context-file discovery agreeing on where the
agent lives. Multi-session gateways can pin a logical cwd via `_SESSION_CWD`.
"""
import logging
@@ -22,18 +19,15 @@ _UNSET: Any = object()
_SESSION_CWD: ContextVar = ContextVar("HERMES_SESSION_CWD", default=_UNSET)
# The Python package/source root (this file lives at <root>/agent/runtime_cwd.py).
# When a backend is launched from, or self-spawns into, this tree (the desktop
# app default), an os.getcwd() fallback would inject this repo's contributor
# AGENTS.md as authoritative project context. Context discovery must never
# resolve here.
# The package/source root (<root>/agent/runtime_cwd.py). A backend launched from
# or self-spawned into this tree (desktop default) must never let an os.getcwd()
# fallback inject this repo's contributor AGENTS.md as project context.
_PACKAGE_ROOT = Path(__file__).resolve().parent.parent
def _is_install_tree(p: Path) -> bool:
# True only when p IS the package root or sits inside it. Ancestors of the
# package root (a user home that happens to contain the checkout, a --user
# site-packages parent) are legitimate workspaces and must not be blocked.
"""True only when ``p`` IS the package root or sits inside it — ancestors
(a home dir containing the checkout) are legitimate workspaces."""
try:
p = p.resolve()
except Exception:
@@ -61,9 +55,9 @@ def _terminal_cwd_env() -> str:
"""Scope-aware TERMINAL_CWD read (tools.terminal_scope.terminal_env).
Under gateway multiplexing the per-turn terminal scope carries the active
profile's cwd; the process-global env var may hold another profile's
value. Only an import failure falls back: an active refusal scope must
raise, not silently resolve the launch profile's cwd.
profile's cwd; the process-global env var may hold another profile's. Only
an ImportError falls back: an active refusal scope must raise, not silently
resolve the launch profile's cwd.
"""
try:
from tools.terminal_scope import terminal_env
@@ -75,51 +69,49 @@ def _terminal_cwd_env() -> str:
def scope_terminal_cwd() -> str:
"""Public wrapper — the scope-aware TERMINAL_CWD value (may be empty).
Shared by agent_init / skill_utils / code_execution_tool so every cwd
consumer reads through the per-turn terminal scope under gateway
multiplexing instead of the process-global env var.
Shared by agent_init / skill_utils / code_execution_tool so every cwd consumer
reads through the per-turn terminal scope under gateway multiplexing.
"""
return _terminal_cwd_env()
def resolve_agent_cwd() -> Path:
def _resolve_configured_cwd(*, override_is_final: bool) -> Path | None:
"""Session override, then TERMINAL_CWD; each validated as a real directory.
``override_is_final``: a set-but-missing session override yields None
instead of falling through to TERMINAL_CWD.
"""
override = _session_cwd_override()
if override:
p = Path(override).expanduser()
if p.is_dir():
return p
logger.warning("configured working directory does not exist: %s", override)
if override_is_final:
return None
raw = _terminal_cwd_env().strip()
if raw:
p = Path(raw).expanduser()
if p.is_dir():
return p
logger.warning("TERMINAL_CWD does not exist: %s", raw)
return Path(os.getcwd())
return None
def resolve_agent_cwd() -> Path:
"""Configured cwd, else the launch dir (os.getcwd() — its OSError on a
deleted cwd deliberately propagates; the caller owns that guard)."""
p = _resolve_configured_cwd(override_is_final=False)
return p if p is not None else Path(os.getcwd())
def resolve_context_cwd() -> Path | None:
# None means "no configured cwd": build_context_files_prompt then falls back
# to the launch dir (os.getcwd()), correct for a local CLI launched inside a
# real project. A configured path is validated here (previously it was passed
# through unchecked, diverging from resolve_agent_cwd). An explicitly
# configured path is otherwise honored verbatim — including the Hermes
# source tree itself, which is a legitimate workspace when the user is
# developing Hermes (per-surface policy for fallback-picked directories
# lives in build_context_files_prompt; see #64590).
override = _session_cwd_override()
if override:
p = Path(override).expanduser()
if not p.is_dir():
logger.warning("configured working directory does not exist: %s", override)
else:
return p
return None
raw = _terminal_cwd_env().strip()
if raw:
p = Path(raw).expanduser()
if not p.is_dir():
logger.warning("TERMINAL_CWD does not exist: %s", raw)
else:
return p
return None
"""Configured cwd for context-file discovery, or None for "no configured cwd".
None makes build_context_files_prompt fall back to the launch dir (correct
for a local CLI launched inside a real project). A configured path is
validated here; an existing one is honored verbatim — including the Hermes
source tree itself, a legitimate workspace when developing Hermes
(fallback-directory policy lives in build_context_files_prompt).
"""
return _resolve_configured_cwd(override_is_final=True)
+50 -196
View File
@@ -1,87 +1,38 @@
"""Skill bundles — aliases that load multiple skills under one slash command.
A skill bundle is a small YAML file that names a set of skills to load
together. Invoking ``/<bundle-name>`` from the CLI or gateway loads every
referenced skill's full content into a single user message, the same way
``/<skill-name>`` does — but for N skills at once.
Storage
-------
Bundles live in ``~/.hermes/skill-bundles/*.yaml`` (and the equivalent
profile-aware directory under ``HERMES_HOME``). Each file looks like::
name: backend-dev
description: Backend feature work — code review, testing, PR workflow.
skills:
- github-code-review
- test-driven-development
- github-pr-workflow
instruction: |
Optional extra guidance to inject above the skill bodies.
The file's stem is treated as a fallback name when ``name:`` is absent, so
dropping a YAML into the directory is enough to register a new bundle.
Conflict resolution
-------------------
If a bundle and a skill share the same slash name, the bundle wins. The
slash command dispatch checks bundles first, then falls back to skills.
This is the intended behavior — a user who names a bundle ``research``
explicitly wants ``/research`` to mean their bundle, not whatever skill
happens to share the slug.
Public API
----------
- :func:`get_skill_bundles` — return ``{"/slug": bundle_info}``
- :func:`resolve_bundle_command_key` — map a user-typed command to its slug
- :func:`build_bundle_invocation_message` — produce the full user message
- :func:`reload_bundles` — re-scan disk and return a diff
- :func:`list_bundles` — return rich info for display (``hermes bundles``)
- :func:`save_bundle` / :func:`delete_bundle` — file-level operations
Bundles are YAML files in ``<HERMES_HOME>/skill-bundles/`` (``name``,
``description``, ``skills: [...]``, optional ``instruction``; the file stem is
the fallback name). ``/<bundle>`` loads every member skill into one user
message. If a bundle and a skill share a slug, the bundle wins — slash dispatch
checks bundles first, on purpose.
"""
from __future__ import annotations
import logging
import os
import re
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
import yaml
from hermes_constants import get_hermes_home
from agent.skill_commands import diff_command_snapshots, slugify_skill_name as _slugify
logger = logging.getLogger(__name__)
# Slug normalization — matches agent/skill_commands.py so a bundle and a
# skill called "Foo Bar" both resolve to "/foo-bar".
_BUNDLE_INVALID_CHARS = re.compile(r"[^a-z0-9-]")
_BUNDLE_MULTI_HYPHEN = re.compile(r"-{2,}")
_bundles_cache: Dict[str, Dict[str, Any]] = {}
_bundles_cache_mtime: Optional[float] = None
def _bundles_dir() -> Path:
"""Return the canonical bundles directory under HERMES_HOME.
Honors ``HERMES_BUNDLES_DIR`` for tests; falls back to
``<HERMES_HOME>/skill-bundles``.
"""
"""Bundles directory: ``HERMES_BUNDLES_DIR`` override (tests) or ``<HERMES_HOME>/skill-bundles``."""
override = os.environ.get("HERMES_BUNDLES_DIR")
if override:
return Path(override).expanduser()
return get_hermes_home() / "skill-bundles"
def _slugify(name: str) -> str:
cmd = name.lower().replace(" ", "-").replace("_", "-")
cmd = _BUNDLE_INVALID_CHARS.sub("", cmd)
cmd = _BUNDLE_MULTI_HYPHEN.sub("-", cmd).strip("-")
return cmd
def _iter_bundle_files() -> List[Path]:
base = _bundles_dir()
if not base.exists():
@@ -93,19 +44,9 @@ def _iter_bundle_files() -> List[Path]:
def _max_mtime(files: List[Path]) -> float:
"""Highest mtime across the bundle files plus the dir itself.
Watching the directory mtime catches deletions; watching individual
files catches edits. Together they're a cheap freshness check.
"""
base = _bundles_dir()
"""Highest mtime across the bundle files plus the dir itself (dir mtime catches deletions)."""
mtimes = []
if base.exists():
try:
mtimes.append(base.stat().st_mtime)
except OSError:
pass
for f in files:
for f in [_bundles_dir(), *files]:
try:
mtimes.append(f.stat().st_mtime)
except OSError:
@@ -114,11 +55,7 @@ def _max_mtime(files: List[Path]) -> float:
def _load_bundle_file(path: Path) -> Optional[Dict[str, Any]]:
"""Parse a single bundle YAML file. Returns ``None`` on any error.
Errors are logged at WARNING level. We don't raise — a broken bundle
shouldn't take down slash command discovery.
"""
"""Parse one bundle YAML; ``None`` (logged) on any error so a broken bundle can't break discovery."""
try:
raw = path.read_text(encoding="utf-8")
except OSError as exc:
@@ -166,12 +103,7 @@ def _load_bundle_file(path: Path) -> Optional[Dict[str, Any]]:
def scan_bundles() -> Dict[str, Dict[str, Any]]:
"""Scan the bundles directory and rebuild the cache.
Returns the same mapping as :func:`get_skill_bundles` — ``"/slug"`` →
bundle info dict. Later bundles with a duplicate slug are skipped with
a warning (first wins, alphabetical order).
"""
"""Rebuild the ``"/slug"`` -> bundle info cache; duplicate slugs keep the first (alphabetical)."""
global _bundles_cache, _bundles_cache_mtime
files = _iter_bundle_files()
out: Dict[str, Dict[str, Any]] = {}
@@ -193,25 +125,15 @@ def scan_bundles() -> Dict[str, Dict[str, Any]]:
def get_skill_bundles() -> Dict[str, Dict[str, Any]]:
"""Return the current bundle mapping, rescanning when disk changed.
Cheap to call repeatedly: only rescans when the bundles directory or
any bundle file's mtime is newer than the cached snapshot.
"""
files = _iter_bundle_files()
current_mtime = _max_mtime(files)
"""Current bundle mapping; rescans only when a bundle file or the dir mtime changed."""
current_mtime = _max_mtime(_iter_bundle_files())
if not _bundles_cache or _bundles_cache_mtime != current_mtime:
scan_bundles()
return _bundles_cache
def resolve_bundle_command_key(command: str) -> Optional[str]:
"""Resolve a user-typed command to its canonical bundle slash key.
Hyphens and underscores are treated interchangeably to mirror the
skill-command behavior (Telegram converts hyphens to underscores in
bot command names).
"""
"""Resolve a user-typed command to its ``/slug`` key (``_`` ≡ ``-``, as Telegram rewrites hyphens)."""
if not command:
return None
cmd_key = f"/{command.replace('_', '-')}"
@@ -219,35 +141,17 @@ def resolve_bundle_command_key(command: str) -> Optional[str]:
def reload_bundles() -> Dict[str, Any]:
"""Re-scan the bundles directory and return a diff.
Mirrors :func:`agent.skill_commands.reload_skills` so callers can use
the same display logic. Returns a dict with ``added``, ``removed``,
``unchanged``, and ``total`` keys.
"""
"""Re-scan and return an ``added``/``removed``/``unchanged``/``total`` diff (same shape as reload_skills)."""
def _snapshot(cmds: Dict[str, Dict[str, Any]]) -> Dict[str, str]:
return {k.lstrip("/"): (v or {}).get("description", "") for k, v in cmds.items()}
before = _snapshot(_bundles_cache)
new = scan_bundles()
after = _snapshot(new)
added_names = sorted(set(after) - set(before))
removed_names = sorted(set(before) - set(after))
unchanged = sorted(set(after) & set(before))
return {
"added": [{"name": n, "description": after[n]} for n in added_names],
"removed": [{"name": n, "description": before[n]} for n in removed_names],
"unchanged": unchanged,
"total": len(after),
}
return diff_command_snapshots(before, _snapshot(scan_bundles()))
def list_bundles() -> List[Dict[str, Any]]:
"""Return a sorted list of bundle info dicts for display."""
bundles = get_skill_bundles()
return sorted(bundles.values(), key=lambda b: b["slug"])
return sorted(get_skill_bundles().values(), key=lambda b: b["slug"])
def build_bundle_invocation_message(
@@ -256,34 +160,20 @@ def build_bundle_invocation_message(
task_id: str | None = None,
platform: str | None = None,
) -> Optional[Tuple[str, List[str], List[str]]]:
"""Build the user message content for a bundle slash command invocation.
"""Build the user message for a bundle invocation.
Returns ``(message, loaded_skill_names, missing_skill_names)`` or
``None`` if the bundle wasn't found.
A bundle that references skills the user doesn't have installed still
loads — the agent gets a note about which ones were skipped. This is
the same forgiving stance ``build_preloaded_skills_prompt`` uses for
``-s`` CLI preloading.
Disabled skills are also skipped: bundles load members via
``_load_skill_payload`` directly, bypassing the scan-time disabled
filter in ``get_skill_commands()``, so the disabled list must be
re-applied here. ``platform`` scopes the check to a specific
platform's ``skills.platform_disabled`` config (gateway dispatch
passes it explicitly because the gateway handles multiple platforms
in one process); when *None*, the platform resolves from session env
vars and the global disabled list still applies. Mirrors the
stacked-skill gate in gateway dispatch (#58888).
Returns ``(message, loaded_skill_names, missing_skill_names)`` or ``None``
if the bundle wasn't found. Uninstalled members are skipped with a note.
Disabled members are skipped too: bundles load via ``_load_skill_payload``,
bypassing the scan-time disabled filter, so the list is re-applied here.
``platform`` scopes that check (gateway passes it; None resolves from env).
"""
bundles = get_skill_bundles()
info = bundles.get(cmd_key)
info = get_skill_bundles().get(cmd_key)
if not info:
return None
# Late import to avoid pulling tools/* at module import time and to
# keep skill_bundles cheap to import in test environments.
from agent.skill_commands import _load_skill_payload, _build_skill_message
# Late import keeps skill_bundles cheap to import (no tools/* at import time).
from agent.skill_commands import _load_skill_payload, _render_skill_block, _scaffold_header
try:
from agent.skill_utils import get_disabled_skill_names
@@ -298,10 +188,8 @@ def build_bundle_invocation_message(
seen: set[str] = set()
bundle_name = info["name"]
skills = info["skills"]
extra_instruction = info.get("instruction") or ""
for skill_id in skills:
for skill_id in info["skills"]:
identifier = (skill_id or "").strip()
if not identifier or identifier in seen:
continue
@@ -311,66 +199,36 @@ def build_bundle_invocation_message(
if not loaded:
missing.append(identifier)
continue
loaded_skill, skill_dir, skill_name = loaded
skill_name = loaded[2]
# Per-platform / global disabled gate. Checked against the loaded
# skill's canonical name (identifiers may be paths or aliases).
# Gate on the loaded skill's canonical name (identifiers may be paths or aliases).
if skill_name in disabled_names or identifier in disabled_names:
disabled.append(skill_name or identifier)
continue
try:
from tools.skill_usage import bump_use
bump_use(skill_name, task_id=task_id)
except Exception:
pass
activation_note = (
f'[Loaded as part of the "{bundle_name}" skill bundle.]'
)
skill_blocks.append(
_build_skill_message(
loaded_skill,
skill_dir,
activation_note,
session_id=task_id,
)
)
skill_blocks.append(_render_skill_block(
loaded,
f'[Loaded as part of the "{bundle_name}" skill bundle.]',
task_id,
))
loaded_names.append(skill_name)
if not skill_blocks:
return None
# Header — tells the agent this is a bundle, lists the skills, and
# provides any author-supplied instruction.
header_lines = [
f'[IMPORTANT: The user has invoked the "{bundle_name}" skill bundle, '
f"loading {len(loaded_names)} skills together. Treat every skill below "
"as active guidance for this turn.]",
"",
f"Bundle: {bundle_name}",
f"Skills loaded: {', '.join(loaded_names)}",
]
if missing:
header_lines.append(f"Skills missing (skipped): {', '.join(missing)}")
if disabled:
header_lines.append(
f"Skills disabled for this platform (skipped): {', '.join(disabled)}"
)
if extra_instruction:
header_lines.extend(["", f"Bundle instruction: {extra_instruction}"])
if user_instruction:
header_lines.extend(
["", f"User instruction: {user_instruction}"]
)
header = "\n".join(header_lines)
header = _scaffold_header(
f'"{bundle_name}" skill bundle',
loaded_names,
lead_lines=[f"Bundle: {bundle_name}"],
missing=missing,
disabled=disabled,
extra_instruction=info.get("instruction") or "",
user_instruction=user_instruction,
)
return ("\n\n".join([header, *skill_blocks]), loaded_names, missing)
# ---------------------------------------------------------------------------
# File-level CRUD helpers — used by `hermes bundles` CLI subcommand.
# ---------------------------------------------------------------------------
# ── File-level CRUD — used by `hermes bundles` ─────────────────────────────
def bundle_path_for(name: str) -> Path:
@@ -388,10 +246,10 @@ def save_bundle(
instruction: str = "",
overwrite: bool = False,
) -> Path:
"""Write a bundle to disk and invalidate the cache.
"""Write a bundle to disk and refresh the cache.
Raises ``FileExistsError`` if the target exists and ``overwrite`` is
False. Raises ``ValueError`` if the inputs are unusable.
Raises ``FileExistsError`` if the target exists and not ``overwrite``;
``ValueError`` for unusable inputs.
"""
name = (name or "").strip()
if not name:
@@ -415,15 +273,12 @@ def save_bundle(
yaml.safe_dump(payload, sort_keys=False, allow_unicode=True),
encoding="utf-8",
)
scan_bundles() # refresh cache
scan_bundles()
return path
def delete_bundle(name: str) -> Path:
"""Delete a bundle by name. Returns the deleted path.
Raises ``FileNotFoundError`` if the bundle doesn't exist.
"""
"""Delete a bundle by name and return its path; ``FileNotFoundError`` if absent."""
path = bundle_path_for(name)
if not path.exists():
raise FileNotFoundError(f"No bundle at {path}")
@@ -434,5 +289,4 @@ def delete_bundle(name: str) -> Path:
def get_bundle(name: str) -> Optional[Dict[str, Any]]:
"""Look up a bundle by name (slug-normalized)."""
slug = _slugify(name)
return get_skill_bundles().get(f"/{slug}")
return get_skill_bundles().get(f"/{_slugify(name)}")
+310 -452
View File
File diff suppressed because it is too large Load Diff
+244 -543
View File
File diff suppressed because it is too large Load Diff
+56 -137
View File
@@ -1,16 +1,11 @@
"""Progressive subdirectory hint discovery.
As the agent navigates into subdirectories via tool calls (read_file, terminal,
search_files, etc.), this module discovers and loads project context files
(AGENTS.md, CLAUDE.md, .cursorrules) from those directories. Discovered hints
are appended to the tool result so the model gets relevant context at the moment
it starts working in a new area of the codebase.
This complements the startup context loading in ``prompt_builder.py`` which only
loads from the CWD. Subdirectory hints are discovered lazily and injected into
the conversation without modifying the system prompt (preserving prompt caching).
Inspired by Block/goose's SubdirectoryHintTracker.
As the agent navigates into subdirectories via tool calls, this module loads
project context files (AGENTS.md, CLAUDE.md, .cursorrules) from those
directories and appends them to the tool result — context arrives without
touching the system prompt (preserving prompt caching). Complements the
startup CWD-only loading in ``prompt_builder.py``. Inspired by goose's
SubdirectoryHintTracker.
"""
import hashlib
@@ -24,32 +19,21 @@ from agent.prompt_builder import _scan_context_content
logger = logging.getLogger(__name__)
# Context files to look for in subdirectories, in priority order.
# Same filenames as prompt_builder.py but we load ALL found (not first-wins)
# since different subdirectories may use different conventions.
# Same filenames as prompt_builder.py, in priority order (first match wins per dir).
_HINT_FILENAMES = [
"AGENTS.override.md",
"AGENTS.md", "agents.md",
"CLAUDE.md", "claude.md",
".cursorrules",
]
# Maximum chars per hint file to prevent context bloat
_MAX_HINT_CHARS = 8_000
# Tool argument keys that typically contain file paths
_PATH_ARG_KEYS = {"path", "file_path", "workdir"}
# Tools that take shell commands where we should extract paths
_COMMAND_TOOLS = {"terminal"}
# How many parent directories to walk up when looking for hints.
# Prevents scanning all the way to / for deeply nested paths.
# Ancestor levels walked per path — bounds the scan for deeply nested paths.
_MAX_ANCESTOR_WALK = 5
# Directory names that never contain authoritative project context.
# Backups, vendored deps, VCS internals, and caches routinely hold *copies* of
# AGENTS.md; loading those duplicates real context and inflates the prompt.
# Directories that hold *copies* of context files (backups, vendored deps,
# VCS internals, caches), never authoritative project context.
_EXCLUDED_DIR_NAMES = frozenset({
"node_modules", "venv", ".venv", "__pycache__",
".git", ".hg", ".svn",
@@ -61,7 +45,7 @@ _EXCLUDED_DIR_NAMES = frozenset({
def _is_ancestor_or_same(a: Path, b: Path) -> bool:
"""Check if *a* is the same as or an ancestor of *b* (parent directory check)."""
"""True if *a* is *b* or one of its ancestors."""
try:
b.relative_to(a)
return True
@@ -72,34 +56,21 @@ def _is_ancestor_or_same(a: Path, b: Path) -> bool:
class SubdirectoryHintTracker:
"""Track which directories the agent visits and load hints on first access.
Usage::
tracker = SubdirectoryHintTracker(working_dir="/path/to/project")
# After each tool call:
hints = tracker.check_tool_call("read_file", {"path": "backend/src/main.py"})
if hints:
tool_result += hints # append to the tool result string
Usage: after each tool call, ``hints = tracker.check_tool_call(name, args)``
and append the returned text to the tool result.
"""
def __init__(self, working_dir: Optional[str] = None):
self.working_dir = Path(working_dir or os.getcwd()).resolve()
self._loaded_dirs: Set[Path] = set()
# Content digests already injected — prevents re-sending the same file
# reachable through symlinks, hardlinks, or duplicated copies.
# The working dir is pre-marked loaded (startup context handles it).
self._loaded_dirs: Set[Path] = {self.working_dir}
# Content digests already injected: the same file reached through
# symlinks/hardlinks/copies is never re-sent.
self._loaded_digests: Set[str] = set()
# Pre-mark the working dir as loaded (startup context handles it)
self._loaded_dirs.add(self.working_dir)
self._seed_working_dir_digest()
def _seed_working_dir_digest(self) -> None:
"""Record the CWD context file's digest so it is never re-injected.
``prompt_builder`` already loads the working directory's context file at
startup. Seeding its digest here means the same content reached through
a different path (a symlink farm, a shared workspace) is recognised as a
duplicate instead of being sent a second time.
"""
"""Record the CWD context file's digest (prompt_builder already loaded it)."""
for filename in _HINT_FILENAMES:
candidate = self.working_dir / filename
try:
@@ -119,23 +90,14 @@ class SubdirectoryHintTracker:
tool_name: str,
tool_args: Dict[str, Any],
) -> Optional[str]:
"""Check tool call arguments for new directories and load any hint files.
Returns formatted hint text to append to the tool result, or None.
"""
dirs = self._extract_directories(tool_name, tool_args)
if not dirs:
return None
"""Return formatted hint text for newly visited directories, or None."""
all_hints = []
for d in dirs:
for d in self._extract_directories(tool_name, tool_args):
hints = self._load_hints_for_directory(d)
if hints:
all_hints.append(hints)
if not all_hints:
return None
return "\n\n" + "\n\n".join(all_hints)
def _extract_directories(
@@ -143,39 +105,30 @@ class SubdirectoryHintTracker:
) -> list:
"""Extract directory paths from tool call arguments."""
candidates: Set[Path] = set()
# Direct path arguments
for key in _PATH_ARG_KEYS:
val = args.get(key)
if isinstance(val, str) and val.strip():
self._add_path_candidate(val, candidates)
# Shell commands — extract path-like tokens
if tool_name in _COMMAND_TOOLS:
cmd = args.get("command", "")
if isinstance(cmd, str):
self._extract_paths_from_command(cmd, candidates)
return list(candidates)
def _add_path_candidate(self, raw_path: str, candidates: Set[Path]):
"""Resolve a raw path and add its directory + ancestors to candidates.
"""Add a raw path's directory and its ancestors to candidates.
Walks up from the resolved directory toward the filesystem root,
stopping at the first directory already in ``_loaded_dirs`` (or after
``_MAX_ANCESTOR_WALK`` levels). This ensures that reading
``project/src/main.py`` discovers ``project/AGENTS.md`` even when
``project/src/`` has no hint files of its own.
Walks up toward the root, stopping at the first already-loaded
directory or after ``_MAX_ANCESTOR_WALK`` levels, so reading
``project/src/main.py`` still discovers ``project/AGENTS.md``.
"""
try:
p = Path(raw_path).expanduser()
if not p.is_absolute():
p = self.working_dir / p
p = p.resolve()
# Use parent if it's a file path (has extension or doesn't exist as dir)
if p.suffix or (p.exists() and p.is_file()):
p = p.parent
# Walk up ancestors — stop at already-loaded or root
for _ in range(_MAX_ANCESTOR_WALK):
if p in self._loaded_dirs:
break
@@ -189,32 +142,34 @@ class SubdirectoryHintTracker:
pass
def _extract_paths_from_command(self, cmd: str, candidates: Set[Path]):
"""Extract path-like tokens from a shell command string."""
"""Extract path-like tokens (contain / or .; not flags or URLs) from a shell command."""
try:
tokens = shlex.split(cmd)
except ValueError:
tokens = cmd.split()
for token in tokens:
# Skip flags
if token.startswith("-"):
continue
# Must look like a path (contains / or .)
if "/" not in token and "." not in token:
continue
# Skip URLs
if token.startswith(("http://", "https://", "git@")):
continue
self._add_path_candidate(token, candidates)
def _is_valid_subdir(self, path: Path) -> bool:
"""Check if path is a valid directory to scan for hints.
def _within_working_dir(self, path: Path) -> bool:
"""Reject paths outside the working-dir tree.
Only allow subdirectories within the working directory tree.
This prevents loading AGENTS.md from outside the active workspace
(e.g. ~/.codex/AGENTS.md, ~/.claude/CLAUDE.md), which causes
cross-agent context contamination and instruction mixup.
Loading ~/.codex/AGENTS.md or ~/.claude/CLAUDE.md would mix another
agent's instructions into this session. ``is_relative_to`` handles
symlinked paths; the ancestor check is a best-effort fallback.
"""
try:
return path.is_relative_to(self.working_dir)
except (OSError, ValueError):
return _is_ancestor_or_same(self.working_dir, path)
def _is_valid_subdir(self, path: Path) -> bool:
"""Directory inside the working-dir tree, not yet loaded, not an excluded copy dir."""
try:
if not path.is_dir():
return False
@@ -222,59 +177,31 @@ class SubdirectoryHintTracker:
return False
if path in self._loaded_dirs:
return False
# Reject paths outside the working directory tree.
# path.resolve() may differ from working_dir.resolve() due to symlinks,
# but path.is_relative_to(working_dir) handles both absolute and
# symlinked paths correctly on Python 3.9+.
try:
if not path.is_relative_to(self.working_dir):
return False
except (OSError, ValueError):
# Older Python or path resolution error — fall back to parent
# check as a best-effort safeguard.
if not _is_ancestor_or_same(self.working_dir, path):
return False
if self._is_excluded(path):
if not self._within_working_dir(path):
return False
return True
return not self._is_excluded(path)
def _is_excluded(self, path: Path) -> bool:
"""True when the path sits inside a directory that holds copies, not context.
"""True when a segment *below* the working dir is an excluded copy dir.
Directories the user is deliberately working inside are never excluded —
if ``working_dir`` is itself under ``vendor/``, that segment is legitimate
and only segments *below* the working dir are screened.
Only segments under ``working_dir`` are screened: a user deliberately
working inside ``vendor/`` keeps that segment legitimate.
"""
try:
rel_parts = path.relative_to(self.working_dir).parts
except ValueError:
# Paths outside the working dir are already rejected by
# _is_valid_subdir before this runs; treat as excluded defensively.
return True
return True # outside the tree — already rejected upstream
return any(part in _EXCLUDED_DIR_NAMES for part in rel_parts)
def _load_hints_for_directory(self, directory: Path) -> Optional[str]:
"""Load hint files from a directory. Returns formatted text or None.
Only loads hints from directories within the working directory tree.
"""
"""Load the first hint file in *directory*; formatted text or None."""
self._loaded_dirs.add(directory)
# Reject paths outside the working directory tree.
try:
if not directory.is_relative_to(self.working_dir):
logger.debug(
"Skipping hint files in %s — outside working_dir %s",
directory, self.working_dir,
)
return None
except (OSError, ValueError):
if not _is_ancestor_or_same(self.working_dir, directory):
logger.debug(
"Skipping hint files in %s — outside working_dir %s",
directory, self.working_dir,
)
return None
if not self._within_working_dir(directory):
logger.debug(
"Skipping hint files in %s — outside working_dir %s",
directory, self.working_dir,
)
return None
found_hints = []
for filename in _HINT_FILENAMES:
@@ -288,10 +215,6 @@ class SubdirectoryHintTracker:
content = hint_path.read_text(encoding="utf-8").strip()
if not content:
continue
# Skip content we've already injected. The same AGENTS.md is
# routinely reachable through several paths (symlinked shared
# workspaces, hardlinks, copied backups); re-sending it burns
# context for zero new information.
digest = hashlib.sha256(content.encode("utf-8")).hexdigest()
if digest in self._loaded_digests:
logger.debug(
@@ -301,14 +224,13 @@ class SubdirectoryHintTracker:
)
break
self._loaded_digests.add(digest)
# Same security scan as startup context loading
# Same security scan as startup context loading.
content = _scan_context_content(content, filename)
if len(content) > _MAX_HINT_CHARS:
content = (
content[:_MAX_HINT_CHARS]
+ f"\n\n[...truncated {filename}: {len(content):,} chars total]"
)
# Best-effort relative path for display
rel_path = str(hint_path)
try:
rel_path = str(hint_path.relative_to(self.working_dir))
@@ -320,20 +242,17 @@ class SubdirectoryHintTracker:
except (ValueError, RuntimeError):
pass # keep absolute
found_hints.append((rel_path, content))
# First match wins per directory (like startup loading)
break
break # first match wins per directory (like startup loading)
except Exception as exc:
logger.debug("Could not read %s: %s", hint_path, exc)
if not found_hints:
return None
sections = []
for rel_path, content in found_hints:
sections.append(
f"[Subdirectory context discovered: {rel_path}]\n{content}"
)
sections = [
f"[Subdirectory context discovered: {rel_path}]\n{content}"
for rel_path, content in found_hints
]
logger.debug(
"Loaded subdirectory hints from %s: %s",
directory,
+92 -226
View File
@@ -1,29 +1,10 @@
"""Stateful scrubber for reasoning/thinking blocks in streamed assistant text.
``run_agent._strip_think_blocks`` is regex-based and correct for a complete
string, but when it runs *per-delta* in ``_fire_stream_delta`` it destroys
the state that downstream consumers (CLI ``_stream_delta``, gateway
``GatewayStreamConsumer._filter_and_accumulate``) rely on.
Concretely, when MiniMax-M2.7 streams
delta1 = "<think>"
delta2 = "Let me check their config"
delta3 = "</think>"
the per-delta regex erases delta1 entirely (case 2: unterminated-open at
boundary matches ``^<think>...``), so the downstream state machine never
sees the open tag, treats delta2 as regular content, and leaks reasoning
to the user. Consumers that don't run their own state machine (ACP,
api_server, TTS) never had any defence at all — they just emitted
whatever survived the upstream regex.
This module centralises the tag-suppression state machine at the
upstream layer so every stream_delta_callback sees text that has
already had reasoning blocks removed. Partial tags at delta
boundaries are held back until the next delta resolves them, and
end-of-stream flushing surfaces any held-back prose that turned out
not to be a real tag.
The regex ``run_agent._strip_think_blocks`` is correct for a complete string but,
run per-delta, erases an opening ``<think>`` that arrives alone in one delta, so
downstream state machines never see the open tag and leak reasoning. This class
centralises tag suppression upstream: partial tags at delta boundaries are held
back until resolved, and ``flush()`` releases held-back prose that was not a tag.
Usage::
@@ -33,25 +14,15 @@ Usage::
if visible:
emit(visible)
tail = scrubber.flush() # at end of stream
if tail:
emit(tail)
The scrubber is re-entrant per agent instance. Call ``reset()`` at
the top of each new turn so a hung block from an interrupted prior
stream cannot taint the next turn's output.
Call ``reset()`` at the top of each turn so an interrupted block cannot taint
the next turn. Tags handled (case-insensitive): ``<think>``, ``<thinking>``,
``<reasoning>``, ``<thought>``, ``<REASONING_SCRATCHPAD>``.
Tag variants handled (case-insensitive):
``<think>``, ``<thinking>``, ``<reasoning>``, ``<thought>``,
``<REASONING_SCRATCHPAD>``.
Block-boundary rule for opens: an opening tag is only treated as a
reasoning-block opener when it appears at the start of the stream,
after a newline (optionally followed by whitespace), or when only
whitespace has been emitted on the current line. This prevents prose
that *mentions* the tag name (e.g. ``"use <think> tags here"``) from
being incorrectly suppressed. Closed pairs (``<think>X</think>``) are
always suppressed regardless of boundary; a closed pair is an
intentional, bounded construct.
Boundary rule: an opening tag only starts a block at a block boundary (stream
start, after a newline, or with only whitespace emitted on the current line), so
prose that *mentions* ``<think>`` is not suppressed. Closed pairs
(``<think>X</think>``) are always suppressed — a closed pair is intentional.
"""
from __future__ import annotations
@@ -64,16 +35,10 @@ __all__ = ["StreamingThinkScrubber"]
class StreamingThinkScrubber:
"""Stateful scrubber for streaming reasoning/thinking blocks.
State machine:
- ``_in_block``: True while inside an opened block, waiting for
a close tag. All text inside is discarded.
- ``_buf``: held-back partial-tag tail. Emitted / discarded on
the next ``feed()`` call or by ``flush()``.
- ``_last_emitted_ended_newline``: True iff the most recent
emission to the consumer ended with ``\\n``, or nothing has
been emitted yet (start-of-stream counts as a boundary). Used
to decide whether an open tag at buffer position 0 is at a
block boundary.
State: ``_in_block`` (inside an open block; text discarded), ``_buf``
(held-back partial-tag tail), ``_last_emitted_ended_newline`` (True iff the
last emission ended with ``\\n`` or nothing has been emitted yet — decides
whether an open tag at buffer position 0 sits at a block boundary).
"""
_OPEN_TAG_NAMES: Tuple[str, ...] = (
@@ -84,31 +49,33 @@ class StreamingThinkScrubber:
"REASONING_SCRATCHPAD",
)
# Materialise literal tag strings so the hot path does string
# operations, not regex compilation per feed().
# Literal tag strings so the hot path does string ops, not regex per feed().
_OPEN_TAGS: Tuple[str, ...] = tuple(f"<{name}>" for name in _OPEN_TAG_NAMES)
_CLOSE_TAGS: Tuple[str, ...] = tuple(f"</{name}>" for name in _OPEN_TAG_NAMES)
# Pre-compute the longest tag (for partial-tag hold-back bound).
_MAX_TAG_LEN: int = max(len(tag) for tag in _OPEN_TAGS + _CLOSE_TAGS)
def __init__(self) -> None:
self.reset()
def reset(self) -> None:
"""Reset all state. Call at the top of every new turn."""
self._in_block: bool = False
self._buf: str = ""
self._last_emitted_ended_newline: bool = True
def reset(self) -> None:
"""Reset all state. Call at the top of every new turn."""
self._in_block = False
self._buf = ""
self._last_emitted_ended_newline = True
def _emit(self, out: list[str], text: str) -> None:
"""Append visible prose to *out* (orphan close tags stripped) and track the newline flag."""
if text:
text = self._strip_orphan_close_tags(text)
if text:
out.append(text)
self._last_emitted_ended_newline = text.endswith("\n")
def feed(self, text: str) -> str:
"""Feed one delta; return the scrubbed visible portion.
May return an empty string when the entire delta is reasoning
content or is being held back pending resolution of a partial
tag at the boundary.
Returns "" when the whole delta is reasoning content or is held back
pending resolution of a partial tag at the boundary.
"""
if not text:
return ""
@@ -118,130 +85,67 @@ class StreamingThinkScrubber:
while buf:
if self._in_block:
# Hunt for the earliest close tag.
close_idx, close_len = self._find_first_tag(
buf, self._CLOSE_TAGS,
)
close_idx, close_len = self._find_first_tag(buf, self._CLOSE_TAGS)
if close_idx == -1:
# No close yet — hold back a potential partial
# close-tag prefix; discard everything else.
# No close yet: hold back a possible partial close-tag prefix, drop the rest.
held = self._max_partial_suffix(buf, self._CLOSE_TAGS)
self._buf = buf[-held:] if held else ""
return "".join(out)
# Found close: discard block content + tag, continue.
buf = buf[close_idx + close_len:]
self._in_block = False
continue
# Priority 1: closed <tag>X</tag> pair anywhere (no boundary gating —
# even inline pairs are almost certainly leaked reasoning).
# Priority 2: unterminated open tag at a block boundary (gated so
# prose that mentions '<think>' isn't over-stripped). Earliest wins.
pair = self._find_earliest_closed_pair(buf)
open_idx, open_len = self._find_open_at_boundary(buf, out)
if pair is not None and (open_idx == -1 or pair[0] <= open_idx):
self._emit(out, buf[:pair[0]])
buf = buf[pair[1]:]
continue
if open_idx != -1:
self._emit(out, buf[:open_idx])
self._in_block = True
buf = buf[open_idx + open_len:]
continue
# No resolvable tag: hold back any partial-tag prefix at the tail
# so a tag split across deltas isn't missed, then emit the rest.
held = max(
self._max_partial_suffix(buf, self._OPEN_TAGS),
self._max_partial_suffix(buf, self._CLOSE_TAGS),
)
if held:
self._emit(out, buf[:-held])
self._buf = buf[-held:]
else:
# Priority 1 — closed <tag>X</tag> pair anywhere in
# buf. Closed pairs are always an intentional,
# bounded construct (even mid-line prose containing
# an open/close pair is almost certainly a model
# leaking reasoning inline), so no boundary gating.
pair = self._find_earliest_closed_pair(buf)
# Priority 2 — unterminated open tag at a block
# boundary. Boundary-gated so prose that mentions
# '<think>' isn't over-stripped.
open_idx, open_len = self._find_open_at_boundary(
buf, out,
)
# Pick whichever match comes earliest in the buffer.
if pair is not None and (
open_idx == -1 or pair[0] <= open_idx
):
start_idx, end_idx = pair
preceding = buf[:start_idx]
if preceding:
preceding = self._strip_orphan_close_tags(preceding)
if preceding:
out.append(preceding)
self._last_emitted_ended_newline = (
preceding.endswith("\n")
)
buf = buf[end_idx:]
continue
if open_idx != -1:
# Unterminated open at boundary — emit preceding,
# enter block, continue loop with remainder.
preceding = buf[:open_idx]
if preceding:
preceding = self._strip_orphan_close_tags(preceding)
if preceding:
out.append(preceding)
self._last_emitted_ended_newline = (
preceding.endswith("\n")
)
self._in_block = True
buf = buf[open_idx + open_len:]
continue
# No resolvable tag structure in buf. Hold back any
# partial-tag prefix at the tail so a split tag
# across deltas isn't missed, then emit the rest.
held = self._max_partial_suffix(buf, self._OPEN_TAGS)
held_close = self._max_partial_suffix(
buf, self._CLOSE_TAGS,
)
held = max(held, held_close)
if held:
emit_text = buf[:-held]
self._buf = buf[-held:]
else:
emit_text = buf
self._buf = ""
if emit_text:
emit_text = self._strip_orphan_close_tags(emit_text)
if emit_text:
out.append(emit_text)
self._last_emitted_ended_newline = (
emit_text.endswith("\n")
)
return "".join(out)
self._emit(out, buf)
return "".join(out)
return "".join(out)
def flush(self) -> str:
"""End-of-stream flush.
If still inside an unterminated block, held-back content is
discarded — leaking partial reasoning is worse than a
truncated answer. Otherwise the held-back partial-tag tail is
emitted verbatim (it turned out not to be a real tag prefix).
Always treats the next ``feed()`` as a fresh stream boundary.
Intra-turn retries (thinking-only prefill, empty-response
retry) flush then stream again without calling ``reset()``;
leaving ``_last_emitted_ended_newline`` False made a new
stream's opening ``<think>`` look mid-line and leak into the
visible reply.
Inside an unterminated block the held-back content is discarded (leaking
partial reasoning is worse than a truncated answer); otherwise the
held-back tail is emitted verbatim. Always resets the boundary flag:
intra-turn retries flush then stream again without ``reset()``, and a
stale False flag made the new stream's opening ``<think>`` look mid-line.
"""
if self._in_block:
self._buf = ""
self._in_block = False
# Next feed() is a new stream — start-of-stream is a boundary.
self._last_emitted_ended_newline = True
return ""
tail = self._buf
tail = "" if self._in_block else self._buf
self._buf = ""
# Same for the non-block path: do NOT derive the boundary flag
# from the flushed tail (e.g. a held-back '<'). End-of-stream
# means the next feed() starts a new model response.
self._in_block = False
self._last_emitted_ended_newline = True
if not tail:
return ""
return self._strip_orphan_close_tags(tail)
return self._strip_orphan_close_tags(tail) if tail else ""
# ── internal helpers ───────────────────────────────────────────────
@staticmethod
def _find_first_tag(
buf: str, tags: Tuple[str, ...],
) -> Tuple[int, int]:
"""Return (earliest_index, tag_length) over *tags*, or (-1, 0).
Case-insensitive match.
"""
def _find_first_tag(buf: str, tags: Tuple[str, ...]) -> Tuple[int, int]:
"""Return (earliest_index, tag_length) over *tags* (case-insensitive), or (-1, 0)."""
buf_lower = buf.lower()
best_idx = -1
best_len = 0
@@ -253,14 +157,10 @@ class StreamingThinkScrubber:
return best_idx, best_len
def _find_earliest_closed_pair(self, buf: str):
"""Return (start_idx, end_idx) of the earliest closed pair, else None.
"""Return (start_idx, end_idx) of the earliest ``<tag>...</tag>`` pair, else None.
A closed pair is ``<tag>...</tag>`` of any variant. Matches are
case-insensitive and non-greedy (the closest close tag after
an open tag wins), matching the regex ``<tag>.*?</tag>``
semantics of ``_strip_think_blocks`` case 1. When two tag
variants could both match, the one whose open tag appears
earlier wins.
Case-insensitive and non-greedy (closest close after the open wins),
matching ``_strip_think_blocks`` case 1; the earliest open tag wins.
"""
buf_lower = buf.lower()
best: "tuple[int, int] | None" = None
@@ -270,23 +170,15 @@ class StreamingThinkScrubber:
open_idx = buf_lower.find(open_lower)
if open_idx == -1:
continue
close_idx = buf_lower.find(
close_lower, open_idx + len(open_lower),
)
close_idx = buf_lower.find(close_lower, open_idx + len(open_lower))
if close_idx == -1:
continue
end_idx = close_idx + len(close_lower)
if best is None or open_idx < best[0]:
best = (open_idx, end_idx)
best = (open_idx, close_idx + len(close_lower))
return best
def _find_open_at_boundary(
self, buf: str, already_emitted: list[str],
) -> Tuple[int, int]:
"""Return the earliest block-boundary open-tag (idx, len).
Returns (-1, 0) if no boundary-legal opener is present.
"""
def _find_open_at_boundary(self, buf: str, already_emitted: list[str]) -> Tuple[int, int]:
"""Return the earliest block-boundary open-tag (idx, len), or (-1, 0)."""
buf_lower = buf.lower()
best_idx = -1
best_len = 0
@@ -305,50 +197,30 @@ class StreamingThinkScrubber:
search_start = idx + 1
return best_idx, best_len
def _is_block_boundary(
self, buf: str, idx: int, already_emitted: list[str],
) -> bool:
def _is_block_boundary(self, buf: str, idx: int, already_emitted: list[str]) -> bool:
"""True iff position *idx* in *buf* is a block boundary.
A block boundary is:
- buf position 0 AND the most recent emission ended with
a newline (or nothing has been emitted yet)
- any position whose preceding text on the current line
(since the last newline in buf) is whitespace-only, AND
if there is no newline in the preceding buf portion, the
most recent prior emission ended with a newline
Boundary = position 0 with the prior emission ending in a newline (or
nothing emitted yet), or any position whose preceding text on the current
line is whitespace-only (when no newline precedes it in *buf*, the prior
emission must also have ended with a newline).
"""
prior_newline = (
already_emitted[-1].endswith("\n") if already_emitted else self._last_emitted_ended_newline
)
if idx == 0:
# Check whether the last already-emitted chunk in THIS
# feed() call ended with a newline, otherwise fall back
# to the cross-feed flag.
if already_emitted:
return already_emitted[-1].endswith("\n")
return self._last_emitted_ended_newline
return prior_newline
preceding = buf[:idx]
last_nl = preceding.rfind("\n")
if last_nl == -1:
# No newline in buf before the tag — boundary only if the
# prior emission ended with a newline AND everything since
# is whitespace.
if already_emitted:
prior_newline = already_emitted[-1].endswith("\n")
else:
prior_newline = self._last_emitted_ended_newline
return prior_newline and preceding.strip() == ""
# Newline present — text between it and the tag must be
# whitespace-only.
return preceding[last_nl + 1:].strip() == ""
@classmethod
def _max_partial_suffix(
cls, buf: str, tags: Tuple[str, ...],
) -> int:
"""Return the longest buf-suffix that is a prefix of any tag.
def _max_partial_suffix(cls, buf: str, tags: Tuple[str, ...]) -> int:
"""Longest buf-suffix that is a strict prefix of any tag (case-insensitive).
Only prefixes strictly shorter than the tag itself count
(full-length suffixes are the tag and are handled as matches,
not held-back partials). Case-insensitive.
Full-length matches are real tags handled elsewhere, not held-back partials.
"""
if not buf:
return 0
@@ -364,12 +236,7 @@ class StreamingThinkScrubber:
@classmethod
def _strip_orphan_close_tags(cls, text: str) -> str:
"""Remove any close tags from *text* (orphan-close handling).
An orphan close tag has no matching open in the current
scrubber state; it's always noise, stripped with any trailing
whitespace so the surrounding prose flows naturally.
"""
"""Remove close tags with no matching open (always noise) plus trailing whitespace."""
if "</" not in text:
return text
text_lower = text.lower()
@@ -382,8 +249,7 @@ class StreamingThinkScrubber:
tag_lower = tag.lower()
tag_len = len(tag_lower)
if text_lower[i:i + tag_len] == tag_lower:
# Skip the tag and any trailing whitespace,
# matching _strip_think_blocks case 3.
# Skip the tag and trailing whitespace (matches _strip_think_blocks case 3).
j = i + tag_len
while j < len(text) and text[j] in " \t\n\r":
j += 1
+3 -13
View File
@@ -111,16 +111,6 @@ class TestTrustGate:
class TestPrecedence:
def test_scan_order_project_first(self, project_env):
_trust(project_env["config"], project_env["repo"])
order = su.get_scan_ordered_skills_dirs()
proj_dirs = {
(project_env["repo"] / ".hermes" / "skills").resolve(),
(project_env["repo"] / ".agents" / "skills").resolve(),
}
assert set(order[:2]) == proj_dirs
assert order[2] == su.get_skills_dir()
def test_project_paths_are_readonly_owned(self, project_env):
_trust(project_env["config"], project_env["repo"])
p = project_env["repo"] / ".hermes" / "skills" / "repo-skill" / "SKILL.md"
@@ -177,9 +167,9 @@ class TestQuarantine:
@pytest.fixture(autouse=True)
def _clear_quarantine_cache(self):
su._project_quarantine_cache_clear()
su._PROJECT_QUARANTINE_CACHE.clear()
yield
su._project_quarantine_cache_clear()
su._PROJECT_QUARANTINE_CACHE.clear()
def _add_malicious_skill(self, repo: Path) -> Path:
d = repo / ".hermes" / "skills" / "evil-skill"
@@ -231,7 +221,7 @@ class TestQuarantine:
(evil_dir / "SKILL.md").write_text(
"---\nname: evil-skill\ndescription: now actually benign\n---\nbody\n"
)
su._project_quarantine_cache_clear()
su._PROJECT_QUARANTINE_CACHE.clear()
assert su.is_quarantined_project_skill(evil_dir / "SKILL.md") is False
def test_scan_cache_outside_repo(self, project_env):
+5 -3
View File
@@ -17,8 +17,8 @@ from unittest.mock import patch
import agent.skill_bundles as skill_bundles
import agent.skill_commands as skill_commands
import tools.skills_tool as skills_tool
import agent.prompt_cache_boundary as prompt_cache_boundary
from agent.prompt_cache_boundary import (
clear_stable_prefixes,
find_stable_prefix,
register_stable_prefix,
)
@@ -36,9 +36,11 @@ SKILL_BODY = "Inspect the report carefully and preserve the stable instructions.
@pytest.fixture(autouse=True)
def _isolated_registry():
clear_stable_prefixes()
with prompt_cache_boundary._lock:
prompt_cache_boundary._prefixes.clear()
yield
clear_stable_prefixes()
with prompt_cache_boundary._lock:
prompt_cache_boundary._prefixes.clear()
def _write_skill(skills_dir, name, body=SKILL_BODY):
+4 -4
View File
@@ -88,7 +88,7 @@ def test_t20880_tool_heavy_native_loop_reproduction():
assert final_tool_marked
assert shared_transaction_endpoint
assert after_exchange.marker_count <= 4
assert _count_cache_markers(after_exchange.messages, after_exchange.tools) <= 4
class TestPromptCachePlan:
@@ -114,7 +114,7 @@ class TestPromptCachePlan:
assert plan.tools is not tools
assert "cache_control" not in tools[-1]
assert plan.tools[-1]["cache_control"] == MARKER
assert plan.marker_count == 4
assert _count_cache_markers(plan.messages, plan.tools) == 4
def test_unmarkable_endpoint_does_not_consume_a_slot(self):
messages = [
@@ -129,7 +129,7 @@ class TestPromptCachePlan:
direct_native_tool_cache=True,
)
assert plan.marker_count == 2
assert _count_cache_markers(plan.messages, plan.tools) == 2
assert "cache_control" not in plan.messages[-1]
def test_static_prefix_equal_to_whole_prompt_emits_no_empty_block(self):
@@ -184,7 +184,7 @@ class TestPromptCachePlan:
native_anthropic=True,
direct_native_tool_cache=True,
)
assert plan.marker_count == 3
assert _count_cache_markers(plan.messages, plan.tools) == 3
assert len(plan.tools) == 0