The pin to "skills-guard-v2" was a change-detector: every rule change that
bumps the scanner version (this PR moves to v3 so cached verdicts re-scan)
turned it red. The intent was "cached v1 verdicts are invalidated", so assert
the contract instead of the current value.
Review-fold from the 3-angle simplify pass:
- sed -Ei / -iE / --in-place now match the shell-critical tier (the
bare '\s-i\b' token missed combined short flags and the GNU long
form); read-only sed stays unflagged. Regression tests added.
- agent_config_contract joins plugin_guard's CODE_EXEMPT_PATTERN_IDS:
content-contract prose in plugin code files (docstrings/comments)
is the same false-positive class the existing agent_config_mod
exemption suppresses. Doc/config files keep the full pattern set.
Efficiency reviewer: 1.24x full-scan cost (+3.4ms/file, install-time
only), worst-case adversarial line 55us — no ReDoS exposure.
Follow-up hardening on top of #92249's tiered scoring:
- Shell-critical tier now also catches tee, and cp/mv with the config
file in destination position (cp/mv reads and .bak backups excluded).
A single '>' redirect must be preceded by a word/quote character so
markdown blockquotes and '->' arrows no longer match.
- Prose tier catches mid-line imperatives behind directive markers
('you must modify...', 'please update...', 'make sure to append...'),
which previously bypassed the line-start anchor.
- Prose instructions aimed at AGENT config files score critical again:
project-skill quarantine acts only on 'dangerous', so high/caution
silently converted 'quarantined' into 'allowed' for exactly the
sentence shape persistence attacks use (concern raised in #88952).
Hermes/other-agent config prose stays high/caution (setup docs
legitimately instruct config.yaml edits).
- New content-contract tier ('AGENTS.md should contain ...') at
high/caution — the shape is shared by authoring guides and attacks.
- .claude/settings and .codex/config gain the same shell-critical tier.
Verified against a 595-skill corpus: 0 skills blocked by these tiers
(main blocked 44 legitimate ones), all mattpocock repro skills from
#92021 install, and 20/20 attack corpus lines keep their verdicts.
The skills-guard-v1 scanner flagged ANY mention of AGENTS.md / CLAUDE.md /
.cursorrules / .clinerules as critical/persistence. Any critical finding
forces a dangerous verdict, and community installs cannot be overridden
with --force — so legitimate meta-skills that merely DISCUSS agent config
files (authoring guides, setup docs, cross-references) were permanently
blocked. Three popular community skills were hit in the wild.
skills-guard-v2 scores the persistence category in three tiers:
- Mechanical persistence (shell redirection or sed -i targeting an agent
config file) stays critical -> dangerous. An unambiguous write path.
- Modification language in imperative position (verb at line/bullet start
within 80 chars of the filename) is high -> caution. Regexes cannot
separate "Edit AGENTS.md to inject instructions" from descriptive prose,
but imperative verbs are the shape real instructions take. Caution keeps
the install confirmable instead of irreversibly blocked.
- Bare references drop to low/informational for auditability without
driving the verdict.
The verb-proximity shape matches the existing convention in
tools/threat_patterns.py, and the tiering mirrors how allowed_tools_field
was already handled. The pattern id agent_config_mod is preserved so
plugin_guard.CODE_EXEMPT_PATTERN_IDS stays valid; hermes_config_mod /
other_agent_config get parallel _shell / _ref splits fixing the whole bug
class. SCANNER_VERSION bumps to v2 so cached v1 dangerous verdicts are
invalidated and re-scanned on next install attempt.