Commit Graph

3587 Commits

Author SHA1 Message Date
Teknium 1ee30352ca fix: background review can now read skills before patching — denial storm ended, cache parity intact (#61521, #39996)
The self-improvement review fork advertises the parent's full tool schema
(deliberate — tools[] must stay byte-identical for prompt-cache parity)
but denied everything except memory/skill tools at dispatch. Models
naturally reach for read_file to inspect a SKILL.md before patching, got
denied, then attempted a blind skill_manage patch which the
read-before-write guard correctly refused. One deployment logged ~142
denials + ~204 refusals over 2 days: the self-improvement loop ran
continuously but almost never landed a skill patch.

Fix is dispatch-side ONLY — zero request-body change, cache untouched:

- Whitelist read_file + search_files on the review fork (reads are
  side-effect-free). Write tools (write_file/patch/terminal) stay denied:
  autonomous maintenance must go through skill_manage's validation.
- read_file now registers full reads with the review fork's
  read-before-write guard (same as skill_view), so the natural
  read_file -> skill_manage(patch) sequence lands. Partial reads
  (offset>1 / truncated) don't count. No-op outside review forks.
- Self-correcting deny message: names skill_view/skill_manage/memory as
  substitutes so one denial redirects the model instead of a storm
  (the actionable half of #61521's proposal 2).

Rejects #39997's alternative (narrow the advertised schema on local
endpoints): local backends have KV/prefix caches too, and re-prefilling
a large snapshot is most expensive exactly there.

Live A/B (real dispatch path, isolated HERMES_HOME): on main,
read_file DENIED -> patch REFUSED (read-before-write); on this branch,
read_file OK -> patch LANDED. tools[] identical in both.
2026-08-29 18:33:41 -07:00
Ayush Nangia 5cace31708 fix(delegation): carry provider request_overrides through the base_url path (#65035) 2026-08-29 18:16:04 -07:00
Teknium b6d535dd88 feat(browser): Brave Origin works for real-profile browsing and default-browser detection
Extends the real-profile machinery (PR #95620) to Brave Origin — Brave's
standalone paid build with a fully separate install identity:

- new canonical key 'brave-origin' in _CHROMIUM_BROWSERS
- Windows: BraveOHTML ProgId -> brave-origin; channel ProgIds BraveOBHTML/
  BraveODHTML/BraveOSHTM fail closed (identifiers from brave-core
  install_static)
- macOS: com.brave.Browser.origin bundle id (exact match); .beta/.dev/
  .nightly channel bundles fail closed; /Applications/Brave Origin.app
- Linux: brave-origin.desktop matched BEFORE the bare 'brave' fragment
  (substring scan would otherwise resolve an Origin default to stable
  Brave and drive the wrong profile — #95549 wrong-principal invariant);
  brave-origin-{beta,nightly,dev} fail closed
- profile dirs: BraveSoftware/Brave-Origin on all three OSes (per
  brave-core kProductPathName + Homebrew cask zap paths)
- /browser connect launch tables: Brave Origin split into its OWN group
  so a 'brave' executable lookup can never resolve to the Origin binary
- user-facing strings/docs/desktop tooltip updated

Tests: progid/bundle/desktop map params + data-dir resolution for all
three OSes; 125 passed in the three browser test files.
2026-08-29 18:13:33 -07:00
Teknium 03e66c8cba polish(tool-search): stub-optimized openers for deferred tools — trigger+verb in the first ~60 chars (the catalog stub is the only ambient hint a deferred tool exists) 2026-08-29 18:13:20 -07:00
nftpoetrist d6a6d87c4a fix(tools): restore setup_mcp's never-hand-edit instruction
9d9f44d638 removed the desktop platform hint's "never hand-edit
mcp_servers config for them" sentence, reasoning it was a "word-for-word
duplicate of the setup_mcp tool schema... taught on every call." The
schema has never contained that instruction — only "never re-ask after
a decline." setup_mcp is desktop_ui-toolset-only and no runtime guard
in agent/file_safety.py covers mcp_servers config, so removing the only
place teaching this left a real gap: a model asked to add/configure an
MCP server could just write_file into mcp_servers config directly,
bypassing the consent-card/OAuth flow the tool exists to enforce.

Restored the instruction directly in SETUP_MCP_SCHEMA's description —
completing the original commit's stated intent (move it to the schema)
rather than reverting to the platform hint, since the schema reaches
every setup_mcp call regardless of platform hint wording changes.

Added a regression test asserting the schema description forbids
hand-editing mcp_servers config, so a future prompt-diet pass can't
silently drop it again without a test failing.
2026-08-29 17:58:51 -07:00
Teknium b1a46e192c fix(tool-search): pull clarify back out of the default defer set — A/B showed structured ask collapses when deferred
Maintainer A/B (288 live runs, 3 model tiers, results in the PR body):
with the clarify schema visible, models used structured ask-the-user
18/18 on ambiguous tasks (score 1.00 all models). Deferred, usage
collapsed to 7/18 (gpt-terra 0/6) — models still asked, but as
plain-text turn-ending questions: no structured choices, no recommended
option, an extra user round-trip. The ask-the-user affordance has to be
ambient to fire; a catalog stub is not enough (~250 tok to keep eager).

- _DEFAULT_DEFERRED_TOOLS: remove clarify (19 -> 18 deferred)
- regression test pins clarify ∉ default defer set AND assembles direct
  while the bridge is active (sabotage-verified: fails with clarify
  in the set)
2026-08-29 17:58:40 -07:00
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium b6bd681e89 feat(todo): nested subtasks via optional parent field
The todo tool now supports hierarchical task lists: an item's optional
'parent' field points at another item's id, making it a subtask.

- tools/todo_tool.py: parent validated (self-ref dropped), dangling refs
  and cycles sanitized; merge mode can set/clear parent; post-compression
  injection renders the tree indented and keeps a finished parent visible
  while any descendant is still active; the in-progress reorder pass is
  skipped for nested lists (a flat move would tear subtasks from parents).
- Schema cost: ~45 tokens added to the cached tool schema (one string
  property + one behavior sentence).
- acp_adapter/tools.py: todo result markdown indents by parent depth.
- Desktop: TodoItem carries parent; todoTree() DFS helper; composer
  status stack renders subtask rows indented (depth-capped), stabilizer
  compares depth.
- Docs: tools-reference todo entry mentions nesting.

Hydration/replay paths (gateway fresh-agent, API-server history) work
unchanged: parent rides inside the same todos array.
2026-08-29 07:25:12 -07:00
kshitijk4poor 154fd10af0 fix: name the preserved snapshot path in the ROLLBACK FAILED payload
Folded from #97748 (the competing fix by @lEWFkRAD): when a rollback
restore fails, the error note now points the operator at the surviving
snapshot directory instead of leaving them to find it in tempdir.
2026-08-29 19:38:31 +05:30
Adolanium 1315e65a52 fix(skills): a failed rollback restore keeps the skill and the snapshots
Rollback removed the live skill directory before restoring its
snapshot. When copytree then failed (disk full, locked file, path too
long on Windows) the except only added a note, and the finally deleted
the snapshot directory too, so nothing survived: the skill was gone
with a success-shaped error payload.

The broken state is now renamed aside first and deleted only after the
snapshot is restored. If the restore still fails, the broken state is
renamed back, so the worst outcome is the half applied batch instead of
no skill at all. When rollback reports any failure the snapshots are
kept on disk and their location is logged, instead of being deleted by
the finally.

Follow-up to #97692, same batch executor.
2026-08-29 19:38:31 +05:30
Clifford Garwood 10e93c6ab9 fix(skills): drop redundant identical-strings guard and its vacuous tests
The guard and the tests around it pinned behavior that already existed.
fuzzy_find_and_replace rejects old_string == new_string at
tools/fuzzy_match.py:69-70, returning "old_string and new_string are
identical" — so main already answered success=False, and the three tests
asserting "identical" in the error passed with the guard deleted.

The earlier claim that this case "silently applies a no-op the model reports
as success" was wrong. Verified against main:

  {"success": false, "error": "old_string and new_string are identical",
   "file_preview": "..."}

The duplicate guard was also strictly worse: it fired before the skill
lookup and returned no file_preview, shadowing the richer message.

Removes the guard, the two identical-strings tests, and the third
parametrize case. What remains is the genuinely new behavior: an actionable
missing-old_string error, reachable through the public tool.
2026-08-28 23:17:35 -07:00
Clifford Garwood 4f4e778db8 fix(skills): make skill_manage patch failures recoverable instead of a dead end
`skill_manage(action='patch')` rejected a missing `old_string` with:

    old_string is required for 'patch'. Provide the text to find.

That is a dead end. The model cannot tell whether it omitted the argument or
supplied text that did not match, so it retries blindly — and then escapes to
the neighbour that always works: `action='write_file'`, which rewrites the
entire skill file and destroys unrelated content. `skill_manage`'s own action
enum puts that destructive path one token away from the failing one.

The error now names the recovery route: `old_string` must be the EXACT text
currently in the file, read the target first (the skill's SKILL.md, or the
file named by `file_path`), copy the snippet verbatim, and do not fall back
to `action='write_file'`.

Validation lives in `_patch_skill` rather than the dispatcher. `skill_manage()`
previously returned its own bare missing-argument error before ever calling
the helper, which would leave the new guidance unreachable through the public
tool. Removing that duplicate makes the helper the single source of truth; its
{"success": False, "error": ...} flows through the same json.dumps path, so
the serialized shape is unchanged, and validation still precedes the skill
lookup — a missing old_string on an unknown skill still reports the argument
error rather than "skill not found".

Fixes #33064
2026-08-28 23:17:35 -07:00
Teknium 217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
Gille 9a1eef7a29 fix(tools): narrow MCP OAuth lock scope 2026-08-28 20:16:26 -07:00
Teknium 62e8126c69 fix(skill_manage): batch failure results carry the failing op's teaching payload (file_preview, hints)
Live A/B eval (old flat vs operations[] on qwen3.8-27b / gpt-5.6-terra /
claude-sonnet-5) caught a real regression: on a fuzzy-match miss the flat
path returns file_preview so the model can self-correct, but the batch
wrapper rebuilt the error dict and dropped every field except error/
failed_index. Sonnet, recovering blind, probed the file by writing and
reverting placeholder patches for 8+ turns (50k tokens vs 15k on the flat
arm). Batch failures now merge through all non-error fields from the
failing op's result.
2026-08-28 19:38:21 -07:00
Teknium dadbfd8990 refactor(patch): V4A mode gated to OpenAI-family mains — base schema is replace-only with real required (365 -> 195 for everyone else, -149; handler accepts both shapes from any model) (#97403) 2026-08-28 12:59:41 -07:00
Teknium 93de1d3430 vision_analyze diet + image routing: explicit aux vision backend becomes the de-facto route (reverses #29135) (#97339)
* refactor(vision_analyze): schema diet — routing mechanics removed (automatic; native path's own result teaches), region flow kept (~271 -> 181 tok/call, -33%)

* feat(image-routing): explicit auxiliary.vision backend is the de-facto image route — reverses #29135 (maintainer decision); native stays default when unset, image_input_mode:native stays absolute
2026-08-28 12:15:24 -07:00
Teknium 72874b0675 feat(skill_manage): operations[] is the call — each op names its skill; atomic with cross-skill rollback (#97295)
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate

* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing

* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal

* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback

* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization

Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
2026-08-28 12:15:18 -07:00
Teknium baa344dee7 refactor(process): schema diet — enum names the verbs, description keeps only non-obvious semantics; write-vs-submit trap teaching emphasized (306 -> 228 tok/call, -25%) (#97279) 2026-08-28 09:23:00 -07:00
Teknium 7b5e1911f8 refactor(todo): schema diet — item shape and merge semantics taught once, by the param schema (323 -> 232 tok/call, -28%) (#97257) 2026-08-28 08:58:12 -07:00
Teknium a9e72f1b58 refactor(read_file): schema diet + bundle anydoc 0.2.4 + typed NeedsOcrError/hosted-OCR wiring (#97195)
* refactor(read_file): capability-gate the anydoc format list; PDF coverage teaching lives in the response-time warning (426 -> 244/291 tok/call)

* feat(read_file): bundle firecrawl-anydoc 0.2.4 in core, typed NeedsOcrError handling, config-gated hosted OCR with local-OCR-first guidance

* refine(read_file): NEEDS-OCR warning hints at checking for an OCR skill without naming one; hosted_ocr knob unadvertised (maintainer-directed)

* simplify(read_file): drop the anydoc schema gate — bundled core dep makes absence a broken install, not a variant; formats stated unconditionally (263 tok/call)

* feat(read_file): PDF wording upgrades to 'scanned or text' when a trusted hosted-OCR route exists (direct key or explicit config; nous gateway excluded until Parse proxy works)

* simplify(read_file): FIRECRAWL_API_KEY is the ONLY hosted-OCR gate — nous gateway route removed (Parse proxy broken), config true no longer unlocks; false still disables
2026-08-28 08:46:11 -07:00
Teknium 95cf7dc9e8 feat: session temp root moves off tmpfs /tmp to ~/.hermes/cache/terminal by default; auto-pruned after 72h
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
2026-08-28 07:50:33 -07:00
rahlquist d7be3f649d feat: expose terminal.temp_dir to redirect session temp root off tmpfs
Some Linux distros (notably RAM-based tmpfs /tmp on several Arch-based
setups) cap the temp directory at a small size, so Hermes runs out of
space for session temp files (background logs/pid/exit files, code-
execution sandboxes). Add a terminal.temp_dir config key that points
these at real storage.

- Add terminal.temp_dir default (empty) in config_defaults.py
- Bridge it to TERMINAL_TEMP_DIR via TERMINAL_CONFIG_ENV_MAP
- Honor TERMINAL_TEMP_DIR first in LocalEnvironment.get_temp_dir(),
  falling through to TMPDIR//tmp//gettempdir when unset/invalid
- Add tests covering override, process-env, missing-dir, and empty
2026-08-28 07:50:33 -07:00
Teknium eff97a8a05 refactor(profiles): retire the cross-profile write guard — profiles are not isolated (maintainer decision); mirror lost-write guards (#32049) survive; patch/write_file schemas drop cross_profile (-83 tok/call) (#97165) 2026-08-28 06:36:22 -07:00
Teknium 9978706e93 refactor(skill_manage): schema diet — patch args defer to the patch tool's semantics, file_path states its skill-dir-relative shape, authoring curriculum compressed (517 -> 427, -17%) (#97152) 2026-08-28 05:48:15 -07:00
Teknium 536adb35c4 refactor(video_generate): capability-gated dynamic schema (~814 → 458/377 tok/call) (#97095)
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests

* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)

* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
2026-08-28 05:10:33 -07:00
Teknium 4029a24f0c fix(vision): degrade image validation gracefully when Pillow is missing
Pillow is an optional dependency in this codebase (every other PIL use in
tools/vision_tools.py imports lazily and falls back). Both salvaged
validators now distinguish 'PIL missing' (pass through, header-only
sniff) from 'decode failed' (reject), so a Pillow-less install keeps
working instead of rejecting every PNG.

Follow-up to salvaged #53307 (@CannibalKush) and #76896 (@HaiyiMei).
2026-08-28 04:58:21 -07:00
CannibalKush 765142d94f fix: validate PNGs at shared image resolver 2026-08-28 04:58:21 -07:00
HaiyiMei ad0e83062f fix(vision): reject truncated images before embedding 2026-08-28 04:58:21 -07:00
liuhao1024 628a414d29 fix(browser): bind CDP binary exemptions to exact method result paths
Second review round: honoring base64Encoded recursively let any nested
dict spoof {"base64Encoded": true, "data": "<secret>"} past the redactor
(Runtime.evaluate returns arbitrary by-value JSON), and two carriers were
missed entirely — Network.streamResourceContent returns unflagged binary
bufferedData, and Network.getRequestPostData's postData was not covered.

Replace the ambient field sets with per-method exact result-path specs:
_CDP_ALWAYS_BINARY_PATHS for declared-binary paths (screenshots, PDFs,
streamResourceContent, beginFrame screenshotData, the nested
CacheStorage.requestCachedResponse.response.body) and
_CDP_FLAGGED_BINARY_PATHS for paths whose carrier object's base64Encoded
sibling gates the exemption (Network/Fetch.getResponseBody body, IO.read
data, getRequestPostData postData). Path suffixes propagate only into the
matching subtree, so base64Encoded is type information solely on trusted
carrier objects — never ambient trust in nested JSON.
2026-08-28 04:57:52 -07:00
liuhao1024 b2a17bfe82 fix(browser): scope the CDP binary-payload exemption to typed fields (#94138)
Architecture-review follow-up: the method-scoped binary_payload flag skipped
redaction for every string anywhere in the result of the two listed methods,
and the same corruption stayed reachable through Network.getResponseBody /
Fetch.getResponseBody / IO.read / Network.streamResourceContent. Make the
exemption field-scoped instead: an explicit schema exempts exactly
Page.captureScreenshot.result.data and Page.printToPDF.result.data (carriers
with no flag of their own), and any dict whose base64Encoded sibling is
exactly True exempts its body/data/bufferedData string (the protocol's own
discriminator — text bodies with base64Encoded: false stay redacted). Every
other string in every result keeps full secret redaction.
2026-08-28 04:57:52 -07:00
liuhao1024 a56885495e fix(browser): keep CDP binary payloads byte-identical through redaction (#94138)
_redact_cdp_output applied redact_sensitive_text(force=True) to every string
in CDP results, including the base64 screenshot/PDF payload of
Page.captureScreenshot and Page.printToPDF. The Fernet pattern (gAAAA + base64
alphabet) matches arbitrary spans inside such payloads wherever gAAAA follows a +
or /, collapsing them to first6...last4: decoded PNGs came out corrupt (valid header,
CRC failures mid-IDAT, no IEND), the persisted full copies in tool_result_storage
were redacted too, and vision_analyze then embedded corrupt images that the provider
rejected with 400 invalid_image - killing resume sessions with a misleading provider
error. Skip redaction for the two binary-payload methods: the payload is binary,
not free text, so there is no secret to protect there. Every other method keeps
full redaction.
2026-08-28 04:57:52 -07:00
Teknium c30ac90a92 feat(compaction): rebuild dynamic tool schemas at the compaction commit boundary — forever-sessions finally pick up config changes (#97073) 2026-08-28 04:01:05 -07:00
Teknium a619db6633 refactor(image_generate): capability-gated dynamic schema (554 → 317 tok/call, −43%) (#97057)
* refactor(image_generate): capability-gated dynamic schema — args render only when the active model honors them (554 -> 317 tok/call, -43%)

* fix(image_gen): fleet-wide supports_upscale declarations — krea (Enhance) + fal plugin (Clarity passthrough) declare it; declaration<->implementation contract-tested across all 7 in-tree providers
2026-08-28 03:58:52 -07:00
Teknium d6a21bc4ed fix(skills-guard): exempt all os.environ.get() reads from the env-dump pattern; os.getenv secret reads score medium
Follow-up to @AIalliAI's #60750 commits: with python_environ_get_secret
downgraded to medium, the high-severity python_os_environ pattern still
fired on the same os.environ.get("...KEY") line, re-escalating the verdict
the downgrade intended to avoid. Exempt every .get() form (non-secret =
config read; secret-shaped = scored medium by the dedicated pattern), and
apply the same medium grade to the sibling os.getenv() secret pattern —
same shape, same rationale (#60709 point 2).
2026-08-28 03:46:21 -07:00
Alli 54909d41b4 fix(skills-guard): handle inline-comment and docstring false positives for os.environ
The original ^(?!\s*#) prefix only skipped full-line comments starting
with '#'. An inline comment like:
  cfg = environ.get('HOME')  # os.environ available
still triggered python_os_environ because the regex matched the code part
before the '#'.

Two complementary fixes:
1. Replace ^(?!\s*#) with ^[^#\n]* in the regex — this rejects any line
   where a '#' comment marker appears anywhere before os.environ.
2. Add _compute_docstring_lines() — a state machine that pre-computes
   lines inside triple-quoted strings (docstrings) and skips them during
   pattern matching. Also handles single-line self-contained docstrings.

6 new regression tests covering: inline comments, multi-line docstrings,
single-line docstrings, full-line comments, and a verification that real
bare dict(os.environ) code still triggers. All 85 tests pass.
2026-08-28 03:46:21 -07:00
AIalliAI 42e6149451 fix(skills-guard): reduce false-positive CRITICAL/HIGH on benign skill patterns
Five targeted fixes for #60709 (reported by @mvanhorn):

1. ruby_env_secret: scope ENV[] to case-sensitive Ruby constant
   ((?-i:ENV)) — no longer matches Python env[key] dict access.

2. python_environ_get_secret: downgrade critical→medium — reading
   an API key via os.environ.get() is normal auth, not exfiltration.

3. python_os_environ: skip comment lines with ^(?!\s*#) — no longer
   flags os.environ references in docstrings or code comments.

4. deception_hide: downgrade critical→high + negative lookahead for
   UX guidance context (unless/except/until/confirm/diagnose/verify).

5. oversized_skill: downgrade high→low + raise cap 1MB→5MB — large
   skills are legitimate; structural size is informational only.

All 80 existing tests pass. 6 new verification tests added for each fix.
2026-08-28 03:46:21 -07:00
kshitijk4poor 8c098e9e81 fix(skills): catch sed flag variants; exempt content-contract prose in plugin code
Review-fold from the 3-angle simplify pass:

- sed -Ei / -iE / --in-place now match the shell-critical tier (the
  bare '\s-i\b' token missed combined short flags and the GNU long
  form); read-only sed stays unflagged. Regression tests added.
- agent_config_contract joins plugin_guard's CODE_EXEMPT_PATTERN_IDS:
  content-contract prose in plugin code files (docstrings/comments)
  is the same false-positive class the existing agent_config_mod
  exemption suppresses. Doc/config files keep the full pattern set.

Efficiency reviewer: 1.24x full-scan cost (+3.4ms/file, install-time
only), worst-case adversarial line 55us — no ReDoS exposure.
2026-08-28 03:24:43 -07:00
kshitijk4poor f2f61e0a45 fix(skills): close shell-write and prose-bypass gaps in agent-config tiers
Follow-up hardening on top of #92249's tiered scoring:

- Shell-critical tier now also catches tee, and cp/mv with the config
  file in destination position (cp/mv reads and .bak backups excluded).
  A single '>' redirect must be preceded by a word/quote character so
  markdown blockquotes and '->' arrows no longer match.
- Prose tier catches mid-line imperatives behind directive markers
  ('you must modify...', 'please update...', 'make sure to append...'),
  which previously bypassed the line-start anchor.
- Prose instructions aimed at AGENT config files score critical again:
  project-skill quarantine acts only on 'dangerous', so high/caution
  silently converted 'quarantined' into 'allowed' for exactly the
  sentence shape persistence attacks use (concern raised in #88952).
  Hermes/other-agent config prose stays high/caution (setup docs
  legitimately instruct config.yaml edits).
- New content-contract tier ('AGENTS.md should contain ...') at
  high/caution — the shape is shared by authoring guides and attacks.
- .claude/settings and .codex/config gain the same shell-critical tier.

Verified against a 595-skill corpus: 0 skills blocked by these tiers
(main blocked 44 legitimate ones), all mattpocock repro skills from
#92021 install, and 20/20 attack corpus lines keep their verdicts.
2026-08-28 03:24:43 -07:00
ClintonEmok 4ade4450bf docs(skills): document programmatic-write scope cut in skills_guard module docstring
Enough1122's review on #92249 asked whether language write APIs
(Python open('w')/write_text/os.replace/shutil, Node fs.writeFileSync/
appendFile) are covered by the agent-config persistence tiers. They are
not: those tiers score shell redirection, sed -i, and imperative prose
only; language-API calls surface just the informational *_ref finding.

Static regexes cannot tie a dynamically-built path to the config-file
destination without executing the skill, so this is a documented scope
cut rather than missing coverage — runtime install gates remain the
backstop. Scoring behavior is unchanged, so SCANNER_VERSION stays at
skills-guard-v2 and cached verdicts remain valid.

Refs #92249
2026-08-28 03:24:43 -07:00
ClintonEmok e22b8b66ce fix(skills): stop agent-config persistence patterns from blocking meta-skills (#92021)
The skills-guard-v1 scanner flagged ANY mention of AGENTS.md / CLAUDE.md /
.cursorrules / .clinerules as critical/persistence. Any critical finding
forces a dangerous verdict, and community installs cannot be overridden
with --force — so legitimate meta-skills that merely DISCUSS agent config
files (authoring guides, setup docs, cross-references) were permanently
blocked. Three popular community skills were hit in the wild.

skills-guard-v2 scores the persistence category in three tiers:

- Mechanical persistence (shell redirection or sed -i targeting an agent
  config file) stays critical -> dangerous. An unambiguous write path.
- Modification language in imperative position (verb at line/bullet start
  within 80 chars of the filename) is high -> caution. Regexes cannot
  separate "Edit AGENTS.md to inject instructions" from descriptive prose,
  but imperative verbs are the shape real instructions take. Caution keeps
  the install confirmable instead of irreversibly blocked.
- Bare references drop to low/informational for auditability without
  driving the verdict.

The verb-proximity shape matches the existing convention in
tools/threat_patterns.py, and the tiering mirrors how allowed_tools_field
was already handled. The pattern id agent_config_mod is preserved so
plugin_guard.CODE_EXEMPT_PATTERN_IDS stays valid; hermes_config_mod /
other_agent_config get parallel _shell / _ref splits fixing the whole bug
class. SCANNER_VERSION bumps to v2 so cached v1 dangerous verdicts are
invalidated and re-scanned on next install attempt.
2026-08-28 03:24:43 -07:00
Teknium 5857231267 fix(execute_code): limits line teaches spillover instead of a bare 50KB cap (#97048)
* fix(execute_code): limits line teaches spillover — big stdout is saved, not lost (follow-up to #97043)

* fix(execute_code): drop the editorializing tail from the limits line (maintainer review)
2026-08-28 03:17:05 -07:00
Teknium ae8c976032 feat(execute_code): stdout spillover — truncated output's full text saved to cache/exec (host) or kernel tmpdir (cells), path + read_file recipe in the result (#97043) 2026-08-28 03:05:03 -07:00
Teknium 2f57cd95b2 refactor(execute_code): schema diet — persistence woven in, not bolted on (712 → 654 tok/call) (#96997)
* refactor(execute_code): integrate kernel persistence into the core description (712 -> 654 tok/call, -8%)

* fix(execute_code): honest interpreter note — Hermes's own python is the common case; project venv only when VIRTUAL_ENV/CONDA_PREFIX is active
2026-08-28 03:04:51 -07:00
Teknium 5f75ec197b feat(code-execution): remote kernel host — session persistence for docker/ssh/modal backends (closes #96873) (#96991) 2026-08-28 01:39:33 -07:00
Teknium 4e7eb39947 refactor(code-execution): session kernels always on — kernel_mode knob retired (#96787)
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)

* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel

* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen

* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
2026-08-27 22:22:39 -07:00
Teknium 31579f781e fix(cron): transient run prompt survives the relay-fronted gateway forward
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
2026-08-27 20:53:02 -07:00
Victor Kyriazakos 5cc47c994b fix(cron): resolve api_server host for the manual-run forward (bind parity)
The forward dialed a hardcoded 127.0.0.1. The api_server adapter binds
extra.host -> API_SERVER_HOST -> 127.0.0.1, so mirror that chain when
dialing. Wildcard binds (0.0.0.0/::) listen on loopback, so keep dialing
loopback for those; bracket bare IPv6 literals.
2026-08-27 19:52:17 -07:00
Victor Kyriazakos 8e8112687b fix(cron): don't reference nonexistent 'hermes cron trigger' in relay-fronted errors
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
2026-08-27 19:52:17 -07:00
Victor Kyriazakos ad0e522362 fix(cron): forward manual run to the gateway for relay-fronted delivery (NS-773)
A manual 'hermes cron run' on a relay-fronted target has no live relay adapter
and no standalone sender, so it now forwards to the running gateway's
POST /api/jobs/{id}/run (marks due for the gateway ticker, which delivers via
the live relay adapter). Gateway unreachable -> the accurate 'start the gateway
or use cron trigger' error. Native topologies are untouched.
2026-08-27 19:52:17 -07:00