Commit Graph

3587 Commits

Author SHA1 Message Date
Teknium 1190825652 fix(bot-mode): DM protocol no longer shell-interpolates message bodies
The Bot Mode teammate-DM protocol told agents to inline the message into a
double-quoted shell argument: quotes truncated the body and $(...)/backticks
executed on the sender's machine. The protocol now writes the message to a
temp file and delivers it via a new 'hermes chat --query-file' flag (or '-'
for stdin); 'hermes peer dm' already accepted stdin and the peer recipe now
uses it. No shell pass touches the body at any point.

Supersedes the tool-based approach in #89077 — same bug, fixed with a CLI
flag + protocol rewrite instead of a new model tool.

Co-authored-by: mehmetkr-31 <mehmetkr-31@users.noreply.github.com>
2026-08-18 22:31:32 -07:00
Teknium a6bada232c feat(mcp): per-server oauth.user_agent for token-endpoint requests (#75576)
Some authorization servers and WAFs reject httpx's default User-Agent on
the OAuth token endpoint. mcp_servers.<name>.oauth.user_agent now stamps a
custom User-Agent onto the two token-endpoint requests (authorization-code
exchange and refresh) on both provider construction paths. Opt-in,
per-server, token requests only — never MCP traffic or discovery, and no
other headers are configurable. Empty/null/non-string values are ignored.

Completes the second half of #75576 (the CIMD half landed via #89566).
2026-08-18 21:57:53 -07:00
Teknium 5dd15872a6 fix: never evict pinned CIMD sockets from the callback reservation FIFO
The _MAX_RESERVED_SOCKETS cap applied to pinned CIMD sockets too, so under
heavy concurrency an ephemeral-reservation churn could close a parked pinned
socket before _wait_for_callback adopted it, silently reopening the
port-stealing window the pin exists to prevent (#22161). Eviction now skips
the pinned range; it is already bounded by _CIMD_PORTS.

Follow-up to the #84050 salvage.
2026-08-18 20:03:18 -07:00
rob-maron 0b588cb3a4 MCP CIMD auth 2026-08-18 20:03:18 -07:00
Brooklyn Nicholson 74f99af470 feat(desktop): agent-applied layout presets — apply_layout joins the desktop_ui toolset
The agent could reveal single panes (focus_pane) but had no way to arrange
the workspace as one act. apply_layout closes that gap: a desktop_ui tool
that emits layout.apply over the existing bridge, resolved in the renderer
against the layouts contribution registry — the same list the layout picker
reads — so core presets (default/focus/terminal-deck/quad), plugin presets,
and user-saved presets are all addressable by id. Active session only, same
as pane.reveal: a background turn never rearranges the user's desktop.
2026-08-18 21:06:18 -05:00
ethernet 66bb77cbf9 docs(clarify): advertise the questions batch in the tool description
The `questions` parameter had a full description, but the top-level
tool description still described three single-question modes and never
mentioned batching. The model decides how to call a tool from that
description, so it kept asking one question per call.

The description now states that 2-5 independent questions can go in
one call and that one batched call is preferred over a chain of
single-question calls. The parameter description also tells the model
to put a short batch title in the still-required top-level question.
Two schema tests pin the contract: the description names the batch
capability, and the questions parameter stays optional with the
MAX_QUESTIONS cap.
2026-08-18 21:28:53 -04:00
ethernet bd8b658a63 feat(clarify): accept a questions batch in the clarify tool core
The clarify tool gets an optional questions parameter (2-5 independent
questions, issue #18450). Batch-capable platform callbacks receive the
normalized list in one call and reply with per-question answers. Legacy
callbacks are looped one question at a time. The loop stops on timeout
so the user is not asked the remaining questions after they walk away.
Locked answers survive a timeout: the result carries them plus a
timed_out flag, and unanswered entries have an empty user_response.

The single-question path is byte-identical to the previous behavior.
2026-08-18 21:28:53 -04:00
Teknium cc421cb697 fix: dashboard console skills commands no longer act on the wrong profile (#65828)
tools/skills_sync.py bound HERMES_HOME / SKILLS_DIR / MANIFEST_FILE at
import time — the third module in the same lineage as skills_tool
(f8723c478) and skill_manager_tool (c6a3d412d). In a long-lived
dashboard/TUI process, console skills commands (reset, diff,
list-modified, opt-in/out, repair-official) dispatched in-process under
_profile_scope's set_hermes_home_override(), but skills_sync's frozen
constants kept resolving against whichever profile was live at import.
Sharpest edge: reset_bundled_skill()'s #48200 rmtree strict-child guard
was computed against the WRONG skills root.

Fix: same call-time accessor pattern as the two prior fixes —
_hermes_home()/_skills_dir()/_manifest_file() honor an explicitly
patched module global (tests, retargeting) and otherwise re-resolve
from the live profile-scoped get_hermes_home() on every call. All 37
call sites migrated; module constants kept for compat.

Also documents in _profile_scope() that skills_sync needs no module
retargeting since the contextvar override now reaches it.

Regression tests (sabotage-verified: all 3 fail on the old binding):
- accessors follow set_hermes_home_override at call time
- explicit module patch still wins over the override
- rmtree guard anchors on the overridden profile's skills root

Fixes #65828
2026-08-18 14:14:42 -07:00
Teknium 0093bc0fcc fix(skills): preserve manifest file mode across atomic rewrites
_write_manifest still used a hand-rolled mkstemp + atomic_replace,
so every sync reset .bundled_manifest to mkstemp's 0600, dropping a
group-readable or shared mode the operator had set. Replace the block
with utils.atomic_write_text(preserve_mode=True) — the same shared
writer and mode-preservation contract PR #86255 applied to the skill
manager's document writes.

This is the remaining half of PR #14410 by @sgaofen, who reported the
manifest mode reset first; the skill-manager half of that PR was
superseded by the atomic_write_text refactor and #86255.
2026-08-18 13:18:55 -07:00
kshitij e88d8831d2 refactor: consolidate owner extraction onto _owner_from_payload
_fetch_owner_handle and the browse catalog walk both inlined the
same owner-handle extraction logic that _owner_from_payload was
extracted to centralize. Replace both with calls to the helper.
2026-08-19 01:23:17 +05:30
xxxigm 696fef6214 fix(skills): keep inspect/install from mixing same-named hub skills
ClawHub treated the last path segment as a slug, so a GitHub-style
id like owner/repo/skills/skillopt fetched a different author's
skillopt. Pair metadata and files from the same source so inspect
cannot show one registry's header and another's SKILL.md.
2026-08-19 01:23:17 +05:30
Vadelma 968853c5b5 fix(skills): preserve document modes during atomic writes
Co-authored-by: Taneli Mielikäinen <taneli.mielikainen@iki.fi>
2026-08-18 12:52:53 -07:00
Teknium 8a770aae0d fix(tools): spillover is the canonical home for oversized results on every backend
Per review: even with an active sandbox env, spilled tool results
belong in $HERMES_HOME/cache/spillover with the other Hermes-owned
caches — not the sandbox temp dir as primary storage.

- Host-side write happens first on every backend; local/no-env
  sessions reference the host path directly (unchanged).
- cache/spillover joins the auto-mount/sync cache-dir list
  (credential_files._CACHE_DIRS), so docker bind-mounts it and
  modal/ssh/daytona file-sync it. Remote references use the
  translated in-sandbox path after a readability probe.
- Probe failure (persistent containers created before spillover
  joined the mount list, translation failures) falls back to the
  previous in-sandbox temp-dir copy, so nothing regresses.
2026-08-18 02:12:13 -07:00
Teknium c91681c69a fix(tools): large tool results persist to HERMES_HOME/cache/spillover instead of truncating when no sandbox env is active
Sessions that never ran a terminal command (MCP-only, cron, gateway)
have no active sandbox environment, so maybe_persist_tool_result()
got env=None and fell through to the inline-truncate fallback --
a 467K MCP result was cut to a ~1.3K preview with no file written
('Full output could not be saved to sandbox').

Now the host-side cases (env=None or the local backend) write the
spill file directly to $HERMES_HOME/cache/spillover/<id>.txt,
alongside the other Hermes-owned caches instead of littering /tmp.
Remote backends (docker/ssh/modal/daytona) keep the in-sandbox
env.execute() write since read_file resolves in-sandbox there.

Cleanup: the gateway housekeeping loop prunes spillover hourly with
the other media caches, and a once-per-process best-effort prune on
first spill covers CLI-only installs.
2026-08-18 02:12:13 -07:00
Hermes Agent b95ec1cb5d fix(delegation): running subagents stay visible to list/steer across parent-agent rebuilds, and child-started process notifications carry delegation attribution
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.

Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.

Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
2026-08-18 00:01:56 -07:00
Teknium 3360590115 fix: address second-round SkillEvaluator review feedback
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):

- Incomplete-validator findings are now PRESERVED as partial evidence;
  only the validator's pass/fail verdict is excluded from the advisory
  verdict. A report with findings from an incomplete check no longer
  reads as clean.
- Clean-report wording is now "no findings from completed checks"
  whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
  config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
  discard-pinning test, plus the completed-checks wording case.
2026-08-17 17:04:40 -07:00
Teknium 2c2697b52e feat: widen Tier 1 advisory scan to license + security checks
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.

- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
  static (no measurable cost); security invokes NVIDIA SkillSpector in
  its keyless static-rules mode (~+1.2s per install). schema/quality
  stay excluded: hygiene signal ("author not specified" is
  high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
  or failing, the security check reports status="incomplete" and the
  adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
  note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
  This also absorbs a live upstream inconsistency: SkillEvaluator's
  anti-tamper cross-check on SkillSpector's risk score currently trips
  on moderate-finding skills (fail verdict with zero findings, e.g.
  github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
  NVIDIA separately; either way an evidence-free fail must not render
  as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
  not-run formatting.

E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
2026-08-17 17:04:40 -07:00
Teknium 183f18d530 feat: advisory NVIDIA SkillEvaluator Tier 1 scan on skill installs
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).

- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
  over the quarantined bundle, parses the JSON report, classifies
  secrets-class findings (private keys, tokens, credentialed connection
  strings) apart from advisory PII findings. Every failure mode
  (binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
  guard's policy decision and before the install confirmation. Findings
  are shown with file:line; secrets-class findings render red with a
  loud warning. Warn-and-continue by design — the built-in skills guard
  remains the only enforcement layer, because the upstream PII scanner
  has known false-positive classes (git@github.com, docs example
  emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
  endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
  optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
  config toggle.

Scanner install (optional):
  uv tool install --python 3.13 \
    "skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"

E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
2026-08-17 17:04:40 -07:00
Teknium 6229683b62 feat(cli): hermes peer — bot-to-bot DMs across machines and gateways
Bots could message teammates on their own machine (hermes -p <bot> chat) and
the desktop could relay user mentions over Connections, but a bot had NO
transport to a bot on another gateway. This adds one, with zero new server
surface: the peer's existing api_server platform is the wire.

- hermes_cli/subcommands/peer.py: `hermes peer add/list/remove/dm`.
  `dm <peer>[/<agent>]` resolves the remote agent's canonical "Bot Chat"
  (list by title, create when missing), runs one synchronous agent turn via
  POST /api/sessions/{id}/chat, and prints the reply on stdout — the exact
  cross-machine twin of the local bot-messaging command, so the Bot Mode
  protocol composes over it unchanged. Named profiles route via the peer's
  /p/<profile>/ multiplex mirror. Peer URLs live in config.yaml
  (`bot_peers`); the peer's API_SERVER_KEY is a credential and lives in
  ~/.hermes/.env as HERMES_PEER_<NAME>_KEY.
- hermes_cli/main.py: parser wiring + fast-path/session-flag command sets.
- tools/bot_mode_probe.py: when peers are registered, the injected Bot Chat
  messaging protocol gains a cross-machine paragraph (peer roster +
  `hermes peer dm` pattern) so agents discover remote teammates on their
  own; peers join the capability fingerprint so registering/removing one
  refreshes eternal Bot Chat prompts on the next message (loud, one-time,
  user-initiated — no per-turn cache drift).
- Docs: Bot Mode guide (bot-initiated DMs across machines) + cli-commands
  reference (`hermes peer` section + summary row).

Tests: tests/hermes_cli/test_peer_cmd.py (target parsing, /p/ scoping,
registry round-trip in isolated config, real-loopback-HTTP dm flow incl.
Bot Chat create-vs-reuse and bearer auth), bot_mode_probe peer-paragraph +
epoch tests. E2E: real `python -m hermes_cli.main peer ...` against a live
fake peer over HTTP with isolated HERMES_HOME (config/.env persistence,
bare + /p/<profile> routing, stdin, --json). 23 passed; ruff clean.
2026-08-17 16:13:30 -07:00
Teknium 6e22d26583 feat: project-skill quarantine + non-interactive trust inheritance
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.

Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).

Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.

Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
2026-08-17 14:06:16 -07:00
Teknium f891d702df feat: project-local skill discovery with per-repo trust gate
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.

Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.

- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
  get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
  project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
  [project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
  skill_view resolves cross-tier collisions in favor of the project tier
  (same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
  commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)

Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
2026-08-17 11:39:13 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
Teknium c86197e607 fix: make the self-repo git guard Windows-only
The live-checkout git mutation guard blocked history-rewriting git ops
(checkout, reset --hard, rebase, cherry-pick, ...) in the running source
checkout and its worktrees on every platform. The hazard it protects
against is only real on Windows, where NTFS locks loaded module files and
an in-place rewrite can corrupt the running process. On POSIX, open file
handles pin the old inodes, so a checkout swap under a running process is
safe, and the guard mostly taxed normal dev/salvage workflows with clone
workarounds.

- tools/self_repo_guard.py: add guard_active() -> os.name == "nt"
- tools/terminal_tool.py: consult guard_active() before running the
  detector; detector logic and block message unchanged for Windows
- tests: wiring tests force the guard on; new tests cover the POSIX
  pass-through and the platform predicate
2026-08-17 10:03:36 -07:00
kshitij 4323c67dcc fix(delegate): disclaim only the fields a failed probe actually left unmeasured
/simplify-code residual. The note hard-coded "'commits' and 'dirty' are
UNKNOWN", but the two probes fail independently: a bad base_commit fails
rev-list while `git status` still succeeds, so `dirty` is a REAL measurement
being reported as unknown. Safety was never affected (the worktree is preserved
either way), but telling the parent a measured value is untrustworthy is its own
kind of misreport — and it would push a human toward re-inspecting something
already proven.

`mark_worktree_payload_unproven()` now takes an `unmeasured` argument, and
finalize tracks which probe actually failed. The raising path still disclaims
both, because which probe raised is unknowable there.

Validation: 22/22 tests/tools/test_subagent_worktree.py; ruff + ty clean. New
guard mutation-checked (hard-coding "commits/dirty" back fails it).
2026-08-17 19:41:32 +05:30
kshitij ce93a398e8 refactor(delegate): extract the unproven-payload factory; drop the source-reading test
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.

- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
  `unproven_worktree_payload()`. Both producers of this schema now call them,
  so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
  hand-building the dict (-16 lines). The re-import is guarded: the outer
  `except` can be entered because the `from tools import subagent_worktree`
  itself failed, in which case the name is unbound -- an inline fallback keeps
  the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
  compares its key set against live `finalize_subagent_worktree()` output. Same
  contract, no source reading, refactor-proof, and it actually executes the
  code.

Also folded from the same review:

- Fail-closed on an unmeasurable commit count. With no `base_commit` the
  rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
  tree still reached `git worktree remove --force` + `git branch -D` -- the
  exact bug class #88113 is about, on a public function that takes a
  caller-supplied dict. Now returns un-inspected instead, with a test driving a
  real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
  names WHICH probe failed, its exit code, and a bounded git stderr tail, so
  the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
  removed the duplicated index-corruption block in favor of the existing
  `_break_git_index()` helper.

Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
2026-08-17 19:41:32 +05:30
kshitij 97c4f9eeec test(delegate): assert the unproven-state contract, not its prose
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.

- The distinguishability test asserted the failure payload was byte-identical
  to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
  That freezes the AMBIGUITY as a required property: emitting `commits: None`
  for "unknown" would improve exactly what #88113 is about and fail the test.
  Now asserts what the parent actually depends on -- both keep the worktree,
  and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
  happy path, forbidding an always-present-but-False flag (a legitimately
  better JSON contract: stable key set for serializers). Now
  `assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
  of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
  was not even a cross-producer contract: delegate_tool's note said "state
  unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
  assert the note names the worktree AND branch -- the actionable part for a
  human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
  before any git call would keep it green while proving nothing). Now checks
  `call_count` and mirrors the branch-survival + note-names-path legs its
  sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
  parent reads, but delegate_tool's fallback -- the second producer -- was
  verified only by reading. It now AST-parses the real fallback dict literal
  and compares against live `finalize_subagent_worktree()` output, so the two
  producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
  commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
  raising, handled in delegate_tool), and the module docstring listed
  `inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
  `_break_git_index()` beside the file's other module-level helpers.

Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
2026-08-17 19:41:32 +05:30
kshitij 38ea711fd0 fix(delegate): tell the parent when a worktree was preserved un-inspected
The preserved worktree is invisible to the only consumer that can act on it.

Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":

  inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
  inspected OK, child produced nothing     -> {commits: 0, dirty: False, pruned: False}

The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.

Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
  naming the worktree/branch, warns, and returns the payload. Both unproven
  exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
  non-numeric rev-list stdout) produced the same unproven payload but logged at
  DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
  exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
  creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
  schema missing commits/dirty/pruned. It now emits the same flagged shape, and
  logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
  affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
  by restoring the unconditional prune and reintroducing this P1.

Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.

Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
  suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
  the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
  work intact on disk; proven-clean still prunes (pruned=true).
2026-08-17 19:41:32 +05:30
liuhao1024 2b490a0513 fix(delegate): keep the worktree when git inspection fails
finalize_subagent_worktree() treated a non-zero exit from its rev-list
or status probes as proof of the payload defaults (commits=0, clean),
then pruned on them: git worktree remove --force plus branch -D
permanently deleted a child's uncommitted work whenever git could not
inspect the tree (e.g. a corrupted index) (#88113).

A destructive cleanup now requires affirmative proof of zero commits
plus a clean tree. Any non-zero inspection result keeps the worktree
and branch for manual review, with a warning naming both.
2026-08-17 19:41:32 +05:30
Teknium 382060f022 feat(mcp): speak the 2026-07-28 stateless protocol
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):

- Protocol-era negotiation (_negotiate_session): per-server `protocol`
  config key — auto (default, handshake-first with server/discover
  fallback on -32022/-32601), stateless (discover-first), legacy
  (handshake only). Auto is handshake-first deliberately: zero extra
  round-trips and zero behavior change for the entire existing server
  fleet, while 2026-07-28-only servers now connect via the fallback.
  All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
  route through the one choke point, so the CLI/desktop probe path
  inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
  during discovery and bound to the lazy-startup schema cache — TTL'd
  entries expire and force a live re-probe; hint-less (pre-2026)
  servers keep the never-expires behavior. Pagination continuation now
  speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
  (config-overridable), with a fallback for 1.x-era metadata models.
  (RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
  native to SDK 2.0's OAuthClientProvider — verified, no client-side
  gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
  Sampling feature as upstream-deprecated (12-month window) — kept
  fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
2026-08-17 03:03:35 -07:00
Xinyu Du bbc894b0ab docs(tools): tighten subprocess env isolation docstrings
Shorter, single-source ownership explanation for
_strip_hermes_owned_pythonpath (the code-level Check comments already
carry the per-branch detail; the docstring only needs the contract).
2026-08-17 02:10:32 -07:00
Xinyu Du 98389e7895 refactor(tools): simplify subprocess env isolation coverage
Same behavior, same coverage, less boilerplate (test file 1691 -> 1512
lines; PR diff unchanged in semantics).

Production (mechanical only):
- Extract _strip_hermes_owned_pythonpath_and_runtime_markers(): the three
  builders (_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env)
  ran the identical strip-then-pop-markers sequence in the same order
  (ordering is load-bearing for VIRTUAL_ENV validation); the helper makes
  that explicit once instead of three times.

Tests:
- Non-owned preservation: 11 single-shape tests -> one parametrized matrix
  (user/Nix/other-version/python2.7/pythonX.Y-contained/raw spelling/empty
  component/empty PYTHONPATH) + one runtime-shaped matrix (other-version SP,
  venv-SP descendant, repo direct child, repo deep child).
- Owned stripping: venv SP, repo root (independent parents[2] computation),
  duplicates, all-owned key removal, mixed ordering -> one matrix.
- Builder integration: _make_run_env/_sanitize_subprocess_env/
  hermes_subprocess_env venv-SP stripping -> one parametrized test;
  same for the four PYTHONHOME builders (incl. build_subprocess_env).
- Junction: same-named non-owned negative control now covers both the
  configured-root location and an unrelated location; shared
  _physical_repo_root helper; profile resolution matrix (root->named,
  profile-shaped->named no nesting, profile-shaped->default, custom root).
- Every independent proof preserved: home-level junction, repo-level
  junction, profile interaction, negative identity control, uv-base lexical
  VIRTUAL_ENV, validated/unrelated VIRTUAL_ENV, no-scrub escape hatch,
  #84500 same-env/external-env composition, PYTHONHOME removal, real
  Windows-only semantics, POSIX fail-closed backslash paths.
2026-08-17 02:10:32 -07:00
Xinyu Du d67cd58e77 fix(tools): recover repo-level junction lexical root via exact identity
Second real-world topology reported and confirmed on native Windows 11:
the repository itself is a cross-drive junction (D:\hermes\hermes-agent ->
C:\...\hermes-agent) under a real HERMES_HOME directory.  The editable
import spelling resolves to the physical location, so _hermes_repo_root is
physical while the launcher writes the lexical spelling into PYTHONPATH.
The home-relative mapping cannot express a cross-drive link (commonpath
raises on different drives), so the lexical repo root survives stripping;
and with the repo alias missing, a lexical VIRTUAL_ENV
(D:\hermes\hermes-agent\venv) also fails _validated_runtime_venv, so the
venv site-packages survives too (uv-base gateway: both entries survive).

Fix: after the existing home/profile-root mapping, try the single
deterministic candidate <lexical root>/<repo dirname> for every trusted
home candidate (configured home, plus the profile root when the configured
home is a profile path) and accept it only when strict resolve proves it is
the exact physical repo root (fail-closed: missing paths, real directories
that are not the known repo, and unrelated spellings are never aliased).
This also re-enables the VIRTUAL_ENV validation for lexical venv spellings,
so uv-base gateway site-packages cleanup follows the repo alias.

Tests: repo-level junction positive + negative control (same-named real
directory preserved), profile-home + repo-level junction combination,
lexical VIRTUAL_ENV validation after recovery (root + site-packages
stripped, user entries kept), and a no-provenance lookalike preserved.
The execute_code composition test now compares composed paths with
os.path.normcase so a Windows case-only spelling difference (resolve() vs
abspath() casing) can never fail the composition contract.
2026-08-17 02:10:32 -07:00
Xinyu Du 57d94dd8dc fix(tools): keep junction lexical root across profile re-home
Confirmed on native Windows 11 with a real junction and the real startup
chain: when the desktop/CLI spawns the backend with HERMES_HOME in the
configured (lexical) spelling and --profile / sticky active_profile is in
play, _apply_profile_override() re-homes HERMES_HOME through
resolve_profile_env(), which resolves the junction under the platform
default and returns the PHYSICAL spelling.  tools.environments.local is
imported after that mutation, so _hermes_repo_root_aliases is built from
the physical home, the lexical repo-root spelling written into PYTHONPATH
by the launcher (D:\hermes\hermes-agent) is not derivable, and the entry
survives stripping (reproduced: cases --profile default / named / sticky
active_profile / cross-drive junction all leave it in place; no-profile
strips it).

Two narrow changes, no heuristics, no new env vars:

- hermes_cli/profiles.py::resolve_profile_env: when HERMES_HOME is set,
  the configured spelling IS the launch root (junction-transparent,
  physically identical dirs); keep it instead of re-deriving the native
  default.  This is the same producer contract _preserve_hermes_home_path
  already follows.
- tools/environments/local.py::_build_hermes_repo_root_aliases: when the
  configured home is a profile home (<root>/profiles/<name>), also derive
  the root spelling lexically (parent of the profiles component, same
  rule get_default_hermes_root uses) and run the exact-ownership mapping
  against it, so the launcher's lexical root is recovered after re-home
  without ever matching arbitrary descendants of HERMES_HOME.

Regression test test_profile_rehome_keeps_junction_lexical_alias covers
junction + profile re-home + inherited lexical PYTHONPATH end to end.
2026-08-17 02:10:32 -07:00
Xinyu Du deb4953776 test(tools): make subprocess env regressions Windows-portable
The PYTHONPATH/PATH sanitization suite was written POSIX-centric and
failed on real Windows 11 (reproduced natively: 4 failures before this
change).  Fix the tests to express the true per-platform contract:

- test_other_major_version_site_packages_preserved /
  test_make_run_env_injects_hermes_bin_dir: build inputs with
  os.pathsep instead of hardcoded ':'.
- test_make_run_env_appends_homebrew_on_minimal_path: split on
  os.pathsep, neutralise Git Bash dir prepending, and assert the
  documented Windows passthrough (_append_missing_sane_path_entries is
  a no-op off POSIX) instead of the Homebrew append.
- test_make_run_env_real_launchd_path_gains_homebrew: mark
  macos_only per repo OS-marker policy (the regression is the macOS
  launchd PATH; the merge is a passthrough on Windows).
- test_configured_home_alias_matches_launcher_output: create the
  configured-home link via a helper that falls back to an unprivileged
  directory junction (cmd /c mklink /J) when symlink creation raises
  WinError 1314, and skips with a clear reason if no mechanism exists.

Also correct a stale comment in execute_code: the child is not always
the same Python as Hermes (project mode can select an external venv),
so the strip is about compatibility, not redundancy.
2026-08-17 02:10:32 -07:00
Xinyu Du 6e9eeb5413 fix(tools): harden subprocess Python runtime ownership 2026-08-17 02:10:32 -07:00
Xinyu Du 73b49f473a fix(tools): tighten Hermes PYTHONPATH ownership semantics
Adversarial review of the previous two commits (and #78917 itself)
found three ownership-boundary issues; this commit addresses them:

1. Repo direct-child over-strip (Finding A)
   No launcher injects <repo>/tools or another direct child as an
   independent PYTHONPATH entry - audited all four producers (Electron
   electron-main.mjs, gateway/run.py::_ensure_windows_gateway_venv_imports,
   cron/scheduler.py::_windows_cron_python_invocation,
   tui_gateway/host_supervisor.py).  The depth<=1 rule deleted user paths
   that merely live under the repo directory; only the EXACT repo root is
   now stripped.

2. Windows junction/symlink alias (Finding B)
   The gateway launcher renders Hermes-owned paths under the configured
   HERMES_HOME spelling (gateway_windows.py::_preserve_hermes_home_path),
   which may be a junction to another drive, so it differs lexically from
   the resolved repo root.  _hermes_repo_root_aliases now carries both the
   resolved and unresolved spellings; both are recognized as Hermes-owned.

3. Stale abstraction rename (Phase 4)
   _strip_mismatched_site_packages -> _strip_hermes_owned_pythonpath:
   the cross-version heuristic is gone, so the old name misdescribes the
   behavior (ownership-based, not version-based).

Tests: direct-child now preserved; junction alias stripped (lexical pair
monkeypatched); Windows-only real-semantics test added (POSIX test remains
a safety test); mixed-ordering, duplicate-Hermes, and no-scrub PYTHONHOME
contract tests added.  Full file: 52 passed / 16 failed (identical failure
set to base, all isolation-venv environment issues).
2026-08-17 02:10:32 -07:00
Xinyu Du 850686a515 fix(tools): sanitize inherited PYTHONHOME (#75018)
The gateway runs inside its own venv; if its PYTHONHOME leaks into
subprocesses (terminal commands, cron no_agent scripts, TTS providers),
any child interpreter redirects its stdlib search to the Hermes venv and
crashes with version-mismatch errors before importing anything.

PYTHONHOME is now part of _ACTIVE_VENV_MARKER_VARS so all env builders
(_make_run_env, _sanitize_subprocess_env, hermes_subprocess_env, and
build_subprocess_env used by cron) drop it, consistent with Hermes'
existing PYTHONHOME handling in managed_uv.py and sqlite_runtime.py.
execute_code already scrubbed it via _SAFE_ENV_PREFIXES.

Tests cover all four builders plus the marker constant.
2026-08-17 02:10:32 -07:00
Xinyu Du ced80b2a20 fix(tools): preserve user PYTHONPATH entries (#74817 follow-up)
Remove the cross-version heuristic from _strip_mismatched_site_packages:
the subprocess env builder cannot know which Python version a child will
run, so judging user PYTHONPATH entries against the backend interpreter's
version deletes legitimate paths meant for a different child Python
(e.g. /custom/lib/python3.13/site-packages while Hermes runs 3.11).

Also fix over-strip: entries merely containing a pythonX.Y path component
(e.g. /opt/tools/python3.13/bin) were stripped even though they are not
site-packages. Hermes-owned entries (repo root, own venv site-packages)
are now identified by path ownership, not by version.

Regression tests cover both cases; user paths with any pythonX.Y
component are preserved.
2026-08-17 02:10:32 -07:00
elphamale 23a86594cc fix(mcp): read the elicitation schema under the SDK's real field name
`ElicitationHandler` read `params.requested_schema`, but on the pinned
`mcp==1.28.1` the model field is spelled `requestedSchema`. The getattr
always missed and returned its `{}` default, so
`_format_elicitation_schema_summary` took its no-properties branch and the
approval prompt collapsed to the generic

    Approval requested by MCP server '<name>'.

for every request. The field names, types, and descriptions the summary
exists to surface never reached the user, so an elicitation asking for a
card number rendered identically to one asking for a nickname — consent
without the substance of what was being consented to.

Read both spellings rather than just correcting to the 1.x name: mcp 2.0
renames this field to `requested_schema` (it renamed every model field to
snake_case and kept camelCase only as a serialization alias, which
pydantic does not expose to attribute access), so a dual read is correct
on either SDK generation and does not go wrong again on the next bump.
Verified against real 1.28.1 and 2.0.0 installs.

Every existing test in tests/tools/test_mcp_elicitation.py builds a
duck-typed `SimpleNamespace` stand-in, which carries whatever field name
the test wrote and therefore cannot detect a mismatch with the real model.
Add one test that constructs the actual `ElicitRequestFormParams` and
asserts the requested field name reaches the consent description; it fails
on the unfixed tree. The cheap stand-ins are left alone elsewhere.

Found while porting the tree to the mcp 2.x SDK in #76736, but independent
of it: this reproduces on the current pin with no other changes, #76736
does not touch this line, and the two branches merge cleanly in either
order.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:56:50 -07:00
Teknium 9adc900ab2 fix(bot-mode): route bot-to-bot sends through --create-if-missing
The teammate-messaging protocol told bots to send with
`chat -c "Bot Chat" -Q -q` and, on 'No session found', fall back to a
manual two-step (send without -c, then sessions rename) — a dance the
CLI already made unnecessary when --create-if-missing landed (#86794).
Profiles that never went through the Bots-panel birth flow (CLI-created,
pre-Bot-Mode, remote-source) hit that miss on every first contact, and
background sends swallowed the error entirely (the original silent-drop
in Hermes-Bot-Mode#48 / #88059).

- tools/bot_mode_probe.py: protocol command gains --create-if-missing
- bundled plugin: Bot Chat prompt section + @mention handoff note use
  the flag; rename-dance instructions deleted

The capability epoch hashes the protocol section, so existing eternal
Bot Chat sessions pick the new instructions up on their next message via
the established once-per-change rebuild — no per-turn cache drift.

Live-verified: fresh profile with zero sessions, protocol command
created 'Bot Chat' and delivered (PONG round-trip); second send resolved
the same session by title (no duplicate); missing-title send WITHOUT the
flag still errors loudly on stderr.
2026-08-16 23:28:13 -07:00
elphamale 77ed1bbf40 fix(mcp): seed MCP-Protocol-Version from the handshake version, not the latest
The HTTP transport seeded `MCP-Protocol-Version` from LATEST_PROTOCOL_VERSION,
which on mcp 2.x is 2026-07-28 — a revision that replaced the `initialize`
handshake with a per-request envelope. But this transport connects through
`ClientSession.initialize()`, which sends LATEST_HANDSHAKE_VERSION (2025-11-25)
in the body. Header and body therefore disagreed by construction, and a
conforming 2.x server honours the header: it routed the request onto its
per-request-envelope ladder and rejected the legacy body with

    params._meta is missing the required envelope key(s):
    io.modelcontextprotocol/protocolVersion,
    io.modelcontextprotocol/clientCapabilities

Observed against a live MCP endpoint, and confirmed by probing the same
endpoint three ways: the header at 2026-07-28 is rejected, at 2025-11-25 it
succeeds, and with no header at all it succeeds.

Third defect in this migration from one cause: the 2.x bump changed what an
existing constant *means* without revisiting its uses. The header seed was
written when LATEST_PROTOCOL_VERSION was 2025-03-26 and was correct then.

Seeded from LATEST_HANDSHAKE_VERSION, imported with a fallback to
LATEST_PROTOCOL_VERSION for SDKs predating the split, where the two are the
same thing and header and body agree either way. An explicitly configured
header still wins — that override is why servers demanding a specific revision
can have one, and a test pins it.
2026-08-16 23:26:10 -07:00
elphamale 2e1d724e3e fix(mcp): accept both SDK generations' streamable-HTTP transport arity
`streamable_http_client` yields `(read, write, get_session_id)` on mcp 1.x and
`(read, write)` on 2.x. `_run_http` unpacked a fixed 3-tuple, so on 2.x every
HTTP and SSE MCP server failed its handshake with `ValueError: not enough
values to unpack (expected 3, got 2)` and parked after exhausting its retry
ladder. Only stdio servers kept working.

This is the same defect as the import gating fixed earlier in this branch, one
layer further in. That fix's own comment claimed reaching
`streamable_http_client` was "the path that does work" — reaching it was
necessary and not sufficient, and the comment asserted the half that was never
exercised. Corrected along with the code.

Unpacked positionally rather than by arity, since this file deliberately
supports both SDK generations and `get_session_id` was never used here.

The reason this survived review is worth the test it now has: the existing
coverage in test_mcp_client_cert.py fakes the transport with a 3-tuple, so it
encoded 1.x's shape into the assertion and passed on 2.x regardless. The new
test drives `_run_http` once per arity the supported SDK range actually yields,
and asserts the streams handed to ClientSession are the first two — positional,
because 1.x's third element is not a stream. Verified it fails on the 2.x case
without this change.

Found while pointing a real HTTP MCP server at a live deployment running this
branch: the server parked at startup and no tool from it ever registered.
2026-08-16 23:26:10 -07:00
elphamale 11a9dcf567 feat(mcp): migrate to the mcp 2.x SDK
mcp 2.0.0 implements MCP revision 2026-07-28 and makes three breaking
changes Hermes sits on top of: `mcp.server.fastmcp` is gone, every model
field is renamed to snake_case (camelCase survives only as a
serialization alias, which pydantic does not expose to attribute
access), and the SDK's own HTTP stack moved from `httpx` to `httpx2`.

Bump the pin across the dev/mcp/computer-use extras and port the tree:

- `mcp_serve.py` and `agent/transports/hermes_tools_mcp_server.py` move
  from `FastMCP` to `mcp.server.MCPServer`, which has the same
  decorator/add_tool surface. The hermes-tools server already
  synthesised `__signature__` from Hermes' JSON Schema, which is exactly
  what 2.0's `add_tool` reads.
- SDK model reads go through `mcp_field(obj, snake, camel)`, which reads
  both spellings. A single-spelling read fails *silently* on the other
  generation — empty tool schemas, dropped structured content, tool
  results vanishing from sampling conversations — and `mcp` is an
  optional extra users install at their own version.
- `sdk_httpx()` resolves the httpx flavour from the SDK's own transport
  module, so objects handed to `streamable_http_client`, the `sse_client`
  factory, and the OAuth metadata helpers come from the module the
  installed SDK actually imports.
- HTTP support is gated on either streamable-HTTP entry point, not just
  the deprecated alias 2.0 removed.
- OAuth: `OAuthClientProvider` lost its `timeout` argument (the
  configured `oauth.timeout` now bounds the callback waiter's own poll
  loop, where the browser round-trip was always awaited), and
  `callback_handler` must return `AuthorizationCodeResult` rather than a
  tuple. 2.0 also validates the RFC 9207 `iss` parameter, so the
  callback handler and paste fallback capture it.

`mcp`/`mcp-types` 2.0.0 are inside the 14-day `exclude-newer` window, so
two narrow `exclude-newer-package` entries unblock `uv lock`, annotated
for removal on or after 2026-08-11. `httpx2` needs no exemption: 2.7.0 is
already outside the window and satisfies mcp's floor.

Refs #69931

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:26:10 -07:00
Yiipu 2824899321 fix(terminal): correct off-by-one in _hermes_repo_root path resolution
_hermes_repo_root used parents[1] which resolves to tools/ instead of
the repository root. The file lives at tools/environments/local.py so it
needs parents[2] to reach the actual repo root that Electron injects
into PYTHONPATH.

The test test_repo_root_stripped reused the module constant under test
as its input. This made it pass regardless of what the constant pointed
at. The test now computes the real repo root independently from the
source file location. It fails with the old parents[1] code and passes
with the fix.

Reported by spfcraze in PR #78917 review.
2026-08-16 23:26:04 -07:00
mcjoys 43c463fa95 fix(terminal): strip Hermes-venv site-packages from terminal subprocess PYTHONPATH to prevent cross-version ABI conflicts
The Desktop Electron process injects the Hermes venv's site-packages path
(e.g. .../python3.11/site-packages) into PYTHONPATH so the Python 3.11
backend can import its packages. When this PYTHONPATH leaks into terminal
subprocesses running a different Python version (e.g. Python 3.13), 3.11
C extension modules appear on sys.path ahead of the correct 3.13 versions
and crash with ImportError (PIL _imaging, cryptography, etc.).

Replace the existing blunt pop of PYTHONPATH from _ACTIVE_VENV_MARKER_VARS
with a surgical Hermes-venv-aware filter:

- Parse each PYTHONPATH entry by path
- Strip only paths under ~/.hermes/hermes-agent/venv/.../site-packages
- Preserve the Hermes source root (needed for import hermes_cli)
- Preserve all user-set PYTHONPATH entries

The same filter is applied in all three env builders:
- _make_run_env (foreground terminal commands)
- _sanitize_subprocess_env (background/PTY spawns)
- PTY env builder

This preserves env_passthrough semantics and never silently discards the
user's own PYTHONPATH configuration.
2026-08-16 23:26:04 -07:00
Teknium 66312aec48 Port from can1357/oh-my-pi#7553: allow quoted shell metacharacters in allowlist matching
command_allowlist glob rules (e.g. 'cargo *') rejected any command whose
quoted arguments contained shell metacharacters — a cargo benchmark
regex filter like '^layer3/write/(a|b)$' disqualified the whole command
even though those characters are literal to the shell.

_has_allowlist_shell_operator is now quote-aware:
- metacharacters inside single/double quotes or behind a backslash are
  treated as literal arguments;
- $ and backtick inside DOUBLE quotes still disqualify (expansion is
  active there);
- quoted/escaped control characters still disqualify when the command
  carries a -c/-e/--command/--eval-style option that hands the payload
  to another interpreter (sh -c '...', git -c alias.x='!...' x);
- unterminated quotes disqualify (shape can't be reasoned about).

Compound commands (unquoted ; & | < > backtick $( newline) are rejected
exactly as before. hermes_cli/approvals_suggest.derive_glob picks up the
same semantics via its existing import.
2026-08-16 22:10:13 -07:00
Teknium 2e4d771c69 Inspired by Factory Droid: accept unique ID prefixes in process tool lookups
Factory Droid v0.175.0 made TaskOutput/TaskStop accept task-ID prefixes so
background tasks can be referenced without pasting the full ID. Hermes'
process tool had the same friction: every action required the exact
proc_<12-hex> session ID.

ProcessRegistry.get() now falls back to unique-prefix resolution when the
exact lookup misses: 'proc_4dae' or bare '4dae' resolves to
proc_4dae56ca81f6 when exactly one running/finished session matches.
Ambiguous or too-short (<4 suffix chars) prefixes still return None, so
callers keep their existing 'No process with ID ...' error and nothing is
ever picked arbitrarily. Exact IDs never pay the scan, and a full ID that
happens to prefix another always wins.

All process actions (poll/log/wait/kill/write/submit/close) route through
get(), so they all gain prefix support from the single change.
2026-08-16 22:09:37 -07:00
Teknium ea29702749 feat(cron): --continuity / --no-continuity flags on hermes cron create/edit
CLI parity for the continuity toggle:

- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
  tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
  a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
  strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
  inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section

E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
2026-08-16 22:09:28 -07:00
Teknium 2e7a46cc27 feat(cron): continuity=true/false flag as the user-facing surface for self-context
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
2026-08-16 22:09:28 -07:00
Teknium 47d7661aa8 Inspired by Amp: cron self-context — context_from='self' gives recurring jobs run-to-run continuity
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.

- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
  continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
  (can't be validated against the store — the job doesn't exist yet at
  create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
2026-08-16 22:09:28 -07:00