Commit Graph

738 Commits

Author SHA1 Message Date
rob-maron 0b588cb3a4 MCP CIMD auth 2026-08-18 20:03:18 -07:00
Teknium 0c5f195ee2 docs: document one-click plugin install links (hermes://plugin/install)
The deeplink-driven plugin install flow shipped in #89464 (salvage of
#82735 by @serefyarar) had no docs. Adds:

- user-guide/features/plugins.md: "One-click install links (Desktop)"
  section under Managing plugins — link forms (repo/enable/force), the
  confirm-first dialog contract (never auto-installs, same install-time
  security scanning as the CLI), hybrid-repo behavior, legacy
  plugin-agent/plugin-desktop routing, hermes-dev:// in dev builds, and
  the no-SDK anchor example. Cross-links the MCP "Add to Hermes link"
  equivalent.
- developer-guide/desktop-plugin-sdk.md: "Distributing with an install
  link" section so plugin authors find the link form next to the
  packaging docs.
2026-08-18 16:42:04 -07:00
Teknium bc76f62c20 feat(cron): configurable media-send timeout + non-empty failure reasons
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):

- Promote the media-send timeout to the standard resolution pattern:
  HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
  cron.media_send_timeout_seconds in config.yaml, then 300s default
  (mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
  (environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
  in delivery_errors (post-#88631 the reason reaches the run status, not
  just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
2026-08-17 17:51:08 -07:00
Teknium 24f7f9a9da docs: reflect the unified Gateways page, settings profile scope, plugins cleanup, Bot Mode group rows, and host.openWorkspace
Update the desktop docs for five just-merged desktop changes:

- Settings → Gateway + Settings → Connections are now one "Gateways" page:
  retitle every reference, describe the Add-connection flow's four kinds
  (Local / Hermes Cloud / Remote gateway / SSH) and the save-time duplicate
  rules (one local; URL-normalized dedupe across remote/cloud; user@host:port
  + remote profile for SSH), and describe the "Per-profile overrides"
  subsection that replaced the page-level Applies to chip row.
- Document the shared "Applies to" profile scope on the config-backed
  settings pages (Model, Workspace, Safety, Memory & Context, Voice, Chat,
  Advanced, Tools & Keys) and the Messaging overlay.
- Agent plugins section: bundled built-ins are hidden (user/git/project/
  pip/portable installs only), Example Plugin is gone, and the section has
  its own Applies to selector backed by plugins.manage's optional profile
  param.
- Bot Mode: group chats are standalone Discord-style roster rows and open
  in the main chat window (older builds fall back to the in-panel view).
- Desktop Plugin SDK: document the new host.openWorkspace(id, { render,
  title, minWidth, onClose }) door, its refresh/re-front semantics, and
  the feature-detection fallback pattern.

Also retitles the Settings → Gateway references in the web-dashboard guide.
No new pages; sidebars.ts unchanged. `npx docusaurus build` passes.
2026-08-17 17:23:25 -07:00
Teknium 3360590115 fix: address second-round SkillEvaluator review feedback
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):

- Incomplete-validator findings are now PRESERVED as partial evidence;
  only the validator's pass/fail verdict is excluded from the advisory
  verdict. A report with findings from an incomplete check no longer
  reads as clean.
- Clean-report wording is now "no findings from completed checks"
  whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
  config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
  discard-pinning test, plus the completed-checks wording case.
2026-08-17 17:04:40 -07:00
Teknium 2c2697b52e feat: widen Tier 1 advisory scan to license + security checks
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.

- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
  static (no measurable cost); security invokes NVIDIA SkillSpector in
  its keyless static-rules mode (~+1.2s per install). schema/quality
  stay excluded: hygiene signal ("author not specified" is
  high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
  or failing, the security check reports status="incomplete" and the
  adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
  note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
  This also absorbs a live upstream inconsistency: SkillEvaluator's
  anti-tamper cross-check on SkillSpector's risk score currently trips
  on moderate-finding skills (fail verdict with zero findings, e.g.
  github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
  NVIDIA separately; either way an evidence-free fail must not render
  as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
  not-run formatting.

E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
2026-08-17 17:04:40 -07:00
Teknium 183f18d530 feat: advisory NVIDIA SkillEvaluator Tier 1 scan on skill installs
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).

- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
  over the quarantined bundle, parses the JSON report, classifies
  secrets-class findings (private keys, tokens, credentialed connection
  strings) apart from advisory PII findings. Every failure mode
  (binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
  guard's policy decision and before the install confirmation. Findings
  are shown with file:line; secrets-class findings render red with a
  loud warning. Warn-and-continue by design — the built-in skills guard
  remains the only enforcement layer, because the upstream PII scanner
  has known false-positive classes (git@github.com, docs example
  emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
  endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
  optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
  config toggle.

Scanner install (optional):
  uv tool install --python 3.13 \
    "skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"

E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
2026-08-17 17:04:40 -07:00
Teknium 9aa1413781 feat(desktop): extend kanban native notifications to blocker/failure events
Builds on @nductien's completion-notify module (PR #87705):

- Notify on the gateway watcher's full terminal set — blocked, gave_up,
  crashed, timed_out, block_loop_detected — not just completed. A worker
  hitting a blocker while the user is away was the original community ask.
- Route all notification copy through the kanban plugin i18n bundles
  (en/ja/zh/zh-hant), with an English-bundle fallback when the translator
  isn't bound yet.
- Wire the ctx.os.notify door so events also fire a NATIVE OS notification
  while the user is away from the Hermes window (host.notify toast covers
  the foreground). OS-door failures are isolated from the toast path.
- Docs: Desktop notifications section in kanban.md, including the
  app-running coverage window.
2026-08-17 17:00:00 -07:00
Teknium 6e22d26583 feat: project-skill quarantine + non-interactive trust inheritance
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.

Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).

Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.

Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
2026-08-17 14:06:16 -07:00
Teknium 481156139d feat: misfire catch-up for external cron providers
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).

fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:

- No-op for the built-in ticker (its tick loop already self-heals
  past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
  own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
  external retry is de-duplicated) and runs fire_claimed in a daemon
  thread, mirroring the webhook admission pattern, so housekeeping
  never blocks for the length of an agent run. Provider re-arm logic
  (Chronos NAS one-shots) runs exactly as for a normal fire.

Docs: cron.md section + cron.misfire_grace_minutes reference.
2026-08-17 11:42:25 -07:00
Teknium f891d702df feat: project-local skill discovery with per-repo trust gate
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.

Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.

- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
  get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
  project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
  [project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
  skill_view resolves cross-tier collisions in favor of the project tier
  (same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
  commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)

Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
2026-08-17 11:39:13 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
kshitij 97c4f9eeec test(delegate): assert the unproven-state contract, not its prose
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.

- The distinguishability test asserted the failure payload was byte-identical
  to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
  That freezes the AMBIGUITY as a required property: emitting `commits: None`
  for "unknown" would improve exactly what #88113 is about and fail the test.
  Now asserts what the parent actually depends on -- both keep the worktree,
  and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
  happy path, forbidding an always-present-but-False flag (a legitimately
  better JSON contract: stable key set for serializers). Now
  `assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
  of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
  was not even a cross-producer contract: delegate_tool's note said "state
  unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
  assert the note names the worktree AND branch -- the actionable part for a
  human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
  before any git call would keep it green while proving nothing). Now checks
  `call_count` and mirrors the branch-survival + note-names-path legs its
  sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
  parent reads, but delegate_tool's fallback -- the second producer -- was
  verified only by reading. It now AST-parses the real fallback dict literal
  and compares against live `finalize_subagent_worktree()` output, so the two
  producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
  commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
  raising, handled in delegate_tool), and the module docstring listed
  `inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
  `_break_git_index()` beside the file's other module-level helpers.

Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
2026-08-17 19:41:32 +05:30
kshitij 38ea711fd0 fix(delegate): tell the parent when a worktree was preserved un-inspected
The preserved worktree is invisible to the only consumer that can act on it.

Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":

  inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
  inspected OK, child produced nothing     -> {commits: 0, dirty: False, pruned: False}

The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.

Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
  naming the worktree/branch, warns, and returns the payload. Both unproven
  exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
  non-numeric rev-list stdout) produced the same unproven payload but logged at
  DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
  exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
  creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
  schema missing commits/dirty/pruned. It now emits the same flagged shape, and
  logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
  affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
  by restoring the unconditional prune and reintroducing this P1.

Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.

Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
  suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
  the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
  work intact on disk; proven-clean still prunes (pruned=true).
2026-08-17 19:41:32 +05:30
Teknium 8800ec66d6 feat(providers): wire CommandCode into doctor/dump/setup surfaces + docs
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):

- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
  (health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
  Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
  are allowed in the generic loop while native anthropic stays skipped

E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
2026-08-17 02:56:17 -07:00
konsisumer 17fa4e2944 fix(cron): direct drift remediation to user-owned pins 2026-08-17 02:27:37 -07:00
Teknium bceda18df0 docs: document MCP sanitization and tool-result annotations from the scout-slate wave
Post-merge docs sweep for the Aug 16 scout slate. Two pages:

- mcp.md: tool-result sanitization section — invisible Unicode TAG chars
  (U+E0000-E007F) stripped from results/resources/descriptions (#80689);
  vendor _meta surfaced to the model minus protocol-reserved
  modelcontextprotocol/mcp prefixes (#80712)
- tools.md: tool result annotations section — signal-death exit notes
  (subprocess -signum definite, shell 128+signum hedged) (#78074); UTF-16
  read_file transcoding with disclosure hint and 10MB cap (#80717)

Security-policy docs (approvals/allowlist) intentionally untouched.
2026-08-17 00:17:05 -07:00
Teknium 0e378e59aa Port from paperclipai/paperclip#10978: skip locally-edited hub skills on update unless --force
paperclip#10978 made destructive replacement an explicit caller choice
in their skill-sync and package-import paths: a rerun must never remove
operator edits by default. Our hub-skill updater had the same hazard --
'hermes skills update' calls do_install(force=True), which rmtree-replaces
the skill directory even when the user edited it after install.

do_update now compares the on-disk content hash against the hash the
lockfile recorded at install time; drifted skills are skipped with a
notice and only overwritten with the new --force flag (CLI + /skills
slash path). Bundled skills already had this protection via the
user-modified manifest in hermes update; this brings hub-installed
skills to parity.

Sabotage-verified: disabling the drift check makes the new skip test fail.
2026-08-16 22:10:37 -07:00
Teknium ea29702749 feat(cron): --continuity / --no-continuity flags on hermes cron create/edit
CLI parity for the continuity toggle:

- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
  tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
  a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
  strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
  inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section

E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
2026-08-16 22:09:28 -07:00
Teknium 2e7a46cc27 feat(cron): continuity=true/false flag as the user-facing surface for self-context
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
2026-08-16 22:09:28 -07:00
Teknium 47d7661aa8 Inspired by Amp: cron self-context — context_from='self' gives recurring jobs run-to-run continuity
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.

- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
  continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
  (can't be validated against the store — the job doesn't exist yet at
  create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
2026-08-16 22:09:28 -07:00
Teknium 07a5179158 Inspired by Poke: nudge review of repeatedly-failing recurring cron jobs
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.

- cron/jobs.py: persist a failure_streak counter in mark_job_run —
  incremented on agent failure, reset on success; delivery failures don't
  count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
  delivered failure summary once a recurring job's streak reaches
  cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
  nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
  failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.

Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
2026-08-16 22:08:59 -07:00
Teknium d44a295492 Inspired by Claude Cowork: security scanning for plugin install/update
Claude Cowork (Aug 6, 2026) added skill & plugin security scanning:
third-party skills and plugins are automatically checked for malicious
content on upload/edit, returning pass/warn/fail. Hermes already scans
hub-installed skills (tools/skills_guard.py), but `hermes plugins
install` cloned and activated arbitrary Git repos completely unscanned —
and plugins run Python in-process, making them the more dangerous
surface.

- tools/plugin_guard.py: plugin-adapted scanner reusing the skills_guard
  pattern engine. Exempts the documented provider-plugin patterns (own
  requires_env API-key reads, HTTP calls with keys) on code files while
  keeping true threat signals (foreign credential-store access, reverse
  shells, destructive/persistence/obfuscation patterns, prompt injection
  in docs). Plugin-sized structural limits; VCS/venv dirs excluded.
- hermes_cli/plugins_cmd.py: scan the temp clone before it is moved into
  ~/.hermes/plugins/. safe=install, caution=confirm (interactive prompt
  or --force), dangerous=blocked (--force does NOT override). Re-scan on
  `hermes plugins update`; a dangerous updated tree is deactivated until
  the user reviews the findings. Dashboard install path returns
  structured scan_blocked/scan_findings.
- Config gate: plugins.scan_on_install (default true) in config.yaml.
- Validated against all 60 bundled plugins: 57 safe, 3 caution (real
  sudo / curl|sh content in their docs), 0 false-positive blocks.
- 15 new tests incl. E2E through _install_plugin_core with real git
  clones.
2026-08-16 22:08:37 -07:00
Teknium a8d5e16ccf Port from earendil-works/pi#7681: support AGENTS.override.md context override
AGENTS.override.md now takes priority over AGENTS.md in both startup
project-context loading (prompt_builder) and progressive subdirectory
hint discovery (subdirectory_hints). Lets developers keep a personal,
typically-gitignored override next to committed project instructions
without editing the tracked file.
2026-08-16 22:07:43 -07:00
Teknium efe41abde0 feat(curator): per-mutation audit ledger + single-edit rollback
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.

- tools/skill_ledger.py: append/list/get, blob store, actor derivation
  (curator|agent|user), single-entry rollback that takes a pre-rollback
  safety entry first and FAILS CLOSED when that capture fails (consistent
  with the whole-run tarball rollback hardening from #63366). Path
  containment check so a hand-edited ledger can't write outside
  HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
  delete intent recorded via absorbed_into/archived evidence),
  archive_skill()/restore_skill(), and curator auto-transitions (tagged
  actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
  Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
  hermes curator rollback <entry-id> (whole-tree snapshot rollback
  unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
  (default 0 = never) + explicit hermes curator purge, recorded in the
  ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
  archive TTL purge.

Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).

Closes #45778, #50875. Tests adapted from #50261 by @yu-xin-c.
2026-08-16 22:06:41 -07:00
Teknium 3b9a963b8e refactor(xai): lift API-key precedence into resolve_xai_http_credentials behind prefer_api_key
Rework of the #88049 inline early-return per review:

- resolve_xai_http_credentials gains an opt-in prefer_api_key flag that
  checks the explicit XAI_API_KEY first and falls back to OAuth. The key
  is read through tools.tool_backend_helpers.resolve_provider_secret
  (config -> profile secret scope -> env/.env -> credential pool) so the
  preferred path enforces the same scope policy as the existing fallback
  branch, including failing closed under a multiplexed gateway turn.
- The preferred path's base URL honors HERMES_XAI_BASE_URL then
  XAI_BASE_URL behind hermes_cli.auth._xai_validate_inference_base_url,
  mirroring the OAuth branch (a foreign origin can't exfiltrate the key).
- x_search's _resolve_xai_bearer now calls the shared resolver with
  prefer_api_key=True instead of re-implementing precedence inline (#88040).
- tools/tts_tool.py _generate_xai_tts converted to the same flag — same
  root cause for /v1/tts 403s (#87045, supersedes the inline shape in
  #87081 by @enwaiax).
- Regression tests retargeted at the tools.xai_http.get_env_value seam and
  the shared resolver; added coverage for the flag's OAuth fallback,
  HERMES_XAI_BASE_URL + origin validation, default-order stability, and a
  profile-scope-only key on the preferred path.
- Docs: x-search authentication section now states the explicit API key
  wins (metered billing implication).
2026-08-16 21:20:29 -07:00
Teknium 86b2057a1b docs(computer-use): note driver contract auto-repair at update and runtime
The runtime-contract repair now also runs during hermes update and once
per session at the first computer_use call (PR #87923); the docs only
mentioned setup and toolset enablement.
2026-08-16 14:08:54 -07:00
f-trycua 12b1f0f83d fix(computer-use): align browser guidance and screenshots 2026-08-16 11:34:40 -07:00
Francesco Bonacci 81af2ef013 fix(computer-use): reconcile existing cua-driver installs 2026-08-16 11:34:40 -07:00
Francesco Bonacci a403fe6f92 feat(computer-use): support Cua Driver 0.20 runtime contracts 2026-08-16 11:34:40 -07:00
Teknium f06c41522e fix(image_gen): disable default-on upscaling everywhere — opt-in only
The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.

Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.

- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
  now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy
2026-08-16 11:12:05 -07:00
Teknium d709d29f19 fix(agent): trim background_review to the enabled switch
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
2026-08-16 10:27:52 -07:00
Ojas Sharma 7095e23eb2 fix(agent): attribute background-review usage and add cost controls
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.

Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
2026-08-16 06:38:38 -07:00
Nikola Hristov d083b85591 feat(hooks): pre_tool_call content transformation via modify directive
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.

- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
  and returns (block_message, modified_args); modify directives
  shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
  {"action": "modify", "args": {...}} and Claude Code-compatible
  {"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
  dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).

Salvaged from PR #28953. Best fix for #18988.
2026-08-15 23:03:52 -07:00
Teknium 20cf326bd1 fix(computer-use): align browser authorization with live-verified cua-driver 0.19.3 contract
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):

- bounded serve flags corrected: the daemon accepts
  --session-policy/--approve-session-policy, not the docs'
  --capability-manifest names (which it rejects). Verified end-to-end:
  a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
  TTY) and its token is a legacy compatibility path disabled by default
  on current drivers (per the live browser_prepare schema). Kept as a
  passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
  with cua-driver's trusted-launcher grant. config opt-in
  computer_use.grant_existing_profile: true appends
  --grant existing-profile to the standard-mode MCP spawn (MCP
  initialize verified accepting the flag). Default false = attachment
  keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
  ladder: config grant > bounded manifest > YOLO; token = legacy.
2026-08-15 15:04:32 -07:00
Teknium 48dd9c87cf feat(computer-use): user-facing authorization for cua-driver browser attachment
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:

- hermes computer-use browser-approve: CLI passthrough that mints
  cua-driver's five-minute single-use attachment token for one exact
  (pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
  browser_route), forwarded only for existing_profile and only as a
  non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
  private per-session embedded daemon launched with
  --capability-manifest/--approve-capability-manifest; missing manifest
  fails loudly. 'unrestricted' is deliberately NOT a config value —
  it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
  rungs and the isolated-profile-first default.

E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
2026-08-15 15:04:32 -07:00
Teknium bb4f680f22 fix(browser): named browser_exec sessions compose with every backend
session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.

Now a named session composes with whatever browser source is configured:

- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
  pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
  too; previously a named daemon ignored it and fell back to scanning
  local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
  keys its own provider browser via the shared _get_session_info cache
  (bu-named-<name>), so each name gets its own cloud browser, the same
  name reuses one across calls and tasks, and unnamed calls keep the
  per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
  (provider resolution would double-session and double-bill).

Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.

E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
2026-08-15 04:05:57 -07:00
Teknium fc8ebff6d8 feat(mcp): unify the desktop MCP suggestion directory into the catalog
The desktop app carried its own hardcoded list of 17 vendor MCP endpoints
(apps/desktop/src/lib/mcp-directory.ts) powering the composer suggestion
pills — a second PR-reviewed vendor list, overlapping and drifting from the
Nous-approved MCP catalog (optional-mcps/).

This makes the catalog the single source of truth:

- manifest schema: optional `suggest:` block (keywords + hosts), parsed,
  validated, and normalized in mcp_catalog.py
- 15 new URL-only hosted-remote catalog entries (atlassian, sentry, datadog,
  notion, stripe, vercel, supabase, netlify, hugging_face, asana, intercom,
  airtable, webflow, paypal, square); figma + linear manifests gain suggest
  blocks
- GET /api/mcp/catalog now serves the suggest metadata
- desktop suggestion provider builds its match index from the catalog;
  the static directory remains only as a compatibility rung for older
  backends without suggest metadata
- setup card source line prefers the catalog entry's transport URL

GitHub stays out of the catalog on purpose: its hosted MCP rejects generic
DCR and the bundled github/* skills (gh CLI) are the stronger integration.
New desktop `github` suggestion provider offers the github-auth skill
instead — gated on a new cached GET /api/git/gh-auth probe so already-
authenticated users never see the pill.
2026-08-15 02:03:24 -07:00
yuric 569b7a34b1 fix(cron): bound post-run cleanup 2026-08-14 21:55:14 -07:00
Jack Lau 6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00
Gille 4542192086 docs: fix voice mode guide link 2026-08-15 09:56:12 +05:30
teknium1 f79440e0f4 feat: /loop — recurring in-session wakeups (Claude Code parity)
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).

Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.

Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.

Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
  process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
  post-turn tick completion, and a supervised loop_wakeup_watcher that
  injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
  notification-poller wakeup driver + post-turn completion in the turn
  dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
  ticks defer until it finishes, pauses, or parks; real user input always
  wins over both

Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
2026-08-14 13:40:19 -07:00
Teknium 44bdcf3a30 docs(dashboard): document memory/disk pressure banner and /api/status resource blocks (NS-656)
Covers the advisory memory and disk blocks added to GET /api/status in
#84965 (thresholds, staleness handling, fail-safe degradation) and the
dashboard resource-pressure banner (trigger precedence, boot-scoped
dismissals).
2026-08-13 22:10:12 -07:00
Teknium c692312704 feat(kanban): GC stale done-task notify subscriptions
Now that subscriptions survive `done` (completion is reversible —
on every 5s notifier tick forever. Add
kanban_db.purge_stale_done_notify_subs(): one DELETE removing subs
whose task has been done with no new events past a retention window
(age = latest task event, falling back to completed_at/created_at, so
any activity exempts the task; a reopened task is exempt by status
alone). The notifier watcher runs it per board once at startup and at
most hourly, re-reading kanban.done_sub_retention_days (config.yaml,
default 30; 0 disables) at each sweep.
2026-08-13 12:21:04 -07:00
kyinhub e5fd1c7b43 fix(kanban): make creator wake turns graph-safe
Carry the worker's completion handoff into the synthetic creator wake
turn and label it as an automatic notification with inspect-the-board /
don't-recreate guidance, so a woken orchestrator doesn't re-decompose
work that already exists (#70752).

Salvaged from PR #71100 by @yinkev; ported onto the restructured wake
region (delivery_mode gating, scope_id, sub chat_id destinations). The
auto_subscribe_on_create config-default half of the original PR was
dropped as already superseded on main.
2026-08-13 12:10:34 -07:00
Teknium 70143def0f fix(docs): remove stray conflict marker in cron.md 2026-08-13 11:20:27 -07:00
S 05a84f205e fix: clarify one-shot cron drift recovery 2026-08-13 11:20:27 -07:00
verybigdog 6e81ce273c feat(kanban): explicit notify/wake delivery modes with faithful wake session routing
Salvage of #37865 by @verybigdog. Adds delivery_mode (notify / notify+wake / wake)
on kanban notify subscriptions, persists chat_type + user_id_alt so a woken turn
reconstructs the creator's real session key, inherits the return path to child
tasks, and keeps wake out of the model-exposed send_message schema.

Original commits were authored under a local placeholder identity
(hermes-agent@users.noreply.local); re-attributed to the contributor's
public email.
2026-08-13 10:47:40 -07:00
Teknium 151e8af932 docs: document nemo_relay session-span segmentation config
Adds an observability/nemo_relay section to the built-in plugins page
(the plugin had no section despite appearing in the shipped table) with
the gateway.telemetry.session_segments keys, defaults-off contract, and
segment metadata; mirrors a summary in the plugin README.
2026-08-13 10:45:15 -07:00
Teknium c41db467fa docs(delegation): document delegate_task action='list'/'steer'/'stop' live orchestration
Companion docs for #85232 — adds the model-facing control section
(list/steer/stop, ownership scoping, spawn-cap exemption) above the
existing TUI/gateway subagent.steer RPC docs.
2026-08-13 10:21:45 -07:00