With memory_enabled: false but user_profile_enabled: true, the memory tool
stays (it backs USER.md) but the full MEMORY_GUIDANCE told the model to save
notes to a MEMORY.md store that does not exist. Split the guidance: a
profile-only block is injected for that configuration, directing writes to
target='user' only.
With memory.memory_enabled and memory.user_profile_enabled both false,
agent_init never builds a MemoryStore -- but check_memory_requirements()
returned True unconditionally and MEMORY_GUIDANCE was gated only on the
tool being present in valid_tool_names. So the tool shipped in every
request's schema while answering "Memory is not available" on every call,
and the system prompt still told the model to save durable facts there.
Gate both on the config flags, using the store predicate for the tool and
the already-resolved agent state for the guidance (config is not re-read
mid-conversation, so the prompt stays byte-stable). Either flag alone
still backs the tool, so only turning both off removes it.
This lets a user running a third-party provider (Hindsight, Mem0, ...)
turn the built-in files off without paying for the dead surface on every
API call. The provider's own tools are unaffected: hiding the built-in
tool moves the decision onto the toolset gate, and listing memory under
agent.disabled_toolsets remains the only switch that takes those down.
- Cross-vendor failover: when Exa's or Parallel's keyless free tier
returns a rate-limit-shaped error, the request retries once on the
other vendor's free endpoint (search + whole-batch extract). Result
notes served_by; a peer pinned to its paid tier is never used;
non-throttle errors never fail over.
- Docs: failover note + Tavily/Firecrawl keyless-when-selected rows.
- Firecrawl keyless test expectations aligned with the keyless tier.
- Updated the Tavily API key description to clarify that it is optional and keyless access is supported.
- Modified the Tavily plugin and provider to handle requests with or without an API key, using Bearer authentication when the key is provided.
- Enhanced documentation to reflect the new keyless functionality and updated environment variable descriptions.
- Added tests to ensure correct behavior for both keyed and keyless requests.
Unpinned zero-credential installs now pick Exa or Parallel by the
parity of the per-process random session id (stable within a process,
even split fleet-wide) instead of always favoring Parallel. An explicit
hermes tools selection (web.backend / per-capability keys) bypasses the
split entirely; the runner-up vendor stays in the walk as fallback.
Live E2E: 6 fresh processes split 3/3 between vendors, each performed
a real keyless search via its picked endpoint; explicit pin verified.
Exa and Parallel now each render as two picker rows in hermes tools —
'Free (keyless)' and 'Paid (API key)'. Selection persists to
web.provider_tier.<name>:
- free: always the anonymous public endpoint, even with a key set
- paid: always the keyed SDK path; missing key errors instead of
silently downgrading to the free tier (is_keyless_available also
returns False so the auto-fallback walk can't route there)
- unset: auto (key present -> paid, else keyless)
Mechanism: get_setup_schema() gains a 'variants' list the picker
flattens into sibling rows sharing one web_backend; selection writes
the tier via both _write_provider_config sites; active-row detection
matches the tier (auto mirrors use_keyless). Routing goes through a
single use_keyless() chokepoint shared by search+extract in both
providers.
Live E2E: tier=free with a fake key present searched keyless OK (a
keyed call would have 401'd); tier=paid without key errored naming
PARALLEL_API_KEY; picker rows verified for both vendors x both tiers.
A 12-request sequential burst from the same IP that earlier saw the
free-tier rate-limit error went 12/12 OK — the limit is a transient
burst/load control, not a tight standing per-IP quota. Soften the docs
and setup-schema wording accordingly (opencode users hit Exa keyless
as their default path in practice without throttling).
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).
- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
after every keyed/importable candidate; check_web_api_key() lights
the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md
E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
Follow-ups on top of the salvaged #82631 surface:
- _select_surface: an unknown model id found in the live /images/models
catalog now ROUTES to the dedicated Image API instead of only logging a
hint — without this, a model picked from the live picker that postdates
the curated snapshot would fall onto chat-completions and fail. Curated
defaults stay pinned to chat (no behaviour change for existing setups);
offline probes still fall back to chat. _HINTED_MODELS removed.
- list_models (OpenRouter): union of the live GET /images/models catalog
(43 models today) and the chat-completions image models, deduped,
defaults first; curated metadata wins for known ids, API names for the
rest. Nous Portal (no /images route) keeps its chat-only catalog.
Offline fallback: static chain + curated Image API snapshot.
- Tests updated/added: unknown-id routing (flipped from the hint-only
pinning test), non-catalog id stays on chat, merged-picker union/dedupe/
order, Nous exclusion.
- Docs: image-generation.md gains the OpenRouter Image API section and an
editing-support row.
Live-verified: picker lists 43 models; generation succeeded through the
dedicated API on google/gemini-3.1-flash-lite-image and on the previously
unreachable black-forest-labs/flux.2-klein-4b (config-selected, no kwarg).
The deeplink-driven plugin install flow shipped in #89464 (salvage of
#82735 by @serefyarar) had no docs. Adds:
- user-guide/features/plugins.md: "One-click install links (Desktop)"
section under Managing plugins — link forms (repo/enable/force), the
confirm-first dialog contract (never auto-installs, same install-time
security scanning as the CLI), hybrid-repo behavior, legacy
plugin-agent/plugin-desktop routing, hermes-dev:// in dev builds, and
the no-SDK anchor example. Cross-links the MCP "Add to Hermes link"
equivalent.
- developer-guide/desktop-plugin-sdk.md: "Distributing with an install
link" section so plugin authors find the link form next to the
packaging docs.
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):
- Promote the media-send timeout to the standard resolution pattern:
HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
cron.media_send_timeout_seconds in config.yaml, then 300s default
(mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
(environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
in delivery_errors (post-#88631 the reason reaches the run status, not
just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
Update the desktop docs for five just-merged desktop changes:
- Settings → Gateway + Settings → Connections are now one "Gateways" page:
retitle every reference, describe the Add-connection flow's four kinds
(Local / Hermes Cloud / Remote gateway / SSH) and the save-time duplicate
rules (one local; URL-normalized dedupe across remote/cloud; user@host:port
+ remote profile for SSH), and describe the "Per-profile overrides"
subsection that replaced the page-level Applies to chip row.
- Document the shared "Applies to" profile scope on the config-backed
settings pages (Model, Workspace, Safety, Memory & Context, Voice, Chat,
Advanced, Tools & Keys) and the Messaging overlay.
- Agent plugins section: bundled built-ins are hidden (user/git/project/
pip/portable installs only), Example Plugin is gone, and the section has
its own Applies to selector backed by plugins.manage's optional profile
param.
- Bot Mode: group chats are standalone Discord-style roster rows and open
in the main chat window (older builds fall back to the in-panel view).
- Desktop Plugin SDK: document the new host.openWorkspace(id, { render,
title, minWidth, onClose }) door, its refresh/re-front semantics, and
the feature-detection fallback pattern.
Also retitles the Settings → Gateway references in the web-dashboard guide.
No new pages; sidebars.ts unchanged. `npx docusaurus build` passes.
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):
- Incomplete-validator findings are now PRESERVED as partial evidence;
only the validator's pass/fail verdict is excluded from the advisory
verdict. A report with findings from an incomplete check no longer
reads as clean.
- Clean-report wording is now "no findings from completed checks"
whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
discard-pinning test, plus the completed-checks wording case.
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.
- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
static (no measurable cost); security invokes NVIDIA SkillSpector in
its keyless static-rules mode (~+1.2s per install). schema/quality
stay excluded: hygiene signal ("author not specified" is
high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
or failing, the security check reports status="incomplete" and the
adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
This also absorbs a live upstream inconsistency: SkillEvaluator's
anti-tamper cross-check on SkillSpector's risk score currently trips
on moderate-finding skills (fail verdict with zero findings, e.g.
github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
NVIDIA separately; either way an evidence-free fail must not render
as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
not-run formatting.
E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).
- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
over the quarantined bundle, parses the JSON report, classifies
secrets-class findings (private keys, tokens, credentialed connection
strings) apart from advisory PII findings. Every failure mode
(binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
guard's policy decision and before the install confirmation. Findings
are shown with file:line; secrets-class findings render red with a
loud warning. Warn-and-continue by design — the built-in skills guard
remains the only enforcement layer, because the upstream PII scanner
has known false-positive classes (git@github.com, docs example
emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
config toggle.
Scanner install (optional):
uv tool install --python 3.13 \
"skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"
E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
Builds on @nductien's completion-notify module (PR #87705):
- Notify on the gateway watcher's full terminal set — blocked, gave_up,
crashed, timed_out, block_loop_detected — not just completed. A worker
hitting a blocker while the user is away was the original community ask.
- Route all notification copy through the kanban plugin i18n bundles
(en/ja/zh/zh-hant), with an English-bundle fallback when the translator
isn't bound yet.
- Wire the ctx.os.notify door so events also fire a NATIVE OS notification
while the user is away from the Hermes window (host.notify toast covers
the foreground). OS-door failures are isolated from the toast path.
- Docs: Desktop notifications section in kanban.md, including the
app-running coverage window.
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.
Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).
Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.
Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).
fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:
- No-op for the built-in ticker (its tick loop already self-heals
past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
external retry is de-duplicated) and runs fire_claimed in a daemon
thread, mirroring the webhook admission pattern, so housekeeping
never blocks for the length of an agent run. Provider re-arm logic
(Chronos NAS one-shots) runs exactly as for a normal fire.
Docs: cron.md section + cron.misfire_grace_minutes reference.
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.
Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.
- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
[project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
skill_view resolves cross-tier collisions in favor of the project tier
(same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)
Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.
- The distinguishability test asserted the failure payload was byte-identical
to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
That freezes the AMBIGUITY as a required property: emitting `commits: None`
for "unknown" would improve exactly what #88113 is about and fail the test.
Now asserts what the parent actually depends on -- both keep the worktree,
and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
happy path, forbidding an always-present-but-False flag (a legitimately
better JSON contract: stable key set for serializers). Now
`assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
was not even a cross-producer contract: delegate_tool's note said "state
unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
assert the note names the worktree AND branch -- the actionable part for a
human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
before any git call would keep it green while proving nothing). Now checks
`call_count` and mirrors the branch-survival + note-names-path legs its
sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
parent reads, but delegate_tool's fallback -- the second producer -- was
verified only by reading. It now AST-parses the real fallback dict literal
and compares against live `finalize_subagent_worktree()` output, so the two
producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
raising, handled in delegate_tool), and the module docstring listed
`inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
`_break_git_index()` beside the file's other module-level helpers.
Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
The preserved worktree is invisible to the only consumer that can act on it.
Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":
inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
inspected OK, child produced nothing -> {commits: 0, dirty: False, pruned: False}
The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.
Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
naming the worktree/branch, warns, and returns the payload. Both unproven
exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
non-numeric rev-list stdout) produced the same unproven payload but logged at
DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
schema missing commits/dirty/pruned. It now emits the same flagged shape, and
logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
by restoring the unconditional prune and reintroducing this P1.
Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.
Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
work intact on disk; proven-clean still prunes (pruned=true).
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):
- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
(health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
are allowed in the generic loop while native anthropic stays skipped
E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
Post-merge docs sweep for the Aug 16 scout slate. Two pages:
- mcp.md: tool-result sanitization section — invisible Unicode TAG chars
(U+E0000-E007F) stripped from results/resources/descriptions (#80689);
vendor _meta surfaced to the model minus protocol-reserved
modelcontextprotocol/mcp prefixes (#80712)
- tools.md: tool result annotations section — signal-death exit notes
(subprocess -signum definite, shell 128+signum hedged) (#78074); UTF-16
read_file transcoding with disclosure hint and 10MB cap (#80717)
Security-policy docs (approvals/allowlist) intentionally untouched.
paperclip#10978 made destructive replacement an explicit caller choice
in their skill-sync and package-import paths: a rerun must never remove
operator edits by default. Our hub-skill updater had the same hazard --
'hermes skills update' calls do_install(force=True), which rmtree-replaces
the skill directory even when the user edited it after install.
do_update now compares the on-disk content hash against the hash the
lockfile recorded at install time; drifted skills are skipped with a
notice and only overwritten with the new --force flag (CLI + /skills
slash path). Bundled skills already had this protection via the
user-modified manifest in hermes update; this brings hub-installed
skills to parity.
Sabotage-verified: disabling the drift check makes the new skip test fail.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.
- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
(can't be validated against the store — the job doesn't exist yet at
create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
Claude Cowork (Aug 6, 2026) added skill & plugin security scanning:
third-party skills and plugins are automatically checked for malicious
content on upload/edit, returning pass/warn/fail. Hermes already scans
hub-installed skills (tools/skills_guard.py), but `hermes plugins
install` cloned and activated arbitrary Git repos completely unscanned —
and plugins run Python in-process, making them the more dangerous
surface.
- tools/plugin_guard.py: plugin-adapted scanner reusing the skills_guard
pattern engine. Exempts the documented provider-plugin patterns (own
requires_env API-key reads, HTTP calls with keys) on code files while
keeping true threat signals (foreign credential-store access, reverse
shells, destructive/persistence/obfuscation patterns, prompt injection
in docs). Plugin-sized structural limits; VCS/venv dirs excluded.
- hermes_cli/plugins_cmd.py: scan the temp clone before it is moved into
~/.hermes/plugins/. safe=install, caution=confirm (interactive prompt
or --force), dangerous=blocked (--force does NOT override). Re-scan on
`hermes plugins update`; a dangerous updated tree is deactivated until
the user reviews the findings. Dashboard install path returns
structured scan_blocked/scan_findings.
- Config gate: plugins.scan_on_install (default true) in config.yaml.
- Validated against all 60 bundled plugins: 57 safe, 3 caution (real
sudo / curl|sh content in their docs), 0 false-positive blocks.
- 15 new tests incl. E2E through _install_plugin_core with real git
clones.
AGENTS.override.md now takes priority over AGENTS.md in both startup
project-context loading (prompt_builder) and progressive subdirectory
hint discovery (subdirectory_hints). Lets developers keep a personal,
typically-gitignored override next to committed project instructions
without editing the tracked file.
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.
- tools/skill_ledger.py: append/list/get, blob store, actor derivation
(curator|agent|user), single-entry rollback that takes a pre-rollback
safety entry first and FAILS CLOSED when that capture fails (consistent
with the whole-run tarball rollback hardening from #63366). Path
containment check so a hand-edited ledger can't write outside
HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
delete intent recorded via absorbed_into/archived evidence),
archive_skill()/restore_skill(), and curator auto-transitions (tagged
actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
hermes curator rollback <entry-id> (whole-tree snapshot rollback
unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
(default 0 = never) + explicit hermes curator purge, recorded in the
ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
archive TTL purge.
Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).
Closes#45778, #50875. Tests adapted from #50261 by @yu-xin-c.
Rework of the #88049 inline early-return per review:
- resolve_xai_http_credentials gains an opt-in prefer_api_key flag that
checks the explicit XAI_API_KEY first and falls back to OAuth. The key
is read through tools.tool_backend_helpers.resolve_provider_secret
(config -> profile secret scope -> env/.env -> credential pool) so the
preferred path enforces the same scope policy as the existing fallback
branch, including failing closed under a multiplexed gateway turn.
- The preferred path's base URL honors HERMES_XAI_BASE_URL then
XAI_BASE_URL behind hermes_cli.auth._xai_validate_inference_base_url,
mirroring the OAuth branch (a foreign origin can't exfiltrate the key).
- x_search's _resolve_xai_bearer now calls the shared resolver with
prefer_api_key=True instead of re-implementing precedence inline (#88040).
- tools/tts_tool.py _generate_xai_tts converted to the same flag — same
root cause for /v1/tts 403s (#87045, supersedes the inline shape in
#87081 by @enwaiax).
- Regression tests retargeted at the tools.xai_http.get_env_value seam and
the shared resolver; added coverage for the flag's OAuth fallback,
HERMES_XAI_BASE_URL + origin validation, default-order stability, and a
profile-scope-only key on the preferred path.
- Docs: x-search authentication section now states the explicit API key
wins (metered billing implication).
The runtime-contract repair now also runs during hermes update and once
per session at the first computer_use call (PR #87923); the docs only
mentioned setup and toolset enablement.
The Aug 8 default-on upscaling policy (66ea4e686) chained the Clarity
Upscaler after every sub-2MP generation. Clarity is an SD1.5 creative
tile-diffusion enhancer (creativity 0.35, "masterpiece" prompt prefix) —
it redraws content, which degraded output on 100% of generations for
models like GPT Image 2 and Ideogram whose value is precise text
rendering, CJK, and photorealistic detail.
Policy now: no model upscales by default, on FAL or Krea. The `upscale`
tool param remains as a per-call opt-in (`upscale: true`); explicit
requests still chain Clarity (FAL) / Krea Enhance as before.
- FAL catalog: all 17 default-on entries flipped to upscale=False
- Krea plugin: medium + medium-turbo per-model defaults flipped off
- Tool schema: upscale param described as opt-in with a fidelity warning
- Tests updated: catalog invariant now pins all-off; default-on cases
now assert no upscaler call
- Docs (en + zh) updated to the opt-in policy
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.