New agents now get a blobatar — a deterministic soft-body face generated
from the bot's name (same name, same face, forever) — as the default
shapes mode, with full manual control:
- Face follows the name live while typing in New Agent
- Randomize re-rolls the seed; Lock face pins the current one so a later
rename can't change it (Unlock returns to name-following)
- Any of the six silhouettes (round/organic/boxy/nub/cloud/sun) can be
pinned via frozen-per-major trait positions while the rest stays
name-derived
- Classic geometric shapes remain one click away, and existing bots keep
their stored looks untouched
Wiring: blobatar@0.2.0 (zero deps, ~3.7KB) exported through the plugin
SDK (blobatarSvg / Blobatar), feature-detected in plugin.js with a
legacy-shape fallback for older desktops. Blob shape strings are
'blobatar[:seed[:kind]]' inside the existing meta.shape field, so
persistence, cross-machine ui_meta sync, and the roster's PNG backfill
(data-bot-face tag preserved) all work unchanged.
Extends ctx.os.notify (the curated plugin OS door from #78685) with icon,
action buttons, and a serializable `activate` target. Body/action clicks
focus the window and navigate to the plugin's screen; activation paths
share one resolver (hermes-open-target.ts) with hermes:// OS deep links,
so `hermes://index-network/intent/1`, `/index-network/intent/1`, and
{ path, params } all land on the same hash-router route. Approval
notifications keep their existing session-scoped channel.
Salvaged from PR #84192 by @serefyarar (net diff of the PR branch applied
onto current main; branch carried merge commits so a single authored
commit preserves attribution).
Follow-up to #89382: the operator runbook, bundled-skill docs page, and the
bundled SKILL.md now cover the organizer-scoped lookup flag and note that
/meet/ short URLs require it while webhook jobs derive the organizer
automatically.
A profile belongs to one gateway, but the Capabilities surface (Skills /
Tools / MCP) always read and wrote through the window's active backend —
scoping to a remote-owned profile silently edited the wrong machine.
- hermes.ts: capability REST helpers accept a ProfileScope
(string | {connectionId, profile}); ambient path now also carries the
active registry connection tag (same contract as the cron helpers,
#87882); profileScopeKey namespaces cache keys per connection.
- SkillsView: scope selector lists (profile, device) rows from the union
agent roster on multi-connection desktops; new fixedConnection prop
pins the whole view to a registered connection (plugin door), with a
probe-able SkillsView.supportsFixedConnection flag.
- MCP tab: live reload.mcp RPC withheld for cross-backend scopes (it
rides the active gateway socket and would reload the wrong machine).
- Bot Mode: remote-target drafts now get the live Capabilities tab
pinned to the target machine via fixedConnection, feature-detected so
older desktops keep the staged checklists.
- Config-record/hub-action stores accept scopes; cache keys fold in the
connection id so two gateways' same-named profiles never share rows.
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
acc614e72 added a raw backslash inside a <code> span in a table row;
MDX reads it as an escape and never finds the closing </code>, failing
docs-site-checks on main and every open PR. Escape it as \.
Live incident 2026-08-17: the source checkout was parked on a stale feature
branch (claude-code-inspired/local-terminal-memory-limit, days behind main),
left there by earlier tooling. 'hermes update' autostashed, refreshed lazy
backends, synced skills, and printed '✓ Code updated!' / '✓ Update complete!'
while the checkout stayed on the stale branch with none of main's new code.
Two sessions burned time on 'the fix is missing' confusion.
- Parked-branch guard: auto-switch back to the update target ONLY when the
parked branch is clean and fully merged (git cherry origin/<target> shows
nothing unmerged); the checkout then STAYS on the target instead of being
re-parked. Otherwise: loud CODE UPDATE SKIPPED block naming the branch,
behind-count, and resolution commands; exit 1; branch untouched.
- The up-to-date (commit_count == 0) path no longer switches back to a
fully-merged parked branch either.
- Post-pull gate additionally refuses to print '✓ Code updated!' when HEAD
ends up attached to a non-target branch.
- Summary lines now carry the actual branch + HEAD short-sha:
'✓ Update complete! [main @ 30fcf9580]' — drift visible at a glance.
- New config toggle updates.auto_switch_parked_branch (default true).
- Real-git-fixture regression tests (init/clone/branch, no subprocess
mocks): clean+merged auto-switch, dirty skip, unmerged skip, cherry-picked
equivalence, config opt-out, unverifiable ref, on-main fast path,
up-to-date no-repark, summary branch/sha assertions.
The unified Gateways settings page (from the recent settings merge) still
carried the legacy per-profile gateway-override machinery: an "Applies to"
profile-chip scope switcher, a scope state machine threaded through load/
save/test/sign-in paths, inherit-mode ModeCard variants, and an SSH
remote-profile mapping row.
The page is machine-level gateway management: it decides which gateway
backends this desktop can connect to, and profiles are discovered FROM the
connected gateways. It must not be profile-scoped.
- Delete the scope chips section, ScopeChip component, and the scope/setScope
state; every scope-conditional collapses to its global (scope === null)
branch. getConnectionConfig/save/apply/test/sign-in are all unscoped now.
- ModeCard local card always renders the local title/desc (inherit variants
gone); SSH remote-profile mapping row removed.
- i18n: drop now-unused gateway keys (appliesTo, allProfiles,
defaultConnection, profileConnection, inheritTitle, inheritDesc,
sshRemoteProfileTitle, sshRemoteProfileDesc) from types.ts and en/zh/
zh-hant/ja/ar in sync; rewrite the gateway intro in each locale to say
connections are machine-level and profiles come from gateways.
- Tests: replace the scope-switching component tests with a machine-level
assertion (loads getConnectionConfig(null), never a profile scope, no
scope UI rendered).
- Docs: update desktop.md and multi-connection-desktop.md wording — gateway
connections are machine-level; per-profile backend routing continues via
the profile rail / session source surfaces, not the settings page.
The electron main-process per-profile override mechanism
(getConnectionConfig(profileName), route map) and the profile-rail connect
flows are intentionally untouched; only the settings page loses the
affordance.
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):
- Promote the media-send timeout to the standard resolution pattern:
HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
cron.media_send_timeout_seconds in config.yaml, then 300s default
(mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
(environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
in delivery_errors (post-#88631 the reason reaches the run status, not
just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
Update the desktop docs for five just-merged desktop changes:
- Settings → Gateway + Settings → Connections are now one "Gateways" page:
retitle every reference, describe the Add-connection flow's four kinds
(Local / Hermes Cloud / Remote gateway / SSH) and the save-time duplicate
rules (one local; URL-normalized dedupe across remote/cloud; user@host:port
+ remote profile for SSH), and describe the "Per-profile overrides"
subsection that replaced the page-level Applies to chip row.
- Document the shared "Applies to" profile scope on the config-backed
settings pages (Model, Workspace, Safety, Memory & Context, Voice, Chat,
Advanced, Tools & Keys) and the Messaging overlay.
- Agent plugins section: bundled built-ins are hidden (user/git/project/
pip/portable installs only), Example Plugin is gone, and the section has
its own Applies to selector backed by plugins.manage's optional profile
param.
- Bot Mode: group chats are standalone Discord-style roster rows and open
in the main chat window (older builds fall back to the in-panel view).
- Desktop Plugin SDK: document the new host.openWorkspace(id, { render,
title, minWidth, onClose }) door, its refresh/re-front semantics, and
the feature-detection fallback pattern.
Also retitles the Settings → Gateway references in the web-dashboard guide.
No new pages; sidebars.ts unchanged. `npx docusaurus build` passes.
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):
- Incomplete-validator findings are now PRESERVED as partial evidence;
only the validator's pass/fail verdict is excluded from the advisory
verdict. A report with findings from an incomplete check no longer
reads as clean.
- Clean-report wording is now "no findings from completed checks"
whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
discard-pinning test, plus the completed-checks wording case.
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.
- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
static (no measurable cost); security invokes NVIDIA SkillSpector in
its keyless static-rules mode (~+1.2s per install). schema/quality
stay excluded: hygiene signal ("author not specified" is
high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
or failing, the security check reports status="incomplete" and the
adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
This also absorbs a live upstream inconsistency: SkillEvaluator's
anti-tamper cross-check on SkillSpector's risk score currently trips
on moderate-finding skills (fail verdict with zero findings, e.g.
github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
NVIDIA separately; either way an evidence-free fail must not render
as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
not-run formatting.
E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).
- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
over the quarantined bundle, parses the JSON report, classifies
secrets-class findings (private keys, tokens, credentialed connection
strings) apart from advisory PII findings. Every failure mode
(binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
guard's policy decision and before the install confirmation. Findings
are shown with file:line; secrets-class findings render red with a
loud warning. Warn-and-continue by design — the built-in skills guard
remains the only enforcement layer, because the upstream PII scanner
has known false-positive classes (git@github.com, docs example
emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
config toggle.
Scanner install (optional):
uv tool install --python 3.13 \
"skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"
E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
Builds on @nductien's completion-notify module (PR #87705):
- Notify on the gateway watcher's full terminal set — blocked, gave_up,
crashed, timed_out, block_loop_detected — not just completed. A worker
hitting a blocker while the user is away was the original community ask.
- Route all notification copy through the kanban plugin i18n bundles
(en/ja/zh/zh-hant), with an English-bundle fallback when the translator
isn't bound yet.
- Wire the ctx.os.notify door so events also fire a NATIVE OS notification
while the user is away from the Hermes window (host.notify toast covers
the foreground). OS-door failures are isolated from the toast path.
- Docs: Desktop notifications section in kanban.md, including the
app-running coverage window.
Bots could message teammates on their own machine (hermes -p <bot> chat) and
the desktop could relay user mentions over Connections, but a bot had NO
transport to a bot on another gateway. This adds one, with zero new server
surface: the peer's existing api_server platform is the wire.
- hermes_cli/subcommands/peer.py: `hermes peer add/list/remove/dm`.
`dm <peer>[/<agent>]` resolves the remote agent's canonical "Bot Chat"
(list by title, create when missing), runs one synchronous agent turn via
POST /api/sessions/{id}/chat, and prints the reply on stdout — the exact
cross-machine twin of the local bot-messaging command, so the Bot Mode
protocol composes over it unchanged. Named profiles route via the peer's
/p/<profile>/ multiplex mirror. Peer URLs live in config.yaml
(`bot_peers`); the peer's API_SERVER_KEY is a credential and lives in
~/.hermes/.env as HERMES_PEER_<NAME>_KEY.
- hermes_cli/main.py: parser wiring + fast-path/session-flag command sets.
- tools/bot_mode_probe.py: when peers are registered, the injected Bot Chat
messaging protocol gains a cross-machine paragraph (peer roster +
`hermes peer dm` pattern) so agents discover remote teammates on their
own; peers join the capability fingerprint so registering/removing one
refreshes eternal Bot Chat prompts on the next message (loud, one-time,
user-initiated — no per-turn cache drift).
- Docs: Bot Mode guide (bot-initiated DMs across machines) + cli-commands
reference (`hermes peer` section + summary row).
Tests: tests/hermes_cli/test_peer_cmd.py (target parsing, /p/ scoping,
registry round-trip in isolated config, real-loopback-HTTP dm flow incl.
Bot Chat create-vs-reuse and bearer auth), bot_mode_probe peer-paragraph +
epoch tests. E2E: real `python -m hermes_cli.main peer ...` against a live
fake peer over HTTP with isolated HERMES_HOME (config/.env persistence,
bare + /p/<profile> routing, stdin, --json). 23 passed; ruff clean.
Bot Mode's group chats spawned one per-member session per room, and those
"Group: ..." rows (plus canonical Bot Chats when the old eye-toggle pref was
off) flooded the global Sessions sidebar — a 6-bot room dumped six identical
rows into recents (reported with screenshot, Aug 17).
Plugin (apps/desktop/src/plugins/hermes-bots/plugin.js):
- session.create now passes hidden:true UNCONDITIONALLY for both canonical
Bot Chats and group-room member sessions; the $hideBotChats pref, its eye
toggle, and its storage hydrate are removed (Bot Mode sessions are plumbing
or plugin-owned forever-chats, never scratch conversations).
- hideOwnedBotSessions(): idempotent reconciliation sweep over every owned
session id (bot meta canonical chats + each room's member sessions) via
session.set_hidden, run on plugin load and on each gateway reconnect, so
rows born visible under the old pref get cleaned up.
- The Bots session browser and canonical-chat recovery scan pass
include_hidden:true so they still see the rows they own.
Gateway (tui_gateway/methods_session.py):
- session.list honors an include_hidden param (default off — the resume
picker and all global callers keep dropping hidden rows).
- session.set_hidden gains a durable fallback: when no LIVE runtime session
matches, resolve the stored session id in the target profile's state.db
(via resolve_session_id) and flip the flag there. The sweep holds stored
ids for chats that aren't live; the old live-only lookup 4001'd them.
Validated E2E with real imports against a temp HERMES_HOME: born-hidden row
(hidden=1), profile-scoped session.list default vs include_hidden (0 vs 1),
and stored-id sweep on a non-live legacy row (hidden=1). Plugin suite
167/167; new RPC regression tests in tests/tui_gateway/test_session_hidden_rpc.py.
Covers the precise pinned-session resolver added in PR #88690: request
param shape, the preferred_session response field, hidden-row and
compression-lineage resolution, and older-gateway behavior.
Community asks (Discord, Aug 17): unclear what happens across cloud vs
desktop, whether every bot replies in group chats, and how to persist
connections to multiple gateways.
Extends the Bot Mode user-guide page (landed on main today) with the new
cross-connection features:
- "Create on" picker: creating an agent on another registered machine, with
the remote-target caveats (clone source, staged capability checklists,
draft discard).
- Group chats: explicit "not every bot replies" explanation of the
round-robin/pass model, and rooms spanning machines with device badges.
- @mentions across machines via the Connections registry (no gateway switch).
- Bots-across-machines section: persistent SSH inventory, last-known rows,
and the stay-in-your-chat interaction model; cloud+desktop recipe.
- desktop.md Bot Mode section links to the full guide; multi-connection
page's Bot Mode reference points at the docs page instead of the old
standalone repo.
Bot Mode ships built into the desktop app (default on) but only had a
short section in desktop.md. This adds a dedicated user-guide page
covering the Bots roster, creating and editing Bots, avatars, routines,
group chats, bot-to-bot messaging (agent.bot_mode_protocol), the
multi-connection roster, and CLI parity.
An inline ::preview widget could render and be clicked, but the click went
nowhere: the sandbox has no channel to the agent, so an interactive chart
was a dead end. Now the frame injects a second script beside the measurer
that gives the page one voice:
window.hermes.send('get-price eth')
<button data-hermes-send="get-price eth">ETH</button> (zero-script form)
The prompt rides postMessage up tagged with the mount token, then goes
through the composer's own send path (requestComposerSubmit -> prompt.submit)
flagged display_kind=hidden — the same row-typing auto-continue and internal
notifications already use. The agent wakes and takes a real turn; the
durable row persists (context, resume, DB audit); but NO bubble renders,
live or on reload. The user clicks ETH and the chart just changes — the
off-screen loop is click -> hidden turn -> agent rewrites the widget file ->
frame hot-swaps.
Trust boundary matches size reports and is tighter where it matters: mount
token required (frames can't forge each other's intents), string-only,
trimmed, capped at 500 chars, throttled to one intent per second per frame.
The gateway whitelists display_kind to "hidden" — the RPC can't mint
arbitrary row types — and the flag threads through both turn paths (inline
and compute-host isolation) so isolated sessions don't resurrect bubbles on
resume.
The desktop platform hint teaches the model to wire interactive widgets
with data-hermes-send and to answer clicks by updating the widget's file
rather than with prose; the SDK doc documents the contract.
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.
Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).
Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.
Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
The SDK doc still described the v1 frame (fixed height attribute, rail card
under the frame) — chrome that no longer exists. And plugin_storage's usage
example used `with plugin_db(...)`, which reads as auto-close but sqlite3's
context manager only scopes transactions; the example now closes explicitly.
The first cut of the core ::preview consumer rendered the classic
preview-attachment card — a button into the right rail we already had, which
made the directive indistinguishable from an ordinary preview link. Now the
directive shows the thing itself: the workspace HTML file renders in a
sandboxed srcdoc iframe inline in the assistant message (opaque origin,
allow-scripts only — no reach into the app, its storage, or the bridge),
with an optional height attribute clamped to 120-1200px and the classic
card kept below as the rail escape hatch.
The frame waits for turn settle before reading the file (mid-stream it is
often mid-write), resolves relative paths against the session's own cwd,
and falls back to the plain card for non-HTML targets and remote gateways
(no local file door there).
Plugins that persist state have been writing into their own install tree
(<hermes home>/plugins/<name>/), which `hermes plugins update` git-pulls and
`hermes plugins remove` deletes — user data dies with the code that wrote it.
plugins/plugin_storage.py is the sanctioned home: plugin_data_dir(name) gives
one data root per plugin under <hermes home>/plugin-data/<name>/ (profile-
aware, created on first use, names validated against traversal), and
plugin_db(name) opens a WAL-mode SQLite database inside it. Secrets stay on
the existing secret-scope path — this is state, not credentials.
hermes-achievements, the in-tree offender, converts with a legacy-file
migration on first read.
The transcript becomes a contribution area (transcript.directives). A plugin
registers a named directive and the model addresses it by emitting
::name{key="value"} as its own paragraph; that leaf renders as the plugin's
component, wrapped in the contribution error boundary. Unclaimed or malformed
directives stay plain prose, so nothing changes for text that merely looks
like a directive (std::vector) or for users with the plugin disabled.
Core ships ::preview{file="..."} as the reference consumer (the existing
preview-attachment card), the desktop platform hint teaches the model the
syntax, and the SDK exports the area + types so runtime plugin.js files get
the surface through the normal plugins API.
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
(after explicit api_mode wins, before provider-name specials) via lazy
import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
usage cache reporting, model-switch override, and config roundtrip; extend
test_model_switch_openai_api_mode with meta cases.
- plugins/model-providers/meta-ai/__init__.py: drop out-of-tree install
instructions from the module docstring (now bundled)
- tests/providers/test_meta_ai_profile.py: port the plugin's test suite
into the repo (registry discovery instead of file-location import)
- website/docs/integrations/providers.md: meta-ai in the first-class
API-key provider list, META_BASE_URL override, contributor-tier
data-training note
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).
fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:
- No-op for the built-in ticker (its tick loop already self-heals
past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
external retry is de-duplicated) and runs fire_claimed in a daemon
thread, mirroring the webhook admission pattern, so housekeeping
never blocks for the length of an agent run. Provider re-arm logic
(Chronos NAS one-shots) runs exactly as for a normal fire.
Docs: cron.md section + cron.misfire_grace_minutes reference.
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.
Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.
- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
[project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
skill_view resolves cross-tier collisions in favor of the project tier
(same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)
Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.
- The distinguishability test asserted the failure payload was byte-identical
to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
That freezes the AMBIGUITY as a required property: emitting `commits: None`
for "unknown" would improve exactly what #88113 is about and fail the test.
Now asserts what the parent actually depends on -- both keep the worktree,
and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
happy path, forbidding an always-present-but-False flag (a legitimately
better JSON contract: stable key set for serializers). Now
`assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
was not even a cross-producer contract: delegate_tool's note said "state
unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
assert the note names the worktree AND branch -- the actionable part for a
human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
before any git call would keep it green while proving nothing). Now checks
`call_count` and mirrors the branch-survival + note-names-path legs its
sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
parent reads, but delegate_tool's fallback -- the second producer -- was
verified only by reading. It now AST-parses the real fallback dict literal
and compares against live `finalize_subagent_worktree()` output, so the two
producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
raising, handled in delegate_tool), and the module docstring listed
`inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
`_break_git_index()` beside the file's other module-level helpers.
Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
The preserved worktree is invisible to the only consumer that can act on it.
Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":
inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
inspected OK, child produced nothing -> {commits: 0, dirty: False, pruned: False}
The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.
Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
naming the worktree/branch, warns, and returns the payload. Both unproven
exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
non-numeric rev-list stdout) produced the same unproven payload but logged at
DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
schema missing commits/dirty/pruned. It now emits the same flagged shape, and
logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
by restoring the unconditional prune and reintroducing this P1.
Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.
Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
work intact on disk; proven-clean still prunes (pruned=true).
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
A blackholed IPv6 path to api.telegram.org never errors, so
_await_with_thread_deadline never fires and connect hangs at
"attempt 1/8". Known A-record IPs connect over IPv4 immediately.
DoH timeout now fail-opens to the seed IPv4 list instead of the
hostname. Hostname stays last for IPv6-only hosts.
Closes#87015
Phase 2 of the MCP 2026-07-28 migration (#69931), on top of the SDK 2.x
migration (#88180):
- Protocol-era negotiation (_negotiate_session): per-server `protocol`
config key — auto (default, handshake-first with server/discover
fallback on -32022/-32601), stateless (discover-first), legacy
(handshake only). Auto is handshake-first deliberately: zero extra
round-trips and zero behavior change for the entire existing server
fleet, while 2026-07-28-only servers now connect via the fallback.
All four transport call sites (stdio, SSE, new HTTP, legacy HTTP)
route through the one choke point, so the CLI/desktop probe path
inherits it too.
- SEP-2549 list caching: tools/list ttlMs/cacheScope hints are captured
during discovery and bound to the lazy-startup schema cache — TTL'd
entries expire and force a live re-probe; hint-less (pre-2026)
servers keep the never-expires behavior. Pagination continuation now
speaks both SDK generations (params= vs cursor=).
- SEP-837: OAuth client metadata declares application_type=native
(config-overridable), with a fallback for 1.x-era metadata models.
(RFC 9207 iss validation and SEP-2352 issuer-keyed credentials are
native to SDK 2.0's OAuthClientProvider — verified, no client-side
gap.)
- SEP-2577 deprecation posture: SamplingHandler docstring marks the
Sampling feature as upstream-deprecated (12-month window) — kept
fully functional, closed to new capability.
- Docs: `protocol` key in the MCP config reference.
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):
- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
(health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
are allowed in the generic loop while native anthropic stays skipped
E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
Post-merge docs sweep for the Aug 16 scout slate. Two pages:
- mcp.md: tool-result sanitization section — invisible Unicode TAG chars
(U+E0000-E007F) stripped from results/resources/descriptions (#80689);
vendor _meta surfaced to the model minus protocol-reserved
modelcontextprotocol/mcp prefixes (#80712)
- tools.md: tool result annotations section — signal-death exit notes
(subprocess -signum definite, shell 128+signum hedged) (#78074); UTF-16
read_file transcoding with disclosure hint and 10MB cap (#80717)
Security-policy docs (approvals/allowlist) intentionally untouched.
Fold the list/mapping parser INSIDE the existing string-typed-value coercion guard (the `not isinstance(_default_value_for_key(key), str)` block from e4ea0a0ed) instead of running it unconditionally, so a genuinely string-typed setting whose value merely starts with '[' or '{' is left untouched while non-string keys get JSON/YAML flow literals parsed to real lists/dicts.
Update website/docs/user-guide/configuring-models.md: the `config set only writes scalar values` note is no longer accurate; document the list/mapping support with a quoted example.
Fixes#40545#50168
paperclip#10978 made destructive replacement an explicit caller choice
in their skill-sync and package-import paths: a rerun must never remove
operator edits by default. Our hub-skill updater had the same hazard --
'hermes skills update' calls do_install(force=True), which rmtree-replaces
the skill directory even when the user edited it after install.
do_update now compares the on-disk content hash against the hash the
lockfile recorded at install time; drifted skills are skipped with a
notice and only overwritten with the new --force flag (CLI + /skills
slash path). Bundled skills already had this protection via the
user-modified manifest in hermes update; this brings hub-installed
skills to parity.
Sabotage-verified: disabling the drift check makes the new skip test fail.
Copilot CLI 1.0.79-3 added /worktree new (start a session in a new
worktree). Hermes already has hermes -w launch-time isolation; this adds
the mid-session counterpart: /worktree new [name] creates a tree under
.worktrees/ (remote-tip base, worktree_sync honored), retargets
TERMINAL_CWD + process cwd, and registers the same keep-if-unpushed exit
cleanup. /worktree shows the active tree; /worktree list lists them.
Named trees skip the hermes- prefix so the startup pruner ages them on
the slower named-tree schedule.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.