Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.
- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
static (no measurable cost); security invokes NVIDIA SkillSpector in
its keyless static-rules mode (~+1.2s per install). schema/quality
stay excluded: hygiene signal ("author not specified" is
high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
or failing, the security check reports status="incomplete" and the
adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
This also absorbs a live upstream inconsistency: SkillEvaluator's
anti-tamper cross-check on SkillSpector's risk score currently trips
on moderate-finding skills (fail verdict with zero findings, e.g.
github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
NVIDIA separately; either way an evidence-free fail must not render
as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
not-run formatting.
E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).
- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
over the quarantined bundle, parses the JSON report, classifies
secrets-class findings (private keys, tokens, credentialed connection
strings) apart from advisory PII findings. Every failure mode
(binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
guard's policy decision and before the install confirmation. Findings
are shown with file:line; secrets-class findings render red with a
loud warning. Warn-and-continue by design — the built-in skills guard
remains the only enforcement layer, because the upstream PII scanner
has known false-positive classes (git@github.com, docs example
emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
config toggle.
Scanner install (optional):
uv tool install --python 3.13 \
"skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"
E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
Bot Mode's bot-to-bot send (`hermes -p <bot> chat --in ~ -c "Bot Chat"
--create-if-missing -Q -q "..."`) runs one turn and exits. When the turn's
in-loop transcript flush failed transiently (state.db write-lock contention
with a multiplex gateway), the one-shot path had no end-of-run durable
retry: the reply reached stdout and agent.log while the resumed titled
session's stored history never changed (#88583). The interactive CLI is
immune — it retries the flush on the next persist point and finalizes the
row on quit — but every one-shot exit path lacked both.
Fix the whole class with cli._flush_one_shot_session_store():
- final _persist_session retry at one-shot exit (idempotent — per-message
persisted-marker stamps mean already-written turns are not re-written)
- drain queued async token-accounting deltas
- end_session(..., "cli_close") so resumed/created titled session rows no
longer dangle open forever after one-shot runs
Wired into _finalize_single_query (quiet -Q -q AND human -q paths, ahead
of memory-provider shutdown so nothing later can lose the turn) and into
the kanban SIGTERM handler before os._exit(0), which skips atexit and the
SessionDB token-drain hook entirely (same gap class as PR #50881).
Handed-off sessions (#88234) and persistence-isolated forks
(_persist_disabled) are skipped.
Fixes#88583🤖 Generated with Hermes Agent
Three cases in TestCreateAgentModelRecovery: the alias is never
cached, a poisoned cache is never served, and legitimate
last-known-good recovery still works.
Salvaged from #72739 (net diff onto current main; the read-side legacy-row
guard now flows through the shared _stored_session_model() helper that
landed in #88751, keeping one resolver for both chat sites).
POST /api/sessions persisted the advertised virtual alias (hermes-agent)
whenever the request omitted model or echoed the alias back; later turns
replayed it upstream as a real model id and every turn on the session
failed. Null the alias in _session_runtime_request_from_body() — shared by
session create, chat, chat-stream, and model-lock — and stop the
create-handler's raw-body fallback from bypassing that normalization.
- Remove the bundled Example Plugin from the desktop renderer plugins
(apps/desktop/src/plugins/example/). Reference/demo plugins live in the
companion hermes-example-plugins repo (already pointed to by
src/plugins/README.md); shipping the counter demo in everyone's Settings
doubled UI noise for no user value.
- Agent plugins section now hides ALL repo-bundled built-ins
(source === 'bundled': browser/browserbase, cron_providers/chronos,
model-providers/deepinfra, platform adapters, image/video backends, …).
The section is the control panel for plugins the user installed
(user/git/project/pip/portable); built-ins ship enabled-by-default and
are configured from their own surfaces. The HIDDEN_KEY_PREFIXES list
stays as a fallback for older backends. Count pill reflects the
filtered list.
- Add an "Applies to" profile scope selector to the Agent plugins section
(same pattern as the Capabilities scope selector): list/toggle any
profile's plugins without switching the whole app. Backend:
plugins.manage now accepts an optional `profile` param via the same
set_hermes_home_override contract as cron.manage / mcp.servers.*;
unscoped calls are unchanged, so older backends keep working.
- i18n: new settings.plugins.agent.appliesTo key (types + en + zh; other
locales fall back through defineLocale).
- Tests: plugins-settings.test.tsx covers bundled hiding, prefix fallback,
count pill, selector visibility, scoped list/toggle payloads; new
tests/test_plugins_manage_profile_scope.py mirrors the cron.manage
profile-scope tests (scoped read, unknown-profile 4064, no override
leak, unscoped contract unchanged).
Two bugs found by a REAL two-gateway live test (two isolated HERMES_HOMEs,
bravo running the api_server platform, alpha's agent autonomously running
`hermes peer dm` from its Bot Chat protocol; reply relayed correctly and
persisted in bravo's canonical Bot Chat):
1. peer dm parsed the session-create response flat, but api_server wraps
the row: {"object": "hermes.session", "session": {...}} — every first DM
to a fresh peer failed with "Peer did not return a session id" (and the
orphaned Bot Chat then 400'd retries with duplicate-title). Parse the
wrapped shape; the test fake now mirrors the real response shape so this
class can't pass green again.
2. api_server: a session created with no model persists the advertised
virtual model ("hermes-agent") on the row; session chat then replayed it
as a REAL model id and the provider 400'd ("hermes-agent is not a valid
model ID"). _request_agent_overrides already filters the virtual model
for per-request bodies — apply the same filter to the stored session
model at both chat sites (sync + stream), so it means "gateway default"
exactly like the request-body path.
Live E2E transcript (bravo's Bot Chat, via /api/sessions/{id}/messages):
user: Message from 🤖 alpha (@alpha): What is your callsign?
assistant: CALLSIGN-BRAVO-7
peer cmd unit suite 10/10 with the corrected fake.
Follow-up to the salvaged #88632 (Jack Lau) fix for #88532, extending the
same repair to the sibling frozen-at-init handle and closing the handle
lifecycle gap the per-path cache introduces.
1. GatewayRunner._session_db had the identical bug class: bound once as
AsyncSessionDB(SessionDB()) in __init__ on the root home, while /resume,
/title, /history and session search all execute inside
_profile_runtime_scope on a multiplexed gateway. Convert it to the same
property-with-pin pattern: per-access resolution of _default_db_path(),
one cached AsyncSessionDB per resolved path under a lock, and assignment
preserved as an explicit pin (many suites install fakes or None).
Construction-time priming keeps the #88235 init-failure broadcast at
startup.
2. Handle lifecycle: the per-path caches accumulate one open SessionDB per
profile served, but the shutdown path closed only store._db /
runner._session_db - which now resolve just the shutdown task's own
(root) scope. Secondary profiles' handles would strand their WAL write
locks until process exit, recreating the abandoned-handle leak
b454e4da76 fixed and breaking --replace restarts with 'database is
locked'. Add close_all_db_handles() / close_all_session_db_handles()
sweeps and call both from the gateway teardown path. SessionDB.close()
is idempotent, so the root handle being closed by both the legacy loop
and the sweep is safe.
Tests: sweep coverage plus a runner-property scope/pin/cache test in
tests/gateway/test_multiplex_session_db_profile_scope.py.
Fixes#88532.
A multiplexed gateway serves every profile from one process, but
SessionStore bound a single SessionDB during __init__:
self._db = SessionDB()
SessionDB(db_path=None) resolves _default_db_path() at call time and
does follow the context-local HERMES_HOME override, so the path
machinery was already correct. The problem was when it ran: at
construction, on the process's own root home, long before any inbound
event enters _profile_runtime_scope. Every profile's rows therefore
landed in the root state.db, even though the scope had redirected
get_hermes_home() correctly for the turn (that helper's own docstring
lists "sessions" among what it scopes).
The rows still carry the right profile_name, stamped from
source.profile by the same handler, so nothing in the data looks wrong.
The only visible symptom is the desktop listing a profile's session
under the default bot: _open_session_db_for_profile opens
profiles/<name>/state.db, which never received the write.
Look the handle up through a property instead, resolving the active
scope per access and caching one handle per resolved path so a hot
inbound path opens SQLite once per profile rather than once per
message. Construction stays under the cache lock so a concurrent first
message on a profile cannot open and then leak a second handle.
Assignment is preserved as an explicit pin, which is what the existing
suites rely on when they install a fake handle or disable the DB with
store._db = None, and a pin keeps winning across scope changes.
Behavior is unchanged when no profile scope is active, so single-profile
gateways resolve exactly the path they did before. This does not migrate
rows that already landed in the root store; those stay where they are.
Follow-up to the #88347 salvage: after release_all_under() force-closes a
profile's shared connection, a store re-created on the same path registers
a fresh entry under the same key. A stale holder that later calls close()
would pop that fresh entry (its refs were transferred nowhere), letting a
third store open a second connection to the same database — exactly the
multi-writer contention the shared registry exists to prevent. close()
now evicts the registry entry only when it is still its own.
The desktop's main serve process opens memory_store.db for every known
profile and nothing closed those connections before delete_profile's
rmtree — on Windows the open SQLite handles make the removal fail with
WinError 32 for both the CLI and the DELETE /api/profiles/<name> route
(#88347). POSIX unlinking of open files hid the same leak.
MemoryStore.close() is refcount-driven, so a live holder keeps the
handle forever; add MemoryStore.release_all_under(directory) to
force-close every shared connection under a directory, and call it in
delete_profile after stopping the profile backends. Inside serve the
handles live in that very process and get released; from the CLI it is
a no-op.
Fixes#88347
Bots could message teammates on their own machine (hermes -p <bot> chat) and
the desktop could relay user mentions over Connections, but a bot had NO
transport to a bot on another gateway. This adds one, with zero new server
surface: the peer's existing api_server platform is the wire.
- hermes_cli/subcommands/peer.py: `hermes peer add/list/remove/dm`.
`dm <peer>[/<agent>]` resolves the remote agent's canonical "Bot Chat"
(list by title, create when missing), runs one synchronous agent turn via
POST /api/sessions/{id}/chat, and prints the reply on stdout — the exact
cross-machine twin of the local bot-messaging command, so the Bot Mode
protocol composes over it unchanged. Named profiles route via the peer's
/p/<profile>/ multiplex mirror. Peer URLs live in config.yaml
(`bot_peers`); the peer's API_SERVER_KEY is a credential and lives in
~/.hermes/.env as HERMES_PEER_<NAME>_KEY.
- hermes_cli/main.py: parser wiring + fast-path/session-flag command sets.
- tools/bot_mode_probe.py: when peers are registered, the injected Bot Chat
messaging protocol gains a cross-machine paragraph (peer roster +
`hermes peer dm` pattern) so agents discover remote teammates on their
own; peers join the capability fingerprint so registering/removing one
refreshes eternal Bot Chat prompts on the next message (loud, one-time,
user-initiated — no per-turn cache drift).
- Docs: Bot Mode guide (bot-initiated DMs across machines) + cli-commands
reference (`hermes peer` section + summary row).
Tests: tests/hermes_cli/test_peer_cmd.py (target parsing, /p/ scoping,
registry round-trip in isolated config, real-loopback-HTTP dm flow incl.
Bot Chat create-vs-reuse and bearer auth), bot_mode_probe peer-paragraph +
epoch tests. E2E: real `python -m hermes_cli.main peer ...` against a live
fake peer over HTTP with isolated HERMES_HOME (config/.env persistence,
bare + /p/<profile> routing, stdin, --json). 23 passed; ruff clean.
Bot Mode's group chats spawned one per-member session per room, and those
"Group: ..." rows (plus canonical Bot Chats when the old eye-toggle pref was
off) flooded the global Sessions sidebar — a 6-bot room dumped six identical
rows into recents (reported with screenshot, Aug 17).
Plugin (apps/desktop/src/plugins/hermes-bots/plugin.js):
- session.create now passes hidden:true UNCONDITIONALLY for both canonical
Bot Chats and group-room member sessions; the $hideBotChats pref, its eye
toggle, and its storage hydrate are removed (Bot Mode sessions are plumbing
or plugin-owned forever-chats, never scratch conversations).
- hideOwnedBotSessions(): idempotent reconciliation sweep over every owned
session id (bot meta canonical chats + each room's member sessions) via
session.set_hidden, run on plugin load and on each gateway reconnect, so
rows born visible under the old pref get cleaned up.
- The Bots session browser and canonical-chat recovery scan pass
include_hidden:true so they still see the rows they own.
Gateway (tui_gateway/methods_session.py):
- session.list honors an include_hidden param (default off — the resume
picker and all global callers keep dropping hidden rows).
- session.set_hidden gains a durable fallback: when no LIVE runtime session
matches, resolve the stored session id in the target profile's state.db
(via resolve_session_id) and flip the flag there. The sweep holds stored
ids for chats that aren't live; the old live-only lookup 4001'd them.
Validated E2E with real imports against a temp HERMES_HOME: born-hidden row
(hidden=1), profile-scoped session.list default vs include_hidden (0 vs 1),
and stored-id sweep on a non-live legacy row (hidden=1). Plugin suite
167/167; new RPC regression tests in tests/tui_gateway/test_session_hidden_rpc.py.
The BOTS sidebar previewed each profile's most recently active session
(last_session) but clicking the row opened the pinned canonical chat —
two different session identities, so the preview described one
conversation and the click landed in another.
- profiles.list gains an optional preferred_session_ids param
({profile: session_id}): an exact, existence-checked per-profile
lookup that resolves hidden rows and compression lineages to the
live tip (the same resolver session.resume uses) and returns a
preferred_session summary alongside the unchanged last_session.
- The hermes-bots plugin sends its canonical-chat pins with each
roster poll and previews preferred_session ?? last_session.
- openBotCanonicalChat verifies pins through the precise resolver
instead of a paginated, hidden-excluding session.list window that
misjudged real hidden pins as gone; transient lookup failures no
longer clear the pin or mint a replacement chat.
- Grandfathering: a bot with history but no pin adopts the previewed
session on first open instead of minting a new empty chat — the
behavior the design comment already promised.
Closes#88200
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.
Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).
Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.
Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
Plugins that persist state have been writing into their own install tree
(<hermes home>/plugins/<name>/), which `hermes plugins update` git-pulls and
`hermes plugins remove` deletes — user data dies with the code that wrote it.
plugins/plugin_storage.py is the sanctioned home: plugin_data_dir(name) gives
one data root per plugin under <hermes home>/plugin-data/<name>/ (profile-
aware, created on first use, names validated against traversal), and
plugin_db(name) opens a WAL-mode SQLite database inside it. Secrets stay on
the existing secret-scope path — this is state, not credentials.
hermes-achievements, the in-tree offender, converts with a legacy-file
migration on first read.
Reconciles the salvaged Responses-API mandate with the bundled meta-ai
plugin that landed in #88565:
- meta-ai profile api_mode -> codex_responses (prompt caching engages
only on /v1/responses; 0% vs 93-99% measured). Custom endpoints with a
non-api.meta.ai base URL still fall through to chat_completions via
the host-driven mandate design.
- cli-config.yaml.example: point the example at MODEL_API_KEY (Meta's
documented env var) and note the bundled provider covers the default
endpoint
- tests/providers/test_meta_ai_profile.py updated for the new wire
Implement Claude Opus review findings for Meta API support:
- Document in agent/agent_init.py that provider="meta" without an api.meta.ai URL falls through to chat_completions by design (URL-driven wire selection).
- Comment on suppression guard in hermes_cli/runtime_provider.py noting api.meta.ai is handled by _detect_api_mode_for_url.
- Replace inline __import__ with top-of-module import in tests/hermes_cli/test_model_switch_openai_api_mode.py.
- Rename test_meta_retention_not_sent_when_overridden -> test_meta_retention_override_wins in tests/agent/transports/test_meta_codex_cache.py.
- Add test in tests/agent/test_meta_agent_init.py for provider="meta" fallback without api.meta.ai URL.
- Add test in tests/agent/test_auxiliary_client.py for prompt_cache_retention: "24h" under _CodexCompletionsAdapter.
Source: Claude Opus review findings for feat/meta-api-support.
Relocate host_mandated_api_mode check from top of api_mode cascade to
fallback else branch so URL-based provider-slug rewrites (e.g.
api.anthropic.com -> provider='anthropic') always run first. Previously
the mandate branch set api_mode for api.anthropic.com without rewriting
provider, leaving provider='' and causing credential_pool_matches_provider
to fail closed and discard anthropic-scoped pools (#63425 regression
introduced in 8f60e8263).
The mandate is now a true fallback for hosts without an elif branch
(api.meta.ai -> codex_responses for 93-99% prompt-cache hits vs 0% on
chat, plus future mandates) with lazy import + try/except preserved.
Add regression tests: provider=None + api.anthropic.com URL implies
provider='anthropic'/api_mode='anthropic_messages' and preserves an
anthropic credential pool; provider=None + api.meta.ai URL implies
codex_responses.
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
(after explicit api_mode wins, before provider-name specials) via lazy
import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
usage cache reporting, model-switch override, and config roundtrip; extend
test_model_switch_openai_api_mode with meta cases.
- plugins/model-providers/meta-ai/__init__.py: drop out-of-tree install
instructions from the module docstring (now bundled)
- tests/providers/test_meta_ai_profile.py: port the plugin's test suite
into the repo (registry discovery instead of file-location import)
- website/docs/integrations/providers.md: meta-ai in the first-class
API-key provider list, META_BASE_URL override, contributor-tier
data-training note
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).
fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:
- No-op for the built-in ticker (its tick loop already self-heals
past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
external retry is de-duplicated) and runs fire_claimed in a daemon
thread, mirroring the webhook admission pattern, so housekeeping
never blocks for the length of an agent run. Provider re-arm logic
(Chronos NAS one-shots) runs exactly as for a normal fire.
Docs: cron.md section + cron.misfire_grace_minutes reference.
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.
Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.
- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
[project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
skill_view resolves cross-tier collisions in favor of the project tier
(same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)
Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
The live-checkout git mutation guard blocked history-rewriting git ops
(checkout, reset --hard, rebase, cherry-pick, ...) in the running source
checkout and its worktrees on every platform. The hazard it protects
against is only real on Windows, where NTFS locks loaded module files and
an in-place rewrite can corrupt the running process. On POSIX, open file
handles pin the old inodes, so a checkout swap under a running process is
safe, and the guard mostly taxed normal dev/salvage workflows with clone
workarounds.
- tools/self_repo_guard.py: add guard_active() -> os.name == "nt"
- tools/terminal_tool.py: consult guard_active() before running the
detector; detector logic and block message unchanged for Windows
- tests: wiring tests force the guard on; new tests cover the POSIX
pass-through and the platform predicate
/simplify-code residual. The note hard-coded "'commits' and 'dirty' are
UNKNOWN", but the two probes fail independently: a bad base_commit fails
rev-list while `git status` still succeeds, so `dirty` is a REAL measurement
being reported as unknown. Safety was never affected (the worktree is preserved
either way), but telling the parent a measured value is untrustworthy is its own
kind of misreport — and it would push a human toward re-inspecting something
already proven.
`mark_worktree_payload_unproven()` now takes an `unmeasured` argument, and
finalize tracks which probe actually failed. The raising path still disclaims
both, because which probe raised is unknowable there.
Validation: 22/22 tests/tools/test_subagent_worktree.py; ruff + ty clean. New
guard mutation-checked (hard-coding "commits/dirty" back fails it).
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.
- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
`unproven_worktree_payload()`. Both producers of this schema now call them,
so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
hand-building the dict (-16 lines). The re-import is guarded: the outer
`except` can be entered because the `from tools import subagent_worktree`
itself failed, in which case the name is unbound -- an inline fallback keeps
the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
compares its key set against live `finalize_subagent_worktree()` output. Same
contract, no source reading, refactor-proof, and it actually executes the
code.
Also folded from the same review:
- Fail-closed on an unmeasurable commit count. With no `base_commit` the
rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
tree still reached `git worktree remove --force` + `git branch -D` -- the
exact bug class #88113 is about, on a public function that takes a
caller-supplied dict. Now returns un-inspected instead, with a test driving a
real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
names WHICH probe failed, its exit code, and a bounded git stderr tail, so
the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
removed the duplicated index-corruption block in favor of the existing
`_break_git_index()` helper.
Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.
- The distinguishability test asserted the failure payload was byte-identical
to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
That freezes the AMBIGUITY as a required property: emitting `commits: None`
for "unknown" would improve exactly what #88113 is about and fail the test.
Now asserts what the parent actually depends on -- both keep the worktree,
and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
happy path, forbidding an always-present-but-False flag (a legitimately
better JSON contract: stable key set for serializers). Now
`assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
was not even a cross-producer contract: delegate_tool's note said "state
unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
assert the note names the worktree AND branch -- the actionable part for a
human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
before any git call would keep it green while proving nothing). Now checks
`call_count` and mirrors the branch-survival + note-names-path legs its
sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
parent reads, but delegate_tool's fallback -- the second producer -- was
verified only by reading. It now AST-parses the real fallback dict literal
and compares against live `finalize_subagent_worktree()` output, so the two
producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
raising, handled in delegate_tool), and the module docstring listed
`inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
`_break_git_index()` beside the file's other module-level helpers.
Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
The preserved worktree is invisible to the only consumer that can act on it.
Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":
inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
inspected OK, child produced nothing -> {commits: 0, dirty: False, pruned: False}
The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.
Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
naming the worktree/branch, warns, and returns the payload. Both unproven
exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
non-numeric rev-list stdout) produced the same unproven payload but logged at
DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
schema missing commits/dirty/pruned. It now emits the same flagged shape, and
logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
by restoring the unconditional prune and reintroducing this P1.
Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.
Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
work intact on disk; proven-clean still prunes (pruned=true).
finalize_subagent_worktree() treated a non-zero exit from its rev-list
or status probes as proof of the payload defaults (commits=0, clean),
then pruned on them: git worktree remove --force plus branch -D
permanently deleted a child's uncommitted work whenever git could not
inspect the tree (e.g. a corrupted index) (#88113).
A destructive cleanup now requires affirmative proof of zero commits
plus a clean tree. Any non-zero inspection result keeps the worktree
and branch for manual review, with a warning naming both.
The rotation path flushes its un-persisted transcript to the parent (#47202)
and only then calls publish_compression_child. The abort handler rolls back
the in-memory transcript and keeps agent.session_id on the parent - its own
comment says "keep the parent live and discard the stale compacted snapshot" -
but the rows the flush just wrote are not part of what it discards. Every
failed rotation therefore leaves the parent transcript longer than it found
it, whatever the failure was.
That is survivable for a one-off failure and pathological for a sticky one.
A parent row carrying ended_at fails the publish on every attempt and nothing
in this path clears it, so each auto-compaction appends another copy of the
current turn to the transcript it was supposed to shrink. Worse, the growth
then satisfies conversation_compression's own len(durable_parent) >
len(messages) check, so the next attempt adopts the inflated snapshot as if it
were genuine concurrent activity and the in-memory transcript doubles too.
Check that one precondition before writing. It is a plain read of the row the
publish is about to read anyway, and it raises the publish's own message, so
split_status=aborted, failure_class=session_split_failed and the rollback path
are all unchanged; a live parent reaches the flush exactly as before.
Deliberately not extended to the compression lease, which is re-acquirable - a
transient miss there would abort a rotation that would otherwise have
committed. old_session_id moves above the flush so a failure raised from here
takes the same in-memory rollback as any other pre-publish failure.
Scope: this fixes the amplification for every abort cause. It does not fix
what marks a live session as ended in the first place (#88197 Bug 1), which
needs a maintainer decision on end-reason taxonomy and is tracked on the
issue; an affected session still aborts every attempt, it just stops making
itself larger while it does.
Refs #88197
/simplify-code finding: only one-shots carry a run_claim, yet the three
dispatch-failure paths called clear_run_claim unconditionally — each call
acquires _jobs_lock (blocking cross-process flock) and does a full
load_jobs read just to return False for any non-'once' job. The trigger
is exactly a failure storm (interpreter shutdown, EMFILE with N due
jobs): N serialized flock+file reads at the moment the process can least
afford I/O, all guaranteed no-ops for the majority job kind.
Gate at the call site on schedule.kind == 'once'; new mutation-checked
test proves recurring dispatch failures skip the claim I/O entirely.
9/9 tests green; ruff clean.
Follow-ups on the #87591 salvage:
- cron/scheduler.py: wrap the three clear_run_claim call sites in a
best-effort helper — clear_run_claim does load_jobs/save_jobs file I/O,
and on the interpreter-shutdown path (or with a corrupt store) it could
itself raise, defeating the skip-cleanly purpose of these early exits.
A claim that can't be cleared simply expires at the TTL, as before.
- tests/cron/test_oneshot_dispatch_failure_run_claim.py (new): 8 tests —
clear_run_claim unit contract (one-shot cleared / already-clear noop /
recurring never touched / unknown id), all three dispatch-failure paths
through a real tick() clear the claim, and a raising clear_run_claim
does not crash the tick. Mutation-verified: reverting the fix makes the
suite fail.
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
Follow-ups on the #87259 salvage:
- cron/scheduler.py: the ledger-terminal reconciliation now requires the
terminal execution row's claimed_at to be >= the in-memory claim's
registration time (_running_since). Without this, the latest terminal
row for a recurring job is usually the PREVIOUS run's outcome — a fresh
claim in the try_register_running_job -> create_execution window (or a
finished run whose worker finally block hasn't released yet) would be
force-released and the job double-dispatched. Unparseable/missing
claimed_at fails closed to the age-based bound.
- cron/scheduler.py: take the _running_job_ids snapshot for the ledger
query under _running_lock — list() over a set concurrently mutated by
try_register/release_running_job can raise RuntimeError.
- tests: existing reconciliation tests updated to the claimed_at contract;
two new race-guard tests (previous-run terminal row never releases a
fresh claim; missing claimed_at fails closed). Mutation-verified:
removing the ownership guard fails both.
The age-only stale-claim sweep (t_3778a491, already on main) force-releases
an in-memory _running_job_ids claim only once it is older than
max(2*interval, 30m). A leaked claim that is YOUNG (inside its allowance)
while the durable executions ledger already proves the last run ended stays
wedged: the job is returned as due every tick, _submit_with_guard short-
circuits on 'already running', and next_run_at keeps fast-forwarding with no
execution — the exact 2026-08-14 recurring-router incident (t_20e23f84),
which survived a gateway restart because the in-memory age bound alone could
not see a run the ledger had already finished.
sweep_stale_inflight now reconciles each in-flight claim against the durable
executions ledger (cron/executions.db): if the job's MOST RECENT execution
row is terminal (completed/failed/unknown), the run provably ended, so the
claim is stale by construction regardless of its in-memory age and is force-
released. This is a persisted-state recovery path: the ledger is written by
the worker that ran the job and read by ANY ticker process (including one
that started AFTER the leak), so a leaked claim is recoverable without
force-run/resume and without depending on which process holds it in memory.
A ledger-terminal release is authoritative — it does not write a synthetic
mark_job_run failure (the ledger already records the outcome).
Added TestLedgerTerminalReconciliation (4 tests): young+terminal -> released
(RED on main, GREEN here), no-ledger-row -> not released, running-row -> not
released, old+terminal -> released once without synthetic failure.
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.
- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
Follow-ups on the #87261 salvage:
- cron/jobs.py: the persisted-error re-arm now respects schedule legality.
Re-arming to `now` fired CRON jobs at times their expression excludes —
a weekday-only 9am job whose Friday run errored would fire on SATURDAY
(croniter measures a 24h cadence on Saturday, so 27h > cadence+grace and
the guard tripped). Cron jobs re-arm to compute_next_run(schedule, now)
— the next LEGAL occurrence — and only when that actually moves
next_run_at earlier; interval jobs (the 2026-08-14 incident class) keep
the immediate now re-arm, which is always legal for intervals.
- cron/jobs.py: cache _schedule_cadence_seconds' croniter measurement per
expr (mirrors scheduler.py's _cron_interval_cache) — it runs inside
_jobs_lock on every tick for every stale-errored job.
- tests/cron/test_persisted_error_rearm_legality.py (new): weekday job
errored Friday re-arms to Monday (not Saturday), correctly-parked cron
value untouched, interval job still due immediately.
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.
cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.
Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.