A claim that expired without a worker ever spawning (worker_pid NULL) was
reclaimed and immediately re-claimed on every dispatcher tick, with
consecutive_failures stuck at 0 — nothing could trip the breaker. Route the
reclaim through _record_task_failure (own txn after the reclaim commit, same
shape as enforce_max_runtime) instead of the salvaged raw counter increment,
so per-task max_retries / kanban.failure_limit and the gave_up event apply
and last_failure_error carries the stale lock. reclaim_task (operator path)
still resets the counter; the live-worker extend path never reaches it.
Trims the salvaged tests to one invariant that walks the breaker to its trip.
The budget in _sweep_agent_cache_under_pressure comes from the gateway's own
cgroup memory.high/memory.max, but read_anon_rss_mb() read /proc/self/status
RssAnon: only the main process. Every child in the unit (execute_code kernels,
terminal commands) is charged against the same limit, so a kernel at 4.7 GiB
pushed the unit to MemoryHigh while the valve saw <1.6 GiB and never fired;
systemd's stop then SIGKILLed the gateway mid-flush (the #80764 signature).
Under a capped cgroup v2, read own memory.stat `anon` (the same scope as the
budget); uncapped, or when the file is unreadable, keep the self reading.
_finite_limit is factored out of _cgroup_limit_bytes so the cap check and
the budget agree on what "unlimited" means.
Fixes#110549
Replace the predicate unit test with an invariant on the real builder: a skill recorded
by a foreground create is in build_learning_graph() with zero uses, and an unmarked
never-used local skill is not. `hermes curator list-unmanaged` prints the actual
marker (created_by:learn) instead of hard-coding created_by:null. Docs: curator.md and
memory.md describe the learn marker and what the journey shows.
Drop the session-only negative (session grants were never consulted on the unattended
path, so it pins pre-existing behaviour rather than the fix). Keep the permanent-key
positive and the Tirith-not-bypassed negative. Docs: command_allowlist rule keys are
honored in cron/-q/unattended sessions.
apply_wal_with_fallback returns via _apply_delete_for_wal_reset_bug on
WAL-reset-vulnerable SQLite builds (Debian 12 / Ubuntu 22.04 system Pythons,
the reporter's pinned image) before the #110848 existing-WAL cross-VM check,
so exactly the deployment class the fix targets still got zero startup signal
while doctor flagged it. Share one helper between both early-return paths and
key the once-per-process dedupe on the DB path instead of the label so a
gateway serving several profiles hears about each database. The docker doc
remedy now uses the image's python3 (it ships libsqlite3 but no sqlite3 shell).
Review finding: vulnerable-SQLite early return skipped the cross-VM ERROR; dedupe per label; sqlite3 shell absent from image.
d8dcdfd620 (v2026.9.14) made apply_wal_with_fallback refuse to ENABLE WAL
on a fresh database whose directory is on a cross-VM bind mount (virtiofs/9p),
but a database that was already WAL on such a mount kept WAL — correctly, we
never live-downgrade under other openers — and emitted nothing. The operator
in #110848 ran exactly that shape (Podman applehv virtiofs bind mount) and got
"database disk image is malformed" within a minute with no prior signal.
- apply_wal_with_fallback: in the on-disk-WAL branch, log a once-per-process
ERROR ('cross_vm_fs_existing_wal') when the DB file is on a cross-VM
filesystem, naming the two remedies (offline PRAGMA journal_mode=DELETE
after stopping every process + database.journal_mode: delete, or move the
database to a native/named volume). The fresh-DB refusal is unchanged.
- hermes doctor: _report_database_journal_modes flags a WAL database on a
cross-VM filesystem with check_warn and the same remedy (ranked above the
WAL-reset exposure warning; the exposure bookkeeping is kept).
- docs: docker.md gains "Filesystem requirements for state.db in containers";
configuration.md's database comment no longer implies operators must set
delete by hand on virtiofs.
Detection stays /proc/self/mountinfo-based (runs inside the Linux container
on macOS/Windows hosts). locking_mode=EXCLUSIVE is deliberately not adopted:
gateway, cron and workers open state.db concurrently.
Fixes#110848
Redaction ran before glob derivation, so a mined `GITHUB_TOKEN=ghp_… git push`
became the pattern `GITHUB_TOKEN=*** git *` and `--apply` persisted it to
config.yaml. In the permanent-allowlist matcher `***` is three fnmatch
wildcards, so that entry pre-approved any `GITHUB_TOKEN=… git …` command
(`sudo git push --force`, `chmod -R 777 /etc git x`) ahead of the dangerous
command detector. Globs now come from the raw normalized command; when the
redacted form would yield a different glob the command is proposed under its
dangerous-class key instead. Redaction stays for example rendering only.
Review finding: masked `***` inside a persisted glob widened the allowlist to arbitrary `KEY=… git …` commands.
- Move the `agent.redact` import out of the per-record loop in build_proposals.
- Keep the two display-boundary invariant tests (rendered `e.g.` line, --json
examples), drop the unit-level duplicate that asserted the same masking.
- website/docs: state that mined examples are masked at display time only;
state.db itself is unchanged.
- contributors/emails: map kokhlo's commit email.
A dashboard/desktop backend (and the per-profile cron ticker) serves sessions of
several profiles through the HERMES_HOME contextvar override while
gateway.multiplex_profiles stays off. _mcp_registry_scope() keyed every MCP
connection by the bare server name in that mode, so the first profile to
discover `zernio` owned the only connection and every later served profile —
including one whose config carries a different Authorization header — called
the server through it and got the other account's data back (#111151).
The registry scope now follows the served home: a routed profile (an override
naming a home other than the process home) gets the same per-profile overlay
the multiplexer uses, so a same-named server with other credentials is a
separate connection, discovery for profile B is a connect candidate instead of
"already connected", and status/tool views stay per profile. Single-profile
processes (no override) keep bare keys, byte-identical to before.
Fixes#111151
Credit: #111158 by @KoNit-K located the inert flag on hermes_cli surfaces; its
fix (activating fail-closed multiplex secret scoping from config.yaml on the
dashboard) is not taken — the connection-key seam, not the secret-scope mode,
is what leaks the connection, and flipping the process-wide mode from the
dashboard would change credential resolution for every code path in it.
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.
`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
faulthandler.register(SIGUSR2, ..., chain=True) writes the stack dump and
then re-raises the signal to its previous handler. SIGUSR2's default
disposition is "terminate", so the diagnostic hook added for #70344 kills the
very process an operator is trying to inspect (rc = -12), which is what the
#110437 reporter hit while introspecting a long-lived gateway. Nothing else
in the gateway installs a SIGUSR2 handler, so there is nothing to chain to.
Test: a child interpreter runs the real _start_install_faulthandler, receives
SIGUSR2, and must still be alive with a dump in gateway_faulthandler.log.
Red on origin/main (rc=-12), green with chain=False. Docs: the stack-dump
signal is now documented next to the event-loop watchdog.
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:
- TurnRunner._clarify_callback_sync: when the bounded wait returns a
sentinel (timeout, /new or run-end clear_session), schedule the retire
with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
the prose is routed as a follow-up (#111019). Lookup is on the adapter
class so MagicMock doubles cannot fabricate the method; no platform ==
SLACK special-case.
The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.
Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
The bridge wrote GATEWAY_ALLOW_ALL_USERS into os.environ only when unset and
never cleared it. In-process restart paths (gateway restart watcher, dashboard
profile actions) copy os.environ into the child, so a config.yaml grant became
a sticky env var: flipping allow_all_users to false and restarting left the
gateway OPEN. The bridge now tracks its own write (module flag), overwrites or
clears it on reload, exports only a truthy grant (presence-based readers such
as the Telegram intake prefilter treated "false" as configured auth), and the
two restart env builders drop the bridge-owned value so the child re-derives
the posture from its own config.yaml. Under multiplex_profiles the DEFAULT
profile's events are authorized inside its secret scope, where gate readers
never fall to os.environ; the bridged grant is now seeded into that profile's
scope mapping only (secondaries never inherit it).
Review finding: bridged GATEWAY_ALLOW_ALL_USERS survives restart and overrides a flipped config.yaml; inert for the default profile under multiplex; "false" exported as configured auth.
`gateway.allow_all_users: true` (and the top-level spelling) in config.yaml
was a silent no-op: `_TOPLEVEL_BRIDGE` never forwarded it, GatewayConfig has
no field, and every allow-all reader (authz mixin default-deny branch,
startup access check, own-policy adapters, Discord/Matrix/Email plugin
gates) consults the GATEWAY_ALLOW_ALL_USERS env var only.
Bridge the YAML key into that env var in `bridge_core_env_settings`, the
one seam every reader already shares, instead of threading a new config
attribute through ten readers:
- first-writer-wins: an explicit env var beats YAML (matches every other
{PLATFORM}_* gate);
- skipped inside a multiplexed secondary profile's scope (#80099 class):
the secondary's config.yaml must not become the default profile's policy,
and the isolation test now asserts GATEWAY_ALLOW_ALL_USERS stays unset;
- a startup warning names config.yaml as the grant source, because the key
was inert until now and a forgotten `true` flips the posture to open.
Tests: both spellings authorize a stranger through `_is_user_authorized`;
`false`, absent key, and env=false over YAML=true all stay denied.
Docs: security guide, env-var reference, gateway internals.
Fixes#110690
Follow-up to the salvaged #110940 commit:
- `cron/scheduler_prompt.py::_CRON_HINT` tells the model the sentinel is a literal
ASCII control token that must never be translated or rephrased — the prompt-side
half of the fix, so a lane answering in any other language is steered to the
canonical token instead of relying on the filter knowing that language.
- Docs: the supported-token list ('Intentional Silence Tokens') gains the zh-Hans
forms in the English page and the zh-Hans mirror gets the section it lacked.
- Tests trimmed to two invariants (translated forms match in every shape the English
ones do; prose that mentions the word is still delivered), proven red on origin/main.
The archived-holder carve-out in _set_session_title lets the next Bot open
claim the "Bot Chat" title from an archived canonical row. That is correct
for a deliberate sidebar archive, but archive_stale_sessions (the opt-in
sessions.auto_archive sweep) could also archive an idle hidden Bot Chat with
end_reason NULL, which unarchive_recoverable_session refuses; the next Bot
click then stripped the old row's title, retiring the bot's whole history
on an idle timer with no way back.
Skip the hidden canonical Bot Chat in the sweep SELECT using the same
hidden + exact-title predicate set_session_pinned already uses to protect
it, so only an explicit archive retires a Bot Chat. Reword the bot-mode doc
so it no longer promises the retired (still hidden) chat is reachable from
the archive view.
Review finding: auto_archive sweep could irreversibly retire the canonical Bot Chat; docs over-promised archive-view reachability.
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.
Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
The ported skill carried the upstream MIT text but the LICENSE file did not
say where it came from, and the frontmatter only pointed at the repo via a
non-standard `homepage:` key. Reviewers asked for proper attribution.
- LICENSE: header naming the upstream repo, the pinned upstream commit
(b1bf517c54a4…) and the copyright holder (s1dashu) above the verbatim MIT text.
- SKILL.md: `metadata.hermes.upstream: <repo> (pinned b1bf517c)` — the same
shape mono-color and pr-lens use — and the adaptation-notes blockquote now
links the upstream repo and commit so the generated docs page links the source.
- Regenerated website/docs/user-guide/skills/optional/creative/creative-ip-as-logo.md
with website/scripts/generate-skill-docs.py (scoped to this skill).
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.
Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
(square aspect, main-prompt constraints mode — no negative_prompt
parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
proposal-round skip for pre-authorized batches, dimensions-reporting
rule when the backend returns only a URL, limbless-subject note
Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.
Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).
Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.
The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:
- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
raw-string check (never resolves — resolving IS the leak trigger).
Wired as the first check in get_read_block_error() and the write
denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
_check_sensitive_path (covers write_file_tool + patch_tool), before
the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
The roster row no longer fans member faces (GroupRow renders the room
image or a single group glyph since 5afa487e9), so describe the picture
as replacing the default glyph. Member sessions are titled by roomId,
not by display name (group-turns.ts), so drop the `Group: <name>`
literal and say "room session" everywhere the doc mentioned it.
Rapid distinct events on the same logical entity (five pushes to one PR,
a burst of ticket edits, a flapping alert) each carry a fresh delivery ID,
so the idempotency cache cannot suppress them and every event wakes a
separate agent run. Roomote solved this for PR review tasks by keeping one
durable review task per PR and superseding stale heads; this ports the
same debounce-and-supersede pattern to the generic webhook adapter.
New opt-in per-route 'coalesce' block: events group by a payload-derived
key, each new event replaces the pending one and re-arms a quiet-window
timer (window_seconds, default 30), bounded by max_wait_seconds (default
300) past the group's first event so a steady stream cannot starve
dispatch. The settled group dispatches ONE agent run on the latest
event's payload/prompt/delivery templates, with a note when earlier
events were superseded. Pending groups flush on disconnect. Startup
validation rejects missing keys, non-positive windows, and the
deliver_only+coalesce combination.
Rebase onto the decomposed webhook adapter (salvage, #92066):
- Coalescing lives in a topical sibling, gateway/platforms/webhook_coalesce.py
(WebhookCoalescer + validate_coalesce_config); webhook.py only wires it in
(__init__, _validate_route, disconnect, _handle_webhook) and splits main's
_dispatch_agent_run into the HTTP-response wrapper plus _spawn_agent_run,
shared by the immediate and coalesced paths.
- Review finding (unresolved key fields collapsed unrelated entities into one
group): an event whose rendered key still contains a {placeholder} is now
dispatched immediately instead of coalesced; documented.
- Review finding (flush-on-disconnect vs process exit): disconnect() awaits
the handoff of flushed runs; the docs claim is scoped to adapter disconnect
and states that a hard kill loses the current window's buffer.
- cron_job + coalesce is rejected like deliver_only + coalesce (cron_job
landed on main after the PR branched).
- Tests trimmed from 17 to 4 (validation parametrized; debounce/supersede/
independent groups/duplicate-first in one behavioural test; max-wait +
unresolved-key; flush-on-disconnect).
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
SKILL.md is restructured to authoring standard 5 (When to Use,
Prerequisites, How to Run, Quick Reference, Procedure, Pitfalls,
Verification) — headings only, upstream body text kept. The asset
references still carried the upstream per-CLI routing tables and command
lines for other agent products; those are replaced with the native
`image_generate` route (product names are allowed only in LICENSE and
credit lines). `WebSearch` residue in recon docs/refscout becomes
`web_search`.
source.mjs wrote its ~2.6MB Google Fonts metadata cache to
$TEMP||$TMPDIR||'.', which is the project cwd on most Linux shells; it
now uses os.tmpdir() and Pitfalls documents the location. Network-access
note now mentions that moodboard.mjs also downloads the image URLs the
search hosts return. Docs page regenerated for this skill only.
Review follow-ups on the port (all verified against the upstream snapshot,
which I re-downloaded and diffed: every scripts/*.mjs and template is the
upstream file byte-for-byte after CRLF→LF, except one `reference/` →
`references/` path fix; the reference docs differ only by Hermes adaptation
notes and the same path fix).
- LICENSE: header naming the upstream repo, pinned commit 9bca227d… and the
copyright holder above the verbatim MIT text.
- SKILL.md frontmatter: `author` credits the upstream human first, Hermes
Agent second (skills/AGENTS.md rule 4); `metadata.hermes.upstream` pin in
the same shape mono-color/pr-lens use; `category: creative`; H1
`# Auteur Skill` with a linked provenance blockquote.
- `platforms` gains `windows`: the declared prerequisites (Node 18+,
Playwright, optional ffmpeg) all run on Windows and no script uses a
POSIX-only primitive (audited: no /tmp, spawn/exec of shells, fcntl, etc).
- Routing examples translated from Russian to English (marked as translated
from upstream) so an English-language skill doesn't carry stray artefacts.
- tests/skills/test_auteur_skill.py: keep the two port-specific invariants
(path annotations, de-Claude residue). Dropped the exact-count tree
snapshot (change detector), the `~/.hermes/hermes-agent` host-dependent
related_skills fallback, and the frontmatter/description checks that
tests/skills/test_authoring_standards.py already enforces repo-wide.
- Regenerated the docs page with website/scripts/generate-skill-docs.py
(scoped to this skill).
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:
- Every successful import registers its source + a content digest of
everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
files changed since the last run (digest compare; unchanged = no-op).
Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
sync; user-created skills under the import category keep conflict
semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
trigger (or leak into) a sync.
Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
User-visible export shape changed with no docs hunk. One paragraph in the JSONL section: what
the block holds (ids/roles/counts/durations, text-free), why complete is always false,
available=false when no message carries a timestamp, and that import ignores it.
When a child's final answer still missed its output_schema after the one
bounded retry, the result entry flipped to status=failed with the error
"Final answer does not satisfy the declared output_schema" — the completion
line printed ✗ and orchestrators read a finished audit as a failure. Five
audits of 413-4103 s were lost this way in the Sep 10-14 retrospective and
the parent had to mine the live transcripts; in four of them the "violation"
was a ```json fence around a valid array, which the candidate extractor
sliced to its first..last object.
Now: status stays completed, `summary` is the child's raw final text,
`schema_valid: false` + `schema_errors` carry the verdict and a `schema_note`
says the text is unvalidated; the sync completion line shows ⚠ with the
reason. The extractor tries the earliest-opening bracket span and keeps the
first that parses (fenced arrays validate). The OUTPUT CONTRACT the child
sees now says "ONLY the JSON value — no prose, no code fence" and what a miss
costs. One bounded retry is unchanged.
HERMES_DELEGATED_CHILD_CONTEXT=1 is deliberately carried into every shell/
execute_code subprocess a delegate_task child spawns (the fence must survive
exec so a grandchild `hermes kanban complete` cannot promote itself). But the
readers treated the bare flag as "fence every Kanban DB": kanban_db_connect
opened ANY board ?mode=ro and write_txn refused ANY mutation. A subagent
running a Kanban reproduction against a scratch HERMES_HOME therefore got a
silently read-only board with a misleading "descendants require an
initialized board" error; only one lane in the retrospective ever discovered
why (deleg_15dac332), every earlier kanban repro ran degraded.
The marker's value is now the fenced board ROOT (kanban_home() at spawn) and
readers deny only paths under that root or the dispatcher-pinned
HERMES_KANBAN_DB (kanban_path_is_fenced). In-process children and a legacy
"1" marker still fence everything; an inherited path marker is never
re-derived, so a grandchild that moved HERMES_HOME cannot unfence the real
board. Owner-gate tests (test_kanban_descendant_scope, cron env isolation,
kanban CLI exit status) are unchanged and green.
A delegated child's execute_code kernel was keyed correctly
(<owner>::child::<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.
A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
3fad83df31 (Aug 11) moved Relay exporter config to a plugins.toml
selected by HERMES_NEMO_RELAY_PLUGINS_TOML. A .env still carrying the legacy
exporter vars and no TOML logs ONE warning and initialises no exporters, so
users who followed the earlier docs lost every trace silently (the
maintainer's stopped Aug 20, noticed Sep 14; five multiplexed profiles on
the same box carry the same eight vars today).
- `hermes_cli/relay_plugin_migrate.py`: build the document from the
`nemo_relay.observability` dataclasses (`ComponentSpec(...).to_dict()`,
so the `type = "file"` sink discriminator is emitted), validate it by
activating it through `nemo_relay.plugin.initialize` + `clear_async`,
write `<home>/relay-plugins.toml` (tomli_w when installed, minimal emitter
otherwise), set HERMES_NEMO_RELAY_PLUGINS_TOML in that .env, and comment
the legacy lines out (never delete). Defaults mirror the removed plugin so
files land where they used to.
- `hermes update` runs it for the default home AND every live named profile
(each writes its own TOML) as a best-effort post-update step, with a loud
notice; `hermes migrate relay [--all-profiles] [--no-validate]` runs it on
demand.
- The runtime WARNING and the `hermes doctor` finding now say "NO traces
are being exported" and name the exact command and file path.
- Docs: environment-variables.md + built-in-plugins.md carry the migration
note and a complete plugins.toml example including `type = "file"`.
Under gateway.multiplex_profiles, `_start_one_profile_adapters` skipped
Platform.RELAY / Platform.WHATSAPP for secondaries with a bare `continue`,
and the startup "not being served" WARNING only covered platforms the
PRIMARY skipped. Four secondaries on one live box had WHATSAPP_ENABLED=true
and nothing in the log, status file, or `hermes gateway status` said the
channel was dead.
- `_note_unserved_secondary_platform`: one INFO per (profile, platform)
naming the reason (shared process-level ingress owned by the default) and
the remedy (enable it on the default profile, or disable it here), plus a
`<profile>:<platform>` runtime-status stamp (state=disabled,
error_code=multiplex_shared_ingress).
- `_start_secondary_profiles` folds those platforms into the loud WARNING
when NO profile (default included) runs them.
- `hermes gateway status --profile X` prints
`whatsapp: not served under multiplex (shared ingress owned by default)`
from that stamp; /api/status excludes `disabled` entries from the
platforms degraded verdict (informational, not a fault).
- Docs: multi-profile-gateways.md gets the shared-ingress rule.
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.
Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.
`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
Port from qwibitai/nanoclaw#3260: adds a third unauthorized_dm_behavior
option, 'decline'. Instead of replying with a pairing code (pair) or
staying silent (ignore), the gateway sends one short, polite decline to
the unknown sender, then stays silent toward that sender for 24 hours.
- gateway/config.py: accept 'decline' in the normalizer; new
unauthorized_dm_decline_message for custom decline text (round-trips
through to_dict/from_dict).
- gateway/pairing.py: persisted decline stamps (_declined.json) on
PairingStore with has_recent_decline/record_decline; stamps are
pruned on write and recorded BEFORE delivery so a send failure can't
become a decline storm (nanoclaw's stamp-first pattern).
- gateway/run.py: decline branch in the unauthorized-sender path;
groups still always silently ignore.
- docs: security.md + configuration.md updated.
Adapted from TypeScript (NanoClaw's pending_sender_approvals 'decline:'
stamp rows) to Hermes' existing PairingStore JSON persistence; the
owner-FYI half of nanoclaw's flow is intentionally not ported — Hermes
logs the unauthorized attempt, and pairing remains the owner-visible
grant path.
Rebase onto the decomposed gateway (salvage, #88028):
- The unauthorized-sender path moved from gateway/run.py to
gateway/run_inbound.py::_hm_admit_event; the decline branch is a sibling
helper _hm_send_unauthorized_decline next to _hm_offer_pairing_code.
- gateway/config.py now validates the enum via _normalize_choice; the
accepted set is the module constant UNAUTHORIZED_DM_BEHAVIORS (used by
both from_dict and get_unauthorized_dm_behavior so a per-platform
`extra.unauthorized_dm_behavior: decline` is honoured too). The default
decline text lives in config as DEFAULT_UNAUTHORIZED_DM_DECLINE_MESSAGE.
- Tests trimmed from 5 to 2 invariant tests (send-once-then-silent through
the real inbound path; config round-trip + real PairingStore stamp
lifecycle with a patched clock instead of rewriting the JSON file).
remove_job() deletes <cron>/output/<job_id>/ together with the record, but the
finishing run then called save_job_output(), re-creating the directory and
writing the final run into it. Every self-removing job leaked an orphan
directory the store no longer knew about, and the docs claim that "only the
job record is gone afterwards" was false. Skip the save when
self_removal_delivery_allowed() is true (the same check that already excuses
the missing record on the delivery and mark paths); delivery composes without
an output_file, as it already does for non-file paths.
The BaseException handler in _run_one_job_body still called mark_job_run on
the missing record after a self-removal crash. Guard it with the same check so
the crash path matches the completion path instead of probing a deleted record.
Docs: state that the record and its output directory are both gone.
Review finding: self-removed run re-creates the rmtree'd output dir (orphan leak); crash path marks a missing record.
Follow-up to the salvaged #111044 commits:
- self_removal_delivery_allowed() now also requires that no record currently
holds the job id. The marker alone said "this run removed its record"; it did
not say the id is still empty. A replacement record (another owner reclaiming
the id) must be treated as a stolen claim, not a self-removal.
- Drop the allow_self_removed kwarg on fire_claim_fence: the fence already has
the job_id and the ContextVar marker, so it can decide on its own; the caller
no longer threads a flag it computed from the same predicate.
- _FireOwnership.lost(): keep the explicit lost event and the no-owner short
circuit ahead of the self-removal check so an interrupted run is still
reported as lost even after it removed its record.
- _finish_completed_run: skip mark_job_run entirely for a self-removed job
(nothing to mark) instead of calling it and then excusing the False.
- Tests trimmed to two invariants, both A/B'd against origin/main: the
self-removing run delivers after a post-removal heartbeat tick (RED on main),
and a self-removal followed by a replacement record is still discarded
(GREEN on main, guards the new predicate).
- Docs: user-guide cron.md notes that a job may remove itself and still report.
The sibling surfaces of the gateway ping rendered the same false claim:
`hermes kanban block` said "needs a human decision", the Desktop toast title
said "needs a decision", the wake status line (locales/*.yaml
gateway.kanban.wake.block_loop_detected) said "needs a decision" and the
docs described the triage route as "for a human decision". A repeated-block
circuit breaker only establishes that orchestration attention is needed.
Surface sweep from PR #111131 (notifier/test hunks dropped in favour of the
typed-kind formatter from PR #111132).
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
stream watchdog sees progress on a refusal-only stream instead of timing it
out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
parts; without it an aux refusal-only turn parsed to content=None and hit the
empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
alongside-content) so there is one test per surface; drop upstream product
references from docstrings (credit stays in the PR body); pass encoding= to
the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
content_filter result, not an empty response to retry.
Authoring standard 4 requires the human first in `author`; the skill was
drafted with Hermes so the tool was credited instead. The intro cited a
third-party product, which is allowed only in LICENSE/credit lines. Also
adds metadata.hermes.category and lowercases tags to match sibling
skills; docs page regenerated for this skill only.
Tests: the two prose tests asserted sentence literals (change-detectors);
they now assert structure — three Procedure phases, a "Done when" per step,
standard headings, and the tool wiring (cronjob/desktop_preview/[SILENT]).