Commit Graph

2673 Commits

Author SHA1 Message Date
teknium1 8bc5894a3a docs(skills): ip-as-logo follows the modern section order; add skill tests
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.

Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
2026-09-15 04:20:15 -07:00
teknium1 febea3e28d docs(skills): pin ip-as-logo upstream attribution to author, repo and commit
The ported skill carried the upstream MIT text but the LICENSE file did not
say where it came from, and the frontmatter only pointed at the repo via a
non-standard `homepage:` key. Reviewers asked for proper attribution.

- LICENSE: header naming the upstream repo, the pinned upstream commit
  (b1bf517c54a4…) and the copyright holder (s1dashu) above the verbatim MIT text.
- SKILL.md: `metadata.hermes.upstream: <repo> (pinned b1bf517c)` — the same
  shape mono-color and pr-lens use — and the adaptation-notes blockquote now
  links the upstream repo and commit so the generated docs page links the source.
- Regenerated website/docs/user-guide/skills/optional/creative/creative-ip-as-logo.md
  with website/scripts/generate-skill-docs.py (scoped to this skill).
2026-09-15 04:20:15 -07:00
Teknium 398279fd6a feat(skills): add ip-as-logo optional skill (minimal cute IP mascot marks)
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.

Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
  (square aspect, main-prompt constraints mode — no negative_prompt
  parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
  own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
  proposal-round skip for pre-authorized batches, dimensions-reporting
  rule when the backend returns only a URL, limbless-subject note

Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.

Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).

Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
2026-09-15 04:20:15 -07:00
teknium1 b531023622 fix(gateway): drop the redundant whole-block lease; one invariant test per atom; document the watchdog env vars
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.

Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.

Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
2026-09-15 04:19:19 -07:00
teknium1 bc3df8a4d5 fix: NT-namespace guard fires before every sibling resolve (checkpoint, ACP bridge, @file:)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.

The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
2026-09-15 04:17:31 -07:00
Teknium faf71eb4c1 Inspired by Claude Code: file tools reject Windows NT-namespace paths (NTLM leak hardening)
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:

- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
  raw-string check (never resolves — resolving IS the leak trigger).
  Wired as the first check in get_read_block_error() and the write
  denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
  _check_sensitive_path (covers write_file_tool + patch_tool), before
  the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
  local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
  forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
2026-09-15 04:17:31 -07:00
teknium1 ce04a6f189 docs: match room picture and member-session wording to the shipped UI
The roster row no longer fans member faces (GroupRow renders the room
image or a single group glyph since 5afa487e9), so describe the picture
as replacing the default glyph. Member sessions are titled by roomId,
not by display name (group-turns.ts), so drop the `Group: <name>`
literal and say "room session" everywhere the doc mentioned it.
2026-09-15 04:15:28 -07:00
Teknium 4f6e8c7345 docs: Bot Mode group chats — editable name and room picture
Documents PR #89371: room picture at creation (upload/generate),
Group settings dialog (rename + picture) after creation, rename
keeps history/sessions and rejects collisions.
2026-09-15 04:15:28 -07:00
teknium1 ac63d0eea5 fix(agent): execution guidance and browser hints drop the web_search stripper; tests assert the invariant
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (3733e4aff5) matched
nothing and were dead; the function now returns the neutral text for
every toolset and its phantom-tool test asserts "no web tool named"
instead of the removed sentence. model_tools ports the PR's hint layer
into main's _DYNAMIC_SCHEMA_REWRITERS table (browser_navigate +
browser_cdp) rather than a second pass after it.

Tests: the two browser_cdp registry tests were re-added by the PR but
main pruned them in 39975613b13b4; replaced with one schema-neutrality
invariant. Exact-wording assertions ("lightweight retrieval tool",
"appropriate permitted retrieval/search tool") were change detectors and
are dropped. tools-reference.md row updated to the new schema text.
2026-09-15 04:13:13 -07:00
teknium1 7020a0081a docs(secrets): hermes update and its probes never resolve external sources 2026-09-15 04:11:11 -07:00
Teknium 3aeb1736c5 Port from RooCodeInc/Roomote#1478: per-route webhook event coalescing
Rapid distinct events on the same logical entity (five pushes to one PR,
a burst of ticket edits, a flapping alert) each carry a fresh delivery ID,
so the idempotency cache cannot suppress them and every event wakes a
separate agent run. Roomote solved this for PR review tasks by keeping one
durable review task per PR and superseding stale heads; this ports the
same debounce-and-supersede pattern to the generic webhook adapter.

New opt-in per-route 'coalesce' block: events group by a payload-derived
key, each new event replaces the pending one and re-arms a quiet-window
timer (window_seconds, default 30), bounded by max_wait_seconds (default
300) past the group's first event so a steady stream cannot starve
dispatch. The settled group dispatches ONE agent run on the latest
event's payload/prompt/delivery templates, with a note when earlier
events were superseded. Pending groups flush on disconnect. Startup
validation rejects missing keys, non-positive windows, and the
deliver_only+coalesce combination.

Rebase onto the decomposed webhook adapter (salvage, #92066):
- Coalescing lives in a topical sibling, gateway/platforms/webhook_coalesce.py
  (WebhookCoalescer + validate_coalesce_config); webhook.py only wires it in
  (__init__, _validate_route, disconnect, _handle_webhook) and splits main's
  _dispatch_agent_run into the HTTP-response wrapper plus _spawn_agent_run,
  shared by the immediate and coalesced paths.
- Review finding (unresolved key fields collapsed unrelated entities into one
  group): an event whose rendered key still contains a {placeholder} is now
  dispatched immediately instead of coalesced; documented.
- Review finding (flush-on-disconnect vs process exit): disconnect() awaits
  the handoff of flushed runs; the docs claim is scoped to adapter disconnect
  and states that a hard kill loses the current window's buffer.
- cron_job + coalesce is rejected like deliver_only + coalesce (cron_job
  landed on main after the PR branched).
- Tests trimmed from 17 to 4 (validation parametrized; debounce/supersede/
  independent groups/duplicate-first in one behavioural test; max-wait +
  unresolved-key; flush-on-disconnect).
2026-09-15 04:08:12 -07:00
teknium1 05fb879609 fix(dashboard): share the fleet's sudo posture; trim to two invariant tests
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
2026-09-15 04:08:00 -07:00
teknium1 8731bb91f5 refactor(gateway): move the supervised-restart handback into gateway_supervised_restart.py
The handback logic was appended to the hermes_cli/gateway.py facade; it now lives in a
topical sibling. Supervisor detection also reads the gateway's own declaration (control
socket `identify` -> supervisor: "external", then the live argv marker, then the argv the
gateway stamped into gateway_state.json) so a gateway whose command line cannot be read via
psutil is still handed back rather than SIGTERMed and shadowed by a foreground run.

Tests trimmed to the two invariants (handback with fresh-PID success; either failure branch
never takes ownership) plus the plain-manual control. Docs: `hermes gateway restart` is now
part of the --external-supervisor contract.
2026-09-15 04:07:13 -07:00
teknium1 7eae49c499 fix(skills): auteur follows the modern section order and routes assets through image_generate
SKILL.md is restructured to authoring standard 5 (When to Use,
Prerequisites, How to Run, Quick Reference, Procedure, Pitfalls,
Verification) — headings only, upstream body text kept. The asset
references still carried the upstream per-CLI routing tables and command
lines for other agent products; those are replaced with the native
`image_generate` route (product names are allowed only in LICENSE and
credit lines). `WebSearch` residue in recon docs/refscout becomes
`web_search`.

source.mjs wrote its ~2.6MB Google Fonts metadata cache to
$TEMP||$TMPDIR||'.', which is the project cwd on most Linux shells; it
now uses os.tmpdir() and Pitfalls documents the location. Network-access
note now mentions that moodboard.mjs also downloads the image URLs the
search hosts return. Docs page regenerated for this skill only.
2026-09-15 04:03:43 -07:00
teknium1 54255f1e9e fix(skills): auteur — proper upstream attribution, windows platform, trimmed tests
Review follow-ups on the port (all verified against the upstream snapshot,
which I re-downloaded and diffed: every scripts/*.mjs and template is the
upstream file byte-for-byte after CRLF→LF, except one `reference/` →
`references/` path fix; the reference docs differ only by Hermes adaptation
notes and the same path fix).

- LICENSE: header naming the upstream repo, pinned commit 9bca227d… and the
  copyright holder above the verbatim MIT text.
- SKILL.md frontmatter: `author` credits the upstream human first, Hermes
  Agent second (skills/AGENTS.md rule 4); `metadata.hermes.upstream` pin in
  the same shape mono-color/pr-lens use; `category: creative`; H1
  `# Auteur Skill` with a linked provenance blockquote.
- `platforms` gains `windows`: the declared prerequisites (Node 18+,
  Playwright, optional ffmpeg) all run on Windows and no script uses a
  POSIX-only primitive (audited: no /tmp, spawn/exec of shells, fcntl, etc).
- Routing examples translated from Russian to English (marked as translated
  from upstream) so an English-language skill doesn't carry stray artefacts.
- tests/skills/test_auteur_skill.py: keep the two port-specific invariants
  (path annotations, de-Claude residue). Dropped the exact-count tree
  snapshot (change detector), the `~/.hermes/hermes-agent` host-dependent
  related_skills fallback, and the frontmatter/description checks that
  tests/skills/test_authoring_standards.py already enforces repo-wide.
- Regenerated the docs page with website/scripts/generate-skill-docs.py
  (scoped to this skill).
2026-09-15 04:03:43 -07:00
Teknium 251ab05000 feat(skills): add auteur optional skill — cinematic web design with executable anti-slop gates
Port of agiwhitelist/auteur (MIT, ~1k stars), snapshot 9bca227d. Three
registers (build / direct / system) on one taste core: commit-sheet-first
art direction, asset generation via image_generate + local CLIs, and
node-based quality gates (slopscan anti-slop linter, motionqa frame-drop
check, systemscan cross-route drift) run through playwright.

- optional-skills/creative/auteur: SKILL.md (de-Clauded, Hermes tool
  framing), LICENSE (upstream MIT), 11 references, 8 verbatim upstream
  .mjs scripts (all pass node --check; slopscan smoke-run verified),
  6 templates. README gallery assets not vendored (size cap).
- tests/skills/test_auteur_skill.py: frontmatter, path-annotation
  invariant, de-Claude residue, related_skills resolution.
- Docs: catalog row, sidebar entry, generated skill page (scoped regen).
2026-09-15 04:03:43 -07:00
teknium1 ea5757a8fe fix(nix): container mode does not linger the host service user; document the cron/linger dependency
The scope-dispatching cron worker runs inside the container there, so a
host user manager would start for nothing (as PR #110641 by @liuhao1024
also gated it). nix-setup.md gains the note operators need when they
declare the user themselves.
2026-09-15 04:03:25 -07:00
kshitijk4poor 288fdc1a4c fix(auth): accept a non-production Portal's own inference host when the operator selected it
A token minted by a non-production Portal is meant to be spent at that environment's own
inference gateway, and the Portal's refresh response names that host. The allowlist applied
to Portal-returned inference URLs was production-only, so the value was refused as "not in
allowlist" and healed to the production host — a token the production Portal never issued,
sent to the production gateway, which 401s it. Every hosted non-production instance hit
this on every gateway turn once #108319 made the deploy-wide NOUS_INFERENCE_BASE_URL
invisible inside a routed profile scope (by design, #65941).

The widening is keyed on the operator's trusted HERMES_PORTAL_BASE_URL override, never on
the stored portal_base_url: when that override names a Portal outside the production
allowlist, any https host under the Nous domain is accepted; otherwise the strict production
set stands. So a poisoned auth.json cannot widen the set, a production-Portal session that
finds a foreign inference URL in its state is still refused and healed, and the bearer can
only ever go to a Nous-owned host. No environment is named in code. Because the override is
read through the profile scope (previous commit), each multiplexed profile decides for
itself.

Validation: 4 invariant tests (accepted only under a non-production override; look-alike
domains, dotless suffix and http still refused; stored portal alone does not widen; the
decision follows the profile scope under multiplex) — the new-behaviour ones red on the
previous commit. Main's existing validation tests are unchanged and green. Live receipt for
the symptom and the fixed chain on a hosted instance: #111589.

Based on #102863 and its rebase onto the decomposed auth_nous.py in #111589, whose
portal-keyed pairing this replaces with the same behaviour and no environment literals.

Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-09-15 16:31:53 +05:30
teknium1 bb745a0e9b fix(goals): failed quality gates re-run every boundary instead of replaying a status fingerprint
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
2026-09-15 03:59:50 -07:00
teknium1 b91ee8c72d docs: remove a stray conflict marker from the optional-skills catalog
47c029927f landed with a leftover ">>>>>>>" line and put the dream-loop
and mono-color rows under autonomous-ai-agents. Move both rows into the
creative table (alphabetical) and drop the marker.
2026-09-15 03:59:20 -07:00
teknium1 ee07fcd4d7 fix(goals): a judge wait_on_pid naming an unobservable pid continues instead of parking
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
2026-09-15 03:59:10 -07:00
Teknium 6fcd011c01 Inspired by ChatGPT Work: keep imported agent setups in sync (hermes import-agent --sync)
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:

- Every successful import registers its source + a content digest of
  everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
  files changed since the last run (digest compare; unchanged = no-op).
  Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
  sync; user-created skills under the import category keep conflict
  semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
  trigger (or leak into) a sync.

Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
2026-09-15 03:58:44 -07:00
teknium1 c164e12bd8 docs(cron): skill-backed jobs receive the skill config block 2026-09-15 03:58:27 -07:00
Teknium e819846b10 Inspired by Amp: relative time bounds (7d/24h/2w) + wrapper forwarding for session_search after/before
Amp's thread feed supports relative time filters (`after:7d`,
`updated_before:7d`) alongside ISO dates. Extend the salvaged
after/before bounds (PR #86067 by @Moodtuner997) the same way:

- `_parse_iso_bound()` now accepts relative durations `Nh`/`Nd`/`Nw`
  (case-insensitive) meaning "now minus N", alongside ISO
  dates/datetimes. Clearer error message names both accepted forms.
- Forward after/before/exclude_session_ids through the public
  `session_search()` wrapper (the PR predates the wrapper/impl split;
  without this the SQL bounds were unreachable from the registry
  handler — same class as the earlier `detail` forwarding fix).
  Appended after `detail` to preserve positional compatibility.
- Tool schema descriptions teach both forms.
- Tests: relative after/before against the discovery shape, unit
  checks for h/d/w math, case-insensitivity, and bad-unit rejection.
- Docs: tools-reference row mentions time bounds + exclude_session_ids.
2026-09-15 03:58:05 -07:00
teknium1 4613f895ab test: give the rewrite-hint fixtures a read baseline; align docs row with schema
The stale-write guard now refuses write_file on an existing file the task
never read in full, so test_write_file_rewrite_hint's overwrite-without-read
fixtures were refused before the hint could be computed. Reading first is
the exact read->whole-file-rewrite pattern the hint exists for.

tools-reference.md's write_file row now mirrors the WRITE_FILE_SCHEMA
description (one-sentence contract + the recovery step) instead of a
longer paraphrase.
2026-09-15 03:57:24 -07:00
Teknium 6569651b87 fix: harden salvage of #65605 — redaction-gated test, sibling test baseline, docs
- test_file_staleness redacted-read case now force-enables redaction
  (matches tests/agent/test_redact.py convention) so it exercises the
  sentinel path in hermetic CI where security.redact_secrets is unset.
- test_write_verification CRLF case establishes a read baseline first
  (the new guard refuses unread existing-file overwrites by design).
- tools-reference.md documents the read-before-overwrite contract.
- contributors/emails mapping for DanSpicyTaco.
2026-09-15 03:57:24 -07:00
teknium1 123db98635 docs(gemini): scope the base-URL normalization claim to the Google host and TTS
The guide said a proxy root like http://localhost:4000/gemini "works the same"
as spelling out /v1beta, but the chat/aux clients only take the native Gemini
adapter when is_native_gemini_base_url() matches the
generativelanguage.googleapis.com host; normalize_gemini_base_url() applies to
the Google host, TTS and the tier probe. Reword the docs to those cases and
tell proxy users to configure an OpenAI-compatible URL. Also note in the
normalize_gemini_base_url docstring that only the last path segment is
inspected and that it does not decide routing.
2026-09-15 03:54:01 -07:00
Teknium f030c03970 Port from cline/cline#13329: normalize host-root Gemini base URLs to /v1beta
A GEMINI_BASE_URL (or tts.gemini.base_url / providers.gemini base_url) set
to a host root — https://generativelanguage.googleapis.com or a proxy root
like http://localhost:4000/gemini — produced native requests to
{base}/models/{model}:generateContent with no API version segment, a
guaranteed 404. Google's own google-genai client treats the base URL as a
host root and appends the version itself, so users reasonably configure it
that way.

normalize_gemini_base_url() appends /v1beta unless the URL already ends
with a version segment (v1, v1beta, v1alpha, ...). Applied at every native
request builder: GeminiNativeClient, probe_gemini_tier, Gemini TTS
(tts_tool.py), and streaming TTS (tts_streaming.py). /openai-suffixed
URLs are untouched (OpenAI-compat path).

Port of cline/cline#13329, which fixed the same bug class after their
ai-sdk migration.
2026-09-15 03:54:01 -07:00
Alan 1657a1ce2d feat(cli): add --format stream-json for structured JSONL output
Adds a --format flag to hermes chat single-query mode. stream-json
emits newline-delimited JSON events (init, text, tool_use, tool_result,
result envelope with token stats + exit code) to stdout for CI
pipelines and external tooling. Session ID stays on stderr.

Salvaged from PR #12278 by @ProDrifterDK onto current main, including
the follow-up commit enforcing the single-query contract (implies
quiet, rejects --tui, emits a final result record with exit code 130
on interrupt).
2026-09-15 03:53:13 -07:00
teknium1 3cad439de7 docs(sessions): describe the timings block in JSONL exports
User-visible export shape changed with no docs hunk. One paragraph in the JSONL section: what
the block holds (ids/roles/counts/durations, text-free), why complete is always false,
available=false when no message carries a timestamp, and that import ignores it.
2026-09-15 03:51:07 -07:00
teknium1 45ab3ad57f fix(delegate_task): return a schema-invalid child's raw text instead of failing the task
When a child's final answer still missed its output_schema after the one
bounded retry, the result entry flipped to status=failed with the error
"Final answer does not satisfy the declared output_schema" — the completion
line printed ✗ and orchestrators read a finished audit as a failure. Five
audits of 413-4103 s were lost this way in the Sep 10-14 retrospective and
the parent had to mine the live transcripts; in four of them the "violation"
was a ```json fence around a valid array, which the candidate extractor
sliced to its first..last object.

Now: status stays completed, `summary` is the child's raw final text,
`schema_valid: false` + `schema_errors` carry the verdict and a `schema_note`
says the text is unvalidated; the sync completion line shows ⚠ with the
reason. The extractor tries the earliest-opening bracket span and keeps the
first that parses (fenced arrays validate). The OUTPUT CONTRACT the child
sees now says "ONLY the JSON value — no prose, no code fence" and what a miss
costs. One bounded retry is unchanged.
2026-09-15 03:45:41 -07:00
teknium1 8a8c3634e8 fix(kanban): scope the delegated-child write fence to the lineage's board root
HERMES_DELEGATED_CHILD_CONTEXT=1 is deliberately carried into every shell/
execute_code subprocess a delegate_task child spawns (the fence must survive
exec so a grandchild `hermes kanban complete` cannot promote itself). But the
readers treated the bare flag as "fence every Kanban DB": kanban_db_connect
opened ANY board ?mode=ro and write_txn refused ANY mutation. A subagent
running a Kanban reproduction against a scratch HERMES_HOME therefore got a
silently read-only board with a misleading "descendants require an
initialized board" error; only one lane in the retrospective ever discovered
why (deleg_15dac332), every earlier kanban repro ran degraded.

The marker's value is now the fenced board ROOT (kanban_home() at spawn) and
readers deny only paths under that root or the dispatcher-pinned
HERMES_KANBAN_DB (kanban_path_is_fenced). In-process children and a legacy
"1" marker still fence everything; an inherited path marker is never
re-derived, so a grandchild that moved HERMES_HOME cannot unfence the real
board. Owner-gate tests (test_kanban_descendant_scope, cron env isolation,
kanban CLI exit status) are unchanged and green.
2026-09-15 03:45:41 -07:00
teknium1 9a49b3c984 fix(execute_code): subagent kernels survive the LRU cap for the child's lifetime
A delegated child's execute_code kernel was keyed correctly
(<owner>::child::<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.

A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
2026-09-15 03:45:41 -07:00
teknium1 081421d838 fix: keep the write-side outcome-uncertain verdict when no server can reconnect
_handle_session_expired_and_retry only reached the at-most-once guard when a
reconnectable server record existed; without one (server torn down, MCP loop
not running) a write-capable call fell through to the generic "MCP call
failed" error, which invites the model to replay a write that may already
have landed. The session-expired classification now runs first and a
write-capable call always gets the outcome_uncertain error; the reconnect is
attempted only when a server can be signalled.

_track_inflight_rpc's teardown RuntimeError said "retry the request on the
rebuilt session" for every op; for a write-capable tools/call it now says the
request may already have been dispatched and must be verified first, so the
wording matches the at-most-once contract the recoverer enforces.

Docs: the readOnlyHint row explains that the same hint gates auto-retry after
a mid-call session expiry, and that unannotated tools on an idle-TTL
Streamable-HTTP server return outcome_uncertain on the first call after idle
instead of being transparently replayed.
2026-09-15 03:45:29 -07:00
teknium1 87ce653d1d feat(relay): migrate legacy HERMES_NEMO_RELAY_ATIF_*/ATOF_* vars into a validated relay-plugins.toml
3fad83df31 (Aug 11) moved Relay exporter config to a plugins.toml
selected by HERMES_NEMO_RELAY_PLUGINS_TOML. A .env still carrying the legacy
exporter vars and no TOML logs ONE warning and initialises no exporters, so
users who followed the earlier docs lost every trace silently (the
maintainer's stopped Aug 20, noticed Sep 14; five multiplexed profiles on
the same box carry the same eight vars today).

- `hermes_cli/relay_plugin_migrate.py`: build the document from the
  `nemo_relay.observability` dataclasses (`ComponentSpec(...).to_dict()`,
  so the `type = "file"` sink discriminator is emitted), validate it by
  activating it through `nemo_relay.plugin.initialize` + `clear_async`,
  write `<home>/relay-plugins.toml` (tomli_w when installed, minimal emitter
  otherwise), set HERMES_NEMO_RELAY_PLUGINS_TOML in that .env, and comment
  the legacy lines out (never delete). Defaults mirror the removed plugin so
  files land where they used to.
- `hermes update` runs it for the default home AND every live named profile
  (each writes its own TOML) as a best-effort post-update step, with a loud
  notice; `hermes migrate relay [--all-profiles] [--no-validate]` runs it on
  demand.
- The runtime WARNING and the `hermes doctor` finding now say "NO traces
  are being exported" and name the exact command and file path.
- Docs: environment-variables.md + built-in-plugins.md carry the migration
  note and a complete plugins.toml example including `type = "file"`.
2026-09-15 03:44:36 -07:00
teknium1 395e4248d0 fix(gateway): surface secondary WhatsApp/Relay skipped under multiplex instead of a silent continue
Under gateway.multiplex_profiles, `_start_one_profile_adapters` skipped
Platform.RELAY / Platform.WHATSAPP for secondaries with a bare `continue`,
and the startup "not being served" WARNING only covered platforms the
PRIMARY skipped. Four secondaries on one live box had WHATSAPP_ENABLED=true
and nothing in the log, status file, or `hermes gateway status` said the
channel was dead.

- `_note_unserved_secondary_platform`: one INFO per (profile, platform)
  naming the reason (shared process-level ingress owned by the default) and
  the remedy (enable it on the default profile, or disable it here), plus a
  `<profile>:<platform>` runtime-status stamp (state=disabled,
  error_code=multiplex_shared_ingress).
- `_start_secondary_profiles` folds those platforms into the loud WARNING
  when NO profile (default included) runs them.
- `hermes gateway status --profile X` prints
  `whatsapp: not served under multiplex (shared ingress owned by default)`
  from that stamp; /api/status excludes `disabled` entries from the
  platforms degraded verdict (informational, not a fault).
- Docs: multi-profile-gateways.md gets the shared-ingress rule.
2026-09-15 03:44:36 -07:00
teknium1 fd303c0137 fix(gateway): 'decline' survives the config.yaml load path; Telegram forwards it; wizard offers it
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.

Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.

`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
2026-09-15 03:44:13 -07:00
Teknium a6e934e0fd feat(gateway): 'decline' unauthorized-DM behavior — one-time polite decline instead of pairing code
Port from qwibitai/nanoclaw#3260: adds a third unauthorized_dm_behavior
option, 'decline'. Instead of replying with a pairing code (pair) or
staying silent (ignore), the gateway sends one short, polite decline to
the unknown sender, then stays silent toward that sender for 24 hours.

- gateway/config.py: accept 'decline' in the normalizer; new
  unauthorized_dm_decline_message for custom decline text (round-trips
  through to_dict/from_dict).
- gateway/pairing.py: persisted decline stamps (_declined.json) on
  PairingStore with has_recent_decline/record_decline; stamps are
  pruned on write and recorded BEFORE delivery so a send failure can't
  become a decline storm (nanoclaw's stamp-first pattern).
- gateway/run.py: decline branch in the unauthorized-sender path;
  groups still always silently ignore.
- docs: security.md + configuration.md updated.

Adapted from TypeScript (NanoClaw's pending_sender_approvals 'decline:'
stamp rows) to Hermes' existing PairingStore JSON persistence; the
owner-FYI half of nanoclaw's flow is intentionally not ported — Hermes
logs the unauthorized attempt, and pairing remains the owner-visible
grant path.

Rebase onto the decomposed gateway (salvage, #88028):
- The unauthorized-sender path moved from gateway/run.py to
  gateway/run_inbound.py::_hm_admit_event; the decline branch is a sibling
  helper _hm_send_unauthorized_decline next to _hm_offer_pairing_code.
- gateway/config.py now validates the enum via _normalize_choice; the
  accepted set is the module constant UNAUTHORIZED_DM_BEHAVIORS (used by
  both from_dict and get_unauthorized_dm_behavior so a per-platform
  `extra.unauthorized_dm_behavior: decline` is honoured too). The default
  decline text lives in config as DEFAULT_UNAUTHORIZED_DM_DECLINE_MESSAGE.
- Tests trimmed from 5 to 2 invariant tests (send-once-then-silent through
  the real inbound path; config round-trip + real PairingStore stamp
  lifecycle with a patched clock instead of rewriting the JSON file).
2026-09-15 03:44:13 -07:00
teknium1 134ef6454d fix(cron): self-removed runs leave no output directory and skip mark_job_run on crash
remove_job() deletes <cron>/output/<job_id>/ together with the record, but the
finishing run then called save_job_output(), re-creating the directory and
writing the final run into it. Every self-removing job leaked an orphan
directory the store no longer knew about, and the docs claim that "only the
job record is gone afterwards" was false. Skip the save when
self_removal_delivery_allowed() is true (the same check that already excuses
the missing record on the delivery and mark paths); delivery composes without
an output_file, as it already does for non-file paths.

The BaseException handler in _run_one_job_body still called mark_job_run on
the missing record after a self-removal crash. Guard it with the same check so
the crash path matches the completion path instead of probing a deleted record.

Docs: state that the record and its output directory are both gone.

Review finding: self-removed run re-creates the rmtree'd output dir (orphan leak); crash path marks a missing record.
2026-09-15 03:42:40 -07:00
teknium1 acb2c45e35 fix(cron): self-removal excuses only a missing record; replacement records stay fail-closed
Follow-up to the salvaged #111044 commits:

- self_removal_delivery_allowed() now also requires that no record currently
  holds the job id. The marker alone said "this run removed its record"; it did
  not say the id is still empty. A replacement record (another owner reclaiming
  the id) must be treated as a stolen claim, not a self-removal.
- Drop the allow_self_removed kwarg on fire_claim_fence: the fence already has
  the job_id and the ContextVar marker, so it can decide on its own; the caller
  no longer threads a flag it computed from the same predicate.
- _FireOwnership.lost(): keep the explicit lost event and the no-owner short
  circuit ahead of the self-removal check so an interrupted run is still
  reported as lost even after it removed its record.
- _finish_completed_run: skip mark_job_run entirely for a self-removed job
  (nothing to mark) instead of calling it and then excusing the False.
- Tests trimmed to two invariants, both A/B'd against origin/main: the
  self-removing run delivers after a post-removal heartbeat tick (RED on main),
  and a self-removal followed by a replacement record is still discarded
  (GREEN on main, guards the new predicate).
- Docs: user-guide cron.md notes that a job may remove itself and still report.
2026-09-15 03:42:40 -07:00
KoNit-K 5c970d9745 fix(kanban): neutral block-loop wording on the CLI, Desktop toast, wake text and docs
The sibling surfaces of the gateway ping rendered the same false claim:
`hermes kanban block` said "needs a human decision", the Desktop toast title
said "needs a decision", the wake status line (locales/*.yaml
gateway.kanban.wake.block_loop_detected) said "needs a decision" and the
docs described the triage route as "for a human decision". A repeated-block
circuit breaker only establishes that orchestration attention is needed.

Surface sweep from PR #111131 (notifier/test hunks dropped in favour of the
typed-kind formatter from PR #111132).
2026-09-15 03:42:00 -07:00
teknium1 88f2844d46 fix(agent): cover the remaining refusal-only surfaces and fold the tests
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
  stream watchdog sees progress on a refusal-only stream instead of timing it
  out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
  parts; without it an aux refusal-only turn parsed to content=None and hit the
  empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
  alongside-content) so there is one test per surface; drop upstream product
  references from docstrings (credit stays in the PR body); pass encoding= to
  the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
  content_filter result, not an empty response to retry.
2026-09-15 03:39:07 -07:00
teknium1 e860b8e4e4 fix(context): compute-host /context and session.context_breakdown carry the per-file manifest; report blocked files
Why: the tui_gateway live formatter (`_format_live_context_output`, used when
the session runs on a compute host) renders its own summary and never got the
"Context files" block, and `session.context_breakdown` had no structured rows,
so Desktop's popover could not show them. The formatter now appends
render_context_file_lines() with the session cwd bound (the RPC thread has no
session context, so the discovery walk would key on the backend's cwd), and
the RPC payload gains a `context_files` list (contract + generated TS/OpenRPC
+ Desktop type). The docs sentence is scoped to the surfaces that render it.

A file whose content _scan_context_content replaces with a BLOCKED marker was
reported "loaded"; the manifest now runs the same scan and reports `blocked`.
The module docstring names the frontmatter-strip / chain-cap approximations
and drops the product-name attribution (credit stays in the PR body).
2026-09-15 03:37:49 -07:00
teknium1 f271ba09b0 fix(context): derive the /context file listing from the builder's own discovery walk
Review follow-up (Enough1122) on the salvaged #91272: the original
list_context_file_sources() hand-mirrored the priority ladder inside
build_context_files_prompt, so the two would drift the moment the builder
gained a context type or changed precedence — misreporting what the prompt
holds is worse than not showing it.

Now prompt_builder exposes one candidate finder per context type
(_CONTEXT_FILE_CANDIDATES → discover_context_files) and BOTH the loaders and
the manifest walk it. The manifest lives in the new sibling
agent/context_file_sources.py (not appended to the facade) and:
- reports empty / unreadable files truthfully instead of "✓ 0 tokens",
- mirrors the install-tree guard ("suppressed") so a Desktop session that
  fell back into the Hermes tree sees why nothing loaded,
- lists every .cursor/rules/*.mdc as loaded, matching the builder which
  concatenates all of them,
- measures truncation on the rendered "## label" section like the builder.

The block now renders on every surface that shows the /context category
table: CLI/TUI (hermes_cli/cli_info_mixin.py) and the messaging gateway
(gateway/slash_commands_status.py). The Desktop popover consumes the raw
session.context_breakdown payload (no text table) and is left as-is.

Tests trimmed to the two invariants: manifest/prompt parity across every
context type at once, and truncated/suppressed follow the builder.
2026-09-15 03:37:49 -07:00
Teknium 5349aa609d Inspired by Copilot CLI: /context now lists each context file with load status and token cost
Copilot CLI 1.0.81-6 shows each user instruction file separately in
/instructions. Hermes loaded AGENTS.md/.hermes.md/CLAUDE.md/.cursorrules/
SOUL.md through a priority ladder but gave the user no visibility into
WHICH files were discovered, which one won, which were shadowed, or how
much context each costs — the /context 'rules' category was one opaque
number.

- agent/prompt_builder.py: list_context_file_sources() — read-only
  manifest mirroring build_context_files_prompt discovery (priority
  ladder, AGENTS.md directory chain with AGENTS.override.md precedence,
  cwd-only CLAUDE.md/.cursorrules, SOUL.md from profile home) with
  per-file chars, est_tokens, and loaded/truncated/shadowed status
- cli.py /context: 'Context files' section rendering the manifest with
  status glyphs and shadowing/truncation notes; zero prompt/cache impact
- docs: reference/slash-commands.md /context row
- tests/agent/test_context_file_sources.py: 11 tests incl. E2E parity
  with build_context_files_prompt shadowing
2026-09-15 03:37:49 -07:00
teknium1 b00ebf6210 fix(skills): live-dashboard credits the human author and drops the product-name intro
Authoring standard 4 requires the human first in `author`; the skill was
drafted with Hermes so the tool was credited instead. The intro cited a
third-party product, which is allowed only in LICENSE/credit lines. Also
adds metadata.hermes.category and lowercases tags to match sibling
skills; docs page regenerated for this skill only.

Tests: the two prose tests asserted sentence literals (change-detectors);
they now assert structure — three Procedure phases, a "Done when" per step,
standard headings, and the tool wiring (cronjob/desktop_preview/[SILENT]).
2026-09-15 03:35:33 -07:00
teknium1 788601358d refactor(skills): live-dashboard becomes an optional skill reconfigured for the Desktop app
Why: the web dashboard is being deprecated in favour of the Electron Desktop
app, and a skill that ships a bespoke cron blueprint should not be bundled by
default. Reconfigure instead of just rebasing:

- Move skills/productivity/live-dashboard -> optional-skills/productivity/
  live-dashboard (install with `hermes skills install
  official/productivity/live-dashboard`); register it like every other
  optional skill: per-skill docs page under user-guide/skills/optional/,
  optional-skills-catalog row, sidebars entry.
- Desktop reality: add a "Show the dashboard" step — when `desktop_preview`
  is in the toolset (Desktop/GUI sessions) render index.html in the in-app
  preview pane after every build/tick and on request; otherwise report the
  absolute file path. Prerequisites section added (Enough1122 review).
- Drop the hard-wired cron/blueprint_catalog.py entry: the curated catalog
  is for bundled skills and would preload a skill that may not be
  installed. Use the skills-pipeline blueprint instead —
  `metadata.hermes.blueprint` on the SKILL.md registers a daily
  all-dashboards sweep as a /suggestions entry at install time (opt-in,
  never auto-scheduled), which is exactly the mechanism main provides for
  optional skills.
- Never hardcode ~/.hermes in prose the agent executes: refer to the Hermes
  home directory's dashboards/<slug>/ and write absolute paths into cron
  prompts (also answers the review's "tick prompt must name the state-file
  path" point).
- Tests follow the skill to optional-skills/, the catalog-blueprint tests
  are replaced by one parse_blueprint/blueprint_to_job_spec invariant and
  one desktop_preview-with-path-fallback invariant.
2026-09-15 03:35:33 -07:00
Teknium 2154458685 Inspired by Energy: one-sentence live dashboards — bundled skill + automation blueprint
Energy (getenergy.com) ships natural-language persistent dashboards:
describe what you want to see in one sentence and the agent builds a
self-updating status page fed by email threads, signed-in websites, and
files. This ports the concept onto Hermes's existing cron + connector
architecture:

- skills/productivity/live-dashboard: setup/tick split skill — pin the
  dashboard contract, verify one live read per source before scheduling,
  keep dashboard.json as source of truth with a self-contained HTML
  projection, stale-read discipline, deliver only on material change.
- cron/blueprint_catalog.py: live-dashboard automation blueprint
  (purpose/sources/time/recurrence/deliver slots) rendering to the
  dashboard form, /blueprint command, and hermes:// deep-link.
- tests/skills/test_live_dashboard_skill.py: skill standards + blueprint
  registration + real fill_blueprint E2E.
- docs: per-skill page, skills catalog row, sidebar entry.
2026-09-15 03:35:33 -07:00
teknium1 0a208b6891 docs(skills): itemize inbox-triage voice calibration and bound its context cost
Why: the calibration step was a single ~120-word paragraph that models follow
less reliably than enumerable rules, and 20-50 full sent messages could
crowd inbox coverage out of context (Enough1122 review). Split sampling /
extraction / record / fallback into bullets, state that truncated excerpts
carry the style facts, add the Sent-folder-naming pitfall (`Sent`, `Sent
Messages`, `[Gmail]/Sent Mail`, localized) so a missing folder name does not
silently trigger the fallback, and compress the provenance parenthetical to
the operative fact. Docs page mirrors the SKILL.md body verbatim.
2026-09-15 03:34:55 -07:00
Teknium 7f2bf68308 Inspired by Energy: evidence-based voice calibration for inbox-triage reply drafts
Energy's cross-app reply agent analyzes ~100 of the user's past replies
before drafting, so drafts land in the user's actual voice instead of
generic-professional AI register. Port the mechanism into the
email-inbox-triage skill's drafting step:

- Step 4 now calibrates on a bounded sample (20-50) of the user's own
  sent replies — greeting/sign-off habits, length, formality, rhythm,
  per-audience differences, how the user pushes back — before drafting,
  with an explicit fallback when Sent is empty or inaccessible.
- New pitfall + verification item pinning the calibration discipline.
- Test locks the evidence-based calibration and fallback language.
2026-09-15 03:34:55 -07:00