Commit Graph

2126 Commits

Author SHA1 Message Date
teknium1 8bc5894a3a docs(skills): ip-as-logo follows the modern section order; add skill tests
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.

Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
2026-09-15 04:20:15 -07:00
teknium1 febea3e28d docs(skills): pin ip-as-logo upstream attribution to author, repo and commit
The ported skill carried the upstream MIT text but the LICENSE file did not
say where it came from, and the frontmatter only pointed at the repo via a
non-standard `homepage:` key. Reviewers asked for proper attribution.

- LICENSE: header naming the upstream repo, the pinned upstream commit
  (b1bf517c54a4…) and the copyright holder (s1dashu) above the verbatim MIT text.
- SKILL.md: `metadata.hermes.upstream: <repo> (pinned b1bf517c)` — the same
  shape mono-color and pr-lens use — and the adaptation-notes blockquote now
  links the upstream repo and commit so the generated docs page links the source.
- Regenerated website/docs/user-guide/skills/optional/creative/creative-ip-as-logo.md
  with website/scripts/generate-skill-docs.py (scoped to this skill).
2026-09-15 04:20:15 -07:00
Teknium 398279fd6a feat(skills): add ip-as-logo optional skill (minimal cute IP mascot marks)
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.

Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
  (square aspect, main-prompt constraints mode — no negative_prompt
  parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
  own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
  proposal-round skip for pre-authorized batches, dimensions-reporting
  rule when the backend returns only a URL, limbless-subject note

Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.

Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).

Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
2026-09-15 04:20:15 -07:00
teknium1 bc3df8a4d5 fix: NT-namespace guard fires before every sibling resolve (checkpoint, ACP bridge, @file:)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.

The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
2026-09-15 04:17:31 -07:00
Teknium faf71eb4c1 Inspired by Claude Code: file tools reject Windows NT-namespace paths (NTLM leak hardening)
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:

- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
  raw-string check (never resolves — resolving IS the leak trigger).
  Wired as the first check in get_read_block_error() and the write
  denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
  _check_sensitive_path (covers write_file_tool + patch_tool), before
  the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
  local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
  forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
2026-09-15 04:17:31 -07:00
teknium1 ce04a6f189 docs: match room picture and member-session wording to the shipped UI
The roster row no longer fans member faces (GroupRow renders the room
image or a single group glyph since 5afa487e9), so describe the picture
as replacing the default glyph. Member sessions are titled by roomId,
not by display name (group-turns.ts), so drop the `Group: <name>`
literal and say "room session" everywhere the doc mentioned it.
2026-09-15 04:15:28 -07:00
Teknium 4f6e8c7345 docs: Bot Mode group chats — editable name and room picture
Documents PR #89371: room picture at creation (upload/generate),
Group settings dialog (rename + picture) after creation, rename
keeps history/sessions and rejects collisions.
2026-09-15 04:15:28 -07:00
Teknium 3aeb1736c5 Port from RooCodeInc/Roomote#1478: per-route webhook event coalescing
Rapid distinct events on the same logical entity (five pushes to one PR,
a burst of ticket edits, a flapping alert) each carry a fresh delivery ID,
so the idempotency cache cannot suppress them and every event wakes a
separate agent run. Roomote solved this for PR review tasks by keeping one
durable review task per PR and superseding stale heads; this ports the
same debounce-and-supersede pattern to the generic webhook adapter.

New opt-in per-route 'coalesce' block: events group by a payload-derived
key, each new event replaces the pending one and re-arms a quiet-window
timer (window_seconds, default 30), bounded by max_wait_seconds (default
300) past the group's first event so a steady stream cannot starve
dispatch. The settled group dispatches ONE agent run on the latest
event's payload/prompt/delivery templates, with a note when earlier
events were superseded. Pending groups flush on disconnect. Startup
validation rejects missing keys, non-positive windows, and the
deliver_only+coalesce combination.

Rebase onto the decomposed webhook adapter (salvage, #92066):
- Coalescing lives in a topical sibling, gateway/platforms/webhook_coalesce.py
  (WebhookCoalescer + validate_coalesce_config); webhook.py only wires it in
  (__init__, _validate_route, disconnect, _handle_webhook) and splits main's
  _dispatch_agent_run into the HTTP-response wrapper plus _spawn_agent_run,
  shared by the immediate and coalesced paths.
- Review finding (unresolved key fields collapsed unrelated entities into one
  group): an event whose rendered key still contains a {placeholder} is now
  dispatched immediately instead of coalesced; documented.
- Review finding (flush-on-disconnect vs process exit): disconnect() awaits
  the handoff of flushed runs; the docs claim is scoped to adapter disconnect
  and states that a hard kill loses the current window's buffer.
- cron_job + coalesce is rejected like deliver_only + coalesce (cron_job
  landed on main after the PR branched).
- Tests trimmed from 17 to 4 (validation parametrized; debounce/supersede/
  independent groups/duplicate-first in one behavioural test; max-wait +
  unresolved-key; flush-on-disconnect).
2026-09-15 04:08:12 -07:00
teknium1 05fb879609 fix(dashboard): share the fleet's sudo posture; trim to two invariant tests
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
2026-09-15 04:08:00 -07:00
teknium1 7eae49c499 fix(skills): auteur follows the modern section order and routes assets through image_generate
SKILL.md is restructured to authoring standard 5 (When to Use,
Prerequisites, How to Run, Quick Reference, Procedure, Pitfalls,
Verification) — headings only, upstream body text kept. The asset
references still carried the upstream per-CLI routing tables and command
lines for other agent products; those are replaced with the native
`image_generate` route (product names are allowed only in LICENSE and
credit lines). `WebSearch` residue in recon docs/refscout becomes
`web_search`.

source.mjs wrote its ~2.6MB Google Fonts metadata cache to
$TEMP||$TMPDIR||'.', which is the project cwd on most Linux shells; it
now uses os.tmpdir() and Pitfalls documents the location. Network-access
note now mentions that moodboard.mjs also downloads the image URLs the
search hosts return. Docs page regenerated for this skill only.
2026-09-15 04:03:43 -07:00
teknium1 54255f1e9e fix(skills): auteur — proper upstream attribution, windows platform, trimmed tests
Review follow-ups on the port (all verified against the upstream snapshot,
which I re-downloaded and diffed: every scripts/*.mjs and template is the
upstream file byte-for-byte after CRLF→LF, except one `reference/` →
`references/` path fix; the reference docs differ only by Hermes adaptation
notes and the same path fix).

- LICENSE: header naming the upstream repo, pinned commit 9bca227d… and the
  copyright holder above the verbatim MIT text.
- SKILL.md frontmatter: `author` credits the upstream human first, Hermes
  Agent second (skills/AGENTS.md rule 4); `metadata.hermes.upstream` pin in
  the same shape mono-color/pr-lens use; `category: creative`; H1
  `# Auteur Skill` with a linked provenance blockquote.
- `platforms` gains `windows`: the declared prerequisites (Node 18+,
  Playwright, optional ffmpeg) all run on Windows and no script uses a
  POSIX-only primitive (audited: no /tmp, spawn/exec of shells, fcntl, etc).
- Routing examples translated from Russian to English (marked as translated
  from upstream) so an English-language skill doesn't carry stray artefacts.
- tests/skills/test_auteur_skill.py: keep the two port-specific invariants
  (path annotations, de-Claude residue). Dropped the exact-count tree
  snapshot (change detector), the `~/.hermes/hermes-agent` host-dependent
  related_skills fallback, and the frontmatter/description checks that
  tests/skills/test_authoring_standards.py already enforces repo-wide.
- Regenerated the docs page with website/scripts/generate-skill-docs.py
  (scoped to this skill).
2026-09-15 04:03:43 -07:00
Teknium 251ab05000 feat(skills): add auteur optional skill — cinematic web design with executable anti-slop gates
Port of agiwhitelist/auteur (MIT, ~1k stars), snapshot 9bca227d. Three
registers (build / direct / system) on one taste core: commit-sheet-first
art direction, asset generation via image_generate + local CLIs, and
node-based quality gates (slopscan anti-slop linter, motionqa frame-drop
check, systemscan cross-route drift) run through playwright.

- optional-skills/creative/auteur: SKILL.md (de-Clauded, Hermes tool
  framing), LICENSE (upstream MIT), 11 references, 8 verbatim upstream
  .mjs scripts (all pass node --check; slopscan smoke-run verified),
  6 templates. README gallery assets not vendored (size cap).
- tests/skills/test_auteur_skill.py: frontmatter, path-annotation
  invariant, de-Claude residue, related_skills resolution.
- Docs: catalog row, sidebar entry, generated skill page (scoped regen).
2026-09-15 04:03:43 -07:00
teknium1 bb745a0e9b fix(goals): failed quality gates re-run every boundary instead of replaying a status fingerprint
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
2026-09-15 03:59:50 -07:00
teknium1 ee07fcd4d7 fix(goals): a judge wait_on_pid naming an unobservable pid continues instead of parking
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
2026-09-15 03:59:10 -07:00
Teknium 6fcd011c01 Inspired by ChatGPT Work: keep imported agent setups in sync (hermes import-agent --sync)
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:

- Every successful import registers its source + a content digest of
  everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
  files changed since the last run (digest compare; unchanged = no-op).
  Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
  sync; user-created skills under the import category keep conflict
  semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
  trigger (or leak into) a sync.

Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
2026-09-15 03:58:44 -07:00
teknium1 c164e12bd8 docs(cron): skill-backed jobs receive the skill config block 2026-09-15 03:58:27 -07:00
teknium1 3cad439de7 docs(sessions): describe the timings block in JSONL exports
User-visible export shape changed with no docs hunk. One paragraph in the JSONL section: what
the block holds (ids/roles/counts/durations, text-free), why complete is always false,
available=false when no message carries a timestamp, and that import ignores it.
2026-09-15 03:51:07 -07:00
teknium1 45ab3ad57f fix(delegate_task): return a schema-invalid child's raw text instead of failing the task
When a child's final answer still missed its output_schema after the one
bounded retry, the result entry flipped to status=failed with the error
"Final answer does not satisfy the declared output_schema" — the completion
line printed ✗ and orchestrators read a finished audit as a failure. Five
audits of 413-4103 s were lost this way in the Sep 10-14 retrospective and
the parent had to mine the live transcripts; in four of them the "violation"
was a ```json fence around a valid array, which the candidate extractor
sliced to its first..last object.

Now: status stays completed, `summary` is the child's raw final text,
`schema_valid: false` + `schema_errors` carry the verdict and a `schema_note`
says the text is unvalidated; the sync completion line shows ⚠ with the
reason. The extractor tries the earliest-opening bracket span and keeps the
first that parses (fenced arrays validate). The OUTPUT CONTRACT the child
sees now says "ONLY the JSON value — no prose, no code fence" and what a miss
costs. One bounded retry is unchanged.
2026-09-15 03:45:41 -07:00
teknium1 8a8c3634e8 fix(kanban): scope the delegated-child write fence to the lineage's board root
HERMES_DELEGATED_CHILD_CONTEXT=1 is deliberately carried into every shell/
execute_code subprocess a delegate_task child spawns (the fence must survive
exec so a grandchild `hermes kanban complete` cannot promote itself). But the
readers treated the bare flag as "fence every Kanban DB": kanban_db_connect
opened ANY board ?mode=ro and write_txn refused ANY mutation. A subagent
running a Kanban reproduction against a scratch HERMES_HOME therefore got a
silently read-only board with a misleading "descendants require an
initialized board" error; only one lane in the retrospective ever discovered
why (deleg_15dac332), every earlier kanban repro ran degraded.

The marker's value is now the fenced board ROOT (kanban_home() at spawn) and
readers deny only paths under that root or the dispatcher-pinned
HERMES_KANBAN_DB (kanban_path_is_fenced). In-process children and a legacy
"1" marker still fence everything; an inherited path marker is never
re-derived, so a grandchild that moved HERMES_HOME cannot unfence the real
board. Owner-gate tests (test_kanban_descendant_scope, cron env isolation,
kanban CLI exit status) are unchanged and green.
2026-09-15 03:45:41 -07:00
teknium1 9a49b3c984 fix(execute_code): subagent kernels survive the LRU cap for the child's lifetime
A delegated child's execute_code kernel was keyed correctly
(<owner>::child::<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.

A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
2026-09-15 03:45:41 -07:00
teknium1 87ce653d1d feat(relay): migrate legacy HERMES_NEMO_RELAY_ATIF_*/ATOF_* vars into a validated relay-plugins.toml
3fad83df31 (Aug 11) moved Relay exporter config to a plugins.toml
selected by HERMES_NEMO_RELAY_PLUGINS_TOML. A .env still carrying the legacy
exporter vars and no TOML logs ONE warning and initialises no exporters, so
users who followed the earlier docs lost every trace silently (the
maintainer's stopped Aug 20, noticed Sep 14; five multiplexed profiles on
the same box carry the same eight vars today).

- `hermes_cli/relay_plugin_migrate.py`: build the document from the
  `nemo_relay.observability` dataclasses (`ComponentSpec(...).to_dict()`,
  so the `type = "file"` sink discriminator is emitted), validate it by
  activating it through `nemo_relay.plugin.initialize` + `clear_async`,
  write `<home>/relay-plugins.toml` (tomli_w when installed, minimal emitter
  otherwise), set HERMES_NEMO_RELAY_PLUGINS_TOML in that .env, and comment
  the legacy lines out (never delete). Defaults mirror the removed plugin so
  files land where they used to.
- `hermes update` runs it for the default home AND every live named profile
  (each writes its own TOML) as a best-effort post-update step, with a loud
  notice; `hermes migrate relay [--all-profiles] [--no-validate]` runs it on
  demand.
- The runtime WARNING and the `hermes doctor` finding now say "NO traces
  are being exported" and name the exact command and file path.
- Docs: environment-variables.md + built-in-plugins.md carry the migration
  note and a complete plugins.toml example including `type = "file"`.
2026-09-15 03:44:36 -07:00
teknium1 395e4248d0 fix(gateway): surface secondary WhatsApp/Relay skipped under multiplex instead of a silent continue
Under gateway.multiplex_profiles, `_start_one_profile_adapters` skipped
Platform.RELAY / Platform.WHATSAPP for secondaries with a bare `continue`,
and the startup "not being served" WARNING only covered platforms the
PRIMARY skipped. Four secondaries on one live box had WHATSAPP_ENABLED=true
and nothing in the log, status file, or `hermes gateway status` said the
channel was dead.

- `_note_unserved_secondary_platform`: one INFO per (profile, platform)
  naming the reason (shared process-level ingress owned by the default) and
  the remedy (enable it on the default profile, or disable it here), plus a
  `<profile>:<platform>` runtime-status stamp (state=disabled,
  error_code=multiplex_shared_ingress).
- `_start_secondary_profiles` folds those platforms into the loud WARNING
  when NO profile (default included) runs them.
- `hermes gateway status --profile X` prints
  `whatsapp: not served under multiplex (shared ingress owned by default)`
  from that stamp; /api/status excludes `disabled` entries from the
  platforms degraded verdict (informational, not a fault).
- Docs: multi-profile-gateways.md gets the shared-ingress rule.
2026-09-15 03:44:36 -07:00
teknium1 fd303c0137 fix(gateway): 'decline' survives the config.yaml load path; Telegram forwards it; wizard offers it
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.

Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.

`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
2026-09-15 03:44:13 -07:00
Teknium a6e934e0fd feat(gateway): 'decline' unauthorized-DM behavior — one-time polite decline instead of pairing code
Port from qwibitai/nanoclaw#3260: adds a third unauthorized_dm_behavior
option, 'decline'. Instead of replying with a pairing code (pair) or
staying silent (ignore), the gateway sends one short, polite decline to
the unknown sender, then stays silent toward that sender for 24 hours.

- gateway/config.py: accept 'decline' in the normalizer; new
  unauthorized_dm_decline_message for custom decline text (round-trips
  through to_dict/from_dict).
- gateway/pairing.py: persisted decline stamps (_declined.json) on
  PairingStore with has_recent_decline/record_decline; stamps are
  pruned on write and recorded BEFORE delivery so a send failure can't
  become a decline storm (nanoclaw's stamp-first pattern).
- gateway/run.py: decline branch in the unauthorized-sender path;
  groups still always silently ignore.
- docs: security.md + configuration.md updated.

Adapted from TypeScript (NanoClaw's pending_sender_approvals 'decline:'
stamp rows) to Hermes' existing PairingStore JSON persistence; the
owner-FYI half of nanoclaw's flow is intentionally not ported — Hermes
logs the unauthorized attempt, and pairing remains the owner-visible
grant path.

Rebase onto the decomposed gateway (salvage, #88028):
- The unauthorized-sender path moved from gateway/run.py to
  gateway/run_inbound.py::_hm_admit_event; the decline branch is a sibling
  helper _hm_send_unauthorized_decline next to _hm_offer_pairing_code.
- gateway/config.py now validates the enum via _normalize_choice; the
  accepted set is the module constant UNAUTHORIZED_DM_BEHAVIORS (used by
  both from_dict and get_unauthorized_dm_behavior so a per-platform
  `extra.unauthorized_dm_behavior: decline` is honoured too). The default
  decline text lives in config as DEFAULT_UNAUTHORIZED_DM_DECLINE_MESSAGE.
- Tests trimmed from 5 to 2 invariant tests (send-once-then-silent through
  the real inbound path; config round-trip + real PairingStore stamp
  lifecycle with a patched clock instead of rewriting the JSON file).
2026-09-15 03:44:13 -07:00
teknium1 134ef6454d fix(cron): self-removed runs leave no output directory and skip mark_job_run on crash
remove_job() deletes <cron>/output/<job_id>/ together with the record, but the
finishing run then called save_job_output(), re-creating the directory and
writing the final run into it. Every self-removing job leaked an orphan
directory the store no longer knew about, and the docs claim that "only the
job record is gone afterwards" was false. Skip the save when
self_removal_delivery_allowed() is true (the same check that already excuses
the missing record on the delivery and mark paths); delivery composes without
an output_file, as it already does for non-file paths.

The BaseException handler in _run_one_job_body still called mark_job_run on
the missing record after a self-removal crash. Guard it with the same check so
the crash path matches the completion path instead of probing a deleted record.

Docs: state that the record and its output directory are both gone.

Review finding: self-removed run re-creates the rmtree'd output dir (orphan leak); crash path marks a missing record.
2026-09-15 03:42:40 -07:00
teknium1 acb2c45e35 fix(cron): self-removal excuses only a missing record; replacement records stay fail-closed
Follow-up to the salvaged #111044 commits:

- self_removal_delivery_allowed() now also requires that no record currently
  holds the job id. The marker alone said "this run removed its record"; it did
  not say the id is still empty. A replacement record (another owner reclaiming
  the id) must be treated as a stolen claim, not a self-removal.
- Drop the allow_self_removed kwarg on fire_claim_fence: the fence already has
  the job_id and the ContextVar marker, so it can decide on its own; the caller
  no longer threads a flag it computed from the same predicate.
- _FireOwnership.lost(): keep the explicit lost event and the no-owner short
  circuit ahead of the self-removal check so an interrupted run is still
  reported as lost even after it removed its record.
- _finish_completed_run: skip mark_job_run entirely for a self-removed job
  (nothing to mark) instead of calling it and then excusing the False.
- Tests trimmed to two invariants, both A/B'd against origin/main: the
  self-removing run delivers after a post-removal heartbeat tick (RED on main),
  and a self-removal followed by a replacement record is still discarded
  (GREEN on main, guards the new predicate).
- Docs: user-guide cron.md notes that a job may remove itself and still report.
2026-09-15 03:42:40 -07:00
KoNit-K 5c970d9745 fix(kanban): neutral block-loop wording on the CLI, Desktop toast, wake text and docs
The sibling surfaces of the gateway ping rendered the same false claim:
`hermes kanban block` said "needs a human decision", the Desktop toast title
said "needs a decision", the wake status line (locales/*.yaml
gateway.kanban.wake.block_loop_detected) said "needs a decision" and the
docs described the triage route as "for a human decision". A repeated-block
circuit breaker only establishes that orchestration attention is needed.

Surface sweep from PR #111131 (notifier/test hunks dropped in favour of the
typed-kind formatter from PR #111132).
2026-09-15 03:42:00 -07:00
teknium1 88f2844d46 fix(agent): cover the remaining refusal-only surfaces and fold the tests
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
  stream watchdog sees progress on a refusal-only stream instead of timing it
  out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
  parts; without it an aux refusal-only turn parsed to content=None and hit the
  empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
  alongside-content) so there is one test per surface; drop upstream product
  references from docstrings (credit stays in the PR body); pass encoding= to
  the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
  content_filter result, not an empty response to retry.
2026-09-15 03:39:07 -07:00
teknium1 b00ebf6210 fix(skills): live-dashboard credits the human author and drops the product-name intro
Authoring standard 4 requires the human first in `author`; the skill was
drafted with Hermes so the tool was credited instead. The intro cited a
third-party product, which is allowed only in LICENSE/credit lines. Also
adds metadata.hermes.category and lowercases tags to match sibling
skills; docs page regenerated for this skill only.

Tests: the two prose tests asserted sentence literals (change-detectors);
they now assert structure — three Procedure phases, a "Done when" per step,
standard headings, and the tool wiring (cronjob/desktop_preview/[SILENT]).
2026-09-15 03:35:33 -07:00
teknium1 788601358d refactor(skills): live-dashboard becomes an optional skill reconfigured for the Desktop app
Why: the web dashboard is being deprecated in favour of the Electron Desktop
app, and a skill that ships a bespoke cron blueprint should not be bundled by
default. Reconfigure instead of just rebasing:

- Move skills/productivity/live-dashboard -> optional-skills/productivity/
  live-dashboard (install with `hermes skills install
  official/productivity/live-dashboard`); register it like every other
  optional skill: per-skill docs page under user-guide/skills/optional/,
  optional-skills-catalog row, sidebars entry.
- Desktop reality: add a "Show the dashboard" step — when `desktop_preview`
  is in the toolset (Desktop/GUI sessions) render index.html in the in-app
  preview pane after every build/tick and on request; otherwise report the
  absolute file path. Prerequisites section added (Enough1122 review).
- Drop the hard-wired cron/blueprint_catalog.py entry: the curated catalog
  is for bundled skills and would preload a skill that may not be
  installed. Use the skills-pipeline blueprint instead —
  `metadata.hermes.blueprint` on the SKILL.md registers a daily
  all-dashboards sweep as a /suggestions entry at install time (opt-in,
  never auto-scheduled), which is exactly the mechanism main provides for
  optional skills.
- Never hardcode ~/.hermes in prose the agent executes: refer to the Hermes
  home directory's dashboards/<slug>/ and write absolute paths into cron
  prompts (also answers the review's "tick prompt must name the state-file
  path" point).
- Tests follow the skill to optional-skills/, the catalog-blueprint tests
  are replaced by one parse_blueprint/blueprint_to_job_spec invariant and
  one desktop_preview-with-path-fallback invariant.
2026-09-15 03:35:33 -07:00
Teknium 2154458685 Inspired by Energy: one-sentence live dashboards — bundled skill + automation blueprint
Energy (getenergy.com) ships natural-language persistent dashboards:
describe what you want to see in one sentence and the agent builds a
self-updating status page fed by email threads, signed-in websites, and
files. This ports the concept onto Hermes's existing cron + connector
architecture:

- skills/productivity/live-dashboard: setup/tick split skill — pin the
  dashboard contract, verify one live read per source before scheduling,
  keep dashboard.json as source of truth with a self-contained HTML
  projection, stale-read discipline, deliver only on material change.
- cron/blueprint_catalog.py: live-dashboard automation blueprint
  (purpose/sources/time/recurrence/deliver slots) rendering to the
  dashboard form, /blueprint command, and hermes:// deep-link.
- tests/skills/test_live_dashboard_skill.py: skill standards + blueprint
  registration + real fill_blueprint E2E.
- docs: per-skill page, skills catalog row, sidebar entry.
2026-09-15 03:35:33 -07:00
teknium1 0a208b6891 docs(skills): itemize inbox-triage voice calibration and bound its context cost
Why: the calibration step was a single ~120-word paragraph that models follow
less reliably than enumerable rules, and 20-50 full sent messages could
crowd inbox coverage out of context (Enough1122 review). Split sampling /
extraction / record / fallback into bullets, state that truncated excerpts
carry the style facts, add the Sent-folder-naming pitfall (`Sent`, `Sent
Messages`, `[Gmail]/Sent Mail`, localized) so a missing folder name does not
silently trigger the fallback, and compress the provenance parenthetical to
the operative fact. Docs page mirrors the SKILL.md body verbatim.
2026-09-15 03:34:55 -07:00
Teknium 7f2bf68308 Inspired by Energy: evidence-based voice calibration for inbox-triage reply drafts
Energy's cross-app reply agent analyzes ~100 of the user's past replies
before drafting, so drafts land in the user's actual voice instead of
generic-professional AI register. Port the mechanism into the
email-inbox-triage skill's drafting step:

- Step 4 now calibrates on a bounded sample (20-50) of the user's own
  sent replies — greeting/sign-off habits, length, formality, rhythm,
  per-audience differences, how the user pushes back — before drafting,
  with an explicit fallback when Sent is empty or inaccessible.
- New pitfall + verification item pinning the calibration discipline.
- Test locks the evidence-based calibration and fallback language.
2026-09-15 03:34:55 -07:00
Carl Taylor 075a256597 fix(skills): auto-load resolves once per agent and dedupes against -s
The system prompt must stay byte-stable for the life of a conversation:
`_auto_load_skills_result` is seeded in `_SESSION_STATE` and filled on
the FIRST prompt build only (HERMES_IGNORE_RULES captured then too), so
model switches, compression and static-prefix restoration reuse the
exact rendered bytes rather than re-reading config or skill files.

CLI: auto_load renders in the existing background `--skills` preload
thread (real session id for ${HERMES_SESSION_ID}), `-s` names dedupe
against the auto-loaded canonical names via
`build_preloaded_skills_prompt(excluded_loaded_names=)`, the activated
skills line shows auto_load first, and the lazily built agent is seeded
with the pre-resolved bytes. `--ignore-rules` skips auto-load with the
rest of the auto-injected context.

Re-implementation of #74060 by @ctaylor86 against current main.
2026-09-15 03:34:20 -07:00
teknium1 923960fa4b fix(kanban): dispatch_profiles is config-only and fail-closed; trim tests
Follow-up to the salvaged #111004 commit, aligning it with the shape agreed
on #110995:

- Drop the HERMES_KANBAN_DISPATCH_PROFILES env bridge: non-secret behaviour
  lives in config.yaml only, like every other kanban.* key.
- Read the key via load_config_readonly() with the same fail-open config
  read as the sibling kanban.* readers (configured_max_in_progress).
- Fail closed when the key is set: the "none" sentinel is gone (an empty
  list already claims nothing), and an assignee that is not a valid
  profile id is never claimable instead of being lower-cased into the
  allowlist.
- Trim the regression file to two invariants (allowlist without `default`
  buckets the card as nonspawnable AND turns has_spawnable_ready off;
  unset key keeps upstream behaviour). Both drive the real dispatch tick
  against a real config.yaml + kanban.db; the first is red on origin/main.
- Docs: move the "Shared boards across homes" section out of the
  gateway-dispatcher paragraph, state that `default` collides by
  construction, add the config-reference row.
2026-09-15 03:33:53 -07:00
Kevin Rajan 2d46af3fa2 fix(kanban): per-home dispatch claim allowlist for shared boards
On a shared kanban.db, every home's profile_exists('default') is
unconditionally True, so any home's dispatcher could claim cards assigned
to 'default'. Wrap the _profile_exists_fn() predicate with an optional
per-home allowlist: kanban.dispatch_profiles (config.yaml, list or
comma-separated string) with a HERMES_KANBAN_DISPATCH_PROFILES env bridge.
Unset preserves upstream behavior; 'none' claims nothing. Foreign
assignees land in the existing skipped_nonspawnable bucket, and the gate
applies to the ready spawn path, _has_spawnable, and review dispatch alike.
Also documents the multi-home default collision in the kanban user guide.

Fixes #110995
2026-09-15 03:33:53 -07:00
teknium1 931387dd42 test(cron): two invariant ESTOP tests for the fire webhook and misfire backstop; docs
Trim the salvaged suite from four tests to the two invariants that were red on
main: (1) with the sentinel engaged POST /api/cron/fire answers 503 +
Retry-After 60 and never calls claim_fire, and the same job is admitted (202,
claimed, fired) once the sentinel is removed; (2) fire_overdue_jobs dispatches
nothing and leaves next_run_at untouched while engaged, and the first sweep
after resume catches the job up through claim_fire. The webhook test lives
beside the other cron-fire webhook tests (test_cron_fire_webhook.py) and uses
their real spy provider instead of a MagicMock resolver; the "verifier crashes
-> 401" case was already covered there.

Docs: cron.md gains a "Pausing everything: hermes pause" section stating that
all three automated doors honour pause, that in-flight runs are never killed,
and that manual runs are an operator override; the CLI reference table lists
hermes pause / hermes resume.
2026-09-15 03:33:13 -07:00
kshitijk4poor 4d55ca9165 docs(supermemory): list the session-switch retry point on the website too
The website key-features bullet (EN + zh-Hans) omitted /reset from the
retry points while the README lists it; also drops a local import in the
test file shadowed by the module-level one.
2026-09-15 11:55:10 +05:30
kshitijk4poor bc1d4776db docs(supermemory): purge stale session-ingest / /v4/conversations references
The per-turn capture rewrite removed the raw urllib /v4/conversations
ingest, but stale references survived outside the diff hunks:

- README "Behavior" still carried the "written once via the conversations
  endpoint" paragraph contradicted by the new bullets right above it.
- website memory-providers (EN + zh-Hans) still listed full-session ingest,
  session-end /v4/conversations ingest, and ingest in the base-url and
  api_timeout rows — the PR had updated one line per file but missed the
  rest of the section.

Now every surface describes per-turn documents.add capture with retry.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu 03627dbf95 fix(supermemory): write turns via documents API instead of session-end conversations ingest
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).

Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.

Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.

Metadata stays type/session_id/timestamp plus the existing sm_source.
2026-09-15 11:55:10 +05:30
salch-cred 101861c7f7 fix(discord): dispatch-side liveness dimension detects an ACKing-but-deaf gateway socket (#109521)
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.

This adds the dispatch-side signal that was requested instead:

- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
  every parsed DISPATCH frame with no debug gate (verified live against
  the real received_message path: 4/4 frames fired with
  enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
  incident report's field-proven operator bound); 0 opts out of this
  dimension ONLY — the #109782 review failure put the knob in
  _start_liveness_probe's all-or-nothing guard, killing the whole
  watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
  stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out

Fixes #109521

(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
2026-09-15 10:48:58 +05:30
teknium1 2de17e5d40 feat(plugin-catalog): default shelf is Desktop, catch-all is General; categorise today's six entries
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
2026-09-14 21:00:29 -07:00
teknium1 55dbd7f6e1 feat(plugin-catalog): shelve the catalog by category (Memory, Desktop, Platforms, …)
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.

/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
2026-09-14 21:00:29 -07:00
teknium1 e80642df60 fix(bot-mode): keep group follow-ups ordered and late answers visible
Serialize room drives through their actual member completion, freeze input
watermarks by retained entry identity, and share the same completion path
with handoff continuations. Stop discards queued work without releasing an
active owner early; rename follows the existing room binding.

Observe stranded replies for the hard-cap duration plus grace after the
foreground wait, and retain unresolved failures in collapsed Activity.
Never automatically retry an ambiguous failed submit within the same drive.

Adapted from the queue and boundary approach in #92041 by @enwaiax and
harvest-budget approach in #107193 by @Finn763; #106502 by @wadib identified
failed-submit watermark consumption. The implementation retains current
numeric watermark storage, room lifecycle bindings and serial round limits.

Related: #92003, #105247, #100026
2026-09-14 17:35:04 -07:00
Teknium 9975fd56be Merge pull request #111320 from NousResearch/docs/plugin-catalog-no-self-update
Plugin catalog: no self-updaters instead of a 2-week pin age, enforced in CI
2026-09-14 17:34:43 -07:00
teknium1 42602bd12a docs(computer-use): skill matches the driver's current vocabulary
The skill described a SOM overlay burned into the screenshot, a driver-side
`capture` tool, and manual symlinking of the cua-driver skill pack; users
who read the driver's own docs then called raw MCP tools (`capture`,
bare `element_index`) and hit "no reviewed risk classification" and
`snapshot_id_required`. State plainly that `computer_use(action=...)` is a
wrapper vocabulary the driver never sees, that `element=N` is translated to
the snapshot token, what a `stale` refusal means, how text-only models get
vision (auxiliary.vision routing / mode=ax), and the Windows WindowsApps
doctor failure. `cua-driver skills install` links into ~/.hermes/skills now.
2026-09-14 17:34:31 -07:00
teknium1 8f7853188f fix(cron): keep deferred delivery exceptions from aborting ticks
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.

Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
2026-09-14 17:29:32 -07:00
teknium1 002ee41cfc fix(cron): keep unowned Bot Chat delivery on its resolved home
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.

Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-14 17:29:32 -07:00
teknium1 3b0fe0cc2b fix(cron): keep deferred Bot Chat delivery bound to admission
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.

Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
2026-09-14 17:29:32 -07:00
teknium1 5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00