Commit Graph

2699 Commits

Author SHA1 Message Date
teknium1 a919c414ce docs(kanban): worker sessions are named after their card, no title model call 2026-09-15 06:08:20 -07:00
teknium1 e89605b4e6 docs(cron): record the due-only occurrence identity in the missed-occurrence contract
Also map the two salvaged contributor emails (kleros109, fangliquanflq).
2026-09-15 06:07:32 -07:00
teknium1 738d63a34d fix(kanban): automatic stale-claim reclaims count toward the failure breaker
A claim that expired without a worker ever spawning (worker_pid NULL) was
reclaimed and immediately re-claimed on every dispatcher tick, with
consecutive_failures stuck at 0 — nothing could trip the breaker. Route the
reclaim through _record_task_failure (own txn after the reclaim commit, same
shape as enforce_max_runtime) instead of the salvaged raw counter increment,
so per-task max_retries / kanban.failure_limit and the gave_up event apply
and last_failure_error carries the stale lock. reclaim_task (operator path)
still resets the counter; the live-worker extend path never reaches it.

Trims the salvaged tests to one invariant that walks the breaker to its trip.
2026-09-15 06:06:32 -07:00
teknium1 7b44106174 fix(gateway): agent-cache pressure valve measures the cgroup's anon charge, not just the gateway's own RSS
The budget in _sweep_agent_cache_under_pressure comes from the gateway's own
cgroup memory.high/memory.max, but read_anon_rss_mb() read /proc/self/status
RssAnon: only the main process. Every child in the unit (execute_code kernels,
terminal commands) is charged against the same limit, so a kernel at 4.7 GiB
pushed the unit to MemoryHigh while the valve saw <1.6 GiB and never fired;
systemd's stop then SIGKILLed the gateway mid-flush (the #80764 signature).

Under a capped cgroup v2, read own memory.stat `anon` (the same scope as the
budget); uncapped, or when the file is unreadable, keep the self reading.
_finite_limit is factored out of _cgroup_limit_bytes so the cap check and
the budget agree on what "unlimited" means.

Fixes #110549
2026-09-15 05:40:15 -07:00
teknium1 238720e928 test(journey): one end-to-end invariant for foreground-created skills; docs + list-unmanaged label
Replace the predicate unit test with an invariant on the real builder: a skill recorded
by a foreground create is in build_learning_graph() with zero uses, and an unmarked
never-used local skill is not. `hermes curator list-unmanaged` prints the actual
marker (created_by:learn) instead of hard-coding created_by:null. Docs: curator.md and
memory.md describe the learn marker and what the journey shows.
2026-09-15 05:38:31 -07:00
fangliquan ebb6dc6e70 fix(stt): prefer cached local whisper models 2026-09-15 05:36:37 -07:00
teknium1 6f87575009 test(approval): keep two invariants for unattended allowlist keys; document rule keys
Drop the session-only negative (session grants were never consulted on the unattended
path, so it pins pre-existing behaviour rather than the fix). Keep the permanent-key
positive and the Tirith-not-bypassed negative. Docs: command_allowlist rule keys are
honored in cron/-q/unattended sessions.
2026-09-15 05:33:41 -07:00
teknium1 3147ccb1ad docs(session-storage): guard raises, not warns; database-location section follows get_hermes_home(); zh-Hans mirror
Corrections on top of #110617: hermes_state._ensure_test_isolation raises RuntimeError (the
draft said 'warning'); the older Database Location section still told readers the default
path is ~/.hermes/state.db, which the new section says never to hard-code. Mirror the new
section into the zh-Hans page.
2026-09-15 05:30:32 -07:00
KoNit-K 71dade7e34 docs(website): document profile state isolation 2026-09-15 05:30:32 -07:00
teknium1 a9a8a3fa2e fix(config): hermes config get masks credentials on every path; --raw opts out
`hermes config get providers`, `config get providers.<p>.api_key`, `config get
<PROVIDER>_API_KEY` (the .env-routed branch) and `config get mcp_servers.<s>.env.X_API_KEY`
all printed the full credential. The agent runs this command from sessions whose transcripts
persist and get forwarded (a Gemini key surfaced in a Discord DM log), so `print` output is a
leak path the logging redactor never sees.

`get_config_value` now applies the structural masker used by `config show` before printing,
honouring `security.redact_secrets` (default on), with a `--raw` flag for operators/scripts
that need the real value. `_is_secret_config_key` extends the exact-name set with the same
`*_API_KEY / *_TOKEN / *_SECRET / *_PASSWORD` suffixes `_is_env_config_key` already routes to
.env, so env-map leaves under `mcp_servers.*.env` mask too, and the `config set` echo uses the
same predicate.

Slim redo of #84153 by @webtecnica (same direction: mask in get_config_value; dropped the
redact_url_query_params re-export and the separate redaction-enabled reader in favour of
agent.redact._redact_enabled, which already resolves the profile-scoped policy).

Fixes #110758
Fixes #84106
2026-09-15 05:08:55 -07:00
teknium1 2034126e0d fix: warn about existing WAL on cross-VM fs on the vulnerable-SQLite path too
apply_wal_with_fallback returns via _apply_delete_for_wal_reset_bug on
WAL-reset-vulnerable SQLite builds (Debian 12 / Ubuntu 22.04 system Pythons,
the reporter's pinned image) before the #110848 existing-WAL cross-VM check,
so exactly the deployment class the fix targets still got zero startup signal
while doctor flagged it. Share one helper between both early-return paths and
key the once-per-process dedupe on the DB path instead of the label so a
gateway serving several profiles hears about each database. The docker doc
remedy now uses the image's python3 (it ships libsqlite3 but no sqlite3 shell).

Review finding: vulnerable-SQLite early return skipped the cross-VM ERROR; dedupe per label; sqlite3 shell absent from image.
2026-09-15 05:02:30 -07:00
teknium1 60d94fd8f4 fix(state): warn when an existing WAL state.db sits on a virtiofs/9p mount; doctor + docs
d8dcdfd620 (v2026.9.14) made apply_wal_with_fallback refuse to ENABLE WAL
on a fresh database whose directory is on a cross-VM bind mount (virtiofs/9p),
but a database that was already WAL on such a mount kept WAL — correctly, we
never live-downgrade under other openers — and emitted nothing. The operator
in #110848 ran exactly that shape (Podman applehv virtiofs bind mount) and got
"database disk image is malformed" within a minute with no prior signal.

- apply_wal_with_fallback: in the on-disk-WAL branch, log a once-per-process
  ERROR ('cross_vm_fs_existing_wal') when the DB file is on a cross-VM
  filesystem, naming the two remedies (offline PRAGMA journal_mode=DELETE
  after stopping every process + database.journal_mode: delete, or move the
  database to a native/named volume). The fresh-DB refusal is unchanged.
- hermes doctor: _report_database_journal_modes flags a WAL database on a
  cross-VM filesystem with check_warn and the same remedy (ranked above the
  WAL-reset exposure warning; the exposure bookkeeping is kept).
- docs: docker.md gains "Filesystem requirements for state.db in containers";
  configuration.md's database comment no longer implies operators must set
  delete by hand on virtiofs.

Detection stays /proc/self/mountinfo-based (runs inside the Linux container
on macOS/Windows hosts). locking_mode=EXCLUSIVE is deliberately not adopted:
gateway, cron and workers open state.db concurrently.

Fixes #110848
2026-09-15 05:02:30 -07:00
teknium1 1f08821bd0 fix: derive suggest globs from the raw command, class-key when redaction hits them
Redaction ran before glob derivation, so a mined `GITHUB_TOKEN=ghp_… git push`
became the pattern `GITHUB_TOKEN=*** git *` and `--apply` persisted it to
config.yaml. In the permanent-allowlist matcher `***` is three fnmatch
wildcards, so that entry pre-approved any `GITHUB_TOKEN=… git …` command
(`sudo git push --force`, `chmod -R 777 /etc git x`) ahead of the dangerous
command detector. Globs now come from the raw normalized command; when the
redacted form would yield a different glob the command is proposed under its
dangerous-class key instead. Redaction stays for example rendering only.

Review finding: masked `***` inside a persisted glob widened the allowlist to arbitrary `KEY=… git …` commands.
2026-09-15 04:57:29 -07:00
Teknium c60fab351f fix(cli): hoist redactor import, trim suggest redaction tests, document masking
- Move the `agent.redact` import out of the per-record loop in build_proposals.
- Keep the two display-boundary invariant tests (rendered `e.g.` line, --json
  examples), drop the unit-level duplicate that asserted the same masking.
- website/docs: state that mined examples are masked at display time only;
  state.db itself is unchanged.
- contributors/emails: map kokhlo's commit email.
2026-09-15 04:57:29 -07:00
teknium1 57c4e1d963 fix(mcp): served profiles keep their own MCP connections without the multiplex flag
A dashboard/desktop backend (and the per-profile cron ticker) serves sessions of
several profiles through the HERMES_HOME contextvar override while
gateway.multiplex_profiles stays off. _mcp_registry_scope() keyed every MCP
connection by the bare server name in that mode, so the first profile to
discover `zernio` owned the only connection and every later served profile —
including one whose config carries a different Authorization header — called
the server through it and got the other account's data back (#111151).

The registry scope now follows the served home: a routed profile (an override
naming a home other than the process home) gets the same per-profile overlay
the multiplexer uses, so a same-named server with other credentials is a
separate connection, discovery for profile B is a connect candidate instead of
"already connected", and status/tool views stay per profile. Single-profile
processes (no override) keep bare keys, byte-identical to before.

Fixes #111151

Credit: #111158 by @KoNit-K located the inert flag on hermes_cli surfaces; its
fix (activating fail-closed multiplex secret scoping from config.yaml on the
dashboard) is not taken — the connection-key seam, not the secret-scope mode,
is what leaks the connection, and flipping the process-wide mode from the
dashboard would change credential resolution for every code path in it.
2026-09-15 04:56:40 -07:00
teknium1 0d1a3e704a docs(vault): manager items fill on every saved website origin 2026-09-15 04:56:01 -07:00
teknium1 7bb52c0b74 feat: profile clone can opt into staying synced with its source (--sync-imports)
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.

`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
2026-09-15 04:44:09 -07:00
teknium1 4da4cc5c37 fix(gateway): SIGUSR2 dumps thread stacks without killing the gateway
faulthandler.register(SIGUSR2, ..., chain=True) writes the stack dump and
then re-raises the signal to its previous handler. SIGUSR2's default
disposition is "terminate", so the diagnostic hook added for #70344 kills the
very process an operator is trying to inspect (rc = -12), which is what the
#110437 reporter hit while introspecting a long-lived gateway. Nothing else
in the gateway installs a SIGUSR2 handler, so there is nothing to chain to.

Test: a child interpreter runs the real _start_install_faulthandler, receives
SIGUSR2, and must still be alive with a dump in gateway_faulthandler.log.
Red on origin/main (rc=-12), green with chain=False. Docs: the stack-dump
signal is now documented next to the event-loop watchdog.
2026-09-15 04:42:20 -07:00
teknium1 71cac9426d fix(gateway): retire native clarify cards on timeout, reset and prose cancel
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:

- TurnRunner._clarify_callback_sync: when the bounded wait returns a
  sentinel (timeout, /new or run-end clear_session), schedule the retire
  with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
  the prose is routed as a follow-up (#111019). Lookup is on the adapter
  class so MagicMock doubles cannot fabricate the method; no platform ==
  SLACK special-case.

The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.

Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
2026-09-15 04:40:04 -07:00
teknium1 08bb2273e4 fix: make the allow_all_users env bridge own what it writes and re-derive it on restart
The bridge wrote GATEWAY_ALLOW_ALL_USERS into os.environ only when unset and
never cleared it. In-process restart paths (gateway restart watcher, dashboard
profile actions) copy os.environ into the child, so a config.yaml grant became
a sticky env var: flipping allow_all_users to false and restarting left the
gateway OPEN. The bridge now tracks its own write (module flag), overwrites or
clears it on reload, exports only a truthy grant (presence-based readers such
as the Telegram intake prefilter treated "false" as configured auth), and the
two restart env builders drop the bridge-owned value so the child re-derives
the posture from its own config.yaml. Under multiplex_profiles the DEFAULT
profile's events are authorized inside its secret scope, where gate readers
never fall to os.environ; the bridged grant is now seeded into that profile's
scope mapping only (secondaries never inherit it).

Review finding: bridged GATEWAY_ALLOW_ALL_USERS survives restart and overrides a flipped config.yaml; inert for the default profile under multiplex; "false" exported as configured auth.
2026-09-15 04:37:41 -07:00
teknium1 5083d5f78e fix(gateway): honour allow_all_users from config.yaml by bridging it to GATEWAY_ALLOW_ALL_USERS
`gateway.allow_all_users: true` (and the top-level spelling) in config.yaml
was a silent no-op: `_TOPLEVEL_BRIDGE` never forwarded it, GatewayConfig has
no field, and every allow-all reader (authz mixin default-deny branch,
startup access check, own-policy adapters, Discord/Matrix/Email plugin
gates) consults the GATEWAY_ALLOW_ALL_USERS env var only.

Bridge the YAML key into that env var in `bridge_core_env_settings`, the
one seam every reader already shares, instead of threading a new config
attribute through ten readers:
- first-writer-wins: an explicit env var beats YAML (matches every other
  {PLATFORM}_* gate);
- skipped inside a multiplexed secondary profile's scope (#80099 class):
  the secondary's config.yaml must not become the default profile's policy,
  and the isolation test now asserts GATEWAY_ALLOW_ALL_USERS stays unset;
- a startup warning names config.yaml as the grant source, because the key
  was inert until now and a forgotten `true` flips the posture to open.

Tests: both spellings authorize a stranger through `_is_user_authorized`;
`false`, absent key, and env=false over YAML=true all stay denied.
Docs: security guide, env-var reference, gateway internals.

Fixes #110690
2026-09-15 04:37:41 -07:00
teknium1 71d229f5cd fix(cron): the silence instruction names [SILENT] as an untranslatable ASCII token; docs + trim tests
Follow-up to the salvaged #110940 commit:

- `cron/scheduler_prompt.py::_CRON_HINT` tells the model the sentinel is a literal
  ASCII control token that must never be translated or rephrased — the prompt-side
  half of the fix, so a lane answering in any other language is steered to the
  canonical token instead of relying on the filter knowing that language.
- Docs: the supported-token list ('Intentional Silence Tokens') gains the zh-Hans
  forms in the English page and the zh-Hans mirror gets the section it lacked.
- Tests trimmed to two invariants (translated forms match in every shape the English
  ones do; prose that mentions the word is still delivered), proven red on origin/main.
2026-09-15 04:30:31 -07:00
teknium1 4db4650a01 fix: exempt the hidden canonical Bot Chat from the idle auto-archive sweep
The archived-holder carve-out in _set_session_title lets the next Bot open
claim the "Bot Chat" title from an archived canonical row. That is correct
for a deliberate sidebar archive, but archive_stale_sessions (the opt-in
sessions.auto_archive sweep) could also archive an idle hidden Bot Chat with
end_reason NULL, which unarchive_recoverable_session refuses; the next Bot
click then stripped the old row's title, retiring the bot's whole history
on an idle timer with no way back.

Skip the hidden canonical Bot Chat in the sweep SELECT using the same
hidden + exact-title predicate set_session_pinned already uses to protect
it, so only an explicit archive retires a Bot Chat. Reword the bot-mode doc
so it no longer promises the retired (still hidden) chat is reachable from
the archive view.

Review finding: auto_archive sweep could irreversibly retire the canonical Bot Chat; docs over-promised archive-view reachability.
2026-09-15 04:27:32 -07:00
teknium1 a4fad54aa7 docs(bot-mode): archiving a Bot Chat retires it; the next open becomes the new canonical chat
Companion to the archived-title release in hermes_state_titles.py (#110871).
2026-09-15 04:27:32 -07:00
kshitijk4poor 569b4242a3 fix(auth): a Portal-returned inference host is accepted only when the operator named it
#111809 accepted any *.nousresearch.com https host once the operator's Portal override
pointed at a non-production Portal. That let a network-provenance value — the Portal's
refresh response — pick any Nous-owned host as the bearer recipient, including hosts that
are not inference gateways. Owning the DNS suffix is not the same as being an authorized
recipient, and the validator's threat model (an injected refresh response) is exactly the
case a suffix rule fails to bound.

The recipient is now the operator's own NOUS_INFERENCE_BASE_URL: a non-production host
returned by the Portal is accepted exactly when it equals that override's host, otherwise
the strict production set stands. The Portal override grants nothing by itself. What the
operator gains over plain use of the override is that the Portal's value is then persisted
and used for the pricing scope and proxy instead of being healed to production, and the
per-turn "refusing inference URL host" warning stops. No environment is named in code.

Raised on #111809 review. Tests: recipient match accepted, unrelated Nous host refused, no
override refused, Portal override alone grants nothing, the match follows the profile scope
under multiplexing; the widening cases are red on main. Docs row for NOUS_INFERENCE_BASE_URL.

Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-09-15 16:55:43 +05:30
wang2 fa12d7556c fix(telegram): allow opt-in CJK rich messages 2026-09-15 04:24:53 -07:00
teknium1 8bc5894a3a docs(skills): ip-as-logo follows the modern section order; add skill tests
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.

Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
2026-09-15 04:20:15 -07:00
teknium1 febea3e28d docs(skills): pin ip-as-logo upstream attribution to author, repo and commit
The ported skill carried the upstream MIT text but the LICENSE file did not
say where it came from, and the frontmatter only pointed at the repo via a
non-standard `homepage:` key. Reviewers asked for proper attribution.

- LICENSE: header naming the upstream repo, the pinned upstream commit
  (b1bf517c54a4…) and the copyright holder (s1dashu) above the verbatim MIT text.
- SKILL.md: `metadata.hermes.upstream: <repo> (pinned b1bf517c)` — the same
  shape mono-color and pr-lens use — and the adaptation-notes blockquote now
  links the upstream repo and commit so the generated docs page links the source.
- Regenerated website/docs/user-guide/skills/optional/creative/creative-ip-as-logo.md
  with website/scripts/generate-skill-docs.py (scoped to this skill).
2026-09-15 04:20:15 -07:00
Teknium 398279fd6a feat(skills): add ip-as-logo optional skill (minimal cute IP mascot marks)
Ports s1dashu/ip-as-logo-skill (MIT, 3.2k stars in 48h, snapshot of
commit b1bf517c) into optional-skills/creative/. Generates extremely
simplified, cute IP mascot characters readable at 32x32 — 3-color
discipline, corner-emergence composition, complexity budget, and a
copy-paste prompt skeleton.

Hermes adaptations (blockquote header + inline edits, upstream body
otherwise intact):
- image path routed through the built-in image_generate tool
  (square aspect, main-prompt constraints mode — no negative_prompt
  parameter exists)
- subagent parallelization mapped to delegate_task, optional
- delivery per platform file conventions; no auto-QA (per upstream's
  own one-pass-draw rules)
- live-test friction fixes folded in: reduced-batch labeling branch,
  proposal-round skip for pre-authorized batches, dimensions-reporting
  rule when the backend returns only a URL, limbless-subject note

Validated via a cold subagent run (2 candidates for a real brief):
both generations succeeded first-draw, verdict SHIP; its three
friction findings are addressed in this commit.

Docs: catalog row + sidebar line + generated skill page (scoped to
this skill only; regen drift for unrelated pages reverted).

Credit: s1dashu (https://github.com/s1dashu/ip-as-logo-skill)
2026-09-15 04:20:15 -07:00
teknium1 b531023622 fix(gateway): drop the redundant whole-block lease; one invariant test per atom; document the watchdog env vars
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.

Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.

Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
2026-09-15 04:19:19 -07:00
teknium1 bc3df8a4d5 fix: NT-namespace guard fires before every sibling resolve (checkpoint, ACP bridge, @file:)
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.

The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
2026-09-15 04:17:31 -07:00
Teknium faf71eb4c1 Inspired by Claude Code: file tools reject Windows NT-namespace paths (NTLM leak hardening)
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:

- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
  raw-string check (never resolves — resolving IS the leak trigger).
  Wired as the first check in get_read_block_error() and the write
  denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
  _check_sensitive_path (covers write_file_tool + patch_tool), before
  the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
  local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
  forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
2026-09-15 04:17:31 -07:00
teknium1 ce04a6f189 docs: match room picture and member-session wording to the shipped UI
The roster row no longer fans member faces (GroupRow renders the room
image or a single group glyph since 5afa487e9), so describe the picture
as replacing the default glyph. Member sessions are titled by roomId,
not by display name (group-turns.ts), so drop the `Group: <name>`
literal and say "room session" everywhere the doc mentioned it.
2026-09-15 04:15:28 -07:00
Teknium 4f6e8c7345 docs: Bot Mode group chats — editable name and room picture
Documents PR #89371: room picture at creation (upload/generate),
Group settings dialog (rename + picture) after creation, rename
keeps history/sessions and rejects collisions.
2026-09-15 04:15:28 -07:00
teknium1 ac63d0eea5 fix(agent): execution guidance and browser hints drop the web_search stripper; tests assert the invariant
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (3733e4aff5) matched
nothing and were dead; the function now returns the neutral text for
every toolset and its phantom-tool test asserts "no web tool named"
instead of the removed sentence. model_tools ports the PR's hint layer
into main's _DYNAMIC_SCHEMA_REWRITERS table (browser_navigate +
browser_cdp) rather than a second pass after it.

Tests: the two browser_cdp registry tests were re-added by the PR but
main pruned them in 39975613b13b4; replaced with one schema-neutrality
invariant. Exact-wording assertions ("lightweight retrieval tool",
"appropriate permitted retrieval/search tool") were change detectors and
are dropped. tools-reference.md row updated to the new schema text.
2026-09-15 04:13:13 -07:00
teknium1 7020a0081a docs(secrets): hermes update and its probes never resolve external sources 2026-09-15 04:11:11 -07:00
Teknium 3aeb1736c5 Port from RooCodeInc/Roomote#1478: per-route webhook event coalescing
Rapid distinct events on the same logical entity (five pushes to one PR,
a burst of ticket edits, a flapping alert) each carry a fresh delivery ID,
so the idempotency cache cannot suppress them and every event wakes a
separate agent run. Roomote solved this for PR review tasks by keeping one
durable review task per PR and superseding stale heads; this ports the
same debounce-and-supersede pattern to the generic webhook adapter.

New opt-in per-route 'coalesce' block: events group by a payload-derived
key, each new event replaces the pending one and re-arms a quiet-window
timer (window_seconds, default 30), bounded by max_wait_seconds (default
300) past the group's first event so a steady stream cannot starve
dispatch. The settled group dispatches ONE agent run on the latest
event's payload/prompt/delivery templates, with a note when earlier
events were superseded. Pending groups flush on disconnect. Startup
validation rejects missing keys, non-positive windows, and the
deliver_only+coalesce combination.

Rebase onto the decomposed webhook adapter (salvage, #92066):
- Coalescing lives in a topical sibling, gateway/platforms/webhook_coalesce.py
  (WebhookCoalescer + validate_coalesce_config); webhook.py only wires it in
  (__init__, _validate_route, disconnect, _handle_webhook) and splits main's
  _dispatch_agent_run into the HTTP-response wrapper plus _spawn_agent_run,
  shared by the immediate and coalesced paths.
- Review finding (unresolved key fields collapsed unrelated entities into one
  group): an event whose rendered key still contains a {placeholder} is now
  dispatched immediately instead of coalesced; documented.
- Review finding (flush-on-disconnect vs process exit): disconnect() awaits
  the handoff of flushed runs; the docs claim is scoped to adapter disconnect
  and states that a hard kill loses the current window's buffer.
- cron_job + coalesce is rejected like deliver_only + coalesce (cron_job
  landed on main after the PR branched).
- Tests trimmed from 17 to 4 (validation parametrized; debounce/supersede/
  independent groups/duplicate-first in one behavioural test; max-wait +
  unresolved-key; flush-on-disconnect).
2026-09-15 04:08:12 -07:00
teknium1 05fb879609 fix(dashboard): share the fleet's sudo posture; trim to two invariant tests
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
2026-09-15 04:08:00 -07:00
teknium1 8731bb91f5 refactor(gateway): move the supervised-restart handback into gateway_supervised_restart.py
The handback logic was appended to the hermes_cli/gateway.py facade; it now lives in a
topical sibling. Supervisor detection also reads the gateway's own declaration (control
socket `identify` -> supervisor: "external", then the live argv marker, then the argv the
gateway stamped into gateway_state.json) so a gateway whose command line cannot be read via
psutil is still handed back rather than SIGTERMed and shadowed by a foreground run.

Tests trimmed to the two invariants (handback with fresh-PID success; either failure branch
never takes ownership) plus the plain-manual control. Docs: `hermes gateway restart` is now
part of the --external-supervisor contract.
2026-09-15 04:07:13 -07:00
teknium1 7eae49c499 fix(skills): auteur follows the modern section order and routes assets through image_generate
SKILL.md is restructured to authoring standard 5 (When to Use,
Prerequisites, How to Run, Quick Reference, Procedure, Pitfalls,
Verification) — headings only, upstream body text kept. The asset
references still carried the upstream per-CLI routing tables and command
lines for other agent products; those are replaced with the native
`image_generate` route (product names are allowed only in LICENSE and
credit lines). `WebSearch` residue in recon docs/refscout becomes
`web_search`.

source.mjs wrote its ~2.6MB Google Fonts metadata cache to
$TEMP||$TMPDIR||'.', which is the project cwd on most Linux shells; it
now uses os.tmpdir() and Pitfalls documents the location. Network-access
note now mentions that moodboard.mjs also downloads the image URLs the
search hosts return. Docs page regenerated for this skill only.
2026-09-15 04:03:43 -07:00
teknium1 54255f1e9e fix(skills): auteur — proper upstream attribution, windows platform, trimmed tests
Review follow-ups on the port (all verified against the upstream snapshot,
which I re-downloaded and diffed: every scripts/*.mjs and template is the
upstream file byte-for-byte after CRLF→LF, except one `reference/` →
`references/` path fix; the reference docs differ only by Hermes adaptation
notes and the same path fix).

- LICENSE: header naming the upstream repo, pinned commit 9bca227d… and the
  copyright holder above the verbatim MIT text.
- SKILL.md frontmatter: `author` credits the upstream human first, Hermes
  Agent second (skills/AGENTS.md rule 4); `metadata.hermes.upstream` pin in
  the same shape mono-color/pr-lens use; `category: creative`; H1
  `# Auteur Skill` with a linked provenance blockquote.
- `platforms` gains `windows`: the declared prerequisites (Node 18+,
  Playwright, optional ffmpeg) all run on Windows and no script uses a
  POSIX-only primitive (audited: no /tmp, spawn/exec of shells, fcntl, etc).
- Routing examples translated from Russian to English (marked as translated
  from upstream) so an English-language skill doesn't carry stray artefacts.
- tests/skills/test_auteur_skill.py: keep the two port-specific invariants
  (path annotations, de-Claude residue). Dropped the exact-count tree
  snapshot (change detector), the `~/.hermes/hermes-agent` host-dependent
  related_skills fallback, and the frontmatter/description checks that
  tests/skills/test_authoring_standards.py already enforces repo-wide.
- Regenerated the docs page with website/scripts/generate-skill-docs.py
  (scoped to this skill).
2026-09-15 04:03:43 -07:00
Teknium 251ab05000 feat(skills): add auteur optional skill — cinematic web design with executable anti-slop gates
Port of agiwhitelist/auteur (MIT, ~1k stars), snapshot 9bca227d. Three
registers (build / direct / system) on one taste core: commit-sheet-first
art direction, asset generation via image_generate + local CLIs, and
node-based quality gates (slopscan anti-slop linter, motionqa frame-drop
check, systemscan cross-route drift) run through playwright.

- optional-skills/creative/auteur: SKILL.md (de-Clauded, Hermes tool
  framing), LICENSE (upstream MIT), 11 references, 8 verbatim upstream
  .mjs scripts (all pass node --check; slopscan smoke-run verified),
  6 templates. README gallery assets not vendored (size cap).
- tests/skills/test_auteur_skill.py: frontmatter, path-annotation
  invariant, de-Claude residue, related_skills resolution.
- Docs: catalog row, sidebar entry, generated skill page (scoped regen).
2026-09-15 04:03:43 -07:00
teknium1 ea5757a8fe fix(nix): container mode does not linger the host service user; document the cron/linger dependency
The scope-dispatching cron worker runs inside the container there, so a
host user manager would start for nothing (as PR #110641 by @liuhao1024
also gated it). nix-setup.md gains the note operators need when they
declare the user themselves.
2026-09-15 04:03:25 -07:00
kshitijk4poor 288fdc1a4c fix(auth): accept a non-production Portal's own inference host when the operator selected it
A token minted by a non-production Portal is meant to be spent at that environment's own
inference gateway, and the Portal's refresh response names that host. The allowlist applied
to Portal-returned inference URLs was production-only, so the value was refused as "not in
allowlist" and healed to the production host — a token the production Portal never issued,
sent to the production gateway, which 401s it. Every hosted non-production instance hit
this on every gateway turn once #108319 made the deploy-wide NOUS_INFERENCE_BASE_URL
invisible inside a routed profile scope (by design, #65941).

The widening is keyed on the operator's trusted HERMES_PORTAL_BASE_URL override, never on
the stored portal_base_url: when that override names a Portal outside the production
allowlist, any https host under the Nous domain is accepted; otherwise the strict production
set stands. So a poisoned auth.json cannot widen the set, a production-Portal session that
finds a foreign inference URL in its state is still refused and healed, and the bearer can
only ever go to a Nous-owned host. No environment is named in code. Because the override is
read through the profile scope (previous commit), each multiplexed profile decides for
itself.

Validation: 4 invariant tests (accepted only under a non-production override; look-alike
domains, dotless suffix and http still refused; stored portal alone does not widen; the
decision follows the profile scope under multiplex) — the new-behaviour ones red on the
previous commit. Main's existing validation tests are unchanged and green. Live receipt for
the symptom and the fixed chain on a hosted instance: #111589.

Based on #102863 and its rebase onto the decomposed auth_nous.py in #111589, whose
portal-keyed pairing this replaces with the same behaviour and no environment literals.

Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-09-15 16:31:53 +05:30
teknium1 bb745a0e9b fix(goals): failed quality gates re-run every boundary instead of replaying a status fingerprint
`_check_gates()` skipped a failed gate whenever sha256(git HEAD + `git status
--porcelain`) matched the last failure. Porcelain sees neither the contents
of an untracked or already-modified file nor inputs outside the repo, so a
repaired input replayed the stale failure and burned retries until the goal
auto-paused (#110649). The gate now runs on every eligible boundary; the
retry cap still bounds a genuinely stuck red suite. `workspace_fingerprint`
and `GoalGate.last_failed_fingerprint` are removed with their only consumer
(old persisted state ignores the extra key on load). Based on the analysis
in #110649 (JsonDaRula69) and PR #110658 (KoNit-K), whose `git diff HEAD`
hash still misses untracked contents and adds a full diff per boundary.
2026-09-15 03:59:50 -07:00
teknium1 b91ee8c72d docs: remove a stray conflict marker from the optional-skills catalog
47c029927f landed with a leftover ">>>>>>>" line and put the dream-loop
and mono-color rows under autonomous-ai-agents. Move both rows into the
creative table (alphabetical) and drop the marker.
2026-09-15 03:59:20 -07:00
teknium1 ee07fcd4d7 fix(goals): a judge wait_on_pid naming an unobservable pid continues instead of parking
`wait_on()` now refuses a dead/remote pid (salvaged from #110829); the
judge path cannot raise there — `_apply_wait_directive` calls it inside
`evaluate_after_turn`, so a ValueError would surface as a turn failure.
Check liveness before the call on that path and fall through to the
normal continue decision: the barrier would otherwise lift ~5 s later,
the judge would see the same remote pid and re-park every turn.
2026-09-15 03:59:10 -07:00
Teknium 6fcd011c01 Inspired by ChatGPT Work: keep imported agent setups in sync (hermes import-agent --sync)
ChatGPT Work's desktop import (Settings > Import, Aug 11 2026 release)
keeps setup imported from Claude Code / Cursor automatically up to date.
This ports the idea to `hermes import-agent`:

- Every successful import registers its source + a content digest of
  everything the importer read in HERMES_HOME/import-sync.json.
- `hermes import-agent --sync` re-imports every registered source whose
  files changed since the last run (digest compare; unchanged = no-op).
  Prompt-free and cron-friendly; `--sync --dry-run` previews.
- Skills previously imported by import-agent are refreshed in place on
  sync; user-created skills under the import category keep conflict
  semantics and are never clobbered.
- Credential files never affect the digest, so token refreshes cannot
  trigger (or leak into) a sync.

Tests: 13 new tests in tests/hermes_cli/test_agent_import.py (61 total
passing), including a sabotage-verified in-place-refresh test; E2E run
against a temp HERMES_HOME exercised register -> no-op sync -> changed
sync through the real command path.
2026-09-15 03:58:44 -07:00
teknium1 c164e12bd8 docs(cron): skill-backed jobs receive the skill config block 2026-09-15 03:58:27 -07:00
Teknium e819846b10 Inspired by Amp: relative time bounds (7d/24h/2w) + wrapper forwarding for session_search after/before
Amp's thread feed supports relative time filters (`after:7d`,
`updated_before:7d`) alongside ISO dates. Extend the salvaged
after/before bounds (PR #86067 by @Moodtuner997) the same way:

- `_parse_iso_bound()` now accepts relative durations `Nh`/`Nd`/`Nw`
  (case-insensitive) meaning "now minus N", alongside ISO
  dates/datetimes. Clearer error message names both accepted forms.
- Forward after/before/exclude_session_ids through the public
  `session_search()` wrapper (the PR predates the wrapper/impl split;
  without this the SQL bounds were unreachable from the registry
  handler — same class as the earlier `detail` forwarding fix).
  Appended after `detail` to preserve positional compatibility.
- Tool schema descriptions teach both forms.
- Tests: relative after/before against the discovery shape, unit
  checks for h/d/w math, case-insensitivity, and bad-unit rejection.
- Docs: tools-reference row mentions time bounds + exclude_session_ids.
2026-09-15 03:58:05 -07:00