Commit Graph

16476 Commits

Author SHA1 Message Date
kshitijk4poor 54041e98cf test(discord): bind event-silence liveness to its config surface and probe gate
The event-silence tests duplicated `_make_adapter`/`_connect` from
test_discord_liveness.py and carried a bespoke handler-wait loop. Reuse
the sibling helpers instead: `_make_adapter` grows an optional
`max_event_silence` that is only written into `extra` when given, so the
sibling's own tests keep the adapter default and stay unchanged.

Two invariants now bind the fix:

- deaf socket: a `None` stamp (no DISPATCH yet) reads healthy across
  several probe intervals, and once armed, transport-green + event
  silence trips `event_silence` through `_liveness_loop`.
- `websocket_event_max_silence_seconds: 0` disables only that
  dimension: with a stale stamp and a stale heartbeat ACK the probe must
  still run and trip on `ack_stale`. This goes red if the knob is moved
  into `_start_liveness_probe`'s all-or-nothing guard (#109782).

The YAML->extra passthrough test also asserts the new key, so dropping
its `_YAML_WEBSOCKET_LIVENESS_KEYS` entry fails.
2026-09-15 10:48:58 +05:30
kshitijk4poor fff69b7dfc test(discord): keep the two binding event-silence invariants
Trim the new event-silence file from 8 tests to the 2 that bind the fix:
the deaf-socket e2e through `_liveness_loop` (transport green, no
DISPATCH → probe trips with the `event_silence` reason) and the
None-stamp window reading healthy (no false trip on quiet reconnects).
The removed cases re-checked knob parsing, defaults and the YAML seed
loop already covered by the sibling liveness knobs' tests, or restated
the two kept invariants from a different angle.
2026-09-15 10:48:58 +05:30
salch-cred 101861c7f7 fix(discord): dispatch-side liveness dimension detects an ACKing-but-deaf gateway socket (#109521)
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.

This adds the dispatch-side signal that was requested instead:

- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
  every parsed DISPATCH frame with no debug gate (verified live against
  the real received_message path: 4/4 frames fired with
  enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
  incident report's field-proven operator bound); 0 opts out of this
  dimension ONLY — the #109782 review failure put the knob in
  _start_liveness_probe's all-or-nothing guard, killing the whole
  watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
  stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out

Fixes #109521

(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
2026-09-15 10:48:58 +05:30
kshitijk4poor 73f808e47f fix(plugins): only mark a timed-out hook worker abandoned while it still holds its token
The timeout branch of _run_hook_callback_bounded unconditionally added
gate_key to _hook_abandoned. A worker that finishes between done.wait()
returning False and the caller taking the lock has already popped its
token via _release_token, so nothing would ever clear that entry: the
callback stayed blocked for every later call id until reload with no
thread behind it. Guard the insert on the worker still being registered.

The new test makes the race deterministic by swapping the module's
threading.Event for one whose wait() lets the worker finish and then
reports a timeout, and asserts a fresh call id still runs.

Also pass tool_call_id inline from terminal_tool_result instead of the
conditional dict plumbing: an empty id is already treated as "no
identity" by _hook_call_identity and unknown fields are withheld from
narrow-signature callbacks (same shape as _fire_approval_hook). Update
the stale "(hook_name, id(cb))" comment above _hook_running_callbacks.
2026-09-15 10:48:49 +05:30
kshitijk4poor f49c1e3bd7 test(plugins): bind the identity gate, not timeout suppression, in the dedupe tests
The same-call negative control fired two sequential calls with a 0.1 s
timeout, so the first call timed out and the second was dropped by the
60 s suppression window; the identity gate itself was never exercised.
Mirror the positive test instead: a 5 s timeout, two threads with the
same tool_call_id while the callback is held on an Event.

Add one test for the abandoned-worker gate: a hung callback followed by
a call with a fresh tool_call_id, with suppression zeroed, must start
exactly one worker. Reverting the gate change makes it fail.
2026-09-15 10:48:49 +05:30
deadczarvc 4121aa295a fix(plugins): gate hook callbacks by call identity, not by tool name alone
Concurrent invocations of the same tool in one session collapsed into a single
busy key (hook_name, id(cb)): the second invocation was reported as 'still
running' and dropped. For pre_tool_call a drop is a fail-closed block, so the
gate silenced itself on an ordinary, healthy callback.

Measured on a busy profile: 3574 skip lines and 0 timeout lines in one hour —
every skip was the 'while still running' branch, i.e. pure key collision, not
slowness.

The gate now keys on the call identity that is already in the payload
(tool_call_id, else turn_id, else none — the last case behaves exactly as
before). Suppression stays keyed coarsely on (hook_name, id(cb)): a hung
callback is a fact about the callback, so its back-off must not be diluted
per call.

Refs #98382. Independent of #107894 (that one releases the slot on timeout;
this one stops healthy concurrency from colliding).

(cherry picked from commit 53b3dacd008418fcdf5fa6dfcadde575a35a776e)
2026-09-15 10:48:49 +05:30
kshitijk4poor 7feaf03883 fix(state): treat ESRCH like ENOENT in deleted-WAL fd identity check
`_fd_is_truly_unlinked` stats `/proc/<pid>/fd/<n>` after the scan has
already read the readlink target. Between those two steps the descriptor
can be closed (ENOENT, handled by the previous commit) or the whole
process can exit (ESRCH). Both mean the descriptor can no longer keep a
retired WAL/SHM generation alive, so neither is evidence of a live holder
and neither should make the guard refuse to open the database.

Match the sibling scan in hermes_state_holders, which already skips both
errnos, by branching on `exc.errno in (ENOENT, ESRCH)`; any other OSError
still fails closed. The existing closed-descriptor test is parametrized
over both errnos, injecting the failure at `os.stat` so the ESRCH path is
exercised on every platform.
2026-09-15 10:48:34 +05:30
KoNit-K 106bf99a0e fix(state): tolerate closed WAL scan descriptors
(cherry picked from commit 299f91bb0c235cd002510034f77c2e627b3985e6)
2026-09-15 10:48:34 +05:30
teknium1 ab0d4735a7 fix(skills-guard): socat only flags a reverse shell when an address spec follows
`\bsocat\b` under IGNORECASE matched "SOCAT", the Surface Ocean CO2 Atlas,
in every oceanography skill of a 2,110-file research bundle (17 critical
findings in one file), burying the bundle's real issues under noise. A real
socat relay always names an address type (TCP:/UDP:/OPENSSL:/EXEC:/SYSTEM:/
PTY:/UNIX-…:), so the pattern now requires one on the same line. `nc -l` /
`ncat -l` are unchanged. Scanner version bumped to v5 so cached verdicts
re-scan.
2026-09-14 21:55:24 -07:00
teknium1 3e2e2c50eb feat(plugin-catalog): rank entries by GitHub stars, probed at most once a day
Catalog entries sort official → stars desc → name, both in browse shelves and
filtered grids, with a ★ pill on each card linking to the repo's stargazers.

Rate-limit discipline is the design constraint: the docs site deploys many
times a day and shares one GitHub App API budget with every other workflow
(tonight's merge train got rate-limited on unrelated uploads). So
website/scripts/fetch-plugin-stars.py first fetches the live site's own
plugin-stars.json (a CDN GET, not the API); if that cache is under 24h old it
is reused verbatim and GitHub is never called. Only a stale cache triggers one
GET /repos/{owner}/{repo} per unique catalog repo, and a 403/429 mid-run keeps
the previous counts instead of zeroing them. extract-plugins.py merges the
cache into plugins.json (`stars`) and plugins-meta.json (`starsFetchedAt`), and
the page footnote says when the ranking was last refreshed.
2026-09-14 21:24:33 -07:00
teknium1 2de17e5d40 feat(plugin-catalog): default shelf is Desktop, catch-all is General; categorise today's six entries
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
2026-09-14 21:00:29 -07:00
teknium1 55dbd7f6e1 feat(plugin-catalog): shelve the catalog by category (Memory, Desktop, Platforms, …)
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.

/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
2026-09-14 21:00:29 -07:00
teknium1 f5a457ad5b fix(tools): one-shot linger waits for a completion that is mid-publish
`ProcessRegistry._move_to_finished` pops the session out of `_running`, then
saves the receipt, releases handles and writes the checkpoint, and only THEN
enqueues the completion and sets `_completion_event`. A quiet one-shot parent
whose turn ends inside that window called `wait_for_pending_completions`,
found nothing in `_running`, drained an empty queue and exited without the
follow-up turn. That is the CI flake in
tests/tools/test_completed_process_results.py::test_headless_terminal_result_survives_cli_exit
(`follow_ups == []`), which also hit unrelated branches.

Consider `_finished` sessions whose event is not yet set as pending too.

Repro: a 6s sleep before the enqueue plus a 2s delay before the parent's first
wait fails the E2E 2/2 on main and passes 2/2 with this change.
2026-09-14 20:44:56 -07:00
Jeffrey Quesnelle d08032655f Merge pull request #111423 from NousResearch/feat/local-engine-update-prompt
feat(local-models): prefer b10964 and add one-click engine updates
2026-09-14 22:25:46 -04:00
emozilla e9e363c856 feat(local-runtime): prefer llama.cpp b10964 2026-09-14 21:49:13 -04:00
teknium1 a4d474777f fix(computer-use): doctor diagnoses a denied cua-driver spawn instead of crashing
On Windows the Hermes venv interpreter cannot CreateProcess a binary under
C:\Program Files\WindowsApps (WinError 5) even though the shell resolves it,
so `hermes computer-use doctor` died with a raw PermissionError traceback
from _open_mcp. Catch the spawn OSError and print what failed, why the tool
may still work (PATH resolves another copy), and the fix (reinstall outside
WindowsApps or HERMES_CUA_DRIVER_CMD), exit 2.
2026-09-14 17:34:31 -07:00
teknium1 45da78d645 test(computer-use): element click carries the token from the live schema alone
Invariant through the real click() path on the modern tools/list shape
(no capabilities[] array, element_token only in inputSchema); red on main.
2026-09-14 17:34:31 -07:00
teknium1 8f7853188f fix(cron): keep deferred delivery exceptions from aborting ticks
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.

Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
2026-09-14 17:29:32 -07:00
teknium1 002ee41cfc fix(cron): keep unowned Bot Chat delivery on its resolved home
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.

Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-14 17:29:32 -07:00
teknium1 3b0fe0cc2b fix(cron): keep deferred Bot Chat delivery bound to admission
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.

Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
2026-09-14 17:29:32 -07:00
teknium1 5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00
liuhao1024 37243bd668 fix(tools): resolve hermes CLI beside the interpreter in bot_mode_dm deliveries
Bot-to-bot message_agent delivery builds both transport argvs (local
teammate chat and peer dm) with a bare "hermes" as argv[0]. Since #96631
the delivery runner spawns under terminal_tool's isolated host-local
environment, which does not inherit the gateway's PATH — so on
docker/service installs (venv at /opt/hermes/.venv) every delivery exits
with FileNotFoundError: 'hermes'.

Resolve the CLI with bot_relay._hermes_cli() (#93590) — the venv sibling
of this interpreter, then shutil.which, then the bare name — at both
argv construction sites. The turn-lock matcher in _delivery_lock()
already matches argv[0] by basename, so absolute paths lock exactly as
before.

Fixes #100662
2026-09-14 17:04:56 -07:00
teknium1 1a990f3062 fix(skills): keep scanning link-shaped arguments inside fenced code blocks
Masking every balanced [x](dest) in a .md file also blanked
`cp [k](../../../.ssh/id_rsa) /tmp` inside a ```sh fence, which main scored
caution and the branch let through as safe. A link inside a code fence is a
command argument, not a hyperlink: toggle masking off between fence markers.
2026-09-14 16:28:49 -07:00
KoNit-K 5f6e3a6fe3 test(skills): trim link-traversal coverage to two invariants
Split the salvaged regression into the two facts the fix must hold:
a 3-level doc link no longer produces a traversal finding or blocks a
community install, and traversal outside a link destination (a shell
script, or prose on the same line as a link) still fires. The shell
script case is taken from #110978 by @KoNit-K; the same-line case is
what distinguishes masking link destinations from gating by extension.
2026-09-14 16:28:49 -07:00
Puvaan Raaj b259f2401d fix(skills): ignore traversal in markdown links
(cherry picked from commit 92b48941b707df4fd3c13029b157a3f28188b4c4)
2026-09-14 16:28:49 -07:00
teknium1 de16ce9d0c fix(threat-scanner): gate ssh_access on every mutating verb, not just copy verbs
Review of the write-verb gate found sed -i, chmod, truncate, curl -o, wget -O,
git clone and a scripted open(...) against ~/.ssh all slipping to no finding,
where the bare path regex on main caught them. Add those verbs and the open(
shape to the gate; read-only mentions stay clear.
2026-09-14 16:16:08 -07:00
teknium1 b1733fd085 fix: keep the ssh_access id, word-bound the verb gate, collapse tests to two invariants
Salvage trim of #89249:
- keep pattern id `ssh_access` (test_memory_tool and callers key on it; renaming buys nothing)
- word-bound the verb alternation: the unanchored form matched `add` inside "address" and
  `dd` inside "middle", so read-only prose still fired; `\b` closes that leak
- fold the second "bare leading redirect" regex into the same alternation (`>>?` branch)
- tests: one parametrized "write shapes still fire" (echo, cat, cp, tee, mv, install,
  printf, dd, scp, rsync, ln, leading redirect, option clusters) and one "read-only
  mention does not fire"; the trade-off/change-detector tests are dropped
- test_memory_tool: the persistence fixture used a read-only mention; use a write shape
2026-09-14 16:16:08 -07:00
Martin Mogis fe68349cd6 fix(threat-scanner): close reviewer-noted ssh_access_write bypass shapes
Extend the write-verb alternation with mv/install/printf/dd/scp/rsync/ln,
allow short option clusters between verb and path, and add a bare-leading-
redirect branch (a > ~/.ssh/... line carries no verb word at all). Prose
that names a write primitive before an SSH path still fails closed — a
false positive costs a review, a false negative is a backdoor.

(cherry picked from commit 4865b96ca3c8661c0ea6f5c479c198ccf68f86c4)
2026-09-14 16:16:08 -07:00
Martin Mogis 34304b722d fix(threat-scanner): require a write verb before SSH paths in the ssh_access strict pattern
The bare \$HOME/.ssh|~/.ssh regex fired on ANY mention of SSH paths in
scanned content, so operational documentation (VPS recovery notes, SSH
configuration write-ups stored in memory or skills) was blocked as a
persistence threat. Require an echo/cat/cp/tee/append/add/write/>>
verb before the path so only backdoor-insertion shapes match.

(cherry picked from commit 8c9b4c8972ac849e4c82fa20d355733ca3e5a231)
2026-09-14 16:16:08 -07:00
teknium1 8c82e93132 fix(gateway): migrate --multiplex rolls back or resumes when the default cannot come up; every installed unit counts
`hermes gateway migrate --multiplex` ran its one fallible step LAST (install +
start the default gateway) with nothing around it. On a fleet whose secondary
ran a system unit as root (#110850) that step raised, leaving the flag on, the
secondary's unit removed and no gateway anywhere, and the re-run hit the
"already multiplexing (flag on)" short-circuit over an empty fleet.

- apply_migration(): the default bring-up runs inside a rollback. On failure the
  manifest written before the first destructive step restores the flag and
  reinstalls every recorded per-profile gateway with its recorded User=.
- MigrationPlan.interrupted: flag on + manifest present + no live default
  gateway is a half-applied migration, not "already multiplexed"; the re-run
  resumes from the manifest (target manager and User= read from it, since the
  units themselves are gone) instead of refusing. Flag off + leftover manifest
  refuses to overwrite it and points at --standalone.
- ProfileGateway.services records EVERY installed unit (user and system) and the
  manifest carries them; apply stops/uninstalls all of them and rollback
  reinstalls all of them, so a second owner is never left live beside the
  multiplexer. The unattended hook treats a two-unit profile as an ambiguous
  topology and refuses (review finding on #110205).
- gateway_identity(): an unresolvable User= on a system unit stays None instead
  of borrowing the profile directory's owner; the unattended hook treats the
  unknown principal as a boundary (review finding on #110205).
- auto_migration_opted_out(): reads the effective config (load_config_readonly
  under the default home), so a managed `false` wins over a user `true` and a
  YAML string "false" is an opt-out, not a truthy value (review finding on
  #110205).

Builds on KoNit-K's #110854 (run_as_user threaded through install, preserved
from the removed system unit).
2026-09-14 16:16:06 -07:00
KoNit-K 18a45464f3 fix(gateway): preserve system service user during migration 2026-09-14 16:16:06 -07:00
KoNit-K 3bb44dc7b6 test(config): reject leftover auto_migrate key
Pin that hermes config set no longer treats gateway.auto_migrate as a known key after DEFAULT_CONFIG uses auto_multiplex_migration.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-14 16:16:06 -07:00
KoNit-K c8a6b3ca94 fix(gateway): keep auto_multiplex_migration spelling
Rename DEFAULT_CONFIG to the documented key so the reader, docs, and defaults agree. Do not reintroduce auto_migrate.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-14 16:16:06 -07:00
teknium1 23ef64081c fix(gateway): bind the session profile once around command.dispatch and bundle routing (#110695)
Follow-up on the salvaged #110698 (which scoped `_dispatch_skill` alone):

- One `_session_home_scope(session)` binding around the whole stage loop in
  `command.dispatch`, so quick commands (`_load_cfg` → `_active_config_path`
  honours the override), bundles (`skill-bundles/` is home-relative) and skills
  all resolve against the SAME profile the routing guard used.
- `slash.exec` resolves `_bundle_key_for(base)` under the same scope, so a
  bundle that exists only in the secondary profile is routed to dispatch at all.
- `_dispatch_skill` uses the home-keyed `get_skill_commands()` (the guard's
  reader) instead of an unconditional `scan_skill_commands()`.
- Test (red on origin/main): a bundle only under profile B's `skill-bundles/`
  is dispatched for a session bound to B.
2026-09-14 16:14:33 -07:00
teknium1 a21747fe4e fix(kanban): an anchorless thread subscription warns once instead of vanishing (#110919)
Follow-up on the salvaged #110928 (CLI `--parent-chat-id` / `--guild-id`):

- `_claim_for_sub` skipped a thread-shaped row that matched no `profile_routes`
  entry at DEBUG on every tick. Legacy rows written before the flags existed can
  never match a channel-level route (no `parent_chat_id`), so the notifier now
  logs ONE WARNING per row naming the task, the thread and the re-subscribe
  command. Still fail-closed: the events stay unclaimed.
- Docs: the kanban user guide explains the anchors and shows the Discord-thread
  subscribe command under `profile_routes`.
- Test (red on origin/main): two collects → exactly one WARNING, events unseen.
2026-09-14 16:14:33 -07:00
teknium1 45b42202fa fix(cron): a raising profile gate ticks nothing; housekeeping respawns a dead ticker (#111010)
Follow-up on the salvaged #111034:

- `_start_multiplex` published the enumerated home list before the gate had
  filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
  #100489) kept the thread alive but ticked every profile UNGATED — racing the
  gateway that owns them for the same cron store. The list is now assigned only
  after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
  thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
  chore that respawns a ticker that ended without a stop request and logs the
  outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
  outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
  `executions.db` (no patched recover) no longer kills the ticker; a raising gate
  keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
  leaves a stopped one alone.

Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
2026-09-14 16:14:33 -07:00
KoNit-K 4c306e9e8c fix(gateway): scope skill dispatch to session profile 2026-09-14 16:14:33 -07:00
KoNit-K 38bda39562 fix(kanban): preserve Discord thread route anchors 2026-09-14 16:14:33 -07:00
Konstantin Khlopkov 8fce4baf8c fix(cron): keep the ticker thread alive through startup recovery and marker writes (#111010)
The gateway runs InProcessCronScheduler on an unsupervised daemon thread, so any exception escaping outside the guarded tick body ends cron silently while the gateway keeps serving. Guard the three remaining escape windows: startup recovery + initial heartbeat in start(), per-cycle profile enumeration/gating in _start_multiplex, and every status-marker store write (heartbeat/error/clear) in both loops. A failing store now degrades to a logged warning and a failed cycle instead of thread death.
2026-09-14 16:14:33 -07:00
teknium1 4ea7c882fc test(skills-guard): cover the wider attribute-access and stdin-password cases
Extend the two salvaged invariant tests instead of adding more: the
`.profile` negative now covers `?.`, `().` and `].` access (the
lookbehind was widened from `[A-Za-z0-9_]` to `[\w)\]?]` during salvage),
the positive covers a bare `.profile` token, and the sudo positive covers
the `-S` (read password from stdin) form, which the gateway's own masked
password prompt relies on.
2026-09-14 16:14:23 -07:00
unsupportedpastels 4f00c456f0 fix(skills-guard): stop flagging sudo.request/sudo.respond event names as sudo usage
`sudo.request` and `sudo.respond` are gateway wire events: the masked
sudo-password prompt the terminal tool raises, which every client surface
(desktop, TUI, any plugin that relays secure prompts to another device)
has to name to forward it. The `sudo_usage` rule matched the bare word,
so any plugin listing those events scored `high` and every install of it
landed on `caution`, which Hermes Desktop cannot confirm past.

A dotted event name is never a shell `sudo` invocation. Exclude exactly
those two names with a negative lookahead; a real `sudo cmd` still fires.

(cherry picked from commit a7a3a31126de8057c2dbb9b7048724aa9d55a7fd)
2026-09-14 16:14:23 -07:00
chenxue 63395922bf fix(skills-guard): stop shell_rc_mod matching attribute access
The shell-startup-file pattern is
`\.(bashrc|zshrc|profile|bash_profile|bash_login|zprofile|zlogin)\b`.
Six of those names are distinctive enough that seeing them after a dot
means the file. `profile` is not: it is also how every language spells
attribute access, so `self.profile`, `user.profile`, and
`request.profile` each score a medium persistence finding.

The cost is signal, not blocking -- `_determine_verdict` treats
medium/low alone as informational -- but a plugin that happens to name a
field `profile` buries the findings a reviewer has to read. A model-
provider plugin whose tests exercise a `profile` object contributed 36
of 38 findings in its scan report, all of them this pattern.

Split `profile` into its own entry anchored on a non-identifier
character before the dot. Real references keep matching in the forms they
actually take (`~/.profile`, `"$HOME/.profile"`, `./.profile`,
bare `.profile`); attribute reads no longer do. The other six names are
untouched. Both entries keep the `shell_rc_mod` id, and scan_file
deduplicates on (pattern_id, line), so a line holding both still yields
one finding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 705d0bddc61a1c55be82956d343b229544350eae)
2026-09-14 16:14:23 -07:00
teknium1 818d781c0b test(tools): trim the salvaged #110632 / #110307 tests to two invariants each
The file-safety salvage carried five tests and the execute_code salvage four;
fold them into the two behaviour contracts per fix: (a) the named-profile scope
exempts the root's direct files and the write lands with no prompt; a child env
sees the ACTIVE profile's home per turn; (b) negatives hold: a checkout's
.hermes/config.yaml stays gated fail-closed, a lookalike profiles/ tree exempts
nothing; a dedicated process with no override is untouched. Both red on base.
Also compress the _hermes_exempt_homes docstring (WHY only; cite #110630).
2026-09-14 16:13:51 -07:00
teknium1 90a65d6e38 fix(gateway): gateway-stop teardown fires memory-provider lifecycle hooks under the owning profile
Under `gateway.multiplex_profiles: true` every gateway-hosted session that ended
inside the gateway process logged `Memory provider 'openviking' on_session_end
failed: get_secret('OPENVIKING_API_KEY') called with no profile secret scope
active` and skipped its end-of-session commit. The provider lifecycle hooks
(`flush_pending` -> `on_session_end` -> provider teardown -> `close`) read
credentials and home at call time; the stop-time finalize pass and the
idle-cache sweep run on the main loop outside any adapter handler, so
`_run_in_executor_with_context` copied an EMPTY scope into the worker.

Route `_cleanup_agent_resources_off_loop` and `_finalize_session_off_loop`
through `_run_release_in_profile_scope` (the seam cache eviction already uses):
a scoped caller keeps its scope; an unscoped caller passes the session key and
the owner's profile scope is entered from it. The stop path passes the key for
both active and idle-cached agents. Cron's post-run cleanup thread is the
sibling seam and is fixed by the cherry-picked #110634.

Live repro (fake OpenViking recording the X-API-Key header, alpha profile,
multiplex on): base - WARNING + traceback, no commit; head - no warning,
`POST /api/v1/sessions/<sid>/commit` arrives with alpha's key.

Fixes #110622.
2026-09-14 16:13:51 -07:00
teknium1 09a7c297ec fix(tools): uncached check_fn probes classify UnscopedSecretError from the live scope
Under `gateway.multiplex_profiles: true` every gateway start logged a WARNING +
traceback from `tools.registry` for `_check_vault_available`. The browser vault
gate is a `no_cache_check_fn`, so `_check_fn_cached` short-circuited into
`_run_check_fn_uncached(fn)` without consulting the cache scope and the default
`unresolved_scope=False` hint reported an EXPECTED boot-time fail-closed read as
"scope was resolved" — the #100697 fix only reached the cached branch.

Derive the verdict from `current_secret_scope()` at the catch site instead of a
branch-derived hint: no scope installed → DEBUG without traceback (expected,
re-probes on the first scoped turn); scope installed but the read still failed
closed → WARNING + traceback (the probe dropped the scope on a bare thread/hop,
a spawn-site bug that must stay loud). Every uncached probe that reads a
credential is covered, vault included.

Fixes #110635. Credit @KoNit-K (#110638) for the report-side diagnosis; that PR
disabled the legacy-cloud probe for the vault gate only, which would have hidden
the vault from legacy Browser Use cloud users and left the class open for every
other uncached probe.
2026-09-14 16:13:51 -07:00
Jony 45aebc11b0 fix(cron): preserve cleanup profile scope 2026-09-14 16:13:51 -07:00
Kevin Rajan 7ef0b98272 fix(tools): honor multiplexed per-turn HERMES_HOME in execute_code child env
Under a multiplexed Desktop/Dashboard connection, one server process serves
multiple profiles, binding a context-local HERMES_HOME override per turn.
_build_child_env scrubbed the server process's os.environ - which carries the
machine-default HERMES_HOME - so skill scripts run via execute_code silently
read/wrote the wrong profile's directory.

Rewrite HERMES_HOME in the child env from the active override when one is
bound (the same per-turn rewrite apply_subprocess_home_env already does for
HOME). With no override (dedicated per-profile process) the inherited value
is left untouched.

Fixes #110303.
2026-09-14 16:13:51 -07:00
moxian ab78e71a9e fix(file-safety): exempt the Hermes ROOT, not just the profile home
Under a named profile (``hermes -p <name>``, i.e. ``HERMES_HOME=<root>/profiles/<name>``)
the protected-instruction gate exempted only the profile dir, so the ROOT's DIRECT
files (LEDGER.md / MEMORY.md / USER.md / SOUL.md / AGENTS.md / DECISIONS.md ...) fell
through to the ``.hermes`` component rule and were treated as project-local
``<repo>/.hermes/config.yaml``. That gate always-asks and fails closed without a human
channel, so every write there was refused headless — it blocked #54 (LEDGER.md edit
could not land). Subdirectories (scripts/, cron/, data/) were unaffected, which is why
#55 sailed through and #54 did not.

``_hermes_exempt_homes()`` now returns the active profile home plus the Hermes ROOT
when that home really is a named profile (``named_profile_home`` validates: ``.hermes``
name, home markers, tombstone, or the resolved default root), so a coincidental
``profiles/`` directory never exempts its parent. ``_get_real_hermes_home()`` keeps its
per-profile semantics — the #107327 multiplex regression test pins that — and the
exemption consumes the new helper instead.

Refs: https://gitlab.yx.netease.com/agent-projects/platforms/agent-mentor/-/issues/60
2026-09-14 16:13:51 -07:00
teknium1 36c7f6c89d refactor(skills): trim the denylist demotion to one owner regex and two invariants
Fold the cherry-picked mechanism into the main file's compact style: one
denylist-owner regex replaces the assignment-target + name regex pair and the
_in_denylist_construct helper; the comment-prefix table loses its dead .css
entry; the denylist demotion lands on high (a confirmable caution) instead of
medium so the finding still gates the install and shell_rc_mod, already
medium, leaves the set. The verb guard widens to stem-prefix matching and
copyfile/copy2/sendfile per the review thread on #92632, closing the
Path.read_text()/shutil.copyfile shape. Tests trimmed to two invariants: a
skill's own denylist is caution (force-overridable), a real write to
authorized_keys stays dangerous.
2026-09-14 16:13:35 -07:00
Jack Lau 52bec9d77e fix(skills): stop scoring a skill's own denylist as an access
`_scan_file` matches every threat pattern against every line with no notion
of what the line is. The path-token patterns — `authorized_keys`, `~/.aws`,
`~/.ssh` — therefore cannot tell `cat ~/.ssh/authorized_keys` from a skill
that spells the path in order to REFUSE to read it. One `critical` becomes
`dangerous` in `_determine_verdict`, and on a community source that blocks the
install with no `--force`, so the skill that bothered to skip credential
files is the one that gets quarantined (#92478).

Demote, do not drop, following the precedent `allowed_tools_field` already
sets in this file: the finding keeps its file, line and matched text so an
auditor still sees the token; it just stops deciding the verdict alone.

Two contexts qualify, and only for the eight path-reference pattern ids:

- a whole-line comment, in a language that HAS comments, drops to `low`.
  Markdown is deliberately excluded: `#` opens a heading there, and Markdown
  prose is the prompt-injection surface itself.
- a line inside a construct NAMED as a denylist (`SKIP_PATTERNS`, `DENY_*`,
  `EXCLUDE_*`), carrying no verb that could touch the path, drops to `medium`.

The issue also suggested demoting any quoted token on a verb-free line. That
is wider than it looks — a fragment in quotes can be interpolated into a
command a line later — so the name test is the primary rule and the verb test
only guards it, because the construct's name is attacker-chosen.

The denylist check is statement-aware rather than per-line: the reported
match sat on a continuation line of a multi-line regex whose name is four
lines up, which a per-line test reads as anonymous.

8 tests. Reverting the demotion fails the two behaviour tests and leaves the
six contract tests green; dropping the verb guard, the Markdown exclusion, or
the statement-awareness each fails exactly its own test.

(cherry picked from commit f50acd34f62375926bd5843125e106f7e7eb6ee8)
2026-09-14 16:13:35 -07:00