Commit Graph

3613 Commits

Author SHA1 Message Date
fangliquanflq f5200a4c10 fix(config): declare shared Docker container key 2026-08-25 04:00:27 -07:00
Teknium 76e306c458 refactor(tools): remove expired BFL FLUX 3 promo core tools (migration v39); FLUX 3 stays via video_gen/FAL for subscribers (#94599)
* refactor(tools): remove expired bfl_flux3_* promo tools; FLUX 3 rides the video_gen provider surface

* test: relay-cutover migration asserts >= v38, not the version literal
2026-08-25 02:45:10 -07:00
Teknium 64a6f42cb3 test: opted-out profile seeding now exercises the essential-only sync subprocess 2026-08-24 20:25:10 -07:00
Teknium 3733e4aff5 fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
2026-08-24 20:25:10 -07:00
Teknium f14059fad2 fix(models): OpenRouter :nitro/:floor routing variants no longer rejected by /model validation
OpenRouter's :nitro, :floor, :exacto, and :online suffixes are request-time
routing modifiers valid on any model id — /models lists only the base model.
validate_requested_model() compared the full suffixed id against the listing,
so a valid variant was either rejected outright or fuzzy-auto-corrected to
the base id, silently stripping the user's routing opt-in.

Now, for OpenRouter only, a recognized variant suffix validates the BASE id
against the live listing (and the curated-catalog soft-accept and static-
catalog fallback paths) while preserving the suffixed id for persistence and
API requests — checked BEFORE fuzzy correction. :free/:batch/:thinking
remain direct catalog SKUs and keep exact-match semantics; unknown suffixes
and unknown bases are still rejected.

Reported by JEB (Jakob's Hermes Agent) via Discord.
2026-08-24 13:36:23 -07:00
Teknium de07bd5fab fix(serve): emit BACKEND_PORT_IN_USE sentinel + exit 75 on port bind conflict (#93608)
A held port made 'hermes serve' print only uvicorn's bare
'ERROR: [Errno 98/10048] error while attempting to bind on address'
and exit 1 — indistinguishable from a broken backend for the desktop
spawn and wrapping scripts.

- Preflight bind probe (matching uvicorn's SO_REUSEADDR bind flags)
  before uvicorn.Server; on conflict print machine-readable
  'BACKEND_PORT_IN_USE port=<port>' + a human hint naming likely
  holders, exit 75 (EX_TEMPFAIL — existing repo convention, see
  gateway/restart.py, kanban_db.py).
- Probe-to-bind race covered: SystemExit(1) from uvicorn's own bind
  failure is re-checked and translated on both POSIX and Windows
  runner paths.
- --port 0 (ephemeral) short-circuits the probe: unchanged behavior.
- HERMES_BACKEND_READY contract untouched.
- Tests: real held-socket repro (sentinel + exit 75, sabotage-proven
  to fail as bare exit 1 without the fix), free-port boot regression,
  ephemeral-port regression, probe/classification units.
- Docs: port-conflict paragraph under 'hermes serve' in
  reference/cli-commands.md.
2026-08-24 09:55:15 -07:00
Brooklyn Nicholson b595fcd5e1 fix(desktop): keep the Linux HUD clickable and recoverable
X11 cannot restore a window that has ignored the mouse, so stay solid
there. Native Wayland keeps click-through via the cursor poll. Add
desktop.ozone_platform_hint so COSMIC users can opt into XWayland for
always-on-top, and a layout reset that restores the default size.

Co-authored-by: Codex Metatron <47930664+BlakeB254@users.noreply.github.com>
Co-authored-by: DeseretSaint <202557515+DeseretSaint@users.noreply.github.com>
Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>
2026-08-24 06:48:10 -05:00
Teknium ec5e369fe6 Merge pull request #93784 from NousResearch/salv/81234-retry-carrier
fix: /retry and /undo no longer replay an older message after compaction (#81233, salvage #81234)
2026-08-24 03:31:29 -07:00
Teknium a87d314e44 fix(auth): malformed OpenRouter env key no longer shadows valid credential-pool key
A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).

- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
  that mismatch a declared prefix, logging a WARNING naming the env var and
  expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
  only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
  are never rejected. Valid env keys still win over the pool (precedence
  unchanged).

Fixes #93593
2026-08-24 03:22:57 -07:00
chelsealong 9c013eaaf8 fix(dashboard): follow scroll on implicit active-session resume (#93518)
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes #93518.
2026-08-24 03:21:49 -07:00
liuhao1024 8d17060249 fix(cli): hard-exit the Windows update hand-off child once work is durable
The re-exec'd venv child spawned by
_reexec_dependency_sync_off_windows_shim completes every update step —
the receipt records success / "completed at command boundary" — but then
hangs in interpreter shutdown on a leftover non-daemon thread, freezing
the PowerShell window for minutes after "Update complete!". On the
hand-off path only (HERMES_UPDATE_REEXEC=1), after the receipt is
finalized, the update lock released, and stdio restored, flush and
os._exit(code) instead of unwinding — the same treatment #79040's cron
workaround applies. SystemExit codes (including early refusals)
propagate to the hard exit; real exceptions keep the normal raise path
so tracebacks still print. Non-hand-off invocations are untouched: the
marker env is set solely when the shim spawns the child.

Fixes #93581
2026-08-24 03:21:28 -07:00
Teknium 93bf6f7225 fix(cli): fail closed on empty fleet probe across all pre-update liveness signals (#93406)
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.

Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
  same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries

Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.

Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.

Builds on RelaxJonh's #93410. Fixes #93406
2026-08-24 03:21:18 -07:00
aniruddhaadak80 1007296ca0 fix(cli): add --reasoning to both top-level value-flag sets
--reasoning takes a value (metavar=LEVEL in _parser.py) but was absent
from _TOP_LEVEL_VALUE_FLAGS (used by _first_positional_argv) and from
_apply_profile_override's value_flags set. Every invocation like
"hermes --reasoning high chat ..." therefore misclassified "high" as the
first positional, and _plugin_cli_discovery_needed() forced full eager
plugin CLI discovery at argparse-setup time - the documented startup
cost paid on every use of the reasoning override.

Add --reasoning to both sets, and add a parser-derived parity regression
test so future drift between the hand-maintained sets and
build_top_level_parser() fails CI instead of silently degrading startup
(the exact drift class AGENTS.md bans).

Fixes #93530

(cherry picked from commit 9280617ab8b308f9a8cf947f1617276a6bc8eb4f)
2026-08-24 03:20:37 -07:00
fangliquanflq 694550e486 fix(cli): derive top-level value flags from parser
(cherry picked from commit f2e5a1388115615fafde49a5f2144e5855a81d89)
2026-08-24 03:20:37 -07:00
fangliquanflq 3963fc6f21 fix(config): stop reporting stripped v15 defaults
(cherry picked from commit 4c6b67ec371b16c15e9ffbb91bfb47a504e913fe)
2026-08-24 03:20:28 -07:00
fyzanshaik 8345effc4e fix(cli): support reliable setup menu navigation
Decode Ghostty/Kitty enhanced selection and cancellation keys, make setup cancellation terminal, and add cross-terminal previous-step navigation to setup and model flows.

Refs #92833
2026-08-24 15:45:52 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Teknium d9a48f656a fix(desktop): scheduled jobs on sleeping profiles keep firing
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
2026-08-24 03:14:30 -07:00
kshitijk4poor 45aa0dc33e fix(cli): widen prompt_toolkit fallback to catch any runtime failure
Widen the exception guard from OSError to Exception (re-raising
KeyboardInterrupt/EOFError first) so any prompt_toolkit runtime
failure degrades to input() — matching the established pattern in
masked_secret_prompt.  ValueError and RuntimeError can arise from
exotic stream wrappers or event-loop issues with the same root cause:
prompt_toolkit cannot attach stdin on the terminal.

Add test_line_input_falls_back_to_input_on_any_prompt_toolkit_failure
covering the ValueError case.
2026-08-24 15:42:58 +05:30
wo-o a88cbe10fb test(cli): add regression test for line_input OSError fallback
Cover the prompt_toolkit runtime-failure path added in the fix commit: a
tty-reporting stdin where prompt_toolkit raises OSError(22) (macOS kqueue
EINVAL on fd 0 under curl|bash installs) must degrade to input() instead
of aborting the setup wizard.
2026-08-24 15:42:58 +05:30
Ben Barclay 9a91f7058c Merge remote-tracking branch 'origin/main' into fix/pkce-samesite-none-salvage 2026-08-24 16:48:06 +10:00
Teknium 03c3554fc2 fixup(curator): align #93002 test stubs with #93149 set_pinned bool contract
Combining both PRs for issue #92993: #93149 makes set_pinned() return a
bool and _cmd_pin/_cmd_unpin exit 1 on a no-op write; #93002's tests
stubbed set_pinned with a None-returning lambda, which the combined
_cmd_pin now reads as failure. The stub reports True (write landed) so
#93002's messaging assertions exercise the intended success path.
2026-08-23 21:12:20 -07:00
liuhao1024 ef882a5595 fix(curator): say what pin actually does on an unmanaged skill
`hermes curator pin` guarded on is_agent_created (a filesystem-shape
check), but the flag only matters when the skill carries the
curator-management marker: curated_report() walks marker-carrying skills
only, so auto-transitions never consider an unmanaged (pre-marker)
skill at all. Pinning one recorded the flag and then printed
"will bypass auto-transitions" — an effect that does not exist.

Keep the write (the flag becomes meaningful after `hermes curator
adopt`) and branch the message on is_curator_managed: unmanaged pins
now say the skill is unmanaged and point at adopt. Unpin gets the
symmetric wording.
2026-08-23 21:12:20 -07:00
beplee dd20c30dec fix(curator): check unpin result, guard status ghost rows, tighten test
Review feedback on #93149:
- _cmd_unpin now checks set_pinned's return (same false-success defect
  existed symmetrically on the unpin path)
- curated_report() pinned-visibility branch requires a local skill dir,
  so stale records for deleted dirs don't render as ghost rows
- test 2 asserts rc==0 unconditionally instead of vacuous-passing
- error message points to list-unmanaged (status doesn't render reasons)
2026-08-23 21:12:20 -07:00
beplee 7caa731e80 fix(curator): report pin failures instead of false success and surface pinned unmanaged skills
`hermes curator pin <skill>` printed success even when the underlying
write never landed. set_pinned() routes through _mutate() with
require_curation_eligible=True, which silently returns None for skills
that pass is_agent_created() but fail is_curation_eligible() — e.g. a
user-created skill named "plan", which PROTECTED_BUILTIN_SKILLS blocks
by name. The CLI then announced a pin that does not exist (#92993).

Also, a pin that DID land on an eligible-but-unmanaged skill (no
created_by marker) was invisible: curated_report() only iterated
list_agent_created_skill_names(), which requires the management marker,
so the skill showed up under 'unmanaged' with no trace of its pin.

- set_pinned() now returns bool write success; _cmd_pin() checks it,
  exits nonzero and explains the refusal when the write did not land
- curated_report() additionally includes curation-eligible skills whose
  usage record carries pinned=true, so their pins are visible in status

Fixes #92993
2026-08-23 21:12:20 -07:00
fangliquanflq 030edf9774 fix(auth): canonicalize configured provider display names 2026-08-23 20:00:53 -07:00
fangliquanflq 3a7c094582 fix(auth): preserve configured provider compatibility 2026-08-23 20:00:53 -07:00
fangliquanflq c527b2c0a4 fix(auth): normalize configured provider pool keys 2026-08-23 20:00:53 -07:00
Teknium b03b8ac51d fix(dashboard): name the exact gate trigger in fail-closed refusals
When the bind is loopback and the only gate trigger is
dashboard.public_url, the startup refusal now says so explicitly and
gives both exits (configure a dashboard auth provider, or remove
dashboard.public_url if the proxy no longer exists). Prevents the
stale-public_url mystery-locked-dashboard upgrade trap.

Adds a truth-table regression suite for should_require_auth and the
fail-closed message shape.
2026-08-23 19:55:17 -07:00
e-macgregor d3df14a7e3 fix(dashboard): secure loopback public URL proxy mode 2026-08-23 19:55:17 -07:00
RickyYii 6b3a7af73d fix(security): cover privilege wrappers and command-string options
Review follow-up on #84203. Both points reproduce; neither was a regression
from the first pass, but both are live bypasses of the same guard.

**Privilege and namespace wrappers were missing.** The allowlist covered the
coreutils-shaped wrappers but not the privilege ones, so each of these ran a
lifecycle script straight past the walk:

    pkexec bash ~/restart.sh
    runuser -u root -- bash ~/restart.sh
    setpriv --reuid=0 -- bash ~/restart.sh
    systemd-run --scope bash ~/restart.sh
    nsenter --target 1 --mount bash ~/restart.sh
    unshare -r bash ~/restart.sh

Added `pkexec`, `su`, `runuser`, `setpriv`, `systemd-run`, `nsenter` and
`unshare`, each with the value-taking options that would otherwise be
mistaken for the command (`nsenter -t 1`, `systemd-run -p X=1`,
`runuser -u root`, …).

**An option can carry a command STRING, not an argv tail.** `env -S` and
`su`/`runuser` `-c` take shell source. The peel treated the operand as an
opaque value and skipped it, so `env -S 'bash ~/restart.sh'` was never
scanned — the string went unread rather than being recursed into.

`_STRING_COMMAND_OPTIONS` now names those options and their values are
re-scanned as shell source, the same treatment `sh -c` payloads already get.
They are read at the ORIGINAL command token, before the transparent-prefix
peel, because peeling past `su`/`env` would discard the very option carrying
the command. `--opt value` and `--opt=value` are both handled.

Scope, stated plainly: this is an enumerated allowlist, not a general
solution to "wrapper that execs its tail". A wrapper outside the set, or a
value-taking option outside these tables, still resolves to no reference —
that fails open, exactly as it did before this PR, and it is a miss rather
than a false block. The reviewer offered "extend the set with tests, or
document that the list is heuristic"; this does the first and states the
second.

Tests: 23 new cases (220 in the file) — every added wrapper against a script
reference including the value-operand option forms, both command-string
option spellings for env/su/runuser, and the same wrappers around ordinary
work (`pkexec systemctl status nginx`, `su -c 'ls -la'`, `env -S 'echo hi'`,
`nsenter -t 1 -m ps aux`) which must stay allowed. 15 fail on the tree
before this commit.

False positives re-checked at scale: the 9,258 command lines from this
repo's own scripts and docs give an identical verdict set before and after —
0 new false positives, 0 lost detections, 0 exceptions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
RickyYii a19e1bae10 fix(cron): stop a relative path from disabling the data-sink exemption
`_mask_data_sink_arguments` exempts lifecycle text living in the arguments of
executables that cannot run them (`grep`, `rg`, `journalctl`, `sqlite3`, …),
so hunting for a restart string in logs is diagnostics rather than a command.
The exemption is dropped when an argument looks like an escape back into
execution — including anything starting with a dot, because sqlite3 spells
its escapes as dot-commands (`.shell`, `.system`).

But `.`, `./x` and `../x` are ordinary path operands, and

    grep -r 'systemctl restart hermes-gateway' .

is the most ordinary recursive search there is. The leading-dot test treated
its `.` operand as a sqlite3 escape, disabled masking for the whole segment,
and blocked the command outright — the exact false-positive class the
exemption exists to prevent, on the shape most likely to hit it. Searching a
relative subdirectory (`./logs`, `../archive`) fails the same way, as does a
relative sqlite3 database path (`sqlite3 ./stats.db "SELECT ..."`).

Require a dot followed by a NAME character (`^\.[A-Za-z]`) so a dot-command
still defeats the exemption while a relative path stays a path. A dotfile
operand (`.env`) still reads as a dot-command — conservative, and unchanged
from today's behavior.

This narrows a security guard in the permissive direction, so the escape
hatches are pinned explicitly: with a relative-path operand present,
`.shell`/`.system`, psql's `\!`, a pipe into `sh`/`bash`/`sudo sh`/`xargs`,
command substitution, and a `;`/`&&` continuation all still block. Only the
segment's own data arguments are masked, and only when nothing in it can
reach execution.

Tests: 18 new cases in tests/hermes_cli/test_gateway_restart_loop.py — the
relative-path shapes that must now be allowed, plus the ten escape-hatch
shapes that must still block. The allow cases fail on the unfixed tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
RickyYii 5921ba8c06 fix(security): see through wrapper prefixes in the gateway lifecycle guards
`sudo`, `env`, `nohup`, `timeout` and friends exec their argument tail, so
the command that actually runs sits further right. Three guards read only the
first token of a segment, saw the wrapper, and never inspected what it runs:

  bash ~/restart.sh                      → blocked
  sudo bash ~/restart.sh                 → allowed
  launchctl submit -l com.x -- helper    → blocked
  sudo launchctl submit -l com.x -- helper → allowed

Same foot-gun, one word of prefix. That reaches both enforcement points —
`cron.jobs.create_job` and `tools/terminal_tool.py` under `_HERMES_GATEWAY=1`
— and defeats the label-independent submit block that #62891 added precisely
because a persistent helper is the indirect route to a restart loop.

`_peel_transparent_prefixes()` walks past a bounded chain of these wrappers,
skipping their own options, their value-taking options (`sudo -u deploy`,
`stdbuf -o0`), `VAR=value` assignments, a `--` end-of-options separator, and
`timeout`'s duration operand, then returns the index of the real command. It
is applied to the referenced-script walk, the `sh -c` payload walk, and the
`launchctl submit`/`bootstrap` block.

In the referenced-script walk the peel is ADDITIVE — the segment is read at
the original token and again at the peeled one — because peeling must never
remove a reference the un-peeled read would have found. A local script named
`./timeout` is a script, not the coreutils wrapper, and consuming it as a
prefix would have silently stopped scanning it. (The other two call sites
need no such care: no wrapper name is also a shell name or `launchctl`, so
peeling there can only add.) That split is why the per-index logic now lives
in `_references_at()`.

This is not a new reading of shell syntax for this module — `_PIPE_TO_INTERPRETER`
already treats `sudo ` as transparent for the pipe case (`... | sudo sh`).
This generalises the same reading to the command position.

Deliberately NOT applied to the data-sink masking in
`_mask_data_sink_arguments`: peeling there would widen an exemption, and the
conservative reading is the safe one.

No false positives: peeling only changes which token is treated as the
command, so a wrapper around ordinary work resolves to a non-shell executable
and yields nothing, exactly as before (`sudo apt-get update`,
`timeout 60 curl ...`, `nice -n 10 make -j4`, a bare `env`).

Tests: 40 new cases in tests/hermes_cli/test_gateway_restart_loop.py — every
wrapper form against a script reference, a dot-source, a nested `sh -c`
payload and `launchctl submit`, plus the benign wrapped commands, a wrapped
clean script, and the `./timeout`-style lookalike names that pin the additive
reading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
王雪帆 c4871226f1 fix(cli): honor target_model when resolving custom providers
resolve_runtime_provider() documents target_model as the explicit model
override for mid-session switches and auxiliary slots, but the custom
provider path (_resolve_named_custom_runtime) never received it and
silently substituted the provider's configured default_model instead.

This made auxiliary slots such as auxiliary.background_review silently
run the provider's default model rather than the configured one — e.g. an
ocx-proxy slot configured for gemini-flash actually executed
cursor/claude-sonnet-5, hitting upstream rate limits.

Pass target_model through to the custom runtime resolver and prefer it
over the provider's default model in both the pooled and non-pooled
credential paths.
2026-08-23 18:25:35 -07:00
Meng Chee 92edb861be fix(cron): close NUL-padded script bypass in lifecycle guard
The #76762 binary check treats any NUL byte in the first chunk as "compiled
binary, nothing to scan":

    if b"\x00" in data:
        return None, False

"Contains a NUL" and "is a compiled binary" are different questions, and the
gap between them is a guard bypass. `bash` executes a *text* script straight
past an embedded NUL, so one pad byte disables the entire scan while the
script still runs:

    #!/bin/bash
    # pad<NUL>
    hermes gateway restart

    scan("bash padded.sh")  -> False   (not blocked)
    bash padded.sh          -> executes the lifecycle command

This shape was blocked before #76762, so the crash fix traded a loud failure
for a silent one.

Keying the check on a leading `#!` is not sufficient: a shebang-less file with
a NUL on any line but the first also executes normally. (A NUL on line 1 of a
shebang-less file is the one shape bash rejects, exit 126 — but that same file
is still executable via `. file`.)

Fix: identify binaries by MAGIC NUMBER — ELF, Mach-O (incl. byte-swapped and
universal/fat), PE/COFF, static archive, gzip, zip — with a shebang always
winning. A NUL-bearing *text* file is scanned with its NULs stripped;
stripping can only splice tokens together, never apart, so it fails closed.
File extensions are deliberately not consulted, so a suffixless shell script
is still scanned.

The size check now runs BEFORE the strip: stripping shrinks the buffer, so
checking afterwards would let an oversized file slip under the threshold and
skip the fail-closed branch. (Caught by
test_oversized_nul_bearing_text_still_fails_closed, which failed on the first
cut of this patch.)

Return values are unchanged, so this does not conflict with the in-flight
crash-class fixes to the same function.

Tests (tests/hermes_cli/test_gateway_restart_loop.py), 3 of which fail on main:

- test_nul_padded_script_is_still_scanned
- test_nul_padded_script_without_shebang_is_scanned
- test_oversized_nul_bearing_text_still_fails_closed
- test_elf_binary_is_not_scanned_as_script       (#76762 stays fixed)
- test_macho_binary_is_not_scanned_as_script     (incl. fat binary)
- test_clean_script_without_lifecycle_command_not_blocked
2026-08-23 18:24:36 -07:00
Meng Chee da30db8e8c fix(cron): scan dot-operator sourced scripts in lifecycle guard
`_iter_referenced_shell_scripts` recognises the `source` builtin so a script
pulled in with `source ./restart.sh` gets scanned for lifecycle commands. The
POSIX dot operator is the same builtin, but it was not caught:

    if executable_name in {".", "source"}:

`executable_name` is `Path(executable).name`, and `Path(".").name` is the
**empty string** -- pathlib normalises "." to the current directory, whose name
is "". So the set membership never matched for `.`, the sourced script was
never added to the reference walk, and its contents were never scanned.

Verified against current main:

    . /tmp/restart.sh        -> not blocked   (script never scanned)
    source /tmp/restart.sh   -> blocked
    bash /tmp/restart.sh     -> blocked

where /tmp/restart.sh contains a `hermes gateway restart` line. Sourcing runs
the script in the current shell, so the dot spelling is not merely equivalent
to `source` -- it is the more common form in practice.

Fix compares the raw token as well as the basename:

    if executable in {".", "source"} or executable_name == "source":

Keeping the `executable_name == "source"` arm preserves the existing behaviour
for a path-qualified spelling, while the raw-token test catches `.` without
relying on pathlib normalisation.

Tests (tests/hermes_cli/test_gateway_restart_loop.py):

- test_dot_operator_sourced_script_is_scanned -- the regression; fails on main
- test_source_builtin_sourced_script_is_scanned -- `source` stays blocked
- test_dot_operator_clean_script_not_blocked -- widening the check must not
  false-block an innocent `. ./activate.sh`

Found while auditing the guard after #76762. Scoped deliberately to this one
defect; the NUL-padded-script bypass I found in the same audit is a separate
PR.
2026-08-23 18:24:36 -07:00
zgqq 9aa0721b23 fix: inert heredoc bodies no longer trip the gateway lifecycle guard (#88336)
Runbook prose inside a quoted-delimiter heredoc feeding a data sink
(cat > file <<'EOF') is documentation, not a command this shell will
execute. Mask provably-inert heredoc bodies (tools/shell_heredoc's
conservative stripper, already used by terminal_tool) before scanning.
Fails open on any ambiguity: executable and unquoted-delimiter heredocs
stay scanned. Salvaged from PR #88336 by @zgqq (the heredoc half; its
Branch D boundary and dir-token halves already landed/were fixed).
2026-08-23 18:10:31 -07:00
Artur Hapantsou b34edd6b01 fix: execute_code and argv-list payloads no longer bypass the gateway lifecycle guard (#68289)
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
2026-08-23 18:01:59 -07:00
KeaneYan 1c791cbfe6 fix(gateway): resolve uninstall lifecycle guard conflict 2026-08-23 18:01:59 -07:00
BotUser 679e07a074 fix(gateway): close order-dependency + missing-verb gap in launchctl lifecycle guards
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.

A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:

    for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
      label=${item%%:*}; plist=${item#*:}
      launchctl bootout "gui/$uid/$label"
      launchctl bootstrap "gui/$uid" "$plist"
    done

The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.

cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).

`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.

We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.

Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.

Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 17:56:36 -07:00
Teknium 20e308fea7 fix: use lookbehind anchor so binary-decoded and remote-read content still scans
The separator-class anchor broke two fail-closed tests (binary bytes
decode to U+FFFD adjacent to the CLI name; remote head-c reads). A
negative lookbehind excluding path/word chars keeps the #77173 fix
while preserving every fail-closed content-scan path.
2026-08-23 17:50:48 -07:00
eaglezzz0522-cloud 180f981125 fix: lifecycle guard Branch A anchors the CLI name at command position (#77173 path false positive)
A file path with embedded spaces (/docs/... with lifecycle words in the
filename) matched Branch A via the path tail and hard-blocked innocent
commands. Anchor the CLI name at command position (start, separator, or
substitution opener). Salvaged from PR #77536 by @eaglezzz0522-cloud,
reapplied onto the current pattern with subshell coverage and tests.
2026-08-23 17:50:48 -07:00
Axel Vanni 6d501c2958 fix(cron): make gateway lifecycle matching shell-token aware (#80269)
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.

contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.

tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.

Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.

Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
webtecnica c595d3564a fix(cron): block profile-flag gateway restart/stop when self-targeting (#78028) 2026-08-23 16:24:03 -07:00
Teknium 45fcaaa54a test: regression coverage for #92372 boundary + quote-aware segmentation
Prose false positives (Branch A trailing boundary, Branch D leading
boundary), quoted-multiline data payloads, and the fail-closed
unbalanced-quote fallback.
2026-08-23 16:18:38 -07:00
fangliquan 4d16a1a73c test(dashboard-auth): cover empty login route behavior 2026-08-23 16:09:14 -07:00
fangliquan bd13d593a9 fix(dashboard-auth): correct authentication docs anchor 2026-08-23 16:09:14 -07:00
fangliquan 7cc92cb134 fix(dashboard-auth): replace stale insecure guidance 2026-08-23 16:09:14 -07:00
Teknium aa9aaaa6cb test: mock ownership probe in stop-guard test instead of raw env var
Follow-up: test_stop_refuses_inside_gateway pinned the old env-var
gate; mock _is_supervised_gateway_process like the sibling helpers.
2026-08-23 15:49:52 -07:00
David Metcalfe 5bb657832d test: regression coverage for inherited-env lifecycle guard bypass (#92560)
Test-helper + regression test from PR #92633 by @DavidMetcalfe:
mock _is_supervised_gateway_process instead of setting the raw env
var, and add a test proving a CLI agent session with inherited
_HERMES_GATEWAY=1 but no PID ownership is no longer blocked.
2026-08-23 15:49:52 -07:00