Commit Graph

2581 Commits

Author SHA1 Message Date
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium f709bd88b6 feat(skills): render the configured create dir in every instruction that names the path
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
2026-09-01 07:32:45 -07:00
Teknium 5a8e8a6b87 fix(terminal): strict Linux-only gating for background-executor systemd scopes (#70716 follow-up)
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):

- Gate every scope-path branch on a new _IS_LINUX constant instead of
  'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
  never touches systemd code — no probe subprocess, no scope argv,
  byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
  legacy argv byte-identical, no unit recorded) and probe-returns-False
  off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
  wired into the on-demand windows-venv-e2e lane): real spawn_local on
  windows-latest asserting jobs run exactly as before — spawned, output
  captured, exit code correct, systemd path never reached even under
  faked gateway identity.

Refs #70716, #71378.
2026-09-01 02:32:53 -07:00
Brooklyn Nicholson b293157f7d feat(tools): withdraw tip and tour when the user switches them off
Off should mean the model is never told the tool exists. A switch that
only made the call fail leaves Hermes offering walkthroughs it cannot
give and promising to point at things it cannot point at, which reads as
a broken agent rather than a respected preference.

Both gate on the switch through a shared desktop_ui.user_enabled helper,
which is the reactions check_fn generalized — same config read, same
reason it has to be config rather than an env var: the switch belongs to
the session's client, and the client may be on another machine.
2026-08-31 21:11:11 -05:00
Brooklyn Nicholson 02458e67ad fix(auxiliary): accept dict and object messages in extract_content_or_reasoning
Compression and some OpenAI-compatible proxies hand us a dict-shaped
response or a bare message, not a ChatCompletion. Reuse the existing
helper instead of a second extractor, and bound an optional reasoning
fallback so a chain-of-thought dump cannot become the summary.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
teknium1 d10ef89ee5 fix(docker): keep forwarded secret values out of world-readable argv
docker run/exec argv previously carried -e KEY=VALUE pairs for every
forwarded/passthrough variable. On Linux /proc/<pid>/cmdline is
world-readable regardless of process owner, so every allowlisted secret
was visible to all local users via plain ps for the duration of every
terminal call.

Emit name-only -e KEY flags and supply values via the docker client
subprocess env instead: the docker CLI resolves valueless --env KEY from
its own environment (documented docker/podman behavior), moving secrets
from /proc/*/cmdline (0444) to /proc/*/environ (0400). Covers the docker
run container-start path, the recreation/recovery path, the init-seeding
exec path, and the per-command runtime exec path.

Reported by @sashalab. Fixes #96268
2026-08-31 11:29:17 -07:00
HexLab98 97f2a571b5 test(terminal): cover hung-wait bound, parent-tid interrupt, and cron inactivity watchdog
Pin that execute() returns at the wall-clock deadline when the inner wait
never returns, that /stop on the tool-worker tid still kills the subprocess,
that the cron inactivity helper fires while the caller thread is blocked,
and that ContextVars plus the activity callback reach the deadline worker.
2026-08-31 10:42:39 -07:00
Teknium 90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
zengzheqing 98e2f110cf fix(curator): restore complete skill packages on ledger rollback (#96962)
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.

The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.

Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
2026-08-31 10:08:13 -07:00
Teknium c8329384a9 test(send_message): drop duplicate buzz UUID target tests
Dispatch cluster (#99431) landed equivalent coverage first; the media
branch's copies shadowed them and tripped
test_no_shadowed_test_definitions.
2026-08-31 10:06:34 -07:00
EmpireOperating fcd34e57a2 fix(buzz): complete media-only delivery reporting 2026-08-31 10:06:34 -07:00
EmpireOperating 9c25704257 fix(buzz): verify live media delivery receipts 2026-08-31 10:06:34 -07:00
EmpireOperating 37c943997b fix(buzz): support media in standalone sends 2026-08-31 10:06:34 -07:00
Andrew Bagrin b7c59bda54 fix(cron): isolate per-execution working directories 2026-08-31 09:59:39 -07:00
kshitijk4poor 936b970e28 fix(browser): lightpanda review follow-ups for #99312
- lightpanda_engine_status: check use_real_profile before the cloud
  provider, matching browser_exec's actual resolution order (real-profile
  resolution runs before backend resolution), so /browser status and
  hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
  (find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
  _using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
2026-08-31 21:59:11 +05:30
Adrià Arrufat e3a85ae5a0 feat(browser): honor browser.engine=lightpanda in Browser Use mode
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.

- browser_use_cli: when the engine is lightpanda and nothing with higher
  precedence claimed the session, get a session from _get_session_info()
  and export its endpoint as BU_CDP_URL; the browser is private to the
  session key, so the own-tab preamble is skipped. The browser_exec
  description gains a Lightpanda header (text-first, new_tab once then
  goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
  --host 127.0.0.1 --port <free>` per session key (new
  tools/browser_lightpanda.py), reusing the session cache, inactivity
  reaper and atexit cleanup; a dead process is respawned on the next call;
  orphans from a crashed Hermes are reaped through per-process records in
  $HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
  reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
  (cloud_provider: local + engine: lightpanda; "Local Browser" resets the
  engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
  is shadowed, the reason.
2026-08-31 21:59:11 +05:30
blunkjamie-dev 1885a40ad3 fix(buzz): preserve literal mentions and exact UUID targets 2026-08-31 09:05:41 -07:00
kshitijk4poor 41dfa2d580 test: update chrome-fallback npx fixture for the get-url response shape
The fallback URL probe now uses 'get url' ({data:{url}}), not
'eval window.location.href' ({data:{result}}); the npx-resolution test's
mocked response predates the #81673 change.
2026-08-31 21:17:37 +05:30
Simone Marzola e355394839 fix(browser): isolate Lightpanda and Chrome fallback flags
Strip Chromium-only launch variables from Lightpanda commands, use a non-recursive Lightpanda URL lookup for Chrome fallback, and share sandbox argument injection with the temporary Chrome path.

Co-authored-by: forjd-hermes-bot <282037251+forjd-hermes-bot@users.noreply.github.com>
2026-08-31 21:17:37 +05:30
Teknium 26f178e5fa fix(terminal): gate the BUZZ_* terminal carve-out on actual Buzz agent context
Compose the two salvaged approaches (#78065 + #78511):

- Keep #78065's terminal-only scrub-path exemption (first-party prefix
  predicate in _make_run_env / _sanitize_subprocess_env, plain env values
  never scope-resolved, snapshot exclusion for cross-profile isolation,
  every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
  an import-time blocklist discard: the blocklist is shared by every
  scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
  execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
  process env (Buzz Desktop buzz-acp harness, #76243) OR the live
  session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
  concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
  session on a host that also runs a Buzz gateway does NOT get the
  signing key in its terminal children (maintainer triage note on
  #76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
  is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
  strips) and positive test (buzz session platform enables carve-out);
  docs updated accordingly.

Closes #78026, closes #76243.
2026-08-31 07:35:47 -07:00
Magnus Hedemark a4c6821218 fix(tools): pass BUZZ_* platform credentials to terminal children
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.

Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.

Follow-up hardening from review:

- First-party matches use the merged env value directly instead of
  _resolve_passthrough_value: under multiplex with no profile secret scope
  installed the resolver raised UnscopedSecretError (fail-closed) at call
  sites like the webhook-filter script runner, a regression where the
  script previously ran without the var. The vars are the process's own
  env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
  shared login-shell snapshot (_additional_profile_scoped_passthrough_names
  override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
  exclusion set (env_passthrough refuses blocklisted names), so without
  this a multiplexed gateway would dump profile A's key into
  hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
  would source it — a cross-profile nsec leak. The names are now excluded
  from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
  like ddgs, computer-use driver, user-script runners) that also receive
  first-party platform vars.

Fixes #78026

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-31 07:35:47 -07:00
Teknium 6fba07cb7e test(terminal): background spawn contract now includes owner_task_id 2026-08-31 07:28:18 -07:00
Teknium 5a4dbdec27 fix(delegation): subagent process notifications stay suppressed when the container key collapses
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.

Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).

Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
2026-08-31 07:28:18 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
itskaism 5ce8f71553 fix(delegation): report schema-invalid child results as failed, not completed
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.

Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.

Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
2026-08-31 01:02:42 -07:00
Teknium d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
Lime-oss-hash f9908e2ed6 fix(bot-mode): avoid inherited stdin on Windows
Query-file DM transports do not consume stdin. Use DEVNULL for both the initial attempt and policy-gated retry so Git Bash cannot pass an invalid pseudo-handle to Windows subprocess creation.
2026-08-30 22:20:06 -07:00
Teknium abdc7b952c test: adapt #97667 rejection-notice tests to the config-reading implementation on main 2026-08-30 22:19:10 -07:00
liuhao1024 8557e0a480 fix(tools): surface config-level model_not_found notices in delegation batch reports
When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).

Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
2026-08-30 22:19:10 -07:00
itskaism 1b6ea1a2c2 fix(delegation): report failed children as failed, not completed
A subagent whose loop gave up on a structured failure (e.g. "API call
failed after 3 retries: HTTP 524") returns that error message as
final_response together with completed=False / failed=True /
failure_reason. _run_single_child derived the batch-entry status from
the summary alone (`elif summary and not _empty_sentinel: status =
"completed"`), so the non-empty error text made the batch report show
the task as "✓ status=completed" — the `failed` flag was never
consulted anywhere in delegate_tool.py. Only the "(empty)" sentinel was
mapped to failed.

Fix, at the single status-determination choke point both the single-task
and batch paths share:

- `failed=True` on the child result now wins over a non-empty summary:
  status = "failed".
- The child's classified failure_reason (rate_limit / billing /
  server_error / ...) is propagated onto the batch entry so the parent
  can tell a quota wall from a real task error without parsing prose.
- exit_reason for a structured failure is "error" instead of falling
  through to "max_iterations" (which also wrongly set truncated=True).

Successful children (completed=True, no failed flag) are untouched —
covered by an explicit control test alongside the regression test,
which is red on the old code and green with the fix.
2026-08-30 22:19:10 -07:00
Teknium cd2bd16057 test: drop duplicate failed-flag regression test superseded by TestDelegateFailedChildStatus 2026-08-30 21:07:04 -07:00
David Metcalfe b4d5174385 fix(delegation): pin failure-status edge cases and document exit_reason enum
Per-finding verdicts from the cross-vendor review of fix/status-fix
(#97655/#97654):

[1] Flash NIT (real, cheap) — FIXED. Added test_error_without_failed_flag_
    marks_failed: an error string with the 'failed' key ABSENT (not False)
    must still be status=failed + exit_reason=error. The branch order
    (result.get('failed') or result.get('error')) already handles this; the
    test pins the error-alone path.

[2] GPT-OSS SHOULD-FIX — PINNED. Added test_empty_error_with_summary_is_
    completed: error='' is falsy so result.get('error') falls through to the
    summary-presence heuristic => status=completed. No code change; the
    existing branch is correct and the new test locks it in.

[3] GPT-OSS SHOULD-FIX — VERIFIED, NO CHANGE. Grepped every delegation
    exit_reason consumer:
      * tools/delegation_live_log.py finalize() prints exit_reason generically
        and only special-cases == 'max_iterations' for a readable suffix.
      * tools/process_registry.py derives truncated as
        (truncated or exit_reason == 'max_iterations') — gated, not exhaustive.
      * tools/async_delegation.py passes exit_reason through generically.
    The gateway/status.py, cron/scheduler.py and run_agent.py 'exit_reason'
    hits are a DIFFERENT field (turn_exit_reason / gateway exit reason), not
    the delegation result's exit_reason. No exhaustive if/elif over the enum
    missing an 'error' case, so nothing to add.

[4] GPT-OSS NIT — DONE. Enriched _run_single_child's docstring to enumerate
    status in {completed, interrupted, failed} and exit_reason in {completed,
    max_iterations, interrupted, error}, and added a compact enum comment at
    the result-entry construction. Verified the process_registry.py renderer
    comment (truncated <= exit_reason == 'max_iterations') still holds — the
    truncation flag is derived exactly that way, so no contradiction.

[5] GPT-OSS NIT — REJECTED. The proposed 'fallback for legacy dicts that
    explicitly set failed=False' is not adopted. No consumer produces a result
    dict with an explicit failed=False and no summary while relying on
    completed semantics: run_agent.py sets failed=True only on genuine failure
    and omits the key on success (no failed=False producer). Also, the
    proposed elif would reintroduce ambiguity (explicit failed=False + no
    summary => 'completed'?) and diverge from the conservative else => 'failed'.
    result.get('failed') is falsy for both explicit-False and absent, so no
    distinction exists to preserve; the else is the correct default.

Tests: 301 passed, 7 skipped (tests/tools -k 'delegate or process_registry').
TestDelegateFailedChildStatus: 6 passed.
2026-08-30 21:07:04 -07:00
David Metcalfe c05d04fffb feat(delegation): surface config-level model_not_found notice in delegation batch reports
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.

Closes #97654.
2026-08-30 21:07:04 -07:00
David Metcalfe ec02d5179a fix(delegation): report provider-failed subagents as failed, not completed/max_iterations
A provider-rejected child (e.g. HTTP 400 "<model> is not a valid model ID")
returns completed=False with failed=True + an error string as its terminal
final_response. _run_single_child keyed status on summary presence alone and
assumed completed=False meant iteration-budget exhaustion, so such a child
was reported status=completed + exit_reason=max_iterations, rendering the
false '"TRUNCATED: hit max_iterations"' banner.

Consult the structured failure fields (failed / error) before falling back to
the summary-presence heuristic, and derive exit_reason honestly: failure ->
'error', interrupted -> 'interrupted', completed -> 'completed', and only
genuine budget exhaustion (completed=False, no failure) -> 'max_iterations'.
The 'truncated' flag stays keyed on exit_reason == 'max_iterations', so it is
now correct automatically. The batch renderer needed no change (the error
field is already plumbed into the result entry for the parent).

Closes #97655
2026-08-30 21:07:04 -07:00
Teknium 5a134383fe fix: failed subagents now surface a clean error to the user (CLI + gateway)
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.

- delegate_tool: result.failed now forces status 'failed' (with the
  error carried on the entry); new shared format_subagent_failure_line()
  renders one clean human-readable line (traceback -> exception message,
  length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
  terminal failure status deliver that line via _deliver_platform_notice
  BEFORE all progress-queue gates; tool_progress_callback is now always
  attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
2026-08-30 20:40:14 -07:00
kshitijk4poor 70370e089c fix(skills): skill_view directory file_path + skill_manage categorized name resolution
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):

1. skill_view(name, file_path='references') returned a raw
   '[Errno 21] Is a directory' OS error. The local-skill branch gated
   on target_file.exists(); a directory passes exists(), fell through
   to read_text(), and raised. The plugin-skill sibling branch already
   gated on is_file() — this aligns the local branch so a directory
   request gets the same helpful not-found payload with
   available_files listing instead of an OS error.

2. skill_manage rejected categorized names ('category/skill-name')
   with 'not found in active profile'. _find_skill matched only the
   bare directory name, while skill_view's own ambiguity hint tells
   the caller to use exactly the categorized form — every call that
   followed the hint failed. _find_skill now also matches the full
   relative path of the skill dir, giving skill_manage resolution
   parity with skill_view across edit/patch/delete/write_file/
   remove_file.

Both fixes are covered by regression tests that fail on main.
2026-08-30 20:04:40 +05:30
Teknium dce2ecb8a9 test: cron ContextVar-masking test now uses a chat platform
test_explicit_blank_masks_leaked_cron_env_for_gateway_classification
used platform=api_server as an arbitrary gateway platform; api_server
is now intentionally excluded from gateway approval contexts
(unattended class). Switch to telegram — the test's subject is the
blank-cron-ContextVar masking, not platform policy.
2026-08-30 07:04:16 -07:00
Teknium ef71f2cad8 fix(approval): widen webhook exclusion to all unattended platforms, deny by default
Builds on liuhao1024's webhook exclusion (#37317): instead of falling
through to auto-approve, unattended programmatic platforms (webhook,
msgraph_webhook, api_server) now resolve approval decisions instantly
via approvals.unattended_mode (default deny), mirroring cron_mode.

- _UNATTENDED_APPROVAL_PLATFORMS set + _is_unattended_platform_approval_context()
- approvals.unattended_mode config key (deny | approve), default deny
- Deny branches in _run_approval_gate, check_all_command_guards (with
  tirith parity), and check_execute_code_guard (#87509 sibling site)
- Docs: security.md; config_defaults.py comment + default

Fixes #37284. Also fixes the api_server half of #87509.
2026-08-30 07:04:16 -07:00
liuhao1024 73f8fb74e0 fix(approval): exclude webhook sessions from gateway approval context
Webhook sessions trigger the gateway approval branch because
HERMES_SESSION_PLATFORM is set, but the webhook adapter has no
send_exec_approval and no way to receive /approve replies.  This
blocks the session for the full approval timeout (60-300 s) with
no human who can resolve it.

Fix: _is_gateway_approval_context() now returns False when the
session platform is 'webhook', falling through to the non-interactive
path (auto-approve with warning, or deny if cron).

Regression tests added for webhook, non-webhook gateway, cron, and
no-platform scenarios.
2026-08-30 07:04:16 -07:00
Teknium 8a8aa850f1 fix(delegate): inherit endpoint-scoped capability map only on the parent's exact route
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
2026-08-30 05:16:10 -07:00
Teknium d5fd2e9366 fix(browser): real-profile browsing runs headless — no focus-stealing window
Real-profile browsing (browser.use_real_profile) is meant to drive a COPY of
the user's profile headlessly in the background so they can keep working while
the agent acts on their behalf. Instead, on any host with a display it launched
the user's real browser binary HEADED, popping a window that grabbed focus on
every turn.

Root cause: the native launch (which bypasses agent-browser's own launcher to
avoid --use-mock-keychain dropping keychain-encrypted cookies) only added
--headless=new when Linux had no DISPLAY/WAYLAND_DISPLAY. On a normal desktop
the guard was false, so Chrome opened a visible window.

Fix: launch headless by default on every platform. Chrome's NEW headless mode
shares the profile's normal cookie store (unlike legacy --headless), and the
cookie drop we guard against comes from --use-mock-keychain, not from headless
— so real-profile auth still loads. Users who want to watch can opt in via the
existing browser.headed / AGENT_BROWSER_HEADED toggle (honored for real-profile
now, same as the rest of the browser stack); display-less hosts stay headless
regardless so the launch doesn't die at startup.

Live A/B on a real X seat: old argv mapped a Chrome window (focus steal),
--headless=new mapped zero windows while still exposing a working CDP port.
Updated the stale test that pinned 'never passes --headless' (a legacy-headless
premise) to positively assert the chrome launch is --headless=new by default.
2026-08-29 20:40:04 -07:00
liuhao1024 6b290b81d5 fix(tools): reduce false positives in exfil_curl/exfil_wget patterns
Anchor env var name matches with \b to avoid matching legitimate
env vars that contain KEY/TOKEN/API as substrings (e.g.,
$TRILLIUM_ETAPI_URL). The patterns now require KEY/TOKEN/SECRET/PASSWORD
to appear at the END of the env var name, reducing false positives on
common API-usage documentation in SOUL.md while still catching actual
exfiltration attempts.

Fixes #63977
2026-08-29 20:39:31 -07:00
Teknium d041ed7ab2 fix(skills-hub): reconcile salvaged install fixes with full-directory fetch
Follow-up reconciling the four cherry-picked contributor fixes with the
full-dir GitHubSource.fetch() that landed in #98246:

- Missing SKILL.md-linked support paths now warn and install without the
  file at all three sources (GitHub full-dir, GitHub fallback, UrlSource) —
  dangling links are prose over-matches or repo-only dev tools, not install
  blockers (#66760/#90081). A referenced path present in the tree as a
  SYMLINK stays a hard rejection.
- Extension requirement dropped from the glob/placeholder filter:
  references/LICENSE is a legitimate support file (82236's tests pin this).
  Truncated prose placeholders (references/type-<name>.md -> 'type-') are
  still rejected via the trailing-separator shape.
- percent-quoted Contents-API path (82236) merged with revision pinning
  (96336) in _fetch_file_bytes.
- Fixture typo fix: four cherry-picked test strings used '\---' where
  '\n---' was meant (DeprecationWarning + frontmatter never parsed).

Validation: 167/167 across tests/tools/{skills_hub,skills_guard,
skill_bundle_provenance} + tests/hermes_cli/test_skills_hub.py; live GitHub
fetches (impeccable 163 files rev-pinned; anthropics frontend-design).
2026-08-29 20:39:09 -07:00
liuhao1024 7e95b67ad8 fix(skills): review follow-up — revision pinning, canonicalization, case-fold guards
Addresses the review on #96336:

- Every GitHub byte fetch in an install (SKILL.md included) now carries
  the resolved tree's SHA as ?ref=, closing the pre-existing TOCTOU
  where /contents floated to the default-branch HEAD and bytes could
  come from a newer revision than the tree the paths were validated
  against. The tree is resolved first (idempotent + cached) so the pin
  covers the whole install.
- Same-dir link targets are canonicalized before validation: query/
  fragment stripped via urlsplit, percent-decoded, leading ./ removed —
  the same normalization the support-dir branch applies.
- A case-variant link to skill.md never ships as a bundle entry, and a
  case-folded collision among accepted siblings (A.md + a.md) drops the
  pair — both would overwrite/collide on case-insensitive filesystems.
2026-08-29 20:39:09 -07:00
liuhao1024 86dda3cd36 fix(skills): fetch explicitly linked same-directory siblings on install (#96310)
_referenced_support_paths only kept links whose first path segment was
one of the five support directories (references/templates/scripts/
assets/examples), so a SKILL.md linking same-directory siblings —
mattpocock/skills' domain-modeling links ./CONTEXT-FORMAT.md and
ADR-FORMAT.md — installed 'successfully' with those files silently
omitted: the bundle came out semantically incomplete with unresolved
links.

A second pass now collects same-directory markdown-link targets
(](./FILE.ext) or ](FILE.ext)) that name an extension-bearing file,
carry no internal slash, and are not SKILL.md itself; a leading '..'
is rejected fail-closed exactly like the support-dir traversal branch,
external URLs/anchors/mailto/site-absolute targets are left to their own
resolution, and every accepted name still runs the bundle path
validator. Unlinked siblings remain excluded — the fetch-minimization
contract is unchanged; only files the document explicitly links ship.
2026-08-29 20:39:09 -07:00
JUNZE f449374f0c skills hub: don't abort installs on missing referenced support files
A SKILL.md that mentions a support-path token which doesn't exist in the
repo (a prose glob like `references/type-*.md`, a placeholder truncated
by the regex, or a repo-only dev script) previously made GitHubSource
and UrlSource fetch fail with 'Could not fetch ... from any source',
even though every real file was present.

- _referenced_support_paths: skip glob/placeholder tokens and bare
  prefix matches that can't name an actual file.
- GitHubSource/UrlSource fetch: warn and continue when a referenced
  support file is missing or unfetchable, bundling what exists, instead
  of aborting the whole install. Non-regular tree entries (e.g. git
  symlinks) are still rejected.

Regression tests for both syntax filtering and the fetch loops.
2026-08-29 20:39:09 -07:00
Fidias Feliciano 972d812429 [verified] fix: ignore glob-shaped skill support paths 2026-08-29 20:39:09 -07:00
PRATHAMESH75 9d813c7352 test(cli): add install-flow E2E for skipped unreachable URL support file (#66760)
Covers the reviewer-requested end-to-end path: a referenced support file
that 404s is skipped, and the URL skill still installs through quarantine,
scan, install, and lock provenance with only the reachable files on disk
and in the lock file.
2026-08-29 20:39:09 -07:00
PRATHAMESH75 7d87bab502 fix(cli): skip unreachable support files instead of aborting URL skill install
UrlSource.fetch() aborted the whole SKILL.md URL install if any single
referenced support file (references/templates/scripts/assets) 404'd or
was otherwise unreachable, even when the SKILL.md itself and most
companion files were fine. Skip the missing file with a warning and
keep the ones that were fetched successfully.

Fixes #66760
2026-08-29 20:39:09 -07:00