Commit Graph

3587 Commits

Author SHA1 Message Date
Andrew Bagrin b7c59bda54 fix(cron): isolate per-execution working directories 2026-08-31 09:59:39 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
kshitijk4poor 936b970e28 fix(browser): lightpanda review follow-ups for #99312
- lightpanda_engine_status: check use_real_profile before the cloud
  provider, matching browser_exec's actual resolution order (real-profile
  resolution runs before backend resolution), so /browser status and
  hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
  (find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
  _using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
2026-08-31 21:59:11 +05:30
Adrià Arrufat e3a85ae5a0 feat(browser): honor browser.engine=lightpanda in Browser Use mode
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.

- browser_use_cli: when the engine is lightpanda and nothing with higher
  precedence claimed the session, get a session from _get_session_info()
  and export its endpoint as BU_CDP_URL; the browser is private to the
  session key, so the own-tab preamble is skipped. The browser_exec
  description gains a Lightpanda header (text-first, new_tab once then
  goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
  --host 127.0.0.1 --port <free>` per session key (new
  tools/browser_lightpanda.py), reusing the session cache, inactivity
  reaper and atexit cleanup; a dead process is respawned on the next call;
  orphans from a crashed Hermes are reaped through per-process records in
  $HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
  reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
  (cloud_provider: local + engine: lightpanda; "Local Browser" resets the
  engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
  is shadowed, the reason.
2026-08-31 21:59:11 +05:30
blunkjamie-dev 1885a40ad3 fix(buzz): preserve literal mentions and exact UUID targets 2026-08-31 09:05:41 -07:00
kshitijk4poor a9eb06a629 fix(browser): log when Chromium-only env flags are stripped for Lightpanda
Review follow-up for the #81673 salvage: the strip was silent, which
would confuse a user whose AGENT_BROWSER_ARGS applies to Chrome commands
but vanishes on Lightpanda ones.
2026-08-31 21:17:37 +05:30
Simone Marzola e355394839 fix(browser): isolate Lightpanda and Chrome fallback flags
Strip Chromium-only launch variables from Lightpanda commands, use a non-recursive Lightpanda URL lookup for Chrome fallback, and share sandbox argument injection with the temporary Chrome path.

Co-authored-by: forjd-hermes-bot <282037251+forjd-hermes-bot@users.noreply.github.com>
2026-08-31 21:17:37 +05:30
Teknium 26f178e5fa fix(terminal): gate the BUZZ_* terminal carve-out on actual Buzz agent context
Compose the two salvaged approaches (#78065 + #78511):

- Keep #78065's terminal-only scrub-path exemption (first-party prefix
  predicate in _make_run_env / _sanitize_subprocess_env, plain env values
  never scope-resolved, snapshot exclusion for cross-profile isolation,
  every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
  an import-time blocklist discard: the blocklist is shared by every
  scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
  execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
  process env (Buzz Desktop buzz-acp harness, #76243) OR the live
  session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
  concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
  session on a host that also runs a Buzz gateway does NOT get the
  signing key in its terminal children (maintainer triage note on
  #76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
  is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
  strips) and positive test (buzz session platform enables carve-out);
  docs updated accordingly.

Closes #78026, closes #76243.
2026-08-31 07:35:47 -07:00
vatevstoil a6192c7ad8 fix(terminal): let Buzz-managed agents inherit BUZZ CLI credentials in terminal env
The bundled buzz platform plugin registers BUZZ_PRIVATE_KEY / BUZZ_RELAY_URL /
BUZZ_AUTH_TAG as messaging credentials, so the terminal env blocklist strips
them unconditionally. But when Hermes runs as a Buzz-ACP managed agent, the
buzz-acp harness hands exactly those vars to the agent process ON PURPOSE:
the platform's reply-delivery contract is the model driving the buzz CLI
(block/buzz#2698 — final session text is not delivered to the channel).
Stripping them made every reply die with 'auth error: BUZZ_PRIVATE_KEY is
required' and the channel stayed empty while activity panels streamed the
answer. Mirror the CLAUDE_CODE_OAUTH_TOKEN exception, gated on
BUZZ_MANAGED_AGENT (set only by the Buzz Desktop harness), so gateway/CLI/
kanban contexts keep stripping the path-3 platform secret.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 065e41c0cceb3a373b5025449301c8e930d8a8f0)
2026-08-31 07:35:47 -07:00
Magnus Hedemark a4c6821218 fix(tools): pass BUZZ_* platform credentials to terminal children
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.

Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.

Follow-up hardening from review:

- First-party matches use the merged env value directly instead of
  _resolve_passthrough_value: under multiplex with no profile secret scope
  installed the resolver raised UnscopedSecretError (fail-closed) at call
  sites like the webhook-filter script runner, a regression where the
  script previously ran without the var. The vars are the process's own
  env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
  shared login-shell snapshot (_additional_profile_scoped_passthrough_names
  override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
  exclusion set (env_passthrough refuses blocklisted names), so without
  this a multiplexed gateway would dump profile A's key into
  hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
  would source it — a cross-profile nsec leak. The names are now excluded
  from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
  like ddgs, computer-use driver, user-script runners) that also receive
  first-party platform vars.

Fixes #78026

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-31 07:35:47 -07:00
Teknium 5a4dbdec27 fix(delegation): subagent process notifications stay suppressed when the container key collapses
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.

Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).

Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
2026-08-31 07:28:18 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
itskaism 5ce8f71553 fix(delegation): report schema-invalid child results as failed, not completed
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.

Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.

Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
2026-08-31 01:02:42 -07:00
Teknium d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
Lime-oss-hash f9908e2ed6 fix(bot-mode): avoid inherited stdin on Windows
Query-file DM transports do not consume stdin. Use DEVNULL for both the initial attempt and policy-gated retry so Git Bash cannot pass an invalid pseudo-handle to Windows subprocess creation.
2026-08-30 22:20:06 -07:00
liuhao1024 8557e0a480 fix(tools): surface config-level model_not_found notices in delegation batch reports
When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).

Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
2026-08-30 22:19:10 -07:00
itskaism 1b6ea1a2c2 fix(delegation): report failed children as failed, not completed
A subagent whose loop gave up on a structured failure (e.g. "API call
failed after 3 retries: HTTP 524") returns that error message as
final_response together with completed=False / failed=True /
failure_reason. _run_single_child derived the batch-entry status from
the summary alone (`elif summary and not _empty_sentinel: status =
"completed"`), so the non-empty error text made the batch report show
the task as "✓ status=completed" — the `failed` flag was never
consulted anywhere in delegate_tool.py. Only the "(empty)" sentinel was
mapped to failed.

Fix, at the single status-determination choke point both the single-task
and batch paths share:

- `failed=True` on the child result now wins over a non-empty summary:
  status = "failed".
- The child's classified failure_reason (rate_limit / billing /
  server_error / ...) is propagated onto the batch entry so the parent
  can tell a quota wall from a real task error without parsing prose.
- exit_reason for a structured failure is "error" instead of falling
  through to "max_iterations" (which also wrongly set truncated=True).

Successful children (completed=True, no failed flag) are untouched —
covered by an explicit control test alongside the regression test,
which is red on the old code and green with the fix.
2026-08-30 22:19:10 -07:00
David Metcalfe b4d5174385 fix(delegation): pin failure-status edge cases and document exit_reason enum
Per-finding verdicts from the cross-vendor review of fix/status-fix
(#97655/#97654):

[1] Flash NIT (real, cheap) — FIXED. Added test_error_without_failed_flag_
    marks_failed: an error string with the 'failed' key ABSENT (not False)
    must still be status=failed + exit_reason=error. The branch order
    (result.get('failed') or result.get('error')) already handles this; the
    test pins the error-alone path.

[2] GPT-OSS SHOULD-FIX — PINNED. Added test_empty_error_with_summary_is_
    completed: error='' is falsy so result.get('error') falls through to the
    summary-presence heuristic => status=completed. No code change; the
    existing branch is correct and the new test locks it in.

[3] GPT-OSS SHOULD-FIX — VERIFIED, NO CHANGE. Grepped every delegation
    exit_reason consumer:
      * tools/delegation_live_log.py finalize() prints exit_reason generically
        and only special-cases == 'max_iterations' for a readable suffix.
      * tools/process_registry.py derives truncated as
        (truncated or exit_reason == 'max_iterations') — gated, not exhaustive.
      * tools/async_delegation.py passes exit_reason through generically.
    The gateway/status.py, cron/scheduler.py and run_agent.py 'exit_reason'
    hits are a DIFFERENT field (turn_exit_reason / gateway exit reason), not
    the delegation result's exit_reason. No exhaustive if/elif over the enum
    missing an 'error' case, so nothing to add.

[4] GPT-OSS NIT — DONE. Enriched _run_single_child's docstring to enumerate
    status in {completed, interrupted, failed} and exit_reason in {completed,
    max_iterations, interrupted, error}, and added a compact enum comment at
    the result-entry construction. Verified the process_registry.py renderer
    comment (truncated <= exit_reason == 'max_iterations') still holds — the
    truncation flag is derived exactly that way, so no contradiction.

[5] GPT-OSS NIT — REJECTED. The proposed 'fallback for legacy dicts that
    explicitly set failed=False' is not adopted. No consumer produces a result
    dict with an explicit failed=False and no summary while relying on
    completed semantics: run_agent.py sets failed=True only on genuine failure
    and omits the key on success (no failed=False producer). Also, the
    proposed elif would reintroduce ambiguity (explicit failed=False + no
    summary => 'completed'?) and diverge from the conservative else => 'failed'.
    result.get('failed') is falsy for both explicit-False and absent, so no
    distinction exists to preserve; the else is the correct default.

Tests: 301 passed, 7 skipped (tests/tools -k 'delegate or process_registry').
TestDelegateFailedChildStatus: 6 passed.
2026-08-30 21:07:04 -07:00
David Metcalfe c05d04fffb feat(delegation): surface config-level model_not_found notice in delegation batch reports
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.

Closes #97654.
2026-08-30 21:07:04 -07:00
David Metcalfe ec02d5179a fix(delegation): report provider-failed subagents as failed, not completed/max_iterations
A provider-rejected child (e.g. HTTP 400 "<model> is not a valid model ID")
returns completed=False with failed=True + an error string as its terminal
final_response. _run_single_child keyed status on summary presence alone and
assumed completed=False meant iteration-budget exhaustion, so such a child
was reported status=completed + exit_reason=max_iterations, rendering the
false '"TRUNCATED: hit max_iterations"' banner.

Consult the structured failure fields (failed / error) before falling back to
the summary-presence heuristic, and derive exit_reason honestly: failure ->
'error', interrupted -> 'interrupted', completed -> 'completed', and only
genuine budget exhaustion (completed=False, no failure) -> 'max_iterations'.
The 'truncated' flag stays keyed on exit_reason == 'max_iterations', so it is
now correct automatically. The batch renderer needed no change (the error
field is already plumbed into the result entry for the parent).

Closes #97655
2026-08-30 21:07:04 -07:00
Teknium 5a134383fe fix: failed subagents now surface a clean error to the user (CLI + gateway)
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.

- delegate_tool: result.failed now forces status 'failed' (with the
  error carried on the entry); new shared format_subagent_failure_line()
  renders one clean human-readable line (traceback -> exception message,
  length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
  terminal failure status deliver that line via _deliver_platform_notice
  BEFORE all progress-queue gates; tool_progress_callback is now always
  attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
2026-08-30 20:40:14 -07:00
Teknium 6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
kshitijk4poor 5cc1369fa2 refactor(skills): clarify _find_skill docstring, narrow _local_root except
Simplify-code pass findings:
- The docstring claimed 'Matching bare-name suffix' but the fast path
  matches the exact directory name (parent.name) — reworded to say
  what the code actually does.
- _local_root() swallowed every Exception silently; narrowed to
  OSError (what resolve() raises) with a logger.debug breadcrumb so a
  recurring resolve failure is diagnosable instead of degrading every
  categorized lookup to a silent not-found.
2026-08-30 20:04:40 +05:30
kshitijk4poor 70220e529a refactor(skills): short-circuit bare-name match before resolve machinery
Review feedback (kokhlo): the categorized-name match ran
resolve().relative_to() for every SKILL.md in the walk even when the
bare-name branch already matched — 50+ resolve calls per invocation on
a bare-name lookup in a large profile.

Restructure so the bare directory-name check stays first and the
resolve/relative_to machinery only runs when the lookup name actually
contains a path separator. The skills root is resolved once, lazily,
only when a categorized lookup happens at all. Also compare the
relative path via as_posix() so 'category/skill' lookups work on
Windows, where str(Path) renders backslashes.
2026-08-30 20:04:40 +05:30
kshitijk4poor 70370e089c fix(skills): skill_view directory file_path + skill_manage categorized name resolution
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):

1. skill_view(name, file_path='references') returned a raw
   '[Errno 21] Is a directory' OS error. The local-skill branch gated
   on target_file.exists(); a directory passes exists(), fell through
   to read_text(), and raised. The plugin-skill sibling branch already
   gated on is_file() — this aligns the local branch so a directory
   request gets the same helpful not-found payload with
   available_files listing instead of an OS error.

2. skill_manage rejected categorized names ('category/skill-name')
   with 'not found in active profile'. _find_skill matched only the
   bare directory name, while skill_view's own ambiguity hint tells
   the caller to use exactly the categorized form — every call that
   followed the hint failed. _find_skill now also matches the full
   relative path of the skill dir, giving skill_manage resolution
   parity with skill_view across edit/patch/delete/write_file/
   remove_file.

Both fixes are covered by regression tests that fail on main.
2026-08-30 20:04:40 +05:30
Teknium ef71f2cad8 fix(approval): widen webhook exclusion to all unattended platforms, deny by default
Builds on liuhao1024's webhook exclusion (#37317): instead of falling
through to auto-approve, unattended programmatic platforms (webhook,
msgraph_webhook, api_server) now resolve approval decisions instantly
via approvals.unattended_mode (default deny), mirroring cron_mode.

- _UNATTENDED_APPROVAL_PLATFORMS set + _is_unattended_platform_approval_context()
- approvals.unattended_mode config key (deny | approve), default deny
- Deny branches in _run_approval_gate, check_all_command_guards (with
  tirith parity), and check_execute_code_guard (#87509 sibling site)
- Docs: security.md; config_defaults.py comment + default

Fixes #37284. Also fixes the api_server half of #87509.
2026-08-30 07:04:16 -07:00
liuhao1024 73f8fb74e0 fix(approval): exclude webhook sessions from gateway approval context
Webhook sessions trigger the gateway approval branch because
HERMES_SESSION_PLATFORM is set, but the webhook adapter has no
send_exec_approval and no way to receive /approve replies.  This
blocks the session for the full approval timeout (60-300 s) with
no human who can resolve it.

Fix: _is_gateway_approval_context() now returns False when the
session platform is 'webhook', falling through to the non-interactive
path (auto-approve with warning, or deny if cron).

Regression tests added for webhook, non-webhook gateway, cron, and
no-platform scenarios.
2026-08-30 07:04:16 -07:00
Teknium 8a8aa850f1 fix(delegate): inherit endpoint-scoped capability map only on the parent's exact route
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
2026-08-30 05:16:10 -07:00
Teknium d5fd2e9366 fix(browser): real-profile browsing runs headless — no focus-stealing window
Real-profile browsing (browser.use_real_profile) is meant to drive a COPY of
the user's profile headlessly in the background so they can keep working while
the agent acts on their behalf. Instead, on any host with a display it launched
the user's real browser binary HEADED, popping a window that grabbed focus on
every turn.

Root cause: the native launch (which bypasses agent-browser's own launcher to
avoid --use-mock-keychain dropping keychain-encrypted cookies) only added
--headless=new when Linux had no DISPLAY/WAYLAND_DISPLAY. On a normal desktop
the guard was false, so Chrome opened a visible window.

Fix: launch headless by default on every platform. Chrome's NEW headless mode
shares the profile's normal cookie store (unlike legacy --headless), and the
cookie drop we guard against comes from --use-mock-keychain, not from headless
— so real-profile auth still loads. Users who want to watch can opt in via the
existing browser.headed / AGENT_BROWSER_HEADED toggle (honored for real-profile
now, same as the rest of the browser stack); display-less hosts stay headless
regardless so the launch doesn't die at startup.

Live A/B on a real X seat: old argv mapped a Chrome window (focus steal),
--headless=new mapped zero windows while still exposing a working CDP port.
Updated the stale test that pinned 'never passes --headless' (a legacy-headless
premise) to positively assert the chrome launch is --headless=new by default.
2026-08-29 20:40:04 -07:00
Teknium 21e52c1fd4 fix(security): widen the exfil substring-suffix fix to the skills-guard sibling patterns
Same bug class as the salvaged terminal-scanner fix: skills_guard's
env_exfil_curl/wget/fetch used unanchored \w*(KEY|TOKEN|...|API)
alternations, so any var with API/KEY/TOKEN mid-name
($TRILLIUM_ETAPI_URL) scored a critical exfiltration finding. Applied
the same \b anchor + plural tolerance, dropped mid-name API (every real
secret it caught already ends in KEY/TOKEN), kept CREDENTIAL, and kept
the loopback exemption from #98246. httpx/requests patterns unchanged —
their (KEY|TOKEN|...) alternation is unanchored-by-design against
argument text, not var-name suffixes.
2026-08-29 20:39:31 -07:00
liuhao1024 6b290b81d5 fix(tools): reduce false positives in exfil_curl/exfil_wget patterns
Anchor env var name matches with \b to avoid matching legitimate
env vars that contain KEY/TOKEN/API as substrings (e.g.,
$TRILLIUM_ETAPI_URL). The patterns now require KEY/TOKEN/SECRET/PASSWORD
to appear at the END of the env var name, reducing false positives on
common API-usage documentation in SOUL.md while still catching actual
exfiltration attempts.

Fixes #63977
2026-08-29 20:39:31 -07:00
Teknium d041ed7ab2 fix(skills-hub): reconcile salvaged install fixes with full-directory fetch
Follow-up reconciling the four cherry-picked contributor fixes with the
full-dir GitHubSource.fetch() that landed in #98246:

- Missing SKILL.md-linked support paths now warn and install without the
  file at all three sources (GitHub full-dir, GitHub fallback, UrlSource) —
  dangling links are prose over-matches or repo-only dev tools, not install
  blockers (#66760/#90081). A referenced path present in the tree as a
  SYMLINK stays a hard rejection.
- Extension requirement dropped from the glob/placeholder filter:
  references/LICENSE is a legitimate support file (82236's tests pin this).
  Truncated prose placeholders (references/type-<name>.md -> 'type-') are
  still rejected via the trailing-separator shape.
- percent-quoted Contents-API path (82236) merged with revision pinning
  (96336) in _fetch_file_bytes.
- Fixture typo fix: four cherry-picked test strings used '\---' where
  '\n---' was meant (DeprecationWarning + frontmatter never parsed).

Validation: 167/167 across tests/tools/{skills_hub,skills_guard,
skill_bundle_provenance} + tests/hermes_cli/test_skills_hub.py; live GitHub
fetches (impeccable 163 files rev-pinned; anthropics frontend-design).
2026-08-29 20:39:09 -07:00
liuhao1024 7e95b67ad8 fix(skills): review follow-up — revision pinning, canonicalization, case-fold guards
Addresses the review on #96336:

- Every GitHub byte fetch in an install (SKILL.md included) now carries
  the resolved tree's SHA as ?ref=, closing the pre-existing TOCTOU
  where /contents floated to the default-branch HEAD and bytes could
  come from a newer revision than the tree the paths were validated
  against. The tree is resolved first (idempotent + cached) so the pin
  covers the whole install.
- Same-dir link targets are canonicalized before validation: query/
  fragment stripped via urlsplit, percent-decoded, leading ./ removed —
  the same normalization the support-dir branch applies.
- A case-variant link to skill.md never ships as a bundle entry, and a
  case-folded collision among accepted siblings (A.md + a.md) drops the
  pair — both would overwrite/collide on case-insensitive filesystems.
2026-08-29 20:39:09 -07:00
liuhao1024 86dda3cd36 fix(skills): fetch explicitly linked same-directory siblings on install (#96310)
_referenced_support_paths only kept links whose first path segment was
one of the five support directories (references/templates/scripts/
assets/examples), so a SKILL.md linking same-directory siblings —
mattpocock/skills' domain-modeling links ./CONTEXT-FORMAT.md and
ADR-FORMAT.md — installed 'successfully' with those files silently
omitted: the bundle came out semantically incomplete with unresolved
links.

A second pass now collects same-directory markdown-link targets
(](./FILE.ext) or ](FILE.ext)) that name an extension-bearing file,
carry no internal slash, and are not SKILL.md itself; a leading '..'
is rejected fail-closed exactly like the support-dir traversal branch,
external URLs/anchors/mailto/site-absolute targets are left to their own
resolution, and every accepted name still runs the bundle path
validator. Unlinked siblings remain excluded — the fetch-minimization
contract is unchanged; only files the document explicitly links ship.
2026-08-29 20:39:09 -07:00
JUNZE f449374f0c skills hub: don't abort installs on missing referenced support files
A SKILL.md that mentions a support-path token which doesn't exist in the
repo (a prose glob like `references/type-*.md`, a placeholder truncated
by the regex, or a repo-only dev script) previously made GitHubSource
and UrlSource fetch fail with 'Could not fetch ... from any source',
even though every real file was present.

- _referenced_support_paths: skip glob/placeholder tokens and bare
  prefix matches that can't name an actual file.
- GitHubSource/UrlSource fetch: warn and continue when a referenced
  support file is missing or unfetchable, bundling what exists, instead
  of aborting the whole install. Non-regular tree entries (e.g. git
  symlinks) are still rejected.

Regression tests for both syntax filtering and the fetch loops.
2026-08-29 20:39:09 -07:00
Fidias Feliciano 972d812429 [verified] fix: ignore glob-shaped skill support paths 2026-08-29 20:39:09 -07:00
PRATHAMESH75 7d87bab502 fix(cli): skip unreachable support files instead of aborting URL skill install
UrlSource.fetch() aborted the whole SKILL.md URL install if any single
referenced support file (references/templates/scripts/assets) 404'd or
was otherwise unreachable, even when the SKILL.md itself and most
companion files were fine. Skip the missing file with a warning and
keep the ones that were fetched successfully.

Fixes #66760
2026-08-29 20:39:09 -07:00
Teknium 52e5e7c034 fix(cron): coerce string repeat values on the UPDATE path too
The create-path coercion (salvaged from #78928) left update_job and the
cronjob tool's update handler comparing/storing raw strings: repeat=
'forever' via update raised TypeError in the tool path and stored the raw
string via update_job, breaking the next mark_job_run ('str' has no
.get). Extract normalize_repeat_value as the shared chokepoint (shape
from #77366 by @andrexibiza, with garbage-rejection semantics) and route
create_job, update_job, and the tool update handler through it.
Completed counters are preserved across repeat updates.

Class: #66824 #64520 #7142 #71987 #95706, update half of #77366.
2026-08-29 20:08:53 -07:00
andrexibiza e8bab87ba7 fix(cron): bare durations are recurring intervals; coerce repeat string forms
Contract bug (2026-08-04): the cronjob tool schema documents '30m' as
'(every 30 minutes)' — recurring — but parse_schedule returned
kind='once' for bare durations, silently creating a one-shot job for a
recurring request (agent passed '30m' for 'every 30 min', job ran once
and died). Bare durations ('30m','2h','1d') now parse as recurring
intervals matching the documented contract; explicit one-shot by
duration is 'in 30m'/'in 2h' (fires once that far from now). ISO
timestamps stay one-shot.

Also fixes the sibling repeat-coercion class (#66824/#64520/#7142):
repeat='forever'/'once'/'N' strings now coerce in create_job instead of
raising "'<=' not supported between instances of 'str' and 'int'".

Tool description rewritten to teach the corrected contract and steer
relative requests to 'in Nm' (no more hand-computed ISO timestamps).
Supersedes the doc-only direction of #53739 while keeping its goal
(relative one-shots must be expressible) via the 'in X' form.

Signed-off-by: andrexibiza <84248988+andrexibiza@users.noreply.github.com>
2026-08-29 19:19:40 -07:00
Teknium 45d9c33d85 feat(skills-hub): impeccable joins the optional-skills catalog, content pulled live from upstream
hermes skills install impeccable (and the docs-page install button) now
installs the impeccable frontend-design skill as an official optional-skills
entry. The local optional-skills/creative/impeccable/ dir is a catalog STUB:
its frontmatter declares metadata.hermes.upstream (repo + path), and
OptionalSkillSource.fetch() pulls the real 163-file bundle live from
pbakaus/impeccable:.hermes/skills/impeccable — the Hermes-native bundle
upstream maintains and verifies. Nothing vendored, never stale.

New mechanism (generic, not impeccable-specific):
- OptionalSkillSource._upstream_pointer(): parses/validates the upstream
  pointer (owner/name repo, clean relative path, traversal rejected).
- _fetch_from_upstream(): delegates to GitHubSource.fetch(), relabels the
  bundle official/<rel> at trust 'trusted' (curated endorsement, but
  third-party content — dangerous scan verdicts still block).
- The live-repo fallback path redirects stubs the same way, so stale local
  checkouts behave identically.

Three real gaps this surfaced, all fixed:
- GitHubSource.fetch() only downloaded SKILL.md plus paths linked from a
  canonical support dir (references/, scripts/, ...). Impeccable keeps its
  playbooks under reference/ (singular) and links scripts only from
  reference files, so fetch shipped 1 of 163 files. fetch() now downloads
  the full skill directory via the git tree (same approach as the
  optional-skills live fetch), still rejecting symlinks/hidden/unsafe paths
  and still failing on a missing SKILL.md-linked references/ path.
- The five env_exfil_* scanner patterns flagged loopback requests as
  critical exfiltration: impeccable's live mode polls
  http://localhost:PORT/status?token=TOKEN and scored two CRITICALs.
  Scheme-anchored loopback exemption added; evil.com/?u=localhost decoys
  still fire (10-case regex matrix in tests).
- unified_search() truncated to limit before ranking, so official catalog
  entries got crowded out by skills.sh mirrors and bare-name installs
  stalled on an ambiguity table. Results now stable-sort by trust rank
  before the cut, and _resolve_short_name prefers a sole official exact
  match over community mirrors.

Also fixes pre-existing test pollution: TestInstallPathSafety's fixture
monkeypatched the PEP 562 dynamic SKILLS_DIR, permanently shadowing dynamic
resolution and breaking the served_repo E2E tests in any combined run
(reproducible on main).

Validation: live E2E do_install("impeccable") against real GitHub —
resolves to official/creative/impeccable, verdict SAFE, 163 files on disk,
skill loads, /impeccable slash command registers. 128/128 targeted tests;
full-dir fetch test sabotage-verified. Docs: optional-skills catalog row,
generated skill page, sidebar.
2026-08-29 19:15:32 -07:00
Teknium 0582ae76d0 fix(cron): repair mangled schedule field in cronjob tool schema 2026-08-29 19:14:51 -07:00
Drexuxux ab9d85287d fix(cron): accept documented "every <weekday> <time>" schedules
parse_schedule's "every " branch passed everything after the prefix
straight to parse_duration(), so documented natural-language schedules
like "every monday 9am" and "every day at 9am" (AGENTS.md, SKILL.md,
cron docs) were rejected with "Invalid duration". Convert weekday and
daily/weekday/weekend phrases to cron expressions before the duration
fallback; "every 30m"/"every 2h" interval parsing is unchanged.
2026-08-29 19:14:51 -07:00
Teknium 0f3fcacd3f feat: /plan graduates from bundled skill to built-in command on every surface
The bundled plan skill's auto-generated slash command fell off the capped
Telegram/Discord command menus for most installs (skills are the only tier
trimmed at the platform caps, alphabetically — 'plan' sat past the cutoff at
index 57 of 82 bundled skills). Converting it to a first-class CommandDef
gives it a guaranteed core-tier menu slot on every platform.

- agent/plan_prompt.py: build_plan_prompt() — plan-mode rules + authoring
  craft distilled from the retired skill; prompt-injection pattern like
  /learn and /init (no engine, no model-tool footprint, cache-safe).
- CLI: _handle_plan_command mixin handler (pending-input injection).
- Gateway: /plan branch rewrites event.text and falls through (role
  alternation preserved).
- TUI: command.dispatch branch ('plan' was already in
  _PENDING_INPUT_COMMANDS).
- Removed skills/software-development/plan/ + docs pages (EN + zh-Hans),
  catalog rows, sidebar entry, related_skills references.
- PROTECTED_BUILTIN_SKILLS is now empty (mechanism kept); dependent
  curator/usage tests moved to monkeypatched sentinels.

Salvages #67292 by @webtecnica (credit: first /plan command submission,
issue #67264); reworked from inline planning prompt to the prompt-injection
pattern with workspace-saved plans. Closes #67264, closes #36821 (empty
/plan infers task from conversation context).
2026-08-29 19:14:15 -07:00
Teknium bacb90fe20 feat(delegation): honor delegation.request_overrides on all three resolution branches with explicit-over-runtime merge precedence
Completes the #90953 salvage on post-#98237 main:

- New _merge_request_overrides helper defines the precedence contract:
  explicit delegation.request_overrides merges OVER runtime/parent-derived
  overrides — explicit top-level keys win; extra_body is deep-merged one
  level so runtime extra_body keys survive unless redefined. Inputs are
  copy.deepcopy'd so transport-side mutation can't leak into config or the
  provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
  provider-alongside-base_url runtime overrides instead of being a separate
  return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
  delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
  (previously only when override_provider was set), enabling the inherit
  branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
  proofs, explicit-over-runtime precedence on the provider-alongside-base_url
  path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
  the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
2026-08-29 19:13:23 -07:00
fabiantax d3bfd2e9b1 feat(delegation): forward delegation.request_overrides on direct-endpoint branch
The direct base_url branch of _resolve_delegation_credentials returned no
request_overrides key, so a direct OpenRouter delegation (provider=custom,
base_url=openrouter.ai/api/v1) could not pass routing hints to its children.
The named-provider branch already forwards runtime request_overrides; this
gives the direct branch the same contract, honouring delegation.request_overrides
from config (dict → forwarded, anything else → None).

Primary use: extra_body.provider = {"sort": "throughput"} so delegation
children route to the fastest OpenRouter provider for their model, per the
fab-swarm throughput work (#901).
2026-08-29 19:13:23 -07:00
Teknium 7b3c9f86ed Merge origin/main — resolve todo-state seam onto the todo_list rename (accept legacy alias) 2026-08-29 18:55:29 -07:00
itsflownium 393af4a310 fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
  returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
  bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
  snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions

The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
2026-08-29 18:40:51 -07:00
Zane Chee b2e24b986f fix(computer-use): stop launching retired browser-grant runtimes 2026-08-29 18:35:17 -07:00
Teknium f8546c2eac fix(browser): real-profile follow-ups — reap launched Chrome, headless display-less Linux, register real_profile_pin default + docs
- _terminate_real_profile_chrome(): directly-launched real browsers are ours
  to reap (agent-browser only attaches); wired into the atexit emergency
  cleanup and both launch-failure paths so orphaned Chrome processes can't
  accumulate.
- Display-less Linux gate: append --headless=new (shares the profile's normal
  cookie store, unlike legacy headless) so the direct-launch path doesn't
  regress servers without DISPLAY/WAYLAND_DISPLAY.
- Register browser.real_profile_pin in config_defaults.py and document the
  new launch model + pin in website/docs/user-guide/features/browser.md.
- Drop unused tempfile import from the cherry-picked commit.
2026-08-29 18:35:12 -07:00
Jason Pollak 8e746668ba fix(browser): real-profile browsing on macOS - launch real binary, kill sqlite hang, normalize profile copy
Four fixes for real-profile browsing (browser.use_real_profile), found and
verified end-to-end on macOS with a live Chrome:

1. _copy_auth_file: sqlite3.connect('file:...?mode=ro') on a live Chrome
   auth DB can block indefinitely inside lock negotiation - the busy
   timeout never fires, so the 'fail fast' path hangs the launch forever.
   Try immutable=1 first (reads instantly, correct for a committed
   snapshot of a file another process owns); mode=ro stays as fallback.

2. Launch shape: agent-browser's own launch injects --use-mock-keychain /
   --password-store=basic / --headless=new. On macOS the mock keychain
   makes Chrome treat every keychain-encrypted cookie as undecryptable
   and drop it - the copied profile launches signed out (~3 anonymous
   cookies instead of the full jar). Launch the user's real browser
   binary directly on the copy (no mock-keychain switches), wait for
   DevToolsActivePort, then attach agent-browser via --cdp.

3. Snapshot copy: Local State was copied verbatim, still naming the
   SOURCE profile (last_used='Profile 2', info_cache listing several)
   while the copy only contains Default. Chrome opens the missing profile
   dir and starts signed out. Normalize the copy's Local State to
   Default-only.

4. CDP resolution: the agent-browser daemon may report the endpoint of a
   browser IT spawned (throwaway temp profile) instead of the real
   browser we launched on the copy. Trust the port our browser wrote to
   DevToolsActivePort.

Also adds browser.real_profile_pin (optional): pin which source Chromium
profile dir is snapshotted instead of following profile.last_used - on a
machine with a work profile and a personal one, last-used roulette can
silently give the agent the wrong identity. A pin naming a missing dir
fails closed (signed out) rather than falling back to last_used.

Tests: 4 new pin tests + 3 launch tests reshaped to the direct-launch
contract (Popen the real binary, agent-browser attaches). 77 passing.
2026-08-29 18:35:12 -07:00