Commit Graph

2660 Commits

Author SHA1 Message Date
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
Teknium 758114bb8d fix(cron): manual run reports delivery_failed as a failed run; docs for the distinct status
A manual cronjob(action='run') derived success from last_status == 'ok'
and read the error from last_error — so a run that now records
delivery_failed came back as success=False with error=None, an unexplained
failure. Surface last_delivery_error as the error in that case (the
#84006 direction, re-applied on the delivery_failed status), and pin the
manual-run completion summary to say 'Result: FAILED' over an undelivered
run. Document the status in the cron user guide.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-02 00:52:58 -07:00
赵桂雄 2f58cbfa7f fix(cron): adapt delivery-notice tests to the return_job claim API
Main grew claim_job_for_fire(job_id, return_job=True) — a claimed
snapshot dict instead of a bool — while this branch sat on an older
base. The merge-ref CI ran the hybrid: the wiring tests still mocked
return_value=True, which fails isinstance(claimed_job, dict) and fell
into the 'already being fired' branch, so every dispatch assert failed.

Mock the claim to return the job snapshot (the API's success shape),
read the summary's deliver from the claimed snapshot the run actually
executes, and keep the dispatch-result failure renderer. Rebased onto
current main; cron suite 710 passed.
2026-09-02 00:52:58 -07:00
赵桂雄 fd387c15eb fix(cron): treat falsy deliver as local in manual-run notice
Review follow-up on the #83993 fix: a stored falsy deliver ("", JSON
null) fell through the local check and produced 'output was delivered
there by the job itself' for a target that does not exist — the exact
false-delivery-claim class the PR removes. Fire time already normalizes
falsy deliver to local (no delivery, output persisted in last_output,
no delivery error), so the summary now canonicalizes with the
scheduler's own _normalize_deliver_value and reads saved-locally.

Whitespace-only deliver is deliberately not folded in: fire time
records 'no delivery target resolved' for it, and the error-driven
FAILED wording must stay visible.
2026-09-02 00:52:58 -07:00
赵桂雄 94e49b82b1 fix(cron): stop manual-run notice from asserting delivery that never happened
The _execute_job_now completion notice unconditionally claimed
"(output was delivered there by the job itself)" for non-local
delivery targets, even when the job record's last_delivery_error
showed the delivery failed (#83993). Derive the note from the
refreshed job record so a failed delivery is reported honestly to
the calling agent.
2026-09-02 00:52:58 -07:00
Teknium f298911467 fix(lazy-deps): pass --compile-bytecode on the uv tier and skip metadata dirs in the warm
Follow-up to the #100829 salvage. uv pip install writes no __pycache__ by
default (pip does), so --compile-bytecode covers the whole install including
transitive deps, which the per-spec warm never sees. Also skip *.dist-info /
*.egg-info roots in _installed_dist_roots — they own no importable code.

Live: fresh cpython-3.12.13 venv, real uv install of anthropic==0.87.0 via
_venv_pip_install: main -> 0 pyc, first import 0.468s; after -> 1212 pyc
(546 anthropic), first import 0.205s.

Refs #100461
2026-09-02 00:03:55 -07:00
joaomarcos d380651a9f fix(lazy-deps): byte-compile lazily installed backends at install time
A pip/uv install writes .py sources and no __pycache__ — and reinstalling
the same version still deletes the cache the previous copy had. Nothing in
Hermes compiles them, so the whole compile is paid by whoever imports the
package next. For a lazily installed backend that is the foreground of a
user request, with nothing printed while it runs.

Measured for anthropic==0.87.0 (541 modules) on cpython-3.12.13: the first
import after an install costs 2.2-2.7s against 0.7-1.0s warm, and 10.5s
under concurrent load. N per-profile daemons cold-starting together each
pay it in full, because none of them has written the cache yet.

Compile the freshly installed distributions in _venv_pip_install instead,
on the success path of both the uv and pip tiers. The caller is already
waiting on an installer there and can see why. Package directories are
resolved from each distribution's own file list, so specs whose import
name differs from their package name (python-telegram-bot -> telegram)
are covered. Best-effort: a compile failure never invalidates an install
that succeeded, and sys.dont_write_bytecode is honored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MhAnkrFktFdmZwf64fUYLE
2026-09-02 00:03:55 -07:00
codexbt 86fa1fcd4f fix(mcp): treat silent ping drop as unsupported rather than dead transport (Closes #97245)
A stdio server that never answers the optional ping (no -32601, no
response at all) produced a bare TimeoutError that _keepalive_probe
classified as a dead transport, tearing down and respawning a healthy
subprocess on every keepalive tick. On a first ping timeout, confirm with
list_tools before declaring death; if it answers, latch _ping_unsupported
and use list_tools from then on. If both fail, propagate as before.
2026-09-01 23:56:41 -07:00
DragonnZhang 9514d354ca fix(tool-search): validate deferred tool_call arguments against the concrete schema before dispatch
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.

Fixes #73175

Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
2026-09-01 23:30:33 -07:00
Teknium 3fb128ea3e test(mcp): child PID snapshot must see a subprocess spawned from another thread 2026-09-01 23:28:39 -07:00
NATHAN Menkin b828624479 fix(mcp): respawn and retry once when a stdio child died
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.

The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.

Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.

The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.

Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
  call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
  exhausted, parked, tools deregistered -- no respawn loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 23:28:39 -07:00
MattMaximo 6f625f7381 fix(tools): propagate caller contextvars in DaemonThreadPoolExecutor.submit
Some bundled CPython runtime builds strip stdlib ThreadPoolExecutor's
copy_context() propagation, so work submitted to the daemon pool runs in a
bare context. Under the multiplexed gateway this dropped the profile
secret scope in pool workers: the context-compression timeout fence
resolved auxiliary provider keys (SURPLUS_API_KEY) with
UnscopedSecretError, silently degrading LLM compression to lossy
deterministic summaries and driving re-read loops in affected sessions.

Restore stdlib semantics in submit() by snapshotting the caller's context
and running the callable inside it (a no-op re-application on runtimes
that already propagate). Mirrors the gateway's
_run_in_executor_with_context pattern.

Tests: daemon pool worker sees caller contextvars; scoped get_secret works
in a daemon-pool worker under multiplex while scoped misses still fail
closed (no env leak).
2026-09-01 22:28:52 -07:00
Teknium 7e310f5037 Merge pull request #97979 from NousResearch/core-tool-deferral
feat(tool-search): core-tool deferral — curated event-triggered set behind the bridge (−6.5K tok/call desktop, −49% schemas)
2026-09-01 22:11:29 -07:00
kshitijk4poor 83b81fc6db test(mcp): pin the non-calling watcher probe + tighten to iscoroutinefunction
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
2026-09-01 22:11:01 -07:00
Teknium aac8d4b9e9 test: sweep two main-side tests onto the renamed gui_tour / process_manage names 2026-09-01 21:50:19 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 4e3feb8bbb feat(tts): speech toggles warm up and unload local TTS engines (#100881)
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.

- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
  acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
  registry; piper/kittentts loaders extracted so warm-up and synthesis
  share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
  failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
  intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.

Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
2026-09-01 21:43:59 -07:00
Teknium 9387bf929c fix(delegate): drain abandoned-worker transports FD-safely on child timeout
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).

Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
  the shared client's pooled sockets (force_close_tcp_sockets — FD release
  stays with the owning worker), abort+poison of the cached per-request
  openai/anthropic wire clients, Codex app-server request_interrupt(), and
  the inline _active_request_abort hook. Never client.close(), never
  socket.close().
- delegate timeout path: after registering the deferred-close callback,
  run one immediate drain plus one 5s re-sweep (covers a connection opened
  between the interrupt and the first sweep). The settled read (EOF/EPIPE)
  lets the worker unwind, which triggers the deferred close on the worker's
  own thread — the only safe FD-release boundary. A worker that still never
  settles retains its resources rather than risking a cross-thread close.

Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.

Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.

Closes #94248
2026-09-01 12:07:52 -07:00
Leandro Piccione aa1d22670e fix(delegate): defer timed-out child teardown 2026-09-01 12:07:52 -07:00
Teknium 7cd91114b4 fix(web): drop tavily from the removed-backend registry after the restore
#100540 added a REMOVED_BACKENDS startup warning keyed on tavily; with the
backend restored, that entry would warn on a working provider. The registry
stays (empty) for future removals; migration tests now pin the machinery via
a synthetic entry plus a guard asserting no live provider is ever listed as
removed.
2026-09-01 10:56:49 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium 043c258ac2 fix(web): stale removed-backend config warns at startup and errors by name
A config still pointing at a web backend that no longer ships in-tree
(web.backend: tavily after the #99199 removal) previously failed silently:
no migration, no startup notice, and only a generic 'no registered web
search provider has that name' at the first tool call (reported by keyed
Tavily users upgrading to v0.21.0, see PR #99731 thread).

- tools/tool_backend_helpers.py: REMOVED_BACKENDS registry +
  removed_backend_note(); selection_error() swaps in the specific
  removal explanation (removed in v0.21.0, keyless alternatives) while
  keeping the uniform remediation contract.
- hermes_cli/config.py: validate_config_structure() checks web.backend /
  search_backend / extract_backend against the registry and emits a
  startup warning (deduped per stale value), surfaced by the existing
  print_config_warnings() path in CLI and gateway.
- tests/tools/test_removed_backend_migration.py: startup warning,
  per-capability keys, dedupe, healthy-config negative, live-backend
  failure text preserved.
2026-09-01 10:43:20 -07:00
Teknium 4789fa1066 test(browser): isolate chrome-fallback sandbox test from the shared tmpdir
The test exercised _run_chrome_fallback_command against the REAL
/tmp agent-browser namespace, so any concurrent hermes/pytest process
running the orphan reaper could rmtree the fresh pidless socket dir
mid-command (FileNotFoundError on _stdout_open — the recurring CI
flake, 3 hits incl. two main runs on 2026-09-01). Route it through a
private tmp_path like every sibling browser test.
2026-09-01 09:27:00 -07:00
Hermes Agent 4cca38be86 fix(browser): preserve starting sessions during orphan reap 2026-09-01 09:27:00 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium f709bd88b6 feat(skills): render the configured create dir in every instruction that names the path
The skill_manage tool schema description, prompt-builder docs, and the
skills docs page now derive the creation path from skills.create_dir
(display_skill_create_dir()) instead of hardcoding ~/.hermes/skills/ —
so pointing the config at e.g. /opt/brain/skills changes what the agent
is told everywhere, with no SOUL.md fights or read-only chmod tricks.
Adds config default + docs section + 16 tests (incl. a read-only
profile-skills-dir scenario).
2026-09-01 07:32:45 -07:00
Teknium 5a8e8a6b87 fix(terminal): strict Linux-only gating for background-executor systemd scopes (#70716 follow-up)
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):

- Gate every scope-path branch on a new _IS_LINUX constant instead of
  'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
  never touches systemd code — no probe subprocess, no scope argv,
  byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
  legacy argv byte-identical, no unit recorded) and probe-returns-False
  off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
  wired into the on-demand windows-venv-e2e lane): real spawn_local on
  windows-latest asserting jobs run exactly as before — spawned, output
  captured, exit code correct, systemd path never reached even under
  faked gateway identity.

Refs #70716, #71378.
2026-09-01 02:32:53 -07:00
Brooklyn Nicholson b293157f7d feat(tools): withdraw tip and tour when the user switches them off
Off should mean the model is never told the tool exists. A switch that
only made the call fail leaves Hermes offering walkthroughs it cannot
give and promising to point at things it cannot point at, which reads as
a broken agent rather than a respected preference.

Both gate on the switch through a shared desktop_ui.user_enabled helper,
which is the reactions check_fn generalized — same config read, same
reason it has to be config rather than an env var: the switch belongs to
the session's client, and the client may be on another machine.
2026-08-31 21:11:11 -05:00
Brooklyn Nicholson 02458e67ad fix(auxiliary): accept dict and object messages in extract_content_or_reasoning
Compression and some OpenAI-compatible proxies hand us a dict-shaped
response or a bare message, not a ChatCompletion. Reuse the existing
helper instead of a second extractor, and bound an optional reasoning
fallback so a chain-of-thought dump cannot become the summary.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
teknium1 d10ef89ee5 fix(docker): keep forwarded secret values out of world-readable argv
docker run/exec argv previously carried -e KEY=VALUE pairs for every
forwarded/passthrough variable. On Linux /proc/<pid>/cmdline is
world-readable regardless of process owner, so every allowlisted secret
was visible to all local users via plain ps for the duration of every
terminal call.

Emit name-only -e KEY flags and supply values via the docker client
subprocess env instead: the docker CLI resolves valueless --env KEY from
its own environment (documented docker/podman behavior), moving secrets
from /proc/*/cmdline (0444) to /proc/*/environ (0400). Covers the docker
run container-start path, the recreation/recovery path, the init-seeding
exec path, and the per-command runtime exec path.

Reported by @sashalab. Fixes #96268
2026-08-31 11:29:17 -07:00
HexLab98 97f2a571b5 test(terminal): cover hung-wait bound, parent-tid interrupt, and cron inactivity watchdog
Pin that execute() returns at the wall-clock deadline when the inner wait
never returns, that /stop on the tool-worker tid still kills the subprocess,
that the cron inactivity helper fires while the caller thread is blocked,
and that ContextVars plus the activity callback reach the deadline worker.
2026-08-31 10:42:39 -07:00
Teknium 90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
zengzheqing 98e2f110cf fix(curator): restore complete skill packages on ledger rollback (#96962)
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.

The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.

Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
2026-08-31 10:08:13 -07:00
Teknium c8329384a9 test(send_message): drop duplicate buzz UUID target tests
Dispatch cluster (#99431) landed equivalent coverage first; the media
branch's copies shadowed them and tripped
test_no_shadowed_test_definitions.
2026-08-31 10:06:34 -07:00
EmpireOperating fcd34e57a2 fix(buzz): complete media-only delivery reporting 2026-08-31 10:06:34 -07:00
EmpireOperating 9c25704257 fix(buzz): verify live media delivery receipts 2026-08-31 10:06:34 -07:00
EmpireOperating 37c943997b fix(buzz): support media in standalone sends 2026-08-31 10:06:34 -07:00
Andrew Bagrin b7c59bda54 fix(cron): isolate per-execution working directories 2026-08-31 09:59:39 -07:00
kshitijk4poor 936b970e28 fix(browser): lightpanda review follow-ups for #99312
- lightpanda_engine_status: check use_real_profile before the cloud
  provider, matching browser_exec's actual resolution order (real-profile
  resolution runs before backend resolution), so /browser status and
  hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
  (find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
  _using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
2026-08-31 21:59:11 +05:30
Adrià Arrufat e3a85ae5a0 feat(browser): honor browser.engine=lightpanda in Browser Use mode
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.

- browser_use_cli: when the engine is lightpanda and nothing with higher
  precedence claimed the session, get a session from _get_session_info()
  and export its endpoint as BU_CDP_URL; the browser is private to the
  session key, so the own-tab preamble is skipped. The browser_exec
  description gains a Lightpanda header (text-first, new_tab once then
  goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
  --host 127.0.0.1 --port <free>` per session key (new
  tools/browser_lightpanda.py), reusing the session cache, inactivity
  reaper and atexit cleanup; a dead process is respawned on the next call;
  orphans from a crashed Hermes are reaped through per-process records in
  $HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
  reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
  (cloud_provider: local + engine: lightpanda; "Local Browser" resets the
  engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
  is shadowed, the reason.
2026-08-31 21:59:11 +05:30
blunkjamie-dev 1885a40ad3 fix(buzz): preserve literal mentions and exact UUID targets 2026-08-31 09:05:41 -07:00
kshitijk4poor 41dfa2d580 test: update chrome-fallback npx fixture for the get-url response shape
The fallback URL probe now uses 'get url' ({data:{url}}), not
'eval window.location.href' ({data:{result}}); the npx-resolution test's
mocked response predates the #81673 change.
2026-08-31 21:17:37 +05:30
Simone Marzola e355394839 fix(browser): isolate Lightpanda and Chrome fallback flags
Strip Chromium-only launch variables from Lightpanda commands, use a non-recursive Lightpanda URL lookup for Chrome fallback, and share sandbox argument injection with the temporary Chrome path.

Co-authored-by: forjd-hermes-bot <282037251+forjd-hermes-bot@users.noreply.github.com>
2026-08-31 21:17:37 +05:30
Teknium 26f178e5fa fix(terminal): gate the BUZZ_* terminal carve-out on actual Buzz agent context
Compose the two salvaged approaches (#78065 + #78511):

- Keep #78065's terminal-only scrub-path exemption (first-party prefix
  predicate in _make_run_env / _sanitize_subprocess_env, plain env values
  never scope-resolved, snapshot exclusion for cross-profile isolation,
  every non-terminal surface sealed).
- Fold #78511's BUZZ_MANAGED_AGENT signal into a context gate instead of
  an import-time blocklist discard: the blocklist is shared by every
  scrub surface, so discarding there would leak BUZZ_PRIVATE_KEY into
  execute_code / hermes_subprocess_env children too.
- New gate _buzz_terminal_context_active(): BUZZ_MANAGED_AGENT in the
  process env (Buzz Desktop buzz-acp harness, #76243) OR the live
  session's platform is buzz (HERMES_SESSION_PLATFORM ContextVar,
  concurrency-safe under a multi-session gateway). A Telegram/CLI/cron
  session on a host that also runs a Buzz gateway does NOT get the
  signing key in its terminal children (maintainer triage note on
  #76243: don't expose the key to unrelated shell commands).
- Snapshot exclusion stays prefix-only (conservative even when the gate
  is inactive).
- Tests updated for the gate + new negative test (non-Buzz session
  strips) and positive test (buzz session platform enables carve-out);
  docs updated accordingly.

Closes #78026, closes #76243.
2026-08-31 07:35:47 -07:00
Magnus Hedemark a4c6821218 fix(tools): pass BUZZ_* platform credentials to terminal children
Buzz platform agents could not use the `buzz` CLI from the terminal tool:
the BUZZ_* vars (BUZZ_PRIVATE_KEY, BUZZ_AUTH_TAG, BUZZ_RELAY_URL, and the
other BUZZ_* names) are added to _HERMES_PROVIDER_ENV_BLOCKLIST from the
buzz plugin.yaml (messaging category), and env_passthrough refuses to
re-allow anything in the blocklist (GHSA-rhgp-j443-p4rf). In the reported
`hermes acp` scenario the Buzz adapter's register() is never invoked, so
the agent runs in-process and its terminal uses _make_run_env directly —
there was no path for the platform credentials to reach terminal children.

Fix: a terminal-only, first-party carve-out in the scrub paths themselves
(not adapter registration). BUZZ_* vars pass through to foreground
(_make_run_env) and background/PTY (_sanitize_subprocess_env) terminal
children via a new prefix predicate (_TERMINAL_FIRST_PARTY_ENV_PREFIXES).
Everything else stays sealed and unchanged: the blocklist itself, the
env_passthrough refusal, execute_code scrubbing, hermes_subprocess_env
(browser/TUI-host/copilot-executor spawns), and docker children. The
GHSA-rhgp-j443-p4rf seal is preserved because no registration path is
opened; skills/config still cannot register these names.

Follow-up hardening from review:

- First-party matches use the merged env value directly instead of
  _resolve_passthrough_value: under multiplex with no profile secret scope
  installed the resolver raised UnscopedSecretError (fail-closed) at call
  sites like the webhook-filter script runner, a regression where the
  script previously ran without the var. The vars are the process's own
  env values and are never scope-resolved.
- LocalEnvironment now excludes first-party terminal env names from the
  shared login-shell snapshot (_additional_profile_scoped_passthrough_names
  override): BUZZ_PRIVATE_KEY can never be in the get_all_passthrough()
  exclusion set (env_passthrough refuses blocklisted names), so without
  this a multiplexed gateway would dump profile A's key into
  hermes-snap-<id>.sh and profile B sharing the collapsed LocalEnvironment
  would source it — a cross-profile nsec leak. The names are now excluded
  from the dump and save/restored per command.
- Docs now name the _sanitize_subprocess_env consumers (search workers
  like ddgs, computer-use driver, user-script runners) that also receive
  first-party platform vars.

Fixes #78026

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-31 07:35:47 -07:00
Teknium 6fba07cb7e test(terminal): background spawn contract now includes owner_task_id 2026-08-31 07:28:18 -07:00
Teknium 5a4dbdec27 fix(delegation): subagent process notifications stay suppressed when the container key collapses
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.

Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).

Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
2026-08-31 07:28:18 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
itskaism 5ce8f71553 fix(delegation): report schema-invalid child results as failed, not completed
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.

Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.

Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
2026-08-31 01:02:42 -07:00