Commit Graph

2641 Commits

Author SHA1 Message Date
Teknium 18700c60d5 Merge branch 'simp/tools' into simp/integration 2026-09-02 15:01:45 -07:00
Teknium ae70536485 Merge branch 'simp/web' into simp/integration 2026-09-02 15:01:45 -07:00
Teknium 2d5e9a0ef1 refactor(tools/delegate): repoint batch-tag test to delegate_tool_progress where _BATCH_ORDINALS is bound 2026-09-02 14:47:47 -07:00
Teknium aed2e156ad refactor(tools/files): finish file_operations wiring to extracted modules; restore WHY comments in file_tools; repoint test 2026-09-02 14:47:03 -07:00
Teknium 20258aad9e refactor(tools/mcp_oauth): extract mcp_oauth_provider; compact oauth manager/dashboard bridge and schema_sanitizer; repoint tests 2026-09-02 14:45:42 -07:00
Teknium 8813eba345 refactor(tools): repoint tests to moved symbols; add voice_mode_transcript module 2026-09-02 14:45:15 -07:00
Teknium 05526b028a refactor(tools/delegate): split delegate_tool into child_run/config/dispatch/progress/registry/results; compact delegation helpers 2026-09-02 14:45:15 -07:00
Teknium 9f6335bc44 refactor(tools/browser): extract eval-policy/lightpanda-fallback/real-profile/snapshot modules from browser_tool; split supervisor dialogs/frames; dedupe camofox/cli 2026-09-02 14:44:15 -07:00
Teknium c1f8af1e86 refactor(tools/mcp): split mcp_tool.py into transport/lifecycle/schema/handlers/... sibling modules; compact watchdog and schema cache 2026-09-02 14:44:15 -07:00
Teknium 606cb2de92 refactor(tools/code_exec): unify code_kernel local/remote helpers, split checkpoint_manager god methods, compact spill helpers 2026-09-02 14:44:15 -07:00
Teknium 5c919161d0 refactor(tools/approval): split approval.py into smart/human-wait/gateway-wait modules; dedupe guards 2026-09-02 14:44:14 -07:00
Teknium ec49ae7c0f refactor(tools): wire file_operations to extracted common/lint/search modules; restore literal security pins in lazy_deps 2026-09-02 14:43:45 -07:00
Teknium 14a64ea985 refactor(tools): repoint tests to moved file_tools guard/path symbols 2026-09-02 14:43:45 -07:00
Teknium d4cec15b47 refactor(tools): first-wave simplification of tools/ (file ops split, lazy_deps, code_exec, approval, browser, delegate, mcp, skills, terminal, voice, media)
Behavior-neutral structural pass over tools/*: god-file extractions into
sibling modules (file_operations_common/lint/search, file_tools_paths/
read_tracking/write, code_execution_env/rpc, tool_search_catalog/names/
validation, tts_command_provider, ...), duplicate helper unification,
if/elif -> dispatch tables, dead-code removal, docstring compaction.
Tool schemas (get_tool_definitions) verified byte-identical to base.
2026-09-02 14:43:45 -07:00
Teknium de109d4097 refactor(web): extract chat-tab WebSocket routes into web_routers/chat_ws.py 2026-09-02 13:32:06 -07:00
Teknium 74edfe04e9 refactor(plugins/spotify): action dispatch table, single request helper 2026-09-02 13:30:10 -07:00
Teknium e43f381b4e refactor(plugins/browser): shared BaseCloudBrowserProvider for browser_use/browserbase/firecrawl 2026-09-02 13:30:10 -07:00
kshitijk4poor ff7233b815 fix(credential_files): apply the same exclusions to the symlink-safe mount copy
_safe_skills_path() is the sibling of iter_skills_files(): when a symlink in
skills/ forces a sanitized copy for mount-based backends (Docker/Singularity),
it rglob-copied the whole tree — .hub, .curator_backups, node_modules and all.
Prune EXCLUDED_SKILL_DIRS before descending, same rule as the sync generator,
so the mounted copy never carries (or walks) the bookkeeping trees either.
2026-09-03 01:33:18 +05:30
Carry00 edac49e473 fix(skills): stop syncing bookkeeping dirs to sandboxes
iter_skills_files() walked the skills tree with a bare rglob("*"), so the
.hub download cache, .archive, curator backups, and any node_modules/.git
under a skill package were uploaded to the sandbox on every sync. The
sandbox never reads them: skill content is resolved host-side.

EXCLUDED_SKILL_DIRS is already the canonical exclusion set, honoured by
discovery and backup. Apply it to the sync path too, across all three
roots iter_skills_files() walks (local, external, project-local), and add
.curator_backups to the set.

Measured on a local install: 900 files / 67.3 MB -> 771 files / 8.4 MB.

This is not just wasted bandwidth on the SSH backend, where the oversized
payload can exceed the 120s _ssh_bulk_upload deadline and surface as the
agent hanging on every tool call.

The filter intentionally does not reuse is_excluded_skill_path(), which
also prunes references/, templates/, assets/ and scripts/ -- those hold
support files and bundled scripts the sandbox does read and execute.
2026-09-03 01:33:18 +05:30
Teknium 7840a0e2d9 feat: delegation batch tags read "set N" instead of a hex id slice
Interleaved subagent fan-outs were tagged with the first 4 hex chars of the
delegation id ([b2ac 3/9]), which is attributable but unreadable. Batches are
now numbered in order of appearance per process: [set 1 · 3/9], [set 2 · 1/7].
Desktop /agents already labels groups "Delegation N", so its duplicate hex
badge is dropped.
2026-09-02 10:12:54 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
joaomarcos 65b0f00002 fix(agent): stop the between-turns tool refresh from forking the cached prefix
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:

* a tool whose `check_fn` merely flapped (headless browser probe, expired
  credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.

Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.

`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.

Refs #100336
2026-09-02 07:22:59 -07:00
Teknium ee0e234a2c fix(gateway): discover and reload MCP servers per profile under multiplex
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.

- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
  served profile inside `_profile_runtime_scope`, carried into the
  executor via `copy_context()` (same shape as
  `_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
  caller (e.g. button-confirm callback) did not; shut down / rediscover /
  report only that profile's servers; refresh only that profile's cached
  agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
  `_server_scope_keys` ownership map; leaves the shared MCP loop running
  while other profiles' servers are live. Unscoped call keeps the full
  historical behavior.
- MCP tools register into the owning profile's registry overlay
  (`registry.register(scope=...)`), and `registry.deregister()` gains a
  matching `scope=` kwarg. Plugin callers still cannot name another
  profile's scope; the plugin-vs-global guard is unchanged for them.

Fixes #95518

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium 4e7aa48716 fix(browser): reap idle multiplexed sessions under their owner profile scope
The inactivity janitor is one process-global thread started by whichever
profile first opens a browser, so under `gateway.multiplex_profiles` it runs
with no secret scope: `cleanup_browser` -> `is_camofox_mode` ->
`get_secret("CAMOFOX_URL")` raises UnscopedSecretError, the session entry is
never removed, and the same failure repeats every 30s while the Chromium
daemon leaks.

- `_update_session_activity` records the owning Hermes home per session;
  `_cleanup_inactive_browser_sessions` re-enters that owner's
  `set_hermes_home_override` + `build_profile_secret_scope` around each
  teardown (`_session_owner_scope`, mirroring `_profile_runtime_scope`).
  copy_context at thread spawn would pin the first profile's secrets onto
  every other profile's teardown; there is no os.environ fallthrough.
- 3 consecutive failures -> `_force_reap_browser_session`, which skips the
  failing `close` round-trips but still closes the cloud provider session
  and kills the local daemon via the shared `_release_session_resources`
  tail (extracted from `_cleanup_single_browser_session`, unchanged).
  An activity touch does not reset the failure budget.

Fixes #86402
Fixes #100738

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-02 07:00:13 -07:00
liuhao1024 7d509657a8 fix(tts): resolve default output dir from the active profile
DEFAULT_OUTPUT_DIR was resolved once at import time, so long-lived
multi-profile runtimes (dashboard console, TUI/Desktop backend, cron,
kanban workers) kept writing synthesized audio into the launch
profile's cache/audio instead of the requesting profile's (#98749).

Same bug class and fix as skills_tool (f8723c478) and skills_sync
(#65828): keep the legacy module attribute for tests and external
patchers, but re-resolve from the live profile-scoped HERMES_HOME on
every synthesis call.
2026-09-02 06:48:31 -07:00
liuhao1024 39fca697ae fix(tools): log boot-time UnscopedSecretError probes at debug, not warning
With multiplexing on, check_fns evaluated before any profile secret scope
exists fail closed by design: get_secret raises UnscopedSecretError and the
tool re-probes on the first scoped turn. _run_check_fn_uncached logged that
expected signal like a crashed check_fn (WARNING + exc_info), so every
multiplexed gateway start printed three full tracebacks that drowned real
check_fn failures.

Split the handler: an unscoped read reported while the profile cache scope
was unresolved logs one debug line without a traceback; the same error with
the scope resolved is a genuinely lost scope and keeps the loud
warning + traceback.

Fixes #100697
2026-09-02 06:48:31 -07:00
Teknium d29a7936e4 fix(bot-mode): DMs to a Desktop-owned Bot Chat land in the live session instead of being dropped (#100523)
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.

bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.

On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.

Closes #100523
Supersedes #100544, #100542

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 06:18:55 -07:00
Teknium c2954c8934 feat(model-catalog): picker catalogs refresh every 20 minutes, gateway keeps them warm
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.

- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
  an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
  disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
  it off-thread every TTL window, so every surface on the machine reads
  a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
2026-09-02 06:16:54 -07:00
Teknium 782dd635fe feat(tts): speech toggles warm/release plugin and command TTS providers
Extends the TTS lease from #100912 (4e3feb8bbb) beyond built-in local
engines: when the configured tts.provider is user-declared, acquiring
the first lease and releasing the last one now reach it, so a
self-hosted TTS server can preload its model when read-aloud / voice
conversation turns on and unload when it turns off (Discord request).

- agent/tts_provider.py: TTSProvider gains concrete no-op warm() /
  release() (not abstract — existing plugins are unaffected).
- tools/tts_tool.py: _signal_user_tts_provider() forwards the lease
  hook; plugin providers get warm()/release(), command providers run
  optional `warm_command` / `release_command` (config.yaml, under
  tts.providers.<name>) through the existing _run_command_tts helper
  on a daemon thread — best-effort, output discarded, failures at
  debug. warm_tts_provider() and release_tts_provider() call it.
- tests/tools/test_tts_lifecycle_leases.py: fake plugin provider and
  fake command provider observe warm/release through acquire/release
  lease (both fail on main with action == "noop").
- docs: features/tts.md — lease section, command-provider optional
  keys table, plugin optional hooks.
2026-09-02 05:35:14 -07:00
muhifni 1cd736ff63 fix(terminal): scope terminal config per turn under profile multiplexing
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.

Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:

- tools/terminal_scope.py: ContextVar holding the routed profile's
  COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
  TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
  resolves ONLY from it - an omitted key yields the defined default,
  never os.environ. Unreadable/malformed policy installs a refusal
  scope; terminal_tool / execute_code refuse instead of running under
  ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
  _profile_runtime_scope, tui_gateway session/build/turn scopes, cron
  per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
  (_get_env_config, _resolve_container_task_id shared key, orphan
  reaper lifetime, degraded mode), gateway/platforms/base.py docker
  media translation (volumes, shared key, persistence), runtime_cwd /
  agent_init / skill_utils / code_execution_tool / file_tools cwd
  anchors, prompt_builder / browser_tool / env_probe backend checks,
  gateway footer, @-refs and slash-command cwd. env_probe resolves the
  backend in the caller's context, since the probe worker thread does
  not inherit the ContextVar.

Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.

Fixes #68559
Fixes #94200
Fixes #101132
Fixes #95470

Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
2026-09-02 05:34:28 -07:00
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
Teknium 758114bb8d fix(cron): manual run reports delivery_failed as a failed run; docs for the distinct status
A manual cronjob(action='run') derived success from last_status == 'ok'
and read the error from last_error — so a run that now records
delivery_failed came back as success=False with error=None, an unexplained
failure. Surface last_delivery_error as the error in that case (the
#84006 direction, re-applied on the delivery_failed status), and pin the
manual-run completion summary to say 'Result: FAILED' over an undelivered
run. Document the status in the cron user guide.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-02 00:52:58 -07:00
赵桂雄 2f58cbfa7f fix(cron): adapt delivery-notice tests to the return_job claim API
Main grew claim_job_for_fire(job_id, return_job=True) — a claimed
snapshot dict instead of a bool — while this branch sat on an older
base. The merge-ref CI ran the hybrid: the wiring tests still mocked
return_value=True, which fails isinstance(claimed_job, dict) and fell
into the 'already being fired' branch, so every dispatch assert failed.

Mock the claim to return the job snapshot (the API's success shape),
read the summary's deliver from the claimed snapshot the run actually
executes, and keep the dispatch-result failure renderer. Rebased onto
current main; cron suite 710 passed.
2026-09-02 00:52:58 -07:00
赵桂雄 fd387c15eb fix(cron): treat falsy deliver as local in manual-run notice
Review follow-up on the #83993 fix: a stored falsy deliver ("", JSON
null) fell through the local check and produced 'output was delivered
there by the job itself' for a target that does not exist — the exact
false-delivery-claim class the PR removes. Fire time already normalizes
falsy deliver to local (no delivery, output persisted in last_output,
no delivery error), so the summary now canonicalizes with the
scheduler's own _normalize_deliver_value and reads saved-locally.

Whitespace-only deliver is deliberately not folded in: fire time
records 'no delivery target resolved' for it, and the error-driven
FAILED wording must stay visible.
2026-09-02 00:52:58 -07:00
赵桂雄 94e49b82b1 fix(cron): stop manual-run notice from asserting delivery that never happened
The _execute_job_now completion notice unconditionally claimed
"(output was delivered there by the job itself)" for non-local
delivery targets, even when the job record's last_delivery_error
showed the delivery failed (#83993). Derive the note from the
refreshed job record so a failed delivery is reported honestly to
the calling agent.
2026-09-02 00:52:58 -07:00
Teknium f298911467 fix(lazy-deps): pass --compile-bytecode on the uv tier and skip metadata dirs in the warm
Follow-up to the #100829 salvage. uv pip install writes no __pycache__ by
default (pip does), so --compile-bytecode covers the whole install including
transitive deps, which the per-spec warm never sees. Also skip *.dist-info /
*.egg-info roots in _installed_dist_roots — they own no importable code.

Live: fresh cpython-3.12.13 venv, real uv install of anthropic==0.87.0 via
_venv_pip_install: main -> 0 pyc, first import 0.468s; after -> 1212 pyc
(546 anthropic), first import 0.205s.

Refs #100461
2026-09-02 00:03:55 -07:00
joaomarcos d380651a9f fix(lazy-deps): byte-compile lazily installed backends at install time
A pip/uv install writes .py sources and no __pycache__ — and reinstalling
the same version still deletes the cache the previous copy had. Nothing in
Hermes compiles them, so the whole compile is paid by whoever imports the
package next. For a lazily installed backend that is the foreground of a
user request, with nothing printed while it runs.

Measured for anthropic==0.87.0 (541 modules) on cpython-3.12.13: the first
import after an install costs 2.2-2.7s against 0.7-1.0s warm, and 10.5s
under concurrent load. N per-profile daemons cold-starting together each
pay it in full, because none of them has written the cache yet.

Compile the freshly installed distributions in _venv_pip_install instead,
on the success path of both the uv and pip tiers. The caller is already
waiting on an installer there and can see why. Package directories are
resolved from each distribution's own file list, so specs whose import
name differs from their package name (python-telegram-bot -> telegram)
are covered. Best-effort: a compile failure never invalidates an install
that succeeded, and sys.dont_write_bytecode is honored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MhAnkrFktFdmZwf64fUYLE
2026-09-02 00:03:55 -07:00
codexbt 86fa1fcd4f fix(mcp): treat silent ping drop as unsupported rather than dead transport (Closes #97245)
A stdio server that never answers the optional ping (no -32601, no
response at all) produced a bare TimeoutError that _keepalive_probe
classified as a dead transport, tearing down and respawning a healthy
subprocess on every keepalive tick. On a first ping timeout, confirm with
list_tools before declaring death; if it answers, latch _ping_unsupported
and use list_tools from then on. If both fail, propagate as before.
2026-09-01 23:56:41 -07:00
DragonnZhang 9514d354ca fix(tool-search): validate deferred tool_call arguments against the concrete schema before dispatch
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.

Fixes #73175

Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
2026-09-01 23:30:33 -07:00
Teknium 3fb128ea3e test(mcp): child PID snapshot must see a subprocess spawned from another thread 2026-09-01 23:28:39 -07:00
NATHAN Menkin b828624479 fix(mcp): respawn and retry once when a stdio child died
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.

The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.

Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.

The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.

Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
  call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
  exhausted, parked, tools deregistered -- no respawn loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 23:28:39 -07:00
MattMaximo 6f625f7381 fix(tools): propagate caller contextvars in DaemonThreadPoolExecutor.submit
Some bundled CPython runtime builds strip stdlib ThreadPoolExecutor's
copy_context() propagation, so work submitted to the daemon pool runs in a
bare context. Under the multiplexed gateway this dropped the profile
secret scope in pool workers: the context-compression timeout fence
resolved auxiliary provider keys (SURPLUS_API_KEY) with
UnscopedSecretError, silently degrading LLM compression to lossy
deterministic summaries and driving re-read loops in affected sessions.

Restore stdlib semantics in submit() by snapshotting the caller's context
and running the callable inside it (a no-op re-application on runtimes
that already propagate). Mirrors the gateway's
_run_in_executor_with_context pattern.

Tests: daemon pool worker sees caller contextvars; scoped get_secret works
in a daemon-pool worker under multiplex while scoped misses still fail
closed (no env leak).
2026-09-01 22:28:52 -07:00
Teknium 7e310f5037 Merge pull request #97979 from NousResearch/core-tool-deferral
feat(tool-search): core-tool deferral — curated event-triggered set behind the bridge (−6.5K tok/call desktop, −49% schemas)
2026-09-01 22:11:29 -07:00
kshitijk4poor 83b81fc6db test(mcp): pin the non-calling watcher probe + tighten to iscoroutinefunction
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
2026-09-01 22:11:01 -07:00
Teknium aac8d4b9e9 test: sweep two main-side tests onto the renamed gui_tour / process_manage names 2026-09-01 21:50:19 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 4e3feb8bbb feat(tts): speech toggles warm up and unload local TTS engines (#100881)
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.

- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
  acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
  registry; piper/kittentts loaders extracted so warm-up and synthesis
  share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
  failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
  intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.

Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
2026-09-01 21:43:59 -07:00
Teknium 9387bf929c fix(delegate): drain abandoned-worker transports FD-safely on child timeout
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).

Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
  the shared client's pooled sockets (force_close_tcp_sockets — FD release
  stays with the owning worker), abort+poison of the cached per-request
  openai/anthropic wire clients, Codex app-server request_interrupt(), and
  the inline _active_request_abort hook. Never client.close(), never
  socket.close().
- delegate timeout path: after registering the deferred-close callback,
  run one immediate drain plus one 5s re-sweep (covers a connection opened
  between the interrupt and the first sweep). The settled read (EOF/EPIPE)
  lets the worker unwind, which triggers the deferred close on the worker's
  own thread — the only safe FD-release boundary. A worker that still never
  settles retains its resources rather than risking a cross-thread close.

Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.

Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.

Closes #94248
2026-09-01 12:07:52 -07:00
Leandro Piccione aa1d22670e fix(delegate): defer timed-out child teardown 2026-09-01 12:07:52 -07:00