Commit Graph

2439 Commits

Author SHA1 Message Date
Dimar Anez 9eb13d07b6 fix(terminal): tolerate macOS TCC PermissionError in _safe_getcwd
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.

_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.

Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.

Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
2026-08-26 04:51:41 -07:00
Teknium bd134d0f30 test: loosen frozen bare-verdict dict in cua_0_9 sibling test to decision contract
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
2026-08-26 04:50:13 -07:00
Teknium 3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
kshitijk4poor d62a05e94c fix(checkpoints): surface skipped_oversize to users and stop misreporting failed deletes as restored
Follow-up to the salvaged #95207 fix, completing the misreport bug class:

- restore() now also drops delete_targets whose unlink failed (OSError
  swallowed) from restored_files — the sibling of the kept-oversize
  misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
  (slash_commands + gateway.rollback.kept_oversize locale key in all 17
  catalogs) now tells the user which files were kept because the size
  cap excluded them from every checkpoint; previously the file was
  correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
2026-08-26 16:44:43 +05:30
RickyYii 595b5ce68a refactor(checkpoints): call the size-cap predicate instead of restating it
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.

The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.

Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.

No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.

Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
2026-08-26 16:44:43 +05:30
RickyYii d28bf79927 fix(checkpoints): stop safe restore deleting files the size cap excluded
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.

Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.

Reproduced on main with shipped defaults:

    corpus.jsonl (2 MB), agent appends to it, then /rollback 1
    safe_restore_plan restore=['corpus.jsonl', 'notes.py']
    restore ok=True restored_files=['corpus.jsonl', 'notes.py']
    notes.py     exists=True   content="v1 = 'original source'"
    corpus.jsonl exists=False  <- deleted, was in no checkpoint

Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.

The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.

The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.

The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.

Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.

Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
2026-08-26 16:44:43 +05:30
kshitijk4poor 365cbc242b test: assert omitted attach_to_session stays absent from formatted list output
Closes the gap flagged in review: the raw store was checked but not the
_format_job surface.
2026-08-26 16:06:41 +05:30
StanleyStetson 5e9adc9e4d fix(cron): forward attach_to_session through cronjob handler
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."

Fixes #84802
2026-08-26 16:06:41 +05:30
Victor Kyriazakos 580daa7b96 fix(cron): mirror continuable-cron briefs for origin-fallback and opted-in explicit targets
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).

The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.

Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
  origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
  gate: origin unchanged; origin_fallback eligible under the same flags
  as origin (per-job attach_to_session wins, else global
  cron.mirror_delivery); explicit platform:chat targets eligible ONLY
  under per-job attach_to_session — the global flag never activates
  them, so it cannot start writing transcript entries into arbitrary
  explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
  keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
  session keys are user-isolated, so a seed without a user_id (origin-
  less job into a shared channel) would create an orphan session no
  reply resolves to — those targets fall back to the plain mirror. DM
  targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.

Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.

15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
2026-08-26 16:06:09 +05:30
Teknium 45db70a80a fix(computer-use): fail closed on unverified CuaDriver.app + background launch
Hardening on top of the TCC daemon-identity salvage:
- _validate_cua_driver_app_signature: codesign -dv gate requiring EXACT
  Identifier=com.trycua.driver and the official team (4YEC26S9KF) before
  /usr/bin/open hands the bundle to LaunchServices — the identity fix must
  not double as a launcher for arbitrary/impostor bundles (suffixed
  identifiers and wrong teams rejected; unsigned dev builds only via
  computer_use.allow_unsigned_driver: true in config.yaml).
- _resolve_cua_driver_app_path: derive the bundle ONLY from the resolved
  driver binary — the /Applications fallback could launch a DIFFERENT
  install than the manifest resolved.
- open -n -g: don't activate/steal focus when launching the daemon.
- 7 new tests incl. sabotage-verified exact-match assertions.

Grafted from #76433's review direction (@Chadmc9889's original fail-closed
validation requirement).

Co-authored-by: Chadmc9889 <Chadmc9889@users.noreply.github.com>
2026-08-26 03:21:37 -07:00
projetsjsl 4746f614be fix(computer-use): preserve macOS TCC daemon identity
Launch private computer-use daemons through CuaDriver.app so Screen
Recording authorization remains attached to its stable bundle identity
instead of Hermes' ad-hoc signature.

Co-Authored-By: GPT-5.6 Codex <noreply@openai.com>
2026-08-26 03:21:37 -07:00
Teknium f0c0c986c4 test: pin overlay policy off in embedded-daemon socket/ack contract test
The embedded spawn now consults the overlay policy (capability probe via
subprocess.run) when _cua_no_overlay() is true — which it is on headless
CI since the Linux X11 default flip. The fixed two-entry run side_effect
in this test didn't budget for the probe call; pin the policy off since
this test pins the socket/ack contract, not overlay behavior.
2026-08-26 00:54:36 -07:00
Teknium afee35700e fix(delegation): suppress subagent-owned process notifications in parent chat by default
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.

- New config key delegation.surface_child_process_notifications
  (default false = suppress). Flag true restores the previous behavior
  exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
  watch_disabled events whose task_id starts with 'sa-' when the flag
  is false, logging at debug with session_id+task_id for diagnosis.
  Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
  events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
  crash the drain loop.
- Docs: delegation.md + configuration.md.
2026-08-25 21:59:04 -07:00
Teknium a935154378 test(tool-search): make salvage-seam tests order- and stem-robust
Two interaction seams between the #92693 salvage (merged as #95050) and
this branch: the source-label indexing test now compares in token space
(the stemmer shortens 'catalogsource' to 'catalogsourc'), and the
unregistered-core-name describe test forces the unregistered condition
via monkeypatch instead of depending on which sibling test file imported
model_tools first.
2026-08-25 16:39:12 -07:00
alt-glitch 00305aec61 test(tool-search): kill the shared-stemmer mutant with cache-missing input
The parallel determinism test warms _stem's lru_cache after ~11 distinct
stems, so almost no iterations reach the underlying stemmer and a shared
(non-thread-local) instance survives it. New test bypasses the cache with
per-iteration unique tokens via _stem.__wrapped__, so thousands of stems
run concurrently: a shared stemmer's mutable parse state fails it within
2,000 calls (verified — the mutant dies 8/8 runs; healthy runs stay green).
2026-08-25 16:39:12 -07:00
alt-glitch 6d37f2b78a test(tool-search): exercise stemmer in parallel 2026-08-25 16:39:12 -07:00
alt-glitch 3b065745ca refactor(tool-search): keep batch caps internal 2026-08-25 16:39:12 -07:00
alt-glitch e35a7bddee perf(tool-search): cache stems and bound result metadata 2026-08-25 16:39:12 -07:00
alt-glitch b09f617722 fix(tool-describe): separate missing and direct names 2026-08-25 16:39:12 -07:00
alt-glitch 1e86a263ed fix(tool-search): preserve exact and per-query ranking semantics 2026-08-25 16:39:12 -07:00
alt-glitch e455e4afd0 feat(tool-search): multi-query search, batched describe, Snowball stemming
tool_search now takes queries: string[] (searched independently against
the same catalog, limit applies per query, default 5 / max 25) and
returns the split shape: per-query groups carry tool names only, one
shared tools map holds each matched tool's source, description (400-char
cap) and required parameter names once. When some queries miss, a single
top-level available_sources + hint block replaces the old per-response
fallback.

tool_describe now takes names: string[] and returns a map keyed by name;
unknown names collect in not_found (with the refresh hint) and
non-deferrable names keep their per-name spelling-check error in errors,
so one bad name no longer fails the whole call. Duplicates dedupe
silently.

The shared tokenizer now applies Snowball stemming (english, exact-pinned
snowballstemmer) at both index and query time, closing the measured
plural/singular miss where 'issues' failed to return create_issue. The
inline BM25 is unchanged. Stemmer instances are thread-local (they carry
mutable parse state and bridge dispatch can run on parallel tool-call
threads).

New config knobs under tools.tool_search: max_queries / max_describe_names
(default 10 each, floor 1, no upper clamp) bound the per-call array
inputs; over-cap calls error so the model repairs in one round-trip.

No backward compatibility with the single query/name shapes, by decision.
scripts/analyze_livetest.py renders both shapes since transcripts on disk
may predate this change.
2026-08-25 16:39:12 -07:00
alt-glitch 62b2d78025 fix(tool-search): bridge batch barrier, listing truncation, source indexing (salvage #92693, part 1)
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):

1. The parallel batch planner now peels the tool_call bridge wrapper and
   decides admission on the underlying tool — supports_parallel_tool_calls
   works again when deferral is active. Unparseable wrappers stay
   sequential barriers; bridged calls get exactly the admission the same
   call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
   version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
   service-name queries reach tools whose own name omits the service; the
   dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).

Salvaged from #92693 by @alt-glitch with authorship preserved.
2026-08-25 15:24:15 -07:00
Gille 7c5c994397 fix(teams): request supported transcript content format 2026-08-26 02:54:35 +05:30
Teknium ba9fc55e16 feat(web): cache_exempt_hosts — always-live fetches for staging/tunnel sites
Sites under active development but tested over the public internet
(Vercel previews, ngrok tunnels, staging domains) are public DNS, so
the local-dev never-cache rule can't catch them. web.cache_exempt_hosts
lists hosts whose pages are always fetched live: exact, "*.wildcard",
or domain-suffix matching (label-boundary aware — mysite.dev covers
preview.mysite.dev but never evilmysite.dev). Checked on both store
and lookup, so adding an exemption takes effect immediately even for
entries cached before the config change.
2026-08-25 04:21:45 -07:00
Teknium f0381ee4ae fix(web): never cache local development URLs in the extract cache
Dev servers, hot-reload builds, and chat-GUI artifact previews live on
localhost/private addresses and change on every save — a 20-minute
cached copy would show a stale build exactly when freshness is the
point of fetching. The extract cache now declines loopback, private,
link-local, *.local, *.localhost, and single-label LAN hostnames on
both put and get. Hostname heuristics only (no DNS) — this is a
freshness carveout, not a security boundary; SSRF enforcement is
unchanged in tools/url_safety.py.

Public URLs keep the full TTL.
2026-08-25 04:21:45 -07:00
Teknium 8adef09be8 fix(web): extract cache serves only after policy + provider gates; rescue and format/provider isolation
Review fixes for #94618 (all three blockers reproduced by the reviewer
through the real web_extract_tool):

1. Cache lookup moved AFTER provider resolution and strict-selection
   validation, and gated per-URL on the website blocklist policy — a
   blocklist-blocked or misconfigured-backend call now behaves exactly
   as it would without a cache instead of serving cached content.
2. Rescue-served extract batches are never cached (mirrors the search
   memo's exclusion), keeping one-shot rescue one-shot.
3. Cache entries now get dedicated per-(url, format, provider) files
   instead of sharing the URL-keyed truncate-store file — html and
   markdown (or two backends') copies of one URL no longer overwrite
   each other, and switching extract backends within the TTL never
   serves the old backend's rendering.

Also from review: per-process index tmp filename (cross-process writers
can no longer truncate each other mid-write) and held flight locks are
never evicted from the bounded lock table (eviction could have allowed
a duplicate paid request).

New regression tests for formats/provider keying; E2E harness extended
with policy-block, strict-selection, rescue-two-call, and dual-format
scenarios — 6/6 pass; original 13/13 still pass.
2026-08-25 04:21:45 -07:00
Teknium 04603fc040 feat(web): TTL result caching for web_search + web_extract
Repeat searches (same normalized query + provider) within a 20-minute
TTL are served from an in-process memo, and concurrent identical
queries are single-flighted so a parallel subagent fan-out pays for
one vendor request instead of N. Requested limits bucket up to
10/20/50/100 so near-identical requests share an entry; callers get
their requested count sliced from the bucket.

Repeat extracts of the same URL are served from the existing
cache/web full-text store (previously written for read_file paging
but never read back), via a small JSON sidecar index. Disk-backed, so
CLI, gateway, cron, and subagents share it. Cached extracts re-run
the normal truncate pipeline, so per-call char_limit still works.

Both caches sit after every safety gate (secret-URL, SSRF, policy,
provider resolution) and directly around the paid vendor call — hits
skip only the network request. Only successful responses cache;
rescue-served responses are never cached (one-shot rescue must stay
one-shot). Config: web.cache_enabled (default on),
web.cache_ttl_minutes (default 20, clamped 1-1440).

Idea credit: query coalescing + num-bucketing pattern observed in
Apodex FrontierAgent (Apache-2.0).
2026-08-25 04:21:45 -07:00
Jan-Stefan Janetzky fb1ec36a4b fix(mcp): treat tools.include: [] as an explicit empty whitelist
_normalize_name_filter([]) returns an empty set, which is falsy, so
_should_register fell through to "no filter" and registered every tool
— the exact opposite of what _apply_tool_selection wrote when the user
unchecked everything in the install checklist ("contributes nothing
until reconfigured"). Whitelist mode is now keyed on the include key
holding a valid filter shape (str/list/tuple/set) rather than on set
truthiness, at both the live-discovery and cached-manifest sites.
Invalid include values keep the old warn-and-ignore behaviour.
2026-08-25 04:21:37 -07:00
Teknium d736f5d53f fix(docker): digest-suffix shared-container identity labels so distinct keys never collide
Review finding on #94633: _sanitize_label_value is lossy ('team/workspace'
and 'team_workspace' both sanitize to 'team_workspace'; >63-char keys
truncate identically), and container reuse is label-keyed — so two teams
with DIFFERENT shared keys could silently attach to one running container
(filesystem, processes, env) while their host sandboxes stayed separate.
Shared-key labels now carry a sha256 digest suffix of the raw key
(deterministic across processes; plain profile labels unchanged for
backward compat). Docs also state the first-creator-wins rule for image/
mounts on a shared container. Adds adversarial collision tests.
2026-08-25 04:00:27 -07:00
Teknium 82b32f32ef feat(terminal): wire shared-container key into profile-scoped resolver and MEDIA delivery
Follow-up on @fangliquanflq's opt-in (#84775): after the profile-scoping fix
(#94560) the container cache key is resolved in _resolve_container_task_id,
so the shared key must unify profiles there too — 'shared:<key>' for every
session of every opted-in profile AND for CLI/no-session runs. Delivery adds
the shared sandbox layout as the first translation candidate. Empty key
keeps strict per-profile isolation; SSH ignores the key entirely.
2026-08-25 04:00:27 -07:00
fangliquanflq 7a67bd07a7 feat(docker): support shared container identities 2026-08-25 04:00:27 -07:00
Teknium 76e306c458 refactor(tools): remove expired BFL FLUX 3 promo core tools (migration v39); FLUX 3 stays via video_gen/FAL for subscribers (#94599)
* refactor(tools): remove expired bfl_flux3_* promo tools; FLUX 3 rides the video_gen provider surface

* test: relay-cutover migration asserts >= v38, not the version literal
2026-08-25 02:45:10 -07:00
Teknium ce9b9a6351 feat(computer_use): guide models from full-screen grabs to interactive lanes
Full-screen captures carry no element tree, so the CaptureResult now has a
'note' field surfaced in the tool summary telling the model to call
capture(app='<AppName>') or capture(app='desktop') when it needs to act on
what it sees. Schema description updated to distinguish app='screen'
(composited full-screen image) from app='desktop' (shell surface with
clickable elements); docs + regression tests (14, sabotage-verified) added.
2026-08-25 02:31:34 -07:00
Teknium 15f7b7293c fix(terminal): persistent Docker containers are profile-scoped, not per-session
Commit a270c4ade's session-key fallback in _resolve_container_task_id was
added to stop cross-profile SSH environment reuse, but it wasn't backend-
gated: persistent Docker silently fragmented into one container per gateway
session, breaking the product contract (one long-lived container per profile,
shared by CLI and every session of that profile). #93950's vanishing MEDIA
attachments were downstream damage.

- persistent Docker (container_persistent: true) now keys to the profile:
  literal 'default' for the default profile (same container as CLI),
  'profile:<name>' for named profiles
- SSH and non-persistent Docker keep session scoping (the original leak fix
  and the #82731 isolation contract are untouched)
- gateway MEDIA translation follows the profile layout and keeps the legacy
  bug-window per-session sandboxes as fallback candidates, trying each until
  the file resolves — old sessions self-heal, no migration
- /root/.hermes credential-surface refusal preserved across all layouts
2026-08-25 02:30:38 -07:00
kshitijk4poor c8c3f4c448 fix(approval): machine-readable outcome parity on the gateway tails + sudo human-wait exclusion (#85125 2e) 2026-08-25 13:27:20 +05:30
Jony 335c60ecdd fix(skills): preserve review marks across contexts 2026-08-24 23:55:39 -07:00
Leegenux 6ce7ab8bfb feat(browser): make snapshot threshold configurable 2026-08-24 21:51:44 -07:00
pierrenode bf8b28f27a fix(tools): route browser snapshot storage through the symlink-safe writer
Today's spill/cache-writer hardening (tools/spill_safety.py,
write_text_exclusive/ensure_spill_dir with O_CREAT|O_EXCL|O_NOFOLLOW)
migrated tools/web_tools.py::_store_full_text() — which writes to the same
cache/web directory with the same content-hash filename scheme — but left
its near-identical sibling, tools/browser_tool.py::_store_full_snapshot(),
on the pre-fix plain open()/write_text() pattern. A pre-planted symlink at
the content-hash path redirected the write onto an arbitrary user-owned
file, same as the sites that commit fixed.

Reproduced live: with a symlink planted at the exact
browser-snapshot-<digest>.txt path (predictable from the snapshot content
hash), the pre-fix write followed the link and overwrote the link's
target with the (secret-redacted but otherwise user/page-controlled)
snapshot content.

Fix mirrors _store_full_text's exact usage: ensure_spill_dir(private=False)
+ write_text_exclusive(private=False, overwrite=True) — not private since
cache/web is bind-mounted into remote backends whose container UID must
read it; overwrite=True because re-snapshotting the same page state
legitimately reuses the same content-hash name (the overwrite path
lstat-unlinks the link itself, never following it to write through).

Added a regression test planting a symlink at the exact digest path and
asserting the link's target is untouched (only the link itself gets
safely replaced by a real file). Mutation-verified: with the fix stashed,
the pre-fix code wrote the snapshot content into the symlink's target
file, reproducing the vulnerability exactly.
2026-08-24 21:45:56 -07:00
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium 48f69e51d3 fix(signal): chunk long standalone sends and cover both delivery paths (salvage #57929 + #67279)
Follow-up to lkz-de's adapter chunking commit: long Signal messages no
longer truncate on ANY delivery path.

- tools/send_message_tool.py: register Signal's 8000-char limit in
  _MAX_LENGTHS (imported from the adapter module so the two paths can't
  drift) so hermes send / cron standalone / MCP sends split via the
  shared truncate_message() pass instead of signal-cli rejecting them.
  Standalone-path idea credited to @5L-hermes01 (#67279).
- tests: regression test proving standalone Signal sends chunk at the
  adapter limit with no truncation footer (fails on pre-fix main).
- docs: Long Messages section on the Signal page (en + zh-Hans).

Both fixes verified by sabotage A/B (tests fail with the respective
half reverted to origin/main) and a real-import E2E: 27k-char message
with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles
in-range UTF-16, lossless reassembly.
2026-08-24 20:03:24 -07:00
kshitij 41447a6d70 Merge pull request #94187 from kshitijk4poor/fix/85125-4b-terminal-treekill
fix(terminal): sweep setsid descendants after local timeout group-kill (#85125 4b)
2026-08-25 03:37:05 +05:30
kshitij 8d29a55bed Merge pull request #94188 from kshitijk4poor/fix/85125-4d-treekill-consolidation
refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
2026-08-25 01:48:09 +05:30
BlackishGreen33 c73d721b1d fix(computer-use): recreate CUA session suspect after MCP timeout (#74799)
An MCP call_tool deadline hit left the cua-driver session wedged for
all later computer-use calls. Mark the session suspect on a
concurrent.futures.TimeoutError and tear down + recreate it before the
next non-lifecycle call; healthy sessions are never restarted.

Fail-closed: the timed-out action may still have taken effect on the
remote screen, so it is never silently replayed — the error result
carries structuredContent.code=timeout_outcome_unknown with
next_step=fresh_state.

Informed by #74877 by BlackishGreen33.

Co-authored-by: BlackishGreen33 <s5460703@gmail.com>
2026-08-25 01:47:21 +05:30
kshitijk4poor 9990bcb8ce fix(terminal): sweep setsid descendants after local timeout group-kill (#85125 4b)
LocalEnvironment._kill_process kills the process GROUP (SIGTERM ->
1s wait -> SIGKILL -> 2s wait), but a descendant that called setsid
escapes the group and survives — the #71148 orphan class, terminal
flavor (issue #84967's local sibling).

Fix: snapshot the descendant set via psutil BEFORE the first SIGTERM
(children reparent to init once the wrapper dies, so a later parent
walk finds nothing — same snapshot-before-signal design as
agent/deadline.py kill_process_tree), then after the existing group
escalation completes, SIGKILL any snapshotted survivor whose pgid is
no longer the (now-dead) group. The TERM->KILL grace window for
in-group members is preserved (interrupts use this path too), the
Windows branch is untouched, and the snapshot is fully guarded — a
broken psutil never breaks the kill path (unit-tested).

Tests: live_system_guard_bypass acceptance test spawning a setsid
grandchild and forcing the timeout path (RED on unmodified file,
GREEN after), plus a psutil-failure unit test.

Docker design note (#84967 open question 1, condensed; full note at
/tmp/4b-docker-design-note.md): the docker backend inherits
base.py:1378 _kill_process, which only proc.kill()s the HOST-side
`docker exec` client — the in-container tree (child of containerd-
shim, not the client) survives every timeout entirely. Option A,
`docker exec <cid> kill -- -<pgid>` with TERM->KILL escalation using
a PGID captured at command start, is surgical and preserves container
state but needs a live container + shell and still misses in-container
setsid escapees. Option B, container restart, is absolute (PID-
namespace teardown kills everything) but destroys all in-container
state mid-session and punishes every other consumer of the shared
persistent container. Recommendation: Option A as a best-effort
_kill_process override (degrade to today's behavior on failure);
reserve restart for the existing container-gone recovery path.
2026-08-25 01:34:56 +05:30
kshitijk4poor 547f985286 refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
Per-site decisions:

1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
   MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
   via a function-local import; keeps the swallow-everything fail-open
   contract and the (proc) -> None signature (agent/shell_hooks.py imports
   it by name; _kill_git_process_tree alias preserved). The old body is kept
   verbatim as _legacy_kill_process_tree and used as fallback when the
   delegation import/call fails. A final proc.kill() is retained on the
   happy path so Popen bookkeeping sees the exit (matches old behavior).

2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
   (delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
   body sent SIGTERM then SIGKILL with zero grace between them; the shared
   primitive sends SIGKILL only. With no grace period the observable effect
   is identical, and the psutil descendant sweep now also reaches
   agent-browser's setsid'd daemon grandchild, which killpg alone missed.
   tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
   the legacy fallback (its assertions describe the fallback's internals).

3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
   MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
   kill when escalate=True) — expressed as two delegated calls:
   kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
   kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
   proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
   body terminated children before the parent; the shared primitive
   signals the group atomically (child is a session leader via
   start_new_session=True) plus an identity-aware descendant sweep —
   strictly wider coverage, same signals.

4. gateway/status.py: KEPT BOTH SITES.
   - terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
     incompatible with the shared primitive — it must RAISE OSError with
     taskkill's stderr on non-zero exit (callers branch on that), falls back
     to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
     single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
     fail-soft primitive would invert the error contract.
   - reap_gateway_children (~l2029): NOT migrated. It operates on a
     pre-snapshotted child list from a parent that is already dead
     (psutil.Process(pid) on the parent would fail), and every signal is
     wrapped in identity/ownership checks the primitive lacks: is_running()
     identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
     plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
     The coupling is the feature; migrating would delete the safety logic.

5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
   Dev tooling that intentionally kills by CAPTURED pgid because the direct
   child is usually already reaped (psutil/pid-based primitive cannot find
   it), and it avoids the psutil import on the test-runner hot path. Its
   docstring already documents why psutil is the wrong tool there.

New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
2026-08-25 01:34:56 +05:30
kshitijk4poor 7dde1b8b0b fix(mcp): resolve tool-call timeouts via the unified deadline layer (#85125 2g)
Both readers of the per-server MCP tool timeout (the connection's run()
and the cache-path registration) read config.get("timeout", 300) as
their own private resolution. Route them through _resolve_tool_timeout:
per-server mcp_servers.<name>.timeout still ALWAYS wins (most specific),
then timeouts.mcp.tool_call from the unified timeouts: section, then
the unchanged 300s default. Values pass through resolve_timeout's
platform clamp; resolution failure falls back to the historical default.

Default-behavior invariance pinned by contract tests (nothing
configured -> exactly 300, per-server beats section, section beats
default, invalid/failed resolution falls back).
2026-08-24 17:12:15 +05:30
Teknium d8d1e18ab9 test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
2026-08-24 03:22:48 -07:00
chelsealong b3730153c3 fix(terminal): cap watch_patterns notifications over a process's lifetime
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
2026-08-24 03:22:48 -07:00
Teknium 42a6d761d2 fix(bot-relay): add shutil.which step to CLI resolution and pin utf-8 decoding on delivery subprocess
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:

- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
  try shutil.which('hermes') before the bare-name fallback, so
  environments with a PATH but no venv sibling resolve exactly what an
  interactive shell would. Platform test switched os.name -> sys.platform
  ('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
  errors='replace' on both subprocess.run sites — without them the
  child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
  Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
  with which=None, and encoding-pin assertions in the deliver transport
  test.

Refs #93590, #93597, #93601
2026-08-24 03:21:37 -07:00
liuhao1024 85cd576b06 test(bot-relay): match delivery CLI by basename in argv filters
CI runners have a real hermes sibling next to the venv python, so
local_delivery_command now resolves an absolute path there — the exact
argv filters in the retry-policy fakes and the relay-methods pins must
match by basename instead of the literal "hermes", mirroring the
_delivery_lock matcher.
2026-08-24 03:21:37 -07:00