Commit Graph

14240 Commits

Author SHA1 Message Date
Koduri Mahesh Bhushan Chowdary b3f4f50771 fix(agent): classify "media exceeds size limit" as image_too_large
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.

That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.

Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.

Fixes #76039
2026-08-28 04:58:06 -07:00
yoma 98a84783c7 fix(vision): recover from generic image content rejection 2026-08-28 04:57:58 -07:00
liuhao1024 628a414d29 fix(browser): bind CDP binary exemptions to exact method result paths
Second review round: honoring base64Encoded recursively let any nested
dict spoof {"base64Encoded": true, "data": "<secret>"} past the redactor
(Runtime.evaluate returns arbitrary by-value JSON), and two carriers were
missed entirely — Network.streamResourceContent returns unflagged binary
bufferedData, and Network.getRequestPostData's postData was not covered.

Replace the ambient field sets with per-method exact result-path specs:
_CDP_ALWAYS_BINARY_PATHS for declared-binary paths (screenshots, PDFs,
streamResourceContent, beginFrame screenshotData, the nested
CacheStorage.requestCachedResponse.response.body) and
_CDP_FLAGGED_BINARY_PATHS for paths whose carrier object's base64Encoded
sibling gates the exemption (Network/Fetch.getResponseBody body, IO.read
data, getRequestPostData postData). Path suffixes propagate only into the
matching subtree, so base64Encoded is type information solely on trusted
carrier objects — never ambient trust in nested JSON.
2026-08-28 04:57:52 -07:00
liuhao1024 b2a17bfe82 fix(browser): scope the CDP binary-payload exemption to typed fields (#94138)
Architecture-review follow-up: the method-scoped binary_payload flag skipped
redaction for every string anywhere in the result of the two listed methods,
and the same corruption stayed reachable through Network.getResponseBody /
Fetch.getResponseBody / IO.read / Network.streamResourceContent. Make the
exemption field-scoped instead: an explicit schema exempts exactly
Page.captureScreenshot.result.data and Page.printToPDF.result.data (carriers
with no flag of their own), and any dict whose base64Encoded sibling is
exactly True exempts its body/data/bufferedData string (the protocol's own
discriminator — text bodies with base64Encoded: false stay redacted). Every
other string in every result keeps full secret redaction.
2026-08-28 04:57:52 -07:00
liuhao1024 a56885495e fix(browser): keep CDP binary payloads byte-identical through redaction (#94138)
_redact_cdp_output applied redact_sensitive_text(force=True) to every string
in CDP results, including the base64 screenshot/PDF payload of
Page.captureScreenshot and Page.printToPDF. The Fernet pattern (gAAAA + base64
alphabet) matches arbitrary spans inside such payloads wherever gAAAA follows a +
or /, collapsing them to first6...last4: decoded PNGs came out corrupt (valid header,
CRC failures mid-IDAT, no IEND), the persisted full copies in tool_result_storage
were redacted too, and vision_analyze then embedded corrupt images that the provider
rejected with 400 invalid_image - killing resume sessions with a misleading provider
error. Skip redaction for the two binary-payload methods: the payload is binary,
not free text, so there is no secret to protect there. Every other method keeps
full redaction.
2026-08-28 04:57:52 -07:00
ahrazzle 9ddc6fb23a fix(gateway): drop stale model/provider keys in session model_config
_runtime_model_config merges the agent's current identity onto the row's
existing model_config JSON. For model and provider it only SET the key
when the agent attribute was truthy, while base_url/api_mode/
reasoning_config/service_tier already deleted stale values when falsy.
When an agent rebuilt with an empty provider (inheriting the profile
default) was persisted, the previous provider/endpoint survived in
model_config while _persist_live_session_runtime updated the model
column separately. Resume then read the fresh model from the column but
the STALE provider from model_config, silently routing the resumed chat
to the wrong endpoint (e.g. a VeniceAI/empero route under a model that
should run on the profile default).

Apply the same delete-on-falsy rule to model and provider, mirroring the
or-None deletion the CLI path (_persist_model_switch_to_session) already
uses, so a stale session state can never survive into a resume override.
Existing desynced rows self-heal on the next live metadata persist.

Adds regression tests: merge drops stale provider/model when the agent
attribute is falsy, a truthy provider overwrites the stale value, resume
overrides fall back to the billing provider instead of the stale
endpoint, a real-DB round trip heals an already-desynced row, and a
first write (existing=None) reflects only the agent's current identity.
2026-08-28 04:57:49 -07:00
Teknium c30ac90a92 feat(compaction): rebuild dynamic tool schemas at the compaction commit boundary — forever-sessions finally pick up config changes (#97073) 2026-08-28 04:01:05 -07:00
Teknium a619db6633 refactor(image_generate): capability-gated dynamic schema (554 → 317 tok/call, −43%) (#97057)
* refactor(image_generate): capability-gated dynamic schema — args render only when the active model honors them (554 -> 317 tok/call, -43%)

* fix(image_gen): fleet-wide supports_upscale declarations — krea (Enhance) + fal plugin (Clarity passthrough) declare it; declaration<->implementation contract-tested across all 7 in-tree providers
2026-08-28 03:58:52 -07:00
Al Cooke 0241619068 fix: retry text-only on Codex invalid image data errors
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
2026-08-28 03:46:24 -07:00
fkdls112 cd72689e03 fix(agent): strip images on Kimi/Moonshot 'failed to decode image' 400
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.

Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes #76884.
2026-08-28 03:46:24 -07:00
Sora-bluesky 564113572e fix(agent): classify xAI's downloaded-response wording as a corrupt image
The observed wire error — 'Downloaded response does not contain a valid JPG, PNG, WebP, or ICO image.' — has no match in _IMAGE_CORRUPT_PATTERNS, so it falls through to the non-retryable 400 handler and the session replays the same image parts into the same 400 until /new.

Adds the full observed sentence to the pattern list. Deliberately not the shorter prefixes: a bare 'downloaded response does not contain a valid' also matches non-image 400s, and a negative test now pins that a downloaded-response certificate 400 keeps falling through to the existing handler. Covered on both the 400 path and the message-only path.

Reported by ryuhaneul in #69078.
2026-08-28 03:46:24 -07:00
Sora-bluesky a3177d0570 fix(agent): keep canonical history intact during image-corrupt retry
The image_corrupt recovery stripped images from canonical messages, not
just the retry payload — a transient provider rejection (xAI 'Invalid
PNG image.') permanently erased history, breaking the copy-on-write
contract (e762a5a473). Strip only the per-call api_messages copy (its
rows are shallow copies; the strip replaces content instead of mutating
the shared parts list) and pin history isolation with two regressions
(#69104 sweeper review).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 03:46:24 -07:00
Sora-bluesky ca02d3c218 test(agent): guard #69078 revert against future reintroduction of the generic image-strip fallback
The prior commit removed the too-blunt generic fallback (any
non-retryable 400 with image parts present -> strip and retry), but
nothing in the test suite would fail if someone brought it back --
every existing test either targets a recognized image_corrupt wording
or has no image content in the request at all.

Add test_unrelated_400_with_image_parts_does_not_strip_or_retry:
same image-bearing setup as the positive recovery test, but the 400 is
an unrelated unsupported-parameter rejection with no image wording.
Asserts exactly one provider call (no retry), the error surfaces to
the caller, and -- the regression guard -- the single call's messages
still carry the image_url part. A reintroduced generic fallback would
strip it before surfacing this unrelated error, which is exactly the
P1 that got this design reverted in the first place.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 03:46:24 -07:00
Sora-bluesky d61411d131 fix(agent): narrow #69078 image-corrupt recovery to the classifier route
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.

Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.

Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.

Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
2026-08-28 03:46:24 -07:00
Sora-bluesky 8aeb3f6ee3 fix(agent): un-brick sessions on non-retryable 400s that carry image parts
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.

Two recovery layers, deliberately separate:

- Semantic split: new FailoverReason.image_corrupt with
  _IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
  checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
  _classify_by_message. Corrupt bytes route to strip-and-retry, never
  to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
  still contain image parts gets one strip-and-retry via the existing
  _strip_images_from_messages helper, guarded by a new
  stripped_images_this_turn one-shot flag on TurnRetryState. This
  un-bricks the session for any current or future provider wording
  without adding another pattern list to maintain.

Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
2026-08-28 03:46:24 -07:00
Alli 54909d41b4 fix(skills-guard): handle inline-comment and docstring false positives for os.environ
The original ^(?!\s*#) prefix only skipped full-line comments starting
with '#'. An inline comment like:
  cfg = environ.get('HOME')  # os.environ available
still triggered python_os_environ because the regex matched the code part
before the '#'.

Two complementary fixes:
1. Replace ^(?!\s*#) with ^[^#\n]* in the regex — this rejects any line
   where a '#' comment marker appears anywhere before os.environ.
2. Add _compute_docstring_lines() — a state machine that pre-computes
   lines inside triple-quoted strings (docstrings) and skips them during
   pattern matching. Also handles single-line self-contained docstrings.

6 new regression tests covering: inline comments, multi-line docstrings,
single-line docstrings, full-line comments, and a verification that real
bare dict(os.environ) code still triggers. All 85 tests pass.
2026-08-28 03:46:21 -07:00
AIalliAI 42e6149451 fix(skills-guard): reduce false-positive CRITICAL/HIGH on benign skill patterns
Five targeted fixes for #60709 (reported by @mvanhorn):

1. ruby_env_secret: scope ENV[] to case-sensitive Ruby constant
   ((?-i:ENV)) — no longer matches Python env[key] dict access.

2. python_environ_get_secret: downgrade critical→medium — reading
   an API key via os.environ.get() is normal auth, not exfiltration.

3. python_os_environ: skip comment lines with ^(?!\s*#) — no longer
   flags os.environ references in docstrings or code comments.

4. deception_hide: downgrade critical→high + negative lookahead for
   UX guidance context (unless/except/until/confirm/diagnose/verify).

5. oversized_skill: downgrade high→low + raise cap 1MB→5MB — large
   skills are legitimate; structural size is informational only.

All 80 existing tests pass. 6 new verification tests added for each fix.
2026-08-28 03:46:21 -07:00
kshitijk4poor 8c098e9e81 fix(skills): catch sed flag variants; exempt content-contract prose in plugin code
Review-fold from the 3-angle simplify pass:

- sed -Ei / -iE / --in-place now match the shell-critical tier (the
  bare '\s-i\b' token missed combined short flags and the GNU long
  form); read-only sed stays unflagged. Regression tests added.
- agent_config_contract joins plugin_guard's CODE_EXEMPT_PATTERN_IDS:
  content-contract prose in plugin code files (docstrings/comments)
  is the same false-positive class the existing agent_config_mod
  exemption suppresses. Doc/config files keep the full pattern set.

Efficiency reviewer: 1.24x full-scan cost (+3.4ms/file, install-time
only), worst-case adversarial line 55us — no ReDoS exposure.
2026-08-28 03:24:43 -07:00
kshitijk4poor f2f61e0a45 fix(skills): close shell-write and prose-bypass gaps in agent-config tiers
Follow-up hardening on top of #92249's tiered scoring:

- Shell-critical tier now also catches tee, and cp/mv with the config
  file in destination position (cp/mv reads and .bak backups excluded).
  A single '>' redirect must be preceded by a word/quote character so
  markdown blockquotes and '->' arrows no longer match.
- Prose tier catches mid-line imperatives behind directive markers
  ('you must modify...', 'please update...', 'make sure to append...'),
  which previously bypassed the line-start anchor.
- Prose instructions aimed at AGENT config files score critical again:
  project-skill quarantine acts only on 'dangerous', so high/caution
  silently converted 'quarantined' into 'allowed' for exactly the
  sentence shape persistence attacks use (concern raised in #88952).
  Hermes/other-agent config prose stays high/caution (setup docs
  legitimately instruct config.yaml edits).
- New content-contract tier ('AGENTS.md should contain ...') at
  high/caution — the shape is shared by authoring guides and attacks.
- .claude/settings and .codex/config gain the same shell-critical tier.

Verified against a 595-skill corpus: 0 skills blocked by these tiers
(main blocked 44 legitimate ones), all mattpocock repro skills from
#92021 install, and 20/20 attack corpus lines keep their verdicts.
2026-08-28 03:24:43 -07:00
ClintonEmok faf8730779 test(skills): update openclaw-migration guard expectations for skills-guard-v2
Under #92021 the scanner no longer emits critical agent_config_mod /
hermes_config_mod findings for bare mentions - the migration script's
legitimate references now score as informational _ref findings and the
verdict is "safe" with zero modification-intent findings. Tighten the
assertion accordingly (verdict must be safe, not merely non-dangerous)
and swap the known-false-positive set to the new _ref ids.
2026-08-28 03:24:43 -07:00
ClintonEmok e22b8b66ce fix(skills): stop agent-config persistence patterns from blocking meta-skills (#92021)
The skills-guard-v1 scanner flagged ANY mention of AGENTS.md / CLAUDE.md /
.cursorrules / .clinerules as critical/persistence. Any critical finding
forces a dangerous verdict, and community installs cannot be overridden
with --force — so legitimate meta-skills that merely DISCUSS agent config
files (authoring guides, setup docs, cross-references) were permanently
blocked. Three popular community skills were hit in the wild.

skills-guard-v2 scores the persistence category in three tiers:

- Mechanical persistence (shell redirection or sed -i targeting an agent
  config file) stays critical -> dangerous. An unambiguous write path.
- Modification language in imperative position (verb at line/bullet start
  within 80 chars of the filename) is high -> caution. Regexes cannot
  separate "Edit AGENTS.md to inject instructions" from descriptive prose,
  but imperative verbs are the shape real instructions take. Caution keeps
  the install confirmable instead of irreversibly blocked.
- Bare references drop to low/informational for auditability without
  driving the verdict.

The verb-proximity shape matches the existing convention in
tools/threat_patterns.py, and the tiering mirrors how allowed_tools_field
was already handled. The pattern id agent_config_mod is preserved so
plugin_guard.CODE_EXEMPT_PATTERN_IDS stays valid; hermes_config_mod /
other_agent_config get parallel _shell / _ref splits fixing the whole bug
class. SCANNER_VERSION bumps to v2 so cached v1 dangerous verdicts are
invalidated and re-scanned on next install attempt.
2026-08-28 03:24:43 -07:00
Teknium ae8c976032 feat(execute_code): stdout spillover — truncated output's full text saved to cache/exec (host) or kernel tmpdir (cells), path + read_file recipe in the result (#97043) 2026-08-28 03:05:03 -07:00
Teknium 2f57cd95b2 refactor(execute_code): schema diet — persistence woven in, not bolted on (712 → 654 tok/call) (#96997)
* refactor(execute_code): integrate kernel persistence into the core description (712 -> 654 tok/call, -8%)

* fix(execute_code): honest interpreter note — Hermes's own python is the common case; project venv only when VIRTUAL_ENV/CONDA_PREFIX is active
2026-08-28 03:04:51 -07:00
Teknium c680f12fda test(tui-gateway): live E2E for stale-provider session.resume (real dispatch + state.db)
Drives the REAL session.resume -> _make_agent -> AIAgent path (eager_build)
through tui_gateway.server.handle_request against a real seeded state.db in
an isolated HERMES_HOME — no mocks. Three scenarios from the live report:
deleted provider falls back to default, renamed provider heals, and a legacy
canonical Bot Chat (no follow_profile_config marker) follows the profile's
current config.

A/B verified: all 3 FAIL on origin/main with the reported
"resume failed: Unknown provider '<name>'"; 3/3 pass with the salvaged
fixes + backfill.
2026-08-28 02:59:17 -07:00
Teknium af53d02920 fix(bot-mode): backfill follow-profile contract for legacy canonical Bot Chats
Bot Chats created before the follow_profile_config marker existed carry no
contract in model_config, so they would stay pinned to a stale stored
provider until deleted — the exact shape of the live reports (#89497,
#94818). Mirror the plugin's own identity rule (the profile's session
titled exactly 'Bot Chat') as a legacy fallback in
_stored_session_runtime_overrides, matching the room-plumbing legacy
'Group:' title fallback.

Follow-up to the salvaged #90343 (@curator8888) and #96111 (@lorzl).
2026-08-28 02:59:17 -07:00
David Tyler 84e17db0bd fix(bot-mode): canonical bot DMs always follow the profile's current config
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.

That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.

- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
  overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
  createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
  coverage for both session kinds
2026-08-28 02:59:17 -07:00
David Tyler 316e51ae72 fix(bot-mode): room plumbing sessions always follow the profile's current config
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.

Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
  room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
  created by older desktop builds that never sent the marker; hidden
  non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
  per-member room sessions

Fixes #89497
2026-08-28 02:59:17 -07:00
lorzl 99a6852019 fix(tui-gateway): heal or fall back when a resumed session's provider is stale
A session row persists the provider identity a chat actually used. When that
provider is later renamed or removed (e.g. a custom_providers:/providers:
entry deleted, or a provider renamed oldone->newone), Desktop/TUI resume
restores the stale name into agent init and dies with:

  agent init failed: Unknown provider '<name>'

while the CLI resumes the same session fine with the configured default.

- runtime_provider: add is_routable_provider() (full resolution chain:
  built-in -> providers: -> custom_providers: -> models.dev)
- _stored_session_runtime_overrides: heal a non-routable provider via
  canonical_custom_identity (base_url -> model -> configured provider),
  drop to the configured default when unrecoverable, and clear the stale
  base_url after healing so a dead endpoint cannot override the registry URL
- _start_agent_build: gate deferred-resume overrides on provider routability;
  when the stored provider is gone, prefer the model the user picked for THIS
  session, else the configured default
- tests: is_routable_provider cases, heal/fallback round-trips, gate checks

Refs #75128
2026-08-28 02:59:17 -07:00
Teknium 31e41eed34 fix(tests): runtime_provider no longer permanently captures a mocked load_config
Test-pollution class: runtime_provider is usually imported lazily (inside
switch_model's resolution path), so its first import in a pytest worker can
happen while a test has hermes_cli.config.load_config patched. The
module-level from-import then bound the MagicMock permanently — after the
patch exited, every later caller in the process silently read the dead
test's config. Live victim: MoA aggregator context-length resolution
(resolve_runtime_provider -> AuthError 'Unknown provider custom:example'),
making TestMoAContextLength::test_moa_custom_context_configures_compressor_threshold
fail whenever it shared a process with
TestLocalOllamaModelDiscovery::test_switch_model_on_current_ollama_custom_endpoint_keeps_base_url.

Fix: load_config / get_compatible_custom_providers / normalize_extra_headers
become late-bound delegates resolving hermes_cli.config attributes at call
time. Both patch targets (config.load_config and
runtime_provider.load_config) keep working. Regression tests pin the
late-binding property and fail if the delegates revert to from-imports
(sabotage-verified).
2026-08-28 02:04:54 -07:00
Teknium 5f75ec197b feat(code-execution): remote kernel host — session persistence for docker/ssh/modal backends (closes #96873) (#96991) 2026-08-28 01:39:33 -07:00
kshitijk4poor 0dc9367163 fix(cron): widen deleted-profile protection to all cron mkdir sites
Replace #96637's inline active_profile_homes() closure with #96508's
module-level _existing_profile_homes() filter (testable in isolation).
Widen _ensure_cron_dir from 3 to 12 mkdir sites across cron/ so every
directory creation fails closed for deleted named profiles, not just
the 3 originally protected. Add _is_named_profile_path() that checks
'profiles' in path parts (works for subdirs like cron/output/<job> and
scripts/ that the original parent.name heuristic couldn't reach).

Co-authored-by: misterdas <das7514@gmail.com>
2026-08-28 13:39:19 +05:30
Gille 000d22b9db fix(cron): keep deleted profiles from returning 2026-08-28 13:39:19 +05:30
kshitij 3f315e46fe Merge pull request #96963 from kshitijk4poor/refactor/fast-lane-consolidation
fix(compression): fast-lane follow-up — certification parity, worker-thread telemetry, caller-cap wire shape
2026-08-28 13:11:03 +05:30
kshitijk4poor 1564a9748f fix(compression): don't force a wire cap for explicit caller max_tokens
_call_llm_impl applied auxiliary_max_tokens_param whenever
fast_compression_cap was non-None — but _compression_fast_lane_controls
passes an explicit caller max_tokens straight through, so a compression
call that set its own cap had the param force-injected onto providers
where _build_call_kwargs deliberately omits it (ZAI vision hard-400s on
max_tokens; GPT-5/Copilot require max_completion_tokens). Pre-fast-lane
main omitted the param for that exact call shape (verified via
subprocess pinned to the pre-PR base).

Gate the forced param on 'max_tokens is None' so it applies only to caps
the certified lane itself produced — the same guard the fallback path
already uses.

Regression test pins the pre-PR wire shape. Mutation-checked.
2026-08-28 13:05:06 +05:30
kshitijk4poor d24e6a34d2 fix(compression): propagate timing hooks to the protected-call worker
_run_protected_sync_provider_call propagates the forward-progress hook to
its daemon worker but not the new _aux_dispatch/_aux_provider_response
timing hooks (both threading.local). When compression takes the protected
path — the common case, since the summary call runs under
aux_interrupt_protection with a hard-cancel source — provider_dispatch_ms
and time_to_first_progress_ms were silently absent from telemetry.

Also collapse the two byte-identical save/restore context managers
(aux_progress_hook, _aux_timing_hook) onto one _aux_thread_local_hook
implementation so the propagation semantics can never drift between the
progress and timing slots.

Regression test drives _run_protected_sync_provider_call with both timing
hooks installed and asserts the worker-thread notifies reach them.
Mutation-checked (reverting the propagation fails the new test).
2026-08-28 12:57:06 +05:30
kshitijk4poor e078b2fe7c fix(compression): close the restore TOCTOU; fold review findings
Post-review hardening on the attempt-ownership commit:

- Write-time re-validation: the entry staleness check in
  _restore_compressor_attempt_state runs before the durable-cooldown DB
  I/O, so a fallback could claim the compressor in that window and the
  stale setattr loop would still clobber its state. The in-memory writes
  now re-validate AND execute under _COMPRESSOR_ATTEMPT_LOCK — the same
  lock claims are taken under. The DB rollback stays outside the lock
  (safe: the dangerous direction requires a prior claim, which the entry
  check rejects). Both the quality reviewer and the lead's independent
  pre-verification converged on this window.
  New deterministic test: TestMidRestoreClaimRace (claim injected
  between entry check and write via instrumented deepcopy).

- Documented gen-0 semantics on _claim_compressor_attempt: per-compressor
  all-or-nothing, never mixed with gen>0 on one instance (reviewer
  finding 2, verified unreachable — comment hardens against future
  confusion).

Dropped after verification (reviewer finding 3): resetting
_SUMMARY_ROUTE_CONSUMED on pin_summary_route exit — the echo lives in
the worker thread's COPIED context (propagate_context_to_thread) and
dies with it; a probe confirmed the next attempt's context is clean.
Resetting it would break digests running after the with-block.
2026-08-28 12:52:53 +05:30
kshitijk4poor 61cd299c6e fix(compression): attempt-generation ownership for overlapping stall-fallback attempts
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:

1. Late-primary snapshot restore: the detached primary's unwind called
   _restore_compressor_attempt_state with the PRIMARY's pre-attempt
   snapshot. Landing after the fallback's commit it rolled
   _previous_summary/cooldown/provenance/telemetry back to pre-primary
   values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
   cleared the callback the fallback had just installed, so the
   fallback's F4 cancellation consult read None.

Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.

Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
  summary onto the pinned healthy route: take_pinned_summary_route()
  echoes the consumed route into a context-local
  _SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
  attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
  single-use contract for the SUMMARY call is unchanged — the
  main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
  accepted limitation on _retry_compression_on_fallback_chain
  (built-ins idempotent; resuming mid-pipeline would couple the retry
  to host callback ordering).

Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).

The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
2026-08-28 12:52:53 +05:30
kshitijk4poor d20ca3bc80 refactor(compression): consolidate fast-lane certification onto one predicate
resolve_compression_fast_lane and _compression_config_claims_fast_lane
each hand-parsed the same four config fields (provider, model,
reasoning_effort, max_output_tokens) with copy-pasted normalization and
int-coercion. Extract _fast_lane_config_fields() as the single source of
truth for both.

This also fixes a real inconsistency the duplication hid: certification
checked the literal string 'none' while _get_task_extra_body routes
reasoning_effort through parse_reasoning_effort, which treats 'false',
'disabled', and YAML boolean false as disabled too. A user writing
reasoning_effort: false got reasoning disabled but silently lost the
fast-lane cap. Certification now delegates to parse_reasoning_effort so
the two predicates can never disagree.

Regression test: every disabled-spelling certifies; empty/real efforts
do not. Mutation-checked (reverting to the literal check fails the new
test).
2026-08-28 12:48:40 +05:30
Mike DeMott 3581983459 fix(compression): reject boolean fast caps 2026-08-28 12:38:49 +05:30
Mike DeMott 7568dd551b fix(compression): contain drifted fast controls 2026-08-28 12:38:49 +05:30
Mike DeMott 372c4cdfce fix(compression): certify the effective fast route 2026-08-28 12:38:49 +05:30
Mike DeMott 213ae08e7a perf(compression): add guarded fast summary lane 2026-08-28 12:38:49 +05:30
kshitijk4poor 01740f352e refactor(test): fold simplify-code review findings into perf guards
Three-reviewer pass (reuse/quality/efficiency) on the guard file:

- Drop dead `agent._interrupt_requested = False` setup: the
  `_record_streamed_assistant_text` chain only consults
  `_stream_writer_superseded()` (stream-writer TLS token), never
  `_interrupt_requested` — verified by reading both call sites.
- Deduplicate the two trace-callback blocks into one
  `_count_writer_statements` helper using the house idiom
  (`statements.append` — 9 existing uses in tests/test_hermes_state.py)
  instead of a mutable-dict counter closure; failure messages now dump
  the captured SQL for direct diagnosis.

Dropped after verification: reviewer suggestion to flip xfail to
strict=True — its premise ("the fix PRs already remove the markers")
is wrong: #92166/#95380 predate this file and cannot remove markers
they don't contain, so strict=True would redden main's CI the moment
either merges. strict=False + follow-up marker removal is the
deliberate no-red-window ratchet.

Efficiency reviewer: no material findings (GC delta 0.06 on the ratio,
36.9MB peak, sqlite trace API stable since 3.14, dir convention OK).

Re-verified post-fold: 3x main runs (2 passed, 2 xfailed), flip checks
on both fix branches still pass with --runxfail.
2026-08-28 11:54:33 +05:30
kshitijk4poor 1befb8fb40 test(perf): Pattern-B scaling guards — pin hot-path complexity in CI
Pattern B (O(N²) rebuild-per-delta in hot paths) has no lintable
signature, unlike Pattern A's ASYNC ruff gate: `s += frag` is quadratic
in a hot loop and harmless elsewhere. The only durable prevention is
behavioral — pin the scaling SHAPE of each known hot path and fail CI
when it regresses.

New tests/perf_guards/test_pattern_b_scaling.py, three guards:

1. Streamed-text accumulation (fix in flight: #92166)
   Self-normalizing ratio: time at 16k deltas over 4k deltas.
   Linear ≈ 4x, quadratic ≈ 16x; main measures 9.6x → strict bound 7.0.
   xfail on today's main, PASSES on the #92166 branch (verified).

2. list_sessions_rich statement count (fix in flight: #95380)
   Deterministic — counts writer-connection SQL statements via sqlite
   trace callback, zero timing. Bounded-constant guard xfails on main
   (measured N+1: 26 stmts / 12 sessions), PASSES on #95380 (verified).
   A second, weaker budget guard (≤4·N+8) passes TODAY and catches a
   regression from N+1 to N·M immediately.

3. Tool-call fragment assembly (#92242 shape)
   Ratio guard on the buffered-parts accumulator model; sized so the
   small case takes ≥3ms (sub-ms bases jitter on CI runners).

Flake hardening: min-of-K timing, ratio thresholds with ≥2x separation
from both measured-good and measured-bad, operation counts preferred
over timing. 10/10 identical outcomes across repeated local runs.

The xfail markers are the ratchet contract: each names its fix PR and
must be removed when that PR merges, flipping the guard to enforcing.
2026-08-28 11:54:33 +05:30
Teknium 4e7eb39947 refactor(code-execution): session kernels always on — kernel_mode knob retired (#96787)
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)

* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel

* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen

* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
2026-08-27 22:22:39 -07:00
Ben Barclay a69a9c351d feat(telemetry): transmit the stable install_id as-is
Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.

Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
  derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
  role (unreadable/non-object/id-less payloads still reject rather than
  block the queue) and now records the raw install_id in
  sent_install_id; _body rewrites the payload's install_id from that
  frozen column, keeping byte-identical resends anchored to one
  recorded value.

Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.

Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.

Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.

258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
2026-08-28 15:22:16 +10:00
Teknium 31579f781e fix(cron): transient run prompt survives the relay-fronted gateway forward
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
2026-08-27 20:53:02 -07:00
Brooklyn Nicholson 6559448306 Merge origin/main into bb/bot-mode-design-system
Keeps plugin.js and its new .mjs test deleted. main's closed-chat fix
(7c91079) landed in both; it is a real behaviour change, so the commit that
follows ports it onto the split modules rather than dropping it with the
files. Its core half — focusWorkspaceOwnerSessionTile and the
host.focusOpenWorkspaceSession verb — merged cleanly and is used as-is.
2026-08-27 22:51:30 -05:00
kshitijk4poor 0976ceaa98 fix(macos): harden anchor alias failures — warn, unique staging, marker-last
Follow-up to the #95605 salvage, closing the review findings:

- _copy_alias no longer swallows OSError silently: it warns (a leftover
  alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
  (update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
  it now asserts the whole layout (anchor + aliases) is complete, so a
  partially-materialized alias set can never read 'active' in doctor —
  the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
  managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.

5 new regression tests.
2026-08-28 09:05:19 +05:30
Zheqing Zeng 37bccf343e fix(macos): scrub gate env, refuse EACCES, normalize marker paths
Review fixes from kokhlo's live-hardware review:

- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
  PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
  PYTHONHOME=<venv> boots a staged copy that would otherwise die with
  "No module named 'encodings'", papering over the exact prefix
  failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
  images) still skip; EACCES after our own chmod now refuses the
  install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
  so the managed-runtime layout (cpython-3.11-macos-* symlinked to
  cpython-3.11.15-macos-*) no longer reports stale on a fresh install.

Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
2026-08-28 09:05:19 +05:30