_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.
When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.
Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate
Closes the sibling gap of fa3ab2ffd / #42405.
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):
- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
tool-result image supersedes it. The reported session opened with a
~200KB poster that previously survived every compaction. The row keeps
a non-empty text placeholder, so the zero-user-turn guard (#58753) and
role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
anchor (newest is kept) and strip (older collapse to their
text_summary via _strip_images_from_tool_msg, which also drops the
stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
input_image, Anthropic-native image) were already matched by
_IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
determinism (double-run is a no-op returning the same object).
Refs #89938, #89965
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.
Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.
User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.
Refs #89938
Follow-up to the tombstone salvage (review point 1): named_profile_home()
treated ANY path whose ancestor's parent dir was literally named
'profiles' as a named profile until a '.hermes' ancestor appeared. An
unrelated custom HERMES_HOME like /srv/profiles/buildcache would resolve
as profile 'buildcache', and setup_logging would raise FileNotFoundError
instead of creating logs/ — a behavior regression for non-profile users.
Recognition now requires the 'profiles' directory's parent to BE a
Hermes home: the classic ~/.hermes layout, a root carrying home marker
files (config.yaml / .env / state.db — covers Docker/custom roots), a
profiles/.deleted tombstone dir (only ever created by profile delete),
or the process's resolved default Hermes root.
Adds regression tests for the /srv/profiles/<x> false-positive shape
(named_profile_home is None, mkdir_under_hermes_home and setup_logging
still create dirs) plus the tombstone-dir and ~/.hermes anchors. Also
adds encoding= to a bare write_text in the salvaged test file
(Windows-footgun gate).
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.
Adds a regression test covering the assistant + tool_calls +
image-only-content case.
Closes#40463
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:
Replaying the pre-rewrite sidecar would resend exactly what the rewrite
removed, so it must be dropped — the cost is one cache boundary miss,
never wrong content.
`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:
agent._vision_supported = False
_imgs_removed = _strip_images_from_messages(messages) # history
if isinstance(api_messages, list):
_strip_images_from_messages(api_messages)
and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.
Reproduced with the real functions:
history content after strip : [{'type': 'text', 'text': 'look'}]
sidecar still present : True
NEXT TURN sends : 'look<IMAGE BYTES SENT LAST TURN>'
This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.
Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.
tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).
Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged
CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
Follow-up for salvaged PR #84397: Test-SystemNodeReady still said
'too old (Hermes requires Node >=22.22.0)' — Node 25 is not too old,
it's an unsupported line.
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests
* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)
* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
Enough1122 review of #94426:
- Legacy Bot tiles persisted before ownerRoute.connectionId existed carry no
connection id, so the local-delete branch (`ownerConnection === 'local'`)
never matched them and they resurrected the deleted profile on relaunch.
Treat the empty connection id as the local connection.
- A route without profile now throws instead of silently falling into the
local-delete branch and dropping nothing remotely owned.
- Pin the 'local' connection spelling with tests covering the legacy
(no id), canonical ('local'), and divergent ('Local') spellings.
- Comment the live-atom filter arms by bucket and unify the duplicated
#94235 call-site comments in sdk/index.ts and delete-profile-dialog.tsx.
Address review findings on #94426:
- Non-route (local profile) deletes now require the tile's owner
connectionId to be 'local' — a same-named bot on another connection
keeps its tile and live conversation.
- Route identity (profile/targetProfile) goes through normalizeProfileKey
like the non-route name, so whitespace-padded identities match on both
paths; profile case stays exact, consistently on both paths.
- Test lint: drop the unused dropTilesForProfile value import (tests call
it via the fresh module namespace) and replace the forbidden typeof
import() type annotation with a type-only namespace import. New tests:
same-name-other-connection survival, whitespace normalization, and
case-exact identity. eslint: 0 errors; typecheck clean.
A Bot Mode bot cloned via 'Clone from profile' resurrects after
deletion: the desktop keeps the bot's chat tile in local storage
(bot-meta cleanup removes the roster entry, but the persisted tile
survives). On relaunch the tile restores, re-dials the deleted
profile's backend, and ensure_hermes_home() re-creates the profile
directory the delete just removed — the empty skeleton with no
config.yaml reported in #94235.
dropTilesForProfile() removes a deleted profile's session-tile
bucket and every Bot Mode tile whose ownerRoute points at it
(exact connection + backend target when a source-scoped route is
given) from memory and local storage, with discard (no Cmd+Shift+T
undo) semantics. Wired into both delete doors: the core
DeleteProfileDialog and the SDK host.deleteProfile path the Bots
plugin uses.
Regression tests cover local and source-scoped routes; verified red
with the cleanup disabled.
The salvaged proactive validation (#53307/#76896) rejects the
b'\x89PNG...' + zero-padding fixtures test_vision_tools.py used.
Replace all 11 sites with a real 1x1 PNG (+ ignored trailing padding
for size-sensitive tests), and pin _resize_image_for_vision passthrough
in the size-limit test so it exercises rejection, not resize recovery.
With both salvaged validators in place, the resolver-boundary verify()
(#53307) catches a truncated PNG before the full-decode gate (#76896)
runs. Either error message means the bytes stayed out of history.
Pillow is an optional dependency in this codebase (every other PIL use in
tools/vision_tools.py imports lazily and falls back). Both salvaged
validators now distinguish 'PIL missing' (pass through, header-only
sniff) from 'decode failed' (reject), so a Pillow-less install keeps
working instead of rejecting every PNG.
Follow-up to salvaged #53307 (@CannibalKush) and #76896 (@HaiyiMei).
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.
That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.
Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.
Fixes#76039
Second review round: honoring base64Encoded recursively let any nested
dict spoof {"base64Encoded": true, "data": "<secret>"} past the redactor
(Runtime.evaluate returns arbitrary by-value JSON), and two carriers were
missed entirely — Network.streamResourceContent returns unflagged binary
bufferedData, and Network.getRequestPostData's postData was not covered.
Replace the ambient field sets with per-method exact result-path specs:
_CDP_ALWAYS_BINARY_PATHS for declared-binary paths (screenshots, PDFs,
streamResourceContent, beginFrame screenshotData, the nested
CacheStorage.requestCachedResponse.response.body) and
_CDP_FLAGGED_BINARY_PATHS for paths whose carrier object's base64Encoded
sibling gates the exemption (Network/Fetch.getResponseBody body, IO.read
data, getRequestPostData postData). Path suffixes propagate only into the
matching subtree, so base64Encoded is type information solely on trusted
carrier objects — never ambient trust in nested JSON.
Architecture-review follow-up: the method-scoped binary_payload flag skipped
redaction for every string anywhere in the result of the two listed methods,
and the same corruption stayed reachable through Network.getResponseBody /
Fetch.getResponseBody / IO.read / Network.streamResourceContent. Make the
exemption field-scoped instead: an explicit schema exempts exactly
Page.captureScreenshot.result.data and Page.printToPDF.result.data (carriers
with no flag of their own), and any dict whose base64Encoded sibling is
exactly True exempts its body/data/bufferedData string (the protocol's own
discriminator — text bodies with base64Encoded: false stay redacted). Every
other string in every result keeps full secret redaction.
_redact_cdp_output applied redact_sensitive_text(force=True) to every string
in CDP results, including the base64 screenshot/PDF payload of
Page.captureScreenshot and Page.printToPDF. The Fernet pattern (gAAAA + base64
alphabet) matches arbitrary spans inside such payloads wherever gAAAA follows a +
or /, collapsing them to first6...last4: decoded PNGs came out corrupt (valid header,
CRC failures mid-IDAT, no IEND), the persisted full copies in tool_result_storage
were redacted too, and vision_analyze then embedded corrupt images that the provider
rejected with 400 invalid_image - killing resume sessions with a misleading provider
error. Skip redaction for the two binary-payload methods: the payload is binary,
not free text, so there is no secret to protect there. Every other method keeps
full redaction.
_runtime_model_config merges the agent's current identity onto the row's
existing model_config JSON. For model and provider it only SET the key
when the agent attribute was truthy, while base_url/api_mode/
reasoning_config/service_tier already deleted stale values when falsy.
When an agent rebuilt with an empty provider (inheriting the profile
default) was persisted, the previous provider/endpoint survived in
model_config while _persist_live_session_runtime updated the model
column separately. Resume then read the fresh model from the column but
the STALE provider from model_config, silently routing the resumed chat
to the wrong endpoint (e.g. a VeniceAI/empero route under a model that
should run on the profile default).
Apply the same delete-on-falsy rule to model and provider, mirroring the
or-None deletion the CLI path (_persist_model_switch_to_session) already
uses, so a stale session state can never survive into a resume override.
Existing desynced rows self-heal on the next live metadata persist.
Adds regression tests: merge drops stale provider/model when the agent
attribute is falsy, a truthy provider overwrites the stale value, resume
overrides fall back to the billing provider instead of the stale
endpoint, a real-DB round trip heals an already-desynced row, and a
first write (existing=None) reflects only the agent's current identity.
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.
Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes#76884.