Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).
The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.
Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):
1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
stray but leaves the declaring assistant message carrying the now
unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
displaced result still exists in the transcript, so the id survives
the set-subtraction, no stub is injected, and the payload 400s.
Fix both layers so every path is order-independent:
- repair_message_sequence: new Pass 2 prunes tool_calls that have no
result in the immediately-following tool run (matching on id or
call_id, same superset rule as Pass 1). If pruning empties the turn
(no content/reasoning left), the whole message is dropped rather than
sending an empty assistant message. Codex interim turns are exempt,
as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
rolling positional walk that drops results not immediately following
their declaring assistant (including results appearing BEFORE their
call) and injects stub results for positionally-uncovered calls even
when a mispositioned result exists elsewhere.
Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.
Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.
Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.
Salvaged from #88960. Fixes#47339.
Pattern-A architectural fix: blocking calls inside async functions freeze
the gateway/uvicorn event loop for every adapter, timer, and health check.
Known incidents: 17-minute getaddrinfo freeze (#91912 class), 10s restart
freeze in start_gateway (#36163).
Fixes at the four unguarded core sites:
- gateway/platforms/webhook.py: `gh pr comment` subprocess (30s timeout)
now runs via asyncio.to_thread — a webhook delivery no longer freezes
every other platform for the duration of a network call.
- gateway/run.py start_gateway --replace: two time.sleep() waits (10s +
5s worst case) become await asyncio.sleep() (re-lands #36163 at current
line numbers, credit AhmetArif0).
- gateway/slash_commands.py /save: session render + file write move off
the loop (scales with transcript size).
- hermes_cli/web_server.py voice TTS: multi-MB audio file read + unlink
move off the loop.
Prevention gate so the bug class cannot re-enter:
- pyproject.toml [tool.ruff.lint] select gains ASYNC210/220/221/251
(blocking HTTP / Popen / subprocess.run / time.sleep in async def).
These run in the existing blocking `ruff check .` CI job.
- Frozen ratchet baseline in per-file-ignores for the remaining legacy
sites (detached restart watchers; router sweep in flight via #84376;
two platform adapters), each documented for burn-down. New files or
new violations fail CI immediately.
- tests/** keeps the relaxation (deliberate sleeps in fixtures).
Verification:
- ruff check . green on this branch; sabotage file with time.sleep +
subprocess.run in async def fails the gate with 2 errors.
- New behavioral test test_webhook_offloop_delivery.py asserts loop
liveness DURING delivery (ticker coroutine): 1 tick on the old
blocking code (fails), 21 ticks off-loop (passes).
- 50 webhook/replace gateway tests + 26 save/export tests pass.
Co-authored-by: AhmetArif0 <147827411+AhmetArif0@users.noreply.github.com>
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
Some Linux distros (notably RAM-based tmpfs /tmp on several Arch-based
setups) cap the temp directory at a small size, so Hermes runs out of
space for session temp files (background logs/pid/exit files, code-
execution sandboxes). Add a terminal.temp_dir config key that points
these at real storage.
- Add terminal.temp_dir default (empty) in config_defaults.py
- Bridge it to TERMINAL_TEMP_DIR via TERMINAL_CONFIG_ENV_MAP
- Honor TERMINAL_TEMP_DIR first in LocalEnvironment.get_temp_dir(),
falling through to TMPDIR//tmp//gettempdir when unset/invalid
- Add tests covering override, process-env, missing-dir, and empty
* fix(desktop): keep This-device Default on the local source
Selecting Default while This device is active took the legacy profile
door, which is also the window-primary key. On a VPS-primary desktop
that activated the remote gateway, so Bots showed the wrong Current
Gateway.
* test(desktop): pin Default on This device away from the window primary
Two sibling readers of raw config.mcp_servers duplicated their own (weaker)
shape guard: mcp-health.ts guarded the map but still passed null entries to
isUrlServer (crash on .url read), and the command palette re-implemented the
map check inline. Both now go through getServers(), the single choke point
that drops malformed entries, so a null entry can't crash the sweep and the
palette lists exactly the servers the MCP tab shows.
Also records the contributor email mapping for the salvaged commits.
Follow-up to the cherry-picked #94338.
`getServers` filters on whole-entry shape only — an entry with a junk field
(`{ command: 42, enabled: 'yes' }`) is still handed to the readers, which
coerce and tolerate. Pin that so a later tightening of `isEntry` can't
quietly start dropping entries that merely carry a bad field.
Raised in review on #94338.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An `mcp_servers` key left without a value parses as `null`, and the MCP tab
reads `enabled` straight off every entry (`serverEnabled`), so a single
dangling key threw during render and took the whole Capabilities workspace
with it — including the screen you would use to fix the config.
Filter non-object entries in `getServers`, the one choke point every reader
goes through, so a hand-edited config degrades to a missing row instead of a
dead pane. The backend already refuses to write such an entry
(`_replace_mcp_servers`), so this only has to survive a config edited
outside the app.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Since 1e9a12a71 (v0.20.6), a remote/cloud/ssh registry PRIMARY makes
globalRemoteActive() true, so the ambient v1 route — and with it every
unpinned Capabilities read — resolves to the remote gateway. The only
road back to this machine is an explicit connectionId:'local' pin, but
capabilityScoped() deliberately DROPPED that pin (pre-#91564, absent id
always meant the local pool), and profileScopeKey() collapsed
'local::<profile>' to the bare profile key.
Consequences on any desktop whose registry primary is remote:
- Capabilities -> MCP showed the REMOTE host's mcp_servers under every
scope; locally configured (and connected) MCP servers vanished from
the UI entirely.
- Picking 'default - This device' in the scope selector was a silent
no-op: the collapsed cache key equaled the current scope key, so
changeScope() early-returned and the selector snapped back.
Fix: capabilityScoped() forwards EVERY non-empty connection id, 'local'
included — Electron's apiRequestRegistryConnectionId/ensureRegistryBackend
already own 'local' pins (forced-local pooled child) and this is the
documented contract there. profileScopeKey() namespaces every explicit
pin ('local::<profile>') so a This-device pick and the ambient path never
share a cache row. profiles.ts profileOwnerScoped, which hand-patched
this exact hole for profile mutations, reduces to a named alias.
Live A/B (Electron + CDP, remote registry primary, local-only servers in
local config): v0.20.6 = local servers invisible, local pick no-op;
fixed = 'This device' lists local-server-alpha/beta, remote scope
unchanged.
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.
Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.
Closes#58357
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:
1. A pre-existing api_content sidecar left stale on the consecutive-
assistant merge. The sidecar takes priority over content at
API-build time, so a merge could silently discard its own freshly
concatenated content on the next call. Only dropped when the merge
actually changes the resulting value (wz-heng, #78063 review) --
content_rewritten compares before/after value, not just whether an
assignment branch fired, so a falsy new_content (e.g. "") that
strips to nothing no longer trips a spurious sidecar drop.
2. sanitize_api_messages never flagged a tool result with a missing/
empty tool_call_id -- its orphan-detection set only ever collected
truthy ids, so an unpaired result with no id passed the final
chokepoint untouched.
Addresses teknium1's rebase request and wz-heng's review findings on
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.
When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.
Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate
Closes the sibling gap of fa3ab2ffd / #42405.
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):
- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
tool-result image supersedes it. The reported session opened with a
~200KB poster that previously survived every compaction. The row keeps
a non-empty text placeholder, so the zero-user-turn guard (#58753) and
role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
anchor (newest is kept) and strip (older collapse to their
text_summary via _strip_images_from_tool_msg, which also drops the
stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
input_image, Anthropic-native image) were already matched by
_IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
determinism (double-run is a no-op returning the same object).
Refs #89938, #89965
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.
Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.
User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.
Refs #89938
Follow-up to the tombstone salvage (review point 1): named_profile_home()
treated ANY path whose ancestor's parent dir was literally named
'profiles' as a named profile until a '.hermes' ancestor appeared. An
unrelated custom HERMES_HOME like /srv/profiles/buildcache would resolve
as profile 'buildcache', and setup_logging would raise FileNotFoundError
instead of creating logs/ — a behavior regression for non-profile users.
Recognition now requires the 'profiles' directory's parent to BE a
Hermes home: the classic ~/.hermes layout, a root carrying home marker
files (config.yaml / .env / state.db — covers Docker/custom roots), a
profiles/.deleted tombstone dir (only ever created by profile delete),
or the process's resolved default Hermes root.
Adds regression tests for the /srv/profiles/<x> false-positive shape
(named_profile_home is None, mkdir_under_hermes_home and setup_logging
still create dirs) plus the tombstone-dir and ~/.hermes anchors. Also
adds encoding= to a bare write_text in the salvaged test file
(Windows-footgun gate).
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.
Adds a regression test covering the assistant + tool_calls +
image-only-content case.
Closes#40463
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:
Replaying the pre-rewrite sidecar would resend exactly what the rewrite
removed, so it must be dropped — the cost is one cache boundary miss,
never wrong content.
`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:
agent._vision_supported = False
_imgs_removed = _strip_images_from_messages(messages) # history
if isinstance(api_messages, list):
_strip_images_from_messages(api_messages)
and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.
Reproduced with the real functions:
history content after strip : [{'type': 'text', 'text': 'look'}]
sidecar still present : True
NEXT TURN sends : 'look<IMAGE BYTES SENT LAST TURN>'
This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.
Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.
tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).
Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged
CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
Follow-up for salvaged PR #84397: Test-SystemNodeReady still said
'too old (Hermes requires Node >=22.22.0)' — Node 25 is not too old,
it's an unsupported line.
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests
* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)
* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
Enough1122 review of #94426:
- Legacy Bot tiles persisted before ownerRoute.connectionId existed carry no
connection id, so the local-delete branch (`ownerConnection === 'local'`)
never matched them and they resurrected the deleted profile on relaunch.
Treat the empty connection id as the local connection.
- A route without profile now throws instead of silently falling into the
local-delete branch and dropping nothing remotely owned.
- Pin the 'local' connection spelling with tests covering the legacy
(no id), canonical ('local'), and divergent ('Local') spellings.
- Comment the live-atom filter arms by bucket and unify the duplicated
#94235 call-site comments in sdk/index.ts and delete-profile-dialog.tsx.
Address review findings on #94426:
- Non-route (local profile) deletes now require the tile's owner
connectionId to be 'local' — a same-named bot on another connection
keeps its tile and live conversation.
- Route identity (profile/targetProfile) goes through normalizeProfileKey
like the non-route name, so whitespace-padded identities match on both
paths; profile case stays exact, consistently on both paths.
- Test lint: drop the unused dropTilesForProfile value import (tests call
it via the fresh module namespace) and replace the forbidden typeof
import() type annotation with a type-only namespace import. New tests:
same-name-other-connection survival, whitespace normalization, and
case-exact identity. eslint: 0 errors; typecheck clean.
A Bot Mode bot cloned via 'Clone from profile' resurrects after
deletion: the desktop keeps the bot's chat tile in local storage
(bot-meta cleanup removes the roster entry, but the persisted tile
survives). On relaunch the tile restores, re-dials the deleted
profile's backend, and ensure_hermes_home() re-creates the profile
directory the delete just removed — the empty skeleton with no
config.yaml reported in #94235.
dropTilesForProfile() removes a deleted profile's session-tile
bucket and every Bot Mode tile whose ownerRoute points at it
(exact connection + backend target when a source-scoped route is
given) from memory and local storage, with discard (no Cmd+Shift+T
undo) semantics. Wired into both delete doors: the core
DeleteProfileDialog and the SDK host.deleteProfile path the Bots
plugin uses.
Regression tests cover local and source-scoped routes; verified red
with the cleanup disabled.