Commit Graph

14240 Commits

Author SHA1 Message Date
Teknium dadbfd8990 refactor(patch): V4A mode gated to OpenAI-family mains — base schema is replace-only with real required (365 -> 195 for everyone else, -149; handler accepts both shapes from any model) (#97403) 2026-08-28 12:59:41 -07:00
Mohamed HAMMANE 6da0ae1cf5 fix(install): preserve project config for locked uv sync (#82446) 2026-08-28 12:36:15 -07:00
Teknium 15eb5caf7a fix(install): Node 24.0–24.10 no longer passes the gates only to die at npm EBADENGINE
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.

- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
  package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
  (install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
  locked dependency's engines.node, and the installer gates must encode
  the same floors as the manifest — so the next babel-style floor bump
  turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
2026-08-28 12:20:40 -07:00
Teknium 93de1d3430 vision_analyze diet + image routing: explicit aux vision backend becomes the de-facto route (reverses #29135) (#97339)
* refactor(vision_analyze): schema diet — routing mechanics removed (automatic; native path's own result teaches), region flow kept (~271 -> 181 tok/call, -33%)

* feat(image-routing): explicit auxiliary.vision backend is the de-facto image route — reverses #29135 (maintainer decision); native stays default when unset, image_input_mode:native stays absolute
2026-08-28 12:15:24 -07:00
Teknium 72874b0675 feat(skill_manage): operations[] is the call — each op names its skill; atomic with cross-skill rollback (#97295)
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate

* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing

* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal

* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback

* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization

Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
2026-08-28 12:15:18 -07:00
69k4xmdfm2-blip fc5fdb8c2a fix(gateway): handoff is broken on multi-profile installs (wrong DB, wrong key, wrong bot)
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.

1. The watcher polls only the ROOT store.
   `_handoff_watcher` resolves `self._session_db` with no profile scope, which
   always yields the root `state.db`. But `/handoff` run under
   `hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
   store. Nothing ever reads it, so the CLI times out after 60s while the
   gateway is alive and connected. The watcher now iterates
   `[(None, None), *secondary_profiles]` and polls each inside
   `_profile_runtime_scope`.

2. The destination session key is built without the profile namespace.
   `_process_handoff` called `build_session_key()` with no `profile=`,
   producing `agent:main:...` while that profile's own adapter routes organic
   inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
   reads.

3. Delivery uses the PRIMARY profile's adapter and config.
   `self.adapters` holds only the default profile's adapters (secondaries live
   in `self._profile_adapters[name]`) and `self.config` only the default's home
   channel. A secondary profile's handoff was therefore sent by the wrong bot,
   to the wrong chat, while persisting the right session key and reporting
   `handoff_state='completed'` — a false positive that looks correct in the
   database and is wrong on the wire.

Two robustness fixes in the same path:

4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
   turn plus delivery; awaiting it inline meant one slow handoff stopped the
   watcher from even polling the other profiles. With the CLI's 60s deadline, a
   valid handoff could time out purely because another profile's was ahead of
   it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
   never claimed twice, and a bounded drain on shutdown.

5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
   one in-process dispatch, so a row still in that state at startup belongs to
   a gateway that died mid-dispatch. It can never reach a terminal state, and
   `request_handoff` refuses new requests unless the state is
   NULL/completed/failed — that session could never hand off again, silently.
   `reclaim_stale_running_handoffs()` now fails those rows once per store at
   watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
   may already have switched the session key and dispatched, so a blind retry
   risks double delivery.

Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.

Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
2026-08-28 11:45:19 -07:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
yoniebans be284cf5c0 fix(update): report when the official repo was not checked on the up-to-date path
Review follow-up on #97052 (helix4u): a fork with no upstream remote whose HEAD matches origin/main used to print plain "Already up to date!" under --yes even though official main was never consulted, so an unattended stale fork looked current. _sync_with_upstream_if_needed now returns whether the official upstream was actually checked, and the commit_count == 0 completion line says "Up to date with your fork (official repo not checked)." when it was not. Skip-as-decline semantics are unchanged: no prompt, no remote mutation, no decline marker. Caller-level regression test added for the fork + no-upstream + --yes + HEAD==origin/main path; helper tests now pin the return contract.
2026-08-28 13:33:56 -04:00
yoniebans b33fa127d2 fix(update): gate the fork-upstream prompt for --yes and non-tty runs
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.

Closes #60240 (prompt half). Supersedes #78678, #92448, #92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
2026-08-28 13:33:56 -04:00
Brooklyn Nicholson 21cc469eb1 Merge remote-tracking branch 'origin/main' into bb/bot-mode-design-system
# Conflicts:
#	apps/desktop/src/plugins/hermes-bots/plugin.js
#	apps/desktop/src/plugins/hermes-bots/tests/group-chat.test.mjs
#	apps/desktop/src/sdk/profile-routing.test.ts
2026-08-28 12:05:11 -05:00
Mariano Nicolini da3c2435e2 fix(nous): only rescue an empty list where emptiness means "filtered out"
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
2026-08-28 13:34:27 -03:00
Teknium baa344dee7 refactor(process): schema diet — enum names the verbs, description keeps only non-obvious semantics; write-vs-submit trap teaching emphasized (306 -> 228 tok/call, -25%) (#97279) 2026-08-28 09:23:00 -07:00
Teknium 7b5e1911f8 refactor(todo): schema diet — item shape and merge semantics taught once, by the param schema (323 -> 232 tok/call, -28%) (#97257) 2026-08-28 08:58:12 -07:00
Teknium a9e72f1b58 refactor(read_file): schema diet + bundle anydoc 0.2.4 + typed NeedsOcrError/hosted-OCR wiring (#97195)
* refactor(read_file): capability-gate the anydoc format list; PDF coverage teaching lives in the response-time warning (426 -> 244/291 tok/call)

* feat(read_file): bundle firecrawl-anydoc 0.2.4 in core, typed NeedsOcrError handling, config-gated hosted OCR with local-OCR-first guidance

* refine(read_file): NEEDS-OCR warning hints at checking for an OCR skill without naming one; hosted_ocr knob unadvertised (maintainer-directed)

* simplify(read_file): drop the anydoc schema gate — bundled core dep makes absence a broken install, not a variant; formats stated unconditionally (263 tok/call)

* feat(read_file): PDF wording upgrades to 'scanned or text' when a trusted hosted-OCR route exists (direct key or explicit config; nous gateway excluded until Parse proxy works)

* simplify(read_file): FIRECRAWL_API_KEY is the ONLY hosted-OCR gate — nous gateway route removed (Parse proxy broken), config true no longer unlocks; false still disables
2026-08-28 08:46:11 -07:00
Mariano Nicolini 04647f15c8 fix(nous): only fall back to the reachable set when the overlap is empty
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.

Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:39:29 -03:00
Mariano Nicolini 117e7fef88 fix(nous): surface allowed models the curated list does not carry
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.

When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.

Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:24:55 -03:00
Teknium 306db2776c test(codex): mid-turn compaction fixtures report realistic anchored usage
The usage anchor (#97206) now trusts provider-reported usage. These two
tests simulated a tool-heavy near-overflow turn while the shared fixture
reported a 12-token prompt — the anchored pressure check honestly
concluded no pressure. Give the scenario 18K anchored prompt tokens so
the tests pin the same compaction decision they always did.
2026-08-28 07:51:31 -07:00
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00
Brian b855f86bc8 fix(agent): 413 recovery measures bytes, not token estimates
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.

Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.

Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.

Salvaged from #88960. Fixes #47339.
2026-08-28 07:51:16 -07:00
kshitijk4poor c0ff25a1f8 fix(gateway): stop blocking the event loop — off-loop hot sites + ASYNC lint ratchet
Pattern-A architectural fix: blocking calls inside async functions freeze
the gateway/uvicorn event loop for every adapter, timer, and health check.
Known incidents: 17-minute getaddrinfo freeze (#91912 class), 10s restart
freeze in start_gateway (#36163).

Fixes at the four unguarded core sites:
- gateway/platforms/webhook.py: `gh pr comment` subprocess (30s timeout)
  now runs via asyncio.to_thread — a webhook delivery no longer freezes
  every other platform for the duration of a network call.
- gateway/run.py start_gateway --replace: two time.sleep() waits (10s +
  5s worst case) become await asyncio.sleep() (re-lands #36163 at current
  line numbers, credit AhmetArif0).
- gateway/slash_commands.py /save: session render + file write move off
  the loop (scales with transcript size).
- hermes_cli/web_server.py voice TTS: multi-MB audio file read + unlink
  move off the loop.

Prevention gate so the bug class cannot re-enter:
- pyproject.toml [tool.ruff.lint] select gains ASYNC210/220/221/251
  (blocking HTTP / Popen / subprocess.run / time.sleep in async def).
  These run in the existing blocking `ruff check .` CI job.
- Frozen ratchet baseline in per-file-ignores for the remaining legacy
  sites (detached restart watchers; router sweep in flight via #84376;
  two platform adapters), each documented for burn-down. New files or
  new violations fail CI immediately.
- tests/** keeps the relaxation (deliberate sleeps in fixtures).

Verification:
- ruff check . green on this branch; sabotage file with time.sleep +
  subprocess.run in async def fails the gate with 2 errors.
- New behavioral test test_webhook_offloop_delivery.py asserts loop
  liveness DURING delivery (ticker coroutine): 1 tick on the old
  blocking code (fails), 21 ticks off-loop (passes).
- 50 webhook/replace gateway tests + 26 save/export tests pass.

Co-authored-by: AhmetArif0 <147827411+AhmetArif0@users.noreply.github.com>
2026-08-28 07:50:51 -07:00
Teknium 95cf7dc9e8 feat: session temp root moves off tmpfs /tmp to ~/.hermes/cache/terminal by default; auto-pruned after 72h
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
2026-08-28 07:50:33 -07:00
rahlquist d7be3f649d feat: expose terminal.temp_dir to redirect session temp root off tmpfs
Some Linux distros (notably RAM-based tmpfs /tmp on several Arch-based
setups) cap the temp directory at a small size, so Hermes runs out of
space for session temp files (background logs/pid/exit files, code-
execution sandboxes). Add a terminal.temp_dir config key that points
these at real storage.

- Add terminal.temp_dir default (empty) in config_defaults.py
- Bridge it to TERMINAL_TEMP_DIR via TERMINAL_CONFIG_ENV_MAP
- Honor TERMINAL_TEMP_DIR first in LocalEnvironment.get_temp_dir(),
  falling through to TMPDIR//tmp//gettempdir when unset/invalid
- Add tests covering override, process-env, missing-dir, and empty
2026-08-28 07:50:33 -07:00
Teknium eff97a8a05 refactor(profiles): retire the cross-profile write guard — profiles are not isolated (maintainer decision); mirror lost-write guards (#32049) survive; patch/write_file schemas drop cross_profile (-83 tok/call) (#97165) 2026-08-28 06:36:22 -07:00
isheng 87cff9d4c1 test(sanitizer): add unit tests for _classify_tool_call_orphans 2026-08-28 06:32:48 -07:00
joaomarcos f0ac2c8f12 fix(agent): drop stale api_content sidecar and unpaired tool results
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:

1. A pre-existing api_content sidecar left stale on the consecutive-
   assistant merge. The sidecar takes priority over content at
   API-build time, so a merge could silently discard its own freshly
   concatenated content on the next call. Only dropped when the merge
   actually changes the resulting value (wz-heng, #78063 review) --
   content_rewritten compares before/after value, not just whether an
   assignment branch fired, so a falsy new_content (e.g. "") that
   strips to nothing no longer trips a spurious sidecar drop.

2. sanitize_api_messages never flagged a tool result with a missing/
   empty tool_call_id -- its orphan-detection set only ever collected
   truthy ids, so an unpaired result with no id passed the final
   chokepoint untouched.

Addresses teknium1's rebase request and wz-heng's review findings on
2026-08-28 06:32:48 -07:00
srojk34 e024bf75ce fix(compression): strip whitespace from tool_call_id in _sanitize_tool_pairs
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.

When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.

Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate

Closes the sibling gap of fa3ab2ffd / #42405.
2026-08-28 06:32:48 -07:00
Teknium 0dce46feb7 fix(compressor): widen compaction-time image aging to first-message and envelope shapes
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):

- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
  tool-result image supersedes it. The reported session opened with a
  ~200KB poster that previously survived every compaction. The row keeps
  a non-empty text placeholder, so the zero-user-turn guard (#58753) and
  role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
  anchor (newest is kept) and strip (older collapse to their
  text_summary via _strip_images_from_tool_msg, which also drops the
  stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
  input_image, Anthropic-native image) were already matched by
  _IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
  determinism (double-run is a no-op returning the same object).

Refs #89938, #89965
2026-08-28 06:32:43 -07:00
Jack Lau b81b599d50 fix(compressor): age out stale tool-result images during compaction
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.

Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.

User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.

Refs #89938
2026-08-28 06:32:43 -07:00
Teknium 9978706e93 refactor(skill_manage): schema diet — patch args defer to the patch tool's semantics, file_path states its skill-dir-relative shape, authoring curriculum compressed (517 -> 427, -17%) (#97152) 2026-08-28 05:48:15 -07:00
Teknium a641644f12 fix(paths): display_hermes_home renders POSIX separators on Windows — kills ~/AppData\Local\hermes chimeras in tool schemas and user-facing messages (#97137) 2026-08-28 05:45:00 -07:00
Teknium 35328345d5 fix(profiles): anchor named-profile detection to real Hermes homes
Follow-up to the tombstone salvage (review point 1): named_profile_home()
treated ANY path whose ancestor's parent dir was literally named
'profiles' as a named profile until a '.hermes' ancestor appeared. An
unrelated custom HERMES_HOME like /srv/profiles/buildcache would resolve
as profile 'buildcache', and setup_logging would raise FileNotFoundError
instead of creating logs/ — a behavior regression for non-profile users.

Recognition now requires the 'profiles' directory's parent to BE a
Hermes home: the classic ~/.hermes layout, a root carrying home marker
files (config.yaml / .env / state.db — covers Docker/custom roots), a
profiles/.deleted tombstone dir (only ever created by profile delete),
or the process's resolved default Hermes root.

Adds regression tests for the /srv/profiles/<x> false-positive shape
(named_profile_home is None, mkdir_under_hermes_home and setup_logging
still create dirs) plus the tombstone-dir and ~/.hermes anchors. Also
adds encoding= to a bare write_text in the salvaged test file
(Windows-footgun gate).
2026-08-28 05:27:15 -07:00
EndeavorYen af2dc685fc fix(profiles): honor tombstones in exists/backfill and tighten named-home detection
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
2026-08-28 05:27:15 -07:00
EndeavorYen 7e34522454 fix(profiles): tombstone deleted named profiles so logging cannot resurrect them
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
2026-08-28 05:27:15 -07:00
Hakan Baysal cb8027afed fix(sanitize): preserve assistant messages with tool_calls when stripping images
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.

Adds a regression test covering the assistant + tool_calls +
image-only-content case.

Closes #40463
2026-08-28 05:17:26 -07:00
Frowtek e1762bd30b fix(agent): drop the api_content sidecar when stripping images from history
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:

    Replaying the pre-rewrite sidecar would resend exactly what the rewrite
    removed, so it must be dropped — the cost is one cache boundary miss,
    never wrong content.

`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:

    agent._vision_supported = False
    _imgs_removed = _strip_images_from_messages(messages)      # history
    if isinstance(api_messages, list):
        _strip_images_from_messages(api_messages)

and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.

Reproduced with the real functions:

    history content after strip : [{'type': 'text', 'text': 'look'}]
    sidecar still present       : True
    NEXT TURN sends             : 'look<IMAGE BYTES SENT LAST TURN>'

This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.

Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.

tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
2026-08-28 05:17:21 -07:00
Teknium a5c7eed5f3 feat(cli/tui): -q now seeds a live interactive session; prompts submit literally
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).

Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged

CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
2026-08-28 05:15:06 -07:00
fangliquan 7d61b6701e test(dashboard): run engine invariants in JavaScript lane 2026-08-28 05:12:33 -07:00
fangliquan 4a7df1d070 fix(dashboard): align supported Node engine lines 2026-08-28 05:12:33 -07:00
Teknium 536adb35c4 refactor(video_generate): capability-gated dynamic schema (~814 → 458/377 tok/call) (#97095)
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests

* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)

* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
2026-08-28 05:10:33 -07:00
Inimitable Mind 4e860c093a fix(tui): harden failed build recovery handoff 2026-08-28 04:58:28 -07:00
Inimitable Mind 258d0283b0 fix(tui): close failed resume recovery races 2026-08-28 04:58:28 -07:00
Inimitable Mind dbb9854591 fix(tui): recover model switch after failed resume 2026-08-28 04:58:28 -07:00
Teknium d48adfa789 test(vision): fixtures use a structurally valid PNG now that decode gates reject magic-bytes-only fakes
The salvaged proactive validation (#53307/#76896) rejects the
b'\x89PNG...' + zero-padding fixtures test_vision_tools.py used.
Replace all 11 sites with a real 1x1 PNG (+ ignored trailing padding
for size-sensitive tests), and pin _resize_image_for_vision passthrough
in the size-limit test so it exercises rejection, not resize recovery.
2026-08-28 04:58:21 -07:00
Teknium c0882093b8 test(vision): truncated-image test accepts rejection at either stacked gate
With both salvaged validators in place, the resolver-boundary verify()
(#53307) catches a truncated PNG before the full-decode gate (#76896)
runs. Either error message means the bytes stayed out of history.
2026-08-28 04:58:21 -07:00
CannibalKush 765142d94f fix: validate PNGs at shared image resolver 2026-08-28 04:58:21 -07:00
HaiyiMei ad0e83062f fix(vision): reject truncated images before embedding 2026-08-28 04:58:21 -07:00
lEWFkRAD b80b9d8271 fix(codex): preserve assistant image slots in replay 2026-08-28 04:58:14 -07:00
lEWFkRAD 8de45940fb fix(codex): drop assistant images from Responses replay
Fixes #96816
2026-08-28 04:58:14 -07:00