Commit Graph

26038 Commits

Author SHA1 Message Date
Teknium 15eb5caf7a fix(install): Node 24.0–24.10 no longer passes the gates only to die at npm EBADENGINE
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.

- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
  package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
  (install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
  locked dependency's engines.node, and the installer gates must encode
  the same floors as the manifest — so the next babel-style floor bump
  turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
2026-08-28 12:20:40 -07:00
Jefferson Nunn c9fa2bba45 fix(install): tier-0 locked sync no longer trips over UV_NO_CONFIG
The installer exports UV_NO_CONFIG=1 at script start (sudo -u hygiene,
#21269). That export also hides the project's own [tool.uv] policy —
exclude-newer and its package exemptions — from uv. The resolver then
runs under a different policy than uv.lock was resolved under, and
--locked makes that mismatch fatal:

  error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.

Every fresh install hit this and fell through to the non-hash-verified
PyPI fallback tiers, defeating the point of Tier 0. Strip the variable
for this one invocation only; it stays exported for every other uv call.
Runtime code already strips UV_NO_CONFIG before its own locked syncs for
the same reason (hermes_cli/managed_uv.py).
2026-08-28 12:20:40 -07:00
Teknium 93de1d3430 vision_analyze diet + image routing: explicit aux vision backend becomes the de-facto route (reverses #29135) (#97339)
* refactor(vision_analyze): schema diet — routing mechanics removed (automatic; native path's own result teaches), region flow kept (~271 -> 181 tok/call, -33%)

* feat(image-routing): explicit auxiliary.vision backend is the de-facto image route — reverses #29135 (maintainer decision); native stays default when unset, image_input_mode:native stays absolute
2026-08-28 12:15:24 -07:00
Teknium 72874b0675 feat(skill_manage): operations[] is the call — each op names its skill; atomic with cross-skill rollback (#97295)
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate

* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing

* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal

* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback

* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization

Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
2026-08-28 12:15:18 -07:00
Teknium 7d1c9aeab7 fix: log swallowed reclaim failures + pin ContextVar dispatch invariant (review follow-up for #91217) 2026-08-28 11:45:19 -07:00
69k4xmdfm2-blip fc5fdb8c2a fix(gateway): handoff is broken on multi-profile installs (wrong DB, wrong key, wrong bot)
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.

1. The watcher polls only the ROOT store.
   `_handoff_watcher` resolves `self._session_db` with no profile scope, which
   always yields the root `state.db`. But `/handoff` run under
   `hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
   store. Nothing ever reads it, so the CLI times out after 60s while the
   gateway is alive and connected. The watcher now iterates
   `[(None, None), *secondary_profiles]` and polls each inside
   `_profile_runtime_scope`.

2. The destination session key is built without the profile namespace.
   `_process_handoff` called `build_session_key()` with no `profile=`,
   producing `agent:main:...` while that profile's own adapter routes organic
   inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
   reads.

3. Delivery uses the PRIMARY profile's adapter and config.
   `self.adapters` holds only the default profile's adapters (secondaries live
   in `self._profile_adapters[name]`) and `self.config` only the default's home
   channel. A secondary profile's handoff was therefore sent by the wrong bot,
   to the wrong chat, while persisting the right session key and reporting
   `handoff_state='completed'` — a false positive that looks correct in the
   database and is wrong on the wire.

Two robustness fixes in the same path:

4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
   turn plus delivery; awaiting it inline meant one slow handoff stopped the
   watcher from even polling the other profiles. With the CLI's 60s deadline, a
   valid handoff could time out purely because another profile's was ahead of
   it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
   never claimed twice, and a bounded drain on shutdown.

5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
   one in-process dispatch, so a row still in that state at startup belongs to
   a gateway that died mid-dispatch. It can never reach a terminal state, and
   `request_handoff` refuses new requests unless the state is
   NULL/completed/failed — that session could never hand off again, silently.
   `reclaim_stale_running_handoffs()` now fails those rows once per store at
   watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
   may already have switched the session key and dispatched, so a blind retry
   risks double delivery.

Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.

Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
2026-08-28 11:45:19 -07:00
hermes-seaeye[bot] 9048530318 fmt(js): npm run fix on merge (#97358)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-28 18:38:38 +00:00
Brooklyn Nicholson a792d0794f fix(desktop): fill session-switch backfill in two frames instead of ten
FIRST_PAINT_BUDGET 20 + BACKFILL_STEP 60 prepended the rest of a 600-unit
page across ~10 visible commits. A 290-unit step keeps the interruptible
commits and removes the strobe.
2026-08-28 13:33:53 -05:00
Brooklyn Nicholson d229648511 fix(desktop): keep the session loader up while known history is empty
Brand-new drafts are empty on purpose. A routed session the list already
knows has messages must not drop the loader just because a runtime id is
bound — that is the blank frame during an unproven warm hold and a cold
switch.
2026-08-28 13:33:53 -05:00
Brooklyn Nicholson b6eb17d01c fix(desktop): hold unproven warm transcripts off the view
A compressed runtime cache is a legal tail, not display history. Publishing
it on session switch then replacing it with the persisted lineage is the
warm-path flicker. Gate that paint on persisted-display provenance and keep
the previous/empty view until REST authority lands.

Co-authored-by: xrbs00 <178640517+xrbs00@users.noreply.github.com>
2026-08-28 13:33:53 -05:00
yoniebans 00bbfc6900 docs(update): --yes help states the fork-upstream prompt is skipped, not accepted
The old text read as if --yes answers yes to every prompt. It accepts the config-migration and stash-restore prompts but skips the fork-upstream prompt without adding a remote (#97052 review); say so.
2026-08-28 13:33:56 -04:00
yoniebans be284cf5c0 fix(update): report when the official repo was not checked on the up-to-date path
Review follow-up on #97052 (helix4u): a fork with no upstream remote whose HEAD matches origin/main used to print plain "Already up to date!" under --yes even though official main was never consulted, so an unattended stale fork looked current. _sync_with_upstream_if_needed now returns whether the official upstream was actually checked, and the commit_count == 0 completion line says "Up to date with your fork (official repo not checked)." when it was not. Skip-as-decline semantics are unchanged: no prompt, no remote mutation, no decline marker. Caller-level regression test added for the fork + no-upstream + --yes + HEAD==origin/main path; helper tests now pin the return contract.
2026-08-28 13:33:56 -04:00
yoniebans b33fa127d2 fix(update): gate the fork-upstream prompt for --yes and non-tty runs
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.

Closes #60240 (prompt half). Supersedes #78678, #92448, #92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
2026-08-28 13:33:56 -04:00
Teknium baa344dee7 refactor(process): schema diet — enum names the verbs, description keeps only non-obvious semantics; write-vs-submit trap teaching emphasized (306 -> 228 tok/call, -25%) (#97279) 2026-08-28 09:23:00 -07:00
Teknium 7b5e1911f8 refactor(todo): schema diet — item shape and merge semantics taught once, by the param schema (323 -> 232 tok/call, -28%) (#97257) 2026-08-28 08:58:12 -07:00
Teknium a9e72f1b58 refactor(read_file): schema diet + bundle anydoc 0.2.4 + typed NeedsOcrError/hosted-OCR wiring (#97195)
* refactor(read_file): capability-gate the anydoc format list; PDF coverage teaching lives in the response-time warning (426 -> 244/291 tok/call)

* feat(read_file): bundle firecrawl-anydoc 0.2.4 in core, typed NeedsOcrError handling, config-gated hosted OCR with local-OCR-first guidance

* refine(read_file): NEEDS-OCR warning hints at checking for an OCR skill without naming one; hosted_ocr knob unadvertised (maintainer-directed)

* simplify(read_file): drop the anydoc schema gate — bundled core dep makes absence a broken install, not a variant; formats stated unconditionally (263 tok/call)

* feat(read_file): PDF wording upgrades to 'scanned or text' when a trusted hosted-OCR route exists (direct key or explicit config; nous gateway excluded until Parse proxy works)

* simplify(read_file): FIRECRAWL_API_KEY is the ONLY hosted-OCR gate — nous gateway route removed (Parse proxy broken), config true no longer unlocks; false still disables
2026-08-28 08:46:11 -07:00
Teknium 306db2776c test(codex): mid-turn compaction fixtures report realistic anchored usage
The usage anchor (#97206) now trusts provider-reported usage. These two
tests simulated a tool-heavy near-overflow turn while the shared fixture
reported a 12-token prompt — the anchored pressure check honestly
concluded no pressure. Give the scenario 18K anchored prompt tokens so
the tests pin the same compaction decision they always did.
2026-08-28 07:51:31 -07:00
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
Teknium c5b44e0756 chore: map contributor emails for TiberiuD and fedebyes 2026-08-28 07:51:23 -07:00
Teknium 5b31602c15 docs: reconcile positional pairing with shared _classify_tool_call_orphans (#97167) — classifier docstring reflects its remaining consumer; empty-id filter note updated 2026-08-28 07:51:23 -07:00
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00
Teknium 3f3ae6850a chore: map brianbaldock contributor email 2026-08-28 07:51:16 -07:00
Brian b855f86bc8 fix(agent): 413 recovery measures bytes, not token estimates
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.

Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.

Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.

Salvaged from #88960. Fixes #47339.
2026-08-28 07:51:16 -07:00
kshitijk4poor c0ff25a1f8 fix(gateway): stop blocking the event loop — off-loop hot sites + ASYNC lint ratchet
Pattern-A architectural fix: blocking calls inside async functions freeze
the gateway/uvicorn event loop for every adapter, timer, and health check.
Known incidents: 17-minute getaddrinfo freeze (#91912 class), 10s restart
freeze in start_gateway (#36163).

Fixes at the four unguarded core sites:
- gateway/platforms/webhook.py: `gh pr comment` subprocess (30s timeout)
  now runs via asyncio.to_thread — a webhook delivery no longer freezes
  every other platform for the duration of a network call.
- gateway/run.py start_gateway --replace: two time.sleep() waits (10s +
  5s worst case) become await asyncio.sleep() (re-lands #36163 at current
  line numbers, credit AhmetArif0).
- gateway/slash_commands.py /save: session render + file write move off
  the loop (scales with transcript size).
- hermes_cli/web_server.py voice TTS: multi-MB audio file read + unlink
  move off the loop.

Prevention gate so the bug class cannot re-enter:
- pyproject.toml [tool.ruff.lint] select gains ASYNC210/220/221/251
  (blocking HTTP / Popen / subprocess.run / time.sleep in async def).
  These run in the existing blocking `ruff check .` CI job.
- Frozen ratchet baseline in per-file-ignores for the remaining legacy
  sites (detached restart watchers; router sweep in flight via #84376;
  two platform adapters), each documented for burn-down. New files or
  new violations fail CI immediately.
- tests/** keeps the relaxation (deliberate sleeps in fixtures).

Verification:
- ruff check . green on this branch; sabotage file with time.sleep +
  subprocess.run in async def fails the gate with 2 errors.
- New behavioral test test_webhook_offloop_delivery.py asserts loop
  liveness DURING delivery (ticker coroutine): 1 tick on the old
  blocking code (fails), 21 ticks off-loop (passes).
- 50 webhook/replace gateway tests + 26 save/export tests pass.

Co-authored-by: AhmetArif0 <147827411+AhmetArif0@users.noreply.github.com>
2026-08-28 07:50:51 -07:00
Teknium 95cf7dc9e8 feat: session temp root moves off tmpfs /tmp to ~/.hermes/cache/terminal by default; auto-pruned after 72h
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
2026-08-28 07:50:33 -07:00
rahlquist d7be3f649d feat: expose terminal.temp_dir to redirect session temp root off tmpfs
Some Linux distros (notably RAM-based tmpfs /tmp on several Arch-based
setups) cap the temp directory at a small size, so Hermes runs out of
space for session temp files (background logs/pid/exit files, code-
execution sandboxes). Add a terminal.temp_dir config key that points
these at real storage.

- Add terminal.temp_dir default (empty) in config_defaults.py
- Bridge it to TERMINAL_TEMP_DIR via TERMINAL_CONFIG_ENV_MAP
- Honor TERMINAL_TEMP_DIR first in LocalEnvironment.get_temp_dir(),
  falling through to TMPDIR//tmp//gettempdir when unset/invalid
- Add tests covering override, process-env, missing-dir, and empty
2026-08-28 07:50:33 -07:00
hermes-seaeye[bot] 0abecf7a93 fmt(js): npm run fix on merge (#97219)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-28 14:28:34 +00:00
xxxigm 76da7ff087 fix(desktop): keep This-device Default on the local source (#97038)
* fix(desktop): keep This-device Default on the local source

Selecting Default while This device is active took the legacy profile
door, which is also the window-primary key. On a VPS-primary desktop
that activated the remote gateway, so Bots showed the wrong Current
Gateway.

* test(desktop): pin Default on This device away from the window primary
2026-08-28 07:21:50 -07:00
Teknium eff97a8a05 refactor(profiles): retire the cross-profile write guard — profiles are not isolated (maintainer decision); mirror lost-write guards (#32049) survive; patch/write_file schemas drop cross_profile (-83 tok/call) (#97165) 2026-08-28 06:36:22 -07:00
Teknium 9f2ab334c8 fix(desktop): route the MCP health sweep and command palette through getServers
Two sibling readers of raw config.mcp_servers duplicated their own (weaker)
shape guard: mcp-health.ts guarded the map but still passed null entries to
isUrlServer (crash on .url read), and the command palette re-implemented the
map check inline. Both now go through getServers(), the single choke point
that drops malformed entries, so a null entry can't crash the sweep and the
palette lists exactly the servers the MCP tab shows.

Also records the contributor email mapping for the salvaged commits.

Follow-up to the cherry-picked #94338.
2026-08-28 06:33:40 -07:00
Nacho Chiappero 49761cebad test(desktop): pin that field-level junk in an mcp_servers entry survives
`getServers` filters on whole-entry shape only — an entry with a junk field
(`{ command: 42, enabled: 'yes' }`) is still handed to the readers, which
coerce and tolerate. Pin that so a later tightening of `isEntry` can't
quietly start dropping entries that merely carry a bad field.

Raised in review on #94338.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 06:33:40 -07:00
Nacho Chiappero 13afe9e9e1 fix(desktop): drop malformed mcp_servers entries instead of crashing
An `mcp_servers` key left without a value parses as `null`, and the MCP tab
reads `enabled` straight off every entry (`serverEnabled`), so a single
dangling key threw during render and took the whole Capabilities workspace
with it — including the screen you would use to fix the config.

Filter non-object entries in `getServers`, the one choke point every reader
goes through, so a hand-edited config degrades to a missing row instead of a
dead pane. The backend already refuses to write such an entry
(`_replace_mcp_servers`), so this only has to survive a config edited
outside the app.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 06:33:40 -07:00
Teknium fe576ba48e fix(desktop): a 'This device' Capabilities pick reaches the local machine again under a remote registry primary
Since 1e9a12a71 (v0.20.6), a remote/cloud/ssh registry PRIMARY makes
globalRemoteActive() true, so the ambient v1 route — and with it every
unpinned Capabilities read — resolves to the remote gateway. The only
road back to this machine is an explicit connectionId:'local' pin, but
capabilityScoped() deliberately DROPPED that pin (pre-#91564, absent id
always meant the local pool), and profileScopeKey() collapsed
'local::<profile>' to the bare profile key.

Consequences on any desktop whose registry primary is remote:
- Capabilities -> MCP showed the REMOTE host's mcp_servers under every
  scope; locally configured (and connected) MCP servers vanished from
  the UI entirely.
- Picking 'default - This device' in the scope selector was a silent
  no-op: the collapsed cache key equaled the current scope key, so
  changeScope() early-returned and the selector snapped back.

Fix: capabilityScoped() forwards EVERY non-empty connection id, 'local'
included — Electron's apiRequestRegistryConnectionId/ensureRegistryBackend
already own 'local' pins (forced-local pooled child) and this is the
documented contract there. profileScopeKey() namespaces every explicit
pin ('local::<profile>') so a This-device pick and the ambient path never
share a cache row. profiles.ts profileOwnerScoped, which hand-patched
this exact hole for profile mutations, reduces to a named alias.

Live A/B (Electron + CDP, remote registry primary, local-only servers in
local config): v0.20.6 = local servers invisible, local pick no-op;
fixed = 'This device' lists local-server-alpha/beta, remote scope
unchanged.
2026-08-28 06:33:35 -07:00
Hermes Agent 225fa13bd3 fix(sanitizer): drop duplicated legacy _classify_tool_call_orphans left by cherry-pick auto-merge 2026-08-28 06:32:48 -07:00
isheng 87cff9d4c1 test(sanitizer): add unit tests for _classify_tool_call_orphans 2026-08-28 06:32:48 -07:00
isheng c6a426e9ad refactor(sanitizer): extract shared _classify_tool_call_orphans to eliminate drift
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.

Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.

Closes #58357
2026-08-28 06:32:48 -07:00
joaomarcos f0ac2c8f12 fix(agent): drop stale api_content sidecar and unpaired tool results
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:

1. A pre-existing api_content sidecar left stale on the consecutive-
   assistant merge. The sidecar takes priority over content at
   API-build time, so a merge could silently discard its own freshly
   concatenated content on the next call. Only dropped when the merge
   actually changes the resulting value (wz-heng, #78063 review) --
   content_rewritten compares before/after value, not just whether an
   assignment branch fired, so a falsy new_content (e.g. "") that
   strips to nothing no longer trips a spurious sidecar drop.

2. sanitize_api_messages never flagged a tool result with a missing/
   empty tool_call_id -- its orphan-detection set only ever collected
   truthy ids, so an unpaired result with no id passed the final
   chokepoint untouched.

Addresses teknium1's rebase request and wz-heng's review findings on
2026-08-28 06:32:48 -07:00
srojk34 e024bf75ce fix(compression): strip whitespace from tool_call_id in _sanitize_tool_pairs
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.

When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.

Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate

Closes the sibling gap of fa3ab2ffd / #42405.
2026-08-28 06:32:48 -07:00
Teknium 0dce46feb7 fix(compressor): widen compaction-time image aging to first-message and envelope shapes
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):

- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
  tool-result image supersedes it. The reported session opened with a
  ~200KB poster that previously survived every compaction. The row keeps
  a non-empty text placeholder, so the zero-user-turn guard (#58753) and
  role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
  anchor (newest is kept) and strip (older collapse to their
  text_summary via _strip_images_from_tool_msg, which also drops the
  stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
  input_image, Anthropic-native image) were already matched by
  _IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
  determinism (double-run is a no-op returning the same object).

Refs #89938, #89965
2026-08-28 06:32:43 -07:00
Jack Lau b81b599d50 fix(compressor): age out stale tool-result images during compaction
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.

Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.

User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.

Refs #89938
2026-08-28 06:32:43 -07:00
Teknium 9978706e93 refactor(skill_manage): schema diet — patch args defer to the patch tool's semantics, file_path states its skill-dir-relative shape, authoring curriculum compressed (517 -> 427, -17%) (#97152) 2026-08-28 05:48:15 -07:00
Teknium a641644f12 fix(paths): display_hermes_home renders POSIX separators on Windows — kills ~/AppData\Local\hermes chimeras in tool schemas and user-facing messages (#97137) 2026-08-28 05:45:00 -07:00
Teknium 35328345d5 fix(profiles): anchor named-profile detection to real Hermes homes
Follow-up to the tombstone salvage (review point 1): named_profile_home()
treated ANY path whose ancestor's parent dir was literally named
'profiles' as a named profile until a '.hermes' ancestor appeared. An
unrelated custom HERMES_HOME like /srv/profiles/buildcache would resolve
as profile 'buildcache', and setup_logging would raise FileNotFoundError
instead of creating logs/ — a behavior regression for non-profile users.

Recognition now requires the 'profiles' directory's parent to BE a
Hermes home: the classic ~/.hermes layout, a root carrying home marker
files (config.yaml / .env / state.db — covers Docker/custom roots), a
profiles/.deleted tombstone dir (only ever created by profile delete),
or the process's resolved default Hermes root.

Adds regression tests for the /srv/profiles/<x> false-positive shape
(named_profile_home is None, mkdir_under_hermes_home and setup_logging
still create dirs) plus the tombstone-dir and ~/.hermes anchors. Also
adds encoding= to a bare write_text in the salvaged test file
(Windows-footgun gate).
2026-08-28 05:27:15 -07:00
EndeavorYen af2dc685fc fix(profiles): honor tombstones in exists/backfill and tighten named-home detection
Treat tombstoned leftover dirs as gone for exists/-p/use, skip them in
env backfill, replace only empty shells on recreate, and stop treating a
default home that merely contains a profiles path segment as named.
2026-08-28 05:27:15 -07:00
EndeavorYen 7e34522454 fix(profiles): tombstone deleted named profiles so logging cannot resurrect them
setup_logging and ensure_hermes_home could mkdir profiles/<name> after
hermes profile delete, so empty shells reappeared in profile list and
Desktop Bot Mode. Write a sibling tombstone, refuse mkdir/bootstrap for
tombstoned homes, and skip them in list/serve.
2026-08-28 05:27:15 -07:00
Teknium 911a41ec97 chore: map contributor email for hakanbaysal 2026-08-28 05:17:26 -07:00
Hakan Baysal cb8027afed fix(sanitize): preserve assistant messages with tool_calls when stripping images
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.

Adds a regression test covering the assistant + tool_calls +
image-only-content case.

Closes #40463
2026-08-28 05:17:26 -07:00
Frowtek e1762bd30b fix(agent): drop the api_content sidecar when stripping images from history
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:

    Replaying the pre-rewrite sidecar would resend exactly what the rewrite
    removed, so it must be dropped — the cost is one cache boundary miss,
    never wrong content.

`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:

    agent._vision_supported = False
    _imgs_removed = _strip_images_from_messages(messages)      # history
    if isinstance(api_messages, list):
        _strip_images_from_messages(api_messages)

and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.

Reproduced with the real functions:

    history content after strip : [{'type': 'text', 'text': 'look'}]
    sidecar still present       : True
    NEXT TURN sends             : 'look<IMAGE BYTES SENT LAST TURN>'

This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.

Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.

tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
2026-08-28 05:17:21 -07:00
Teknium a5c7eed5f3 feat(cli/tui): -q now seeds a live interactive session; prompts submit literally
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).

Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged

CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
2026-08-28 05:15:06 -07:00