Commit Graph

14240 Commits

Author SHA1 Message Date
Teknium be597fc730 fix: extend corrupt-config fail-closed guard to gateway, serve, and cron surfaces
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
  construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
  on every surface
2026-09-01 07:00:22 -07:00
embwl0x 6f85df97fd fix(cli): keep quiet prompts interactive 2026-09-01 07:00:22 -07:00
embwl0x 55e7ecd260 test(cli): cover config guard edge cases 2026-09-01 07:00:22 -07:00
embwl0x c335dc734a fix(cli): reject corrupt config in noninteractive runs 2026-09-01 07:00:22 -07:00
joaomarcos 9d3f1de994 fix(anthropic): alias session_search/memory OAuth billing-classifier triggers
Anthropic subscription OAuth (claude_code credential) misroutes Hermes
sessions carrying the session_search or memory toolset into the
extra-usage lane, surfacing as HTTP 400 "You're out of extra usage" on
a valid subscription. Live-verified via the
anthropic-ratelimit-unified-representative-claim response header
(deterministic lane oracle, no dependency on the laggy usage counter):
tool schemas are innocent, the trigger is three specific system-prompt
sentences (session_search recall + two skill_manage sentences),
required jointly — breaking any one clears the classifier.

Two independent layers:
1. OAuth wire alias (anthropic_adapter.py, transports/anthropic.py):
   session_search -> chat_history_lookup, memory -> context_notes in
   tool name, description, and (session_search only) system-prompt
   prose, with wire-collision guarding and a normalize_response
   reverse-map that keeps GH-25255 registered-tool precedence. Also
   routes named tool_choice through the same normalizer, closing a gap
   where a forced tool_choice would leak the raw trigger string and
   stop matching tools[].
2. Prompt-preserving reword (prompt_builder.py): rewords the two
   triggering SKILLS_GUIDANCE sentences while keeping the same
   meaning, still naming skill_manage, and leaving the Skill Safety
   Rule section untouched. Applies to all auth paths since it's a
   prompt-copy change, not a wire-level transform.

Two layers rather than one because the three-sentence AND-condition
means a classifier tightening could start firing on either remaining
leg alone.

Fixes #65365
2026-09-01 16:27:20 +05:30
kshitijk4poor 2755d3dd27 test: give the base FakeReviewAgent the release_clients cleanup seam
The base fake still stubbed the OLD cleanup surface (shutdown_memory_provider
/ close). Production now calls release_clients(); on fakes without it the
AttributeError is swallowed by the cleanup's except-Exception, so those
tests silently stopped exercising the cleanup path. The inner recording
fake in the renamed test keeps its close() stub deliberately — it's the
tripwire proving close() is never called on the shared session.
2026-09-01 16:27:13 +05:30
konsisumer fcd5bbb3fb fix(agent): preserve foreground resources after background review 2026-09-01 16:27:13 +05:30
Teknium 5a8e8a6b87 fix(terminal): strict Linux-only gating for background-executor systemd scopes (#70716 follow-up)
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):

- Gate every scope-path branch on a new _IS_LINUX constant instead of
  'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
  never touches systemd code — no probe subprocess, no scope argv,
  byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
  legacy argv byte-identical, no unit recorded) and probe-returns-False
  off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
  wired into the on-demand windows-venv-e2e lane): real spawn_local on
  windows-latest asserting jobs run exactly as before — spawned, output
  captured, exit code correct, systemd path never reached even under
  faked gateway identity.

Refs #70716, #71378.
2026-09-01 02:32:53 -07:00
lesseradmin 779482598f fix(desktop): pass --disable-setuid-sandbox on the userns launch path
When chrome-sandbox is present but not root-owned 4755, Chromium can still
abort via setuid_sandbox_host even though the namespace sandbox works.
After the userns probe skips sudo, append --disable-setuid-sandbox so
.desktop/no-TTY launches keep the namespace sandbox without a privilege
prompt. Does not add --no-sandbox.

Fixes #51327
2026-09-01 02:32:39 -07:00
4dlt 3a7f2234a6 fix(cli): use Chromium's namespace sandbox when userns is available on Linux
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).

On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.

Fixes #88032
Fixes #51327
Fixes #58593

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 02:32:39 -07:00
Yong Li f41ed09b51 fix(gemini): strip call ids on insert, name the realignments in the log
Review follow-ups:

- Strip the tool_call id when populating the call-name map so it matches the
  stripped lookup (and pass 1's result_call_ids). A padded id previously
  skipped realignment silently.
- Log which names were rewritten, not just how many.
- Note in the comment that a result whose assistant call frame was pruned is
  already dropped by the orphan pass, so it cannot reach the provider with a
  stale name; cover that with a test.
- Rename test_sanitize_leaves_matching_and_unpaired_tool_result_names_alone,
  which only ever exercised the matching case, and add the padded-id case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
Yong Li 2e9435d2f6 fix(gemini): echo bridged tool_call name on the OpenAI-compatible path
Google matches functionResponse.name against functionCall.name and rejects
a mismatch with HTTP 400 INVALID_ARGUMENT. #72089 fixed this for the native
Gemini adapter, where _translate_tool_result_to_gemini() now prefers
tool_name_by_call_id over the result message's internal name.

Requests that reach Gemini through an OpenAI-compatible gateway (OpenRouter,
Vertex/LiteLLM proxies) never run that translation, so they still put the
unwrapped internal tool name on the wire: the model calls the tool_search
bridge tool `tool_call`, make_tool_result_message() labels the result
`mcp__strava__get_recent_activities`, and the next turn 400s with a bare
"Provider returned error". The bad pair stays in the transcript, so every
later request in that session fails too.

Hold the same invariant at the final pre-API chokepoint instead of in the
OpenAI-compat serializer: Gemini arrives under many model strings and base
URLs, so sniffing for "is this really Google?" is unreliable, while every
other provider either ignores the field or already agrees with the call
name. Only a name that is present and disagrees is rewritten, so clean
transcripts still pass through byte-identical for prompt caching, and the
rewrite lands on the per-call copy so the stored trajectory keeps the real
tool name for the session DB and UI. No-op for the native Gemini path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
Teknium 530aa7b10f test(cache): fix stale Studio construction-order paragraph in the bridge witness docstring
The module intro still described the pre-b449d7c1 order (row first, agent
second); _reply_scope and numbered item 2 already state the correct one.
2026-09-01 02:14:35 -07:00
joaomarcos 41ec9d3591 fix(cache,tests): durable generation docstring + studio-bridge order (PR #98811, refs #96811)
agent/prompt_cache_scope.py: durable per-(source,session_key) monotonic generation; list _RESET_END_REASONS; note never garbage-collected. tests/agent/test_studio_bridge_affinity.py: build AIAgent before db.create_session; docstring corrected.
2026-09-01 02:14:35 -07:00
joaomarcos 5abba03a99 test(cache): pin the affinity contract on Studio's real bridge construction
The reproduction reported on #96811 is Hermes Studio's group chat, and it is
the one half of the issue that cannot be closed from inside this repository.
This states why in executable form, and pins the contract the host adoption
depends on so a later refactor cannot quietly break it.

Root cause, traced to the host. Studio reaches Hermes as a LIBRARY, not
through the gateway. Its bridge mints groupRuntimeSessionId(room, profile,
name) -- a gc_run_ prefix truncated to 96 characters plus a fresh UUID4 hex --
for every reply, writes the session row itself, and constructs AIAgent(...)
with platform / session_id / session_db and no routing identity of any kind:
no gateway_session_key, no chat_id, no user_id, no parent_session_id. So the
declared scope is unreachable, the row is a lineage root, and
resolve_prompt_cache_scope() correctly falls through to the physical id, which
moves on every reply. No stable carrier crosses the boundary. Hermes must not
recover one from the id's syntax -- that is #79017's collision class, and the
two negative controls merged via #97704 exist to keep it from trying.

What does cross the boundary is a value Studio already has:
groupBridgeSessionId(room, profile, name, sessionSeed, runtimeConfig) is
stable for one conversation in one room, already carries the room, profile,
agent name and the room-owned sessionSeed, and is already hashed and
length-bounded. Passing it as gateway_session_key is the entire adoption, and
it is a Studio-side change; this PR keeps Refs #96811 for exactly that reason.

What this suite adds is the half that IS reachable here. The existing suites
simulate Studio's id SHAPE against synthetic agent doubles; none of them runs
the host's construction path, so nothing today would fail if that path stopped
honouring a declaration. These tests build a real AIAgent in the bridge's own
order -- row written first, agent second, new physical id per reply -- and
state what the adoption buys:

- three replies with distinct physical ids hold ONE affinity scope, and the
  scope carries neither the room nor the member name;
- two members of one room, and two rooms, never share a bucket;
- a new sessionSeed rotates the scope, which is Studio's own conversation
  boundary and needs no reset observed by Hermes;
- equal keys under different row sources never collapse, because the row's
  source -- not agent.platform -- is the identity the peer queries match on;
- a tool child and a background-review fork on the same declared key keep
  their own scope, so #79161 survives this construction path too;
- and an undeclared bridge is byte-identical: per-reply scope, and a
  compression rotation still walks its lineage.

Against origin/main the two declared-conversation assertions fail --
"assert 3 == 1", and the raw gc_run_ id leaking as the routing key -- which is
the reproduction; the remaining eight pass there and here, because they
describe behaviour that must not change.

Reported by @cervantesh, whose re-check of Studio main@86d0c95375 located the
missing consumer on the real path and asked for exactly this witness.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos 2bb1d80bd9 test(api): cover the declared-conversation precedence through the real handlers
The review asked for "real handler + DB coverage for both precedence paths".
The first pass did not deliver that: TestBindFollowsPrecedence restated the
gate expression inside the test, so it asserted a copy of the rule rather than
the rule, and would have stayed green if the handlers stopped applying it.

These drive POST /v1/responses and POST /v1/runs over real routes with a real
adapter and a real SessionDB:

- the declared key selects and records the conversation, and the row it
  produces carries the key;
- three replies on one declared key land on one session id;
- an undeclared request keeps a per-request id and records nothing;
- a request carrying conversation A's previous_response_id plus a foreign
  header key settles on A, records nothing, leaves A's own key intact, and the
  foreign key still cannot recover A -- the end-to-end shape of the blocker;
- /v1/runs, which owns its agent lifecycle rather than routing through
  _run_agent, settles on the declared conversation, and an explicit body
  session_id outranks the header key without rebinding it.

The _run_agent stand-in creates the session row the way
AIAgent._ensure_db_session does and performs the bind the way _run_agent's
finally block does, so the assertions land on real rows instead of on a mock's
call args. /v1/runs is captured at _create_agent for the same reason.

The restated-gate tests are kept as the cheap unit layer beneath these.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos d63e5d8a10 fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee.

1. The generation lookup was not in the same identity domain as recovery.
   latest_conversation_boundary() selected on session_key alone, while
   _declared_conversation_session() is qualified by (source, session_key).
   X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
   API conversation may legally carry the same key as a Telegram row in one
   database -- a /new over there rotated this conversation's gwk_ generation
   while recovery correctly refused to cross the same line, moving the affinity
   identity out from under a physical identity that had not moved.

   The boundary read now takes (session_key, source), and the carrier is
   'source|key|generation' rather than 'key|generation' -- keying on the string
   alone would also collapse two same-key conversations from different sources
   onto one routing key, since this value leaves the process verbatim as
   OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
   from the agent's own session row, falling back to the platform the row will
   be created with before it lands.

2. The declared key's stated lower precedence did not survive settlement. Both
   handlers let stored_session_id / an explicit body session_id win, then called
   _bind_declared_conversation() unconditionally. record_gateway_session_peer()
   does SET session_key = ? across compression ancestors, so a request carrying
   conversation A's chain plus header key B silently rebound A to B: A could no
   longer be recovered by its own key, and B recovered A's session.

   Recording is now gated on the declared key having actually selected or
   minted the session, on both paths. Behind that gate the bind itself refuses
   to overwrite a row already bound to a different key, so a future caller
   cannot reintroduce the same defect by opting in wrongly.

test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.

Refs #96811

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos e5bce4df4b fix(cache): make the conversation generation survive a backwards clock
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an
NTP correction between two resets writes a SMALLER boundary, MAX keeps
returning the older one, and the next conversation silently reuses the
previous generation -- two conversations on one routing key, which is the
defect this PR exists to remove.

latest_conversation_boundary now returns (count, latest_ended_at) and the
marker is 'count:ended_at'. The two halves fail under different conditions --
a backwards clock defeats the timestamp, retention pruning of an old ended row
decrements the count -- so a generation repeats only if both happen at once.
The pair is deliberately biased toward changing: a spurious change costs one
cold prompt-cache bucket, a repeat would merge two conversations.

Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites
the second boundary to land before the first and asserts three conversations
still resolve to three distinct scopes.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 3739cf3b86 fix(api): resolve the declared conversation instead of minting a session per request
POST /v1/responses and POST /v1/runs parse and authenticate the client's
X-Hermes-Session-Key, pass it downstream for memory scoping, and then mint a
throwaway physical session id anyway whenever the client manages its own
history (no previous_response_id chain to carry one forward).

Every conversation-affinity hint Hermes sends is derived from that physical
id, so all four re-keyed on every single reply: prompt_cache_key on both
OpenAI-wire transports, the OpenRouter and Nous sticky session_id, and xAI's
x-grok-conv-id. The conversation never landed back on a warm prefix.

Fix the identity rather than the four consumers. The declared key resolves to
its live session through find_latest_gateway_session_for_peer -- the same
reset-fenced recovery every native gateway platform already uses -- and the
turn records the row it ended on through record_gateway_session_peer, which
AIAgent._ensure_db_session never did (it knows the key and writes the row
unkeyed, so the mapping the next reply needs did not exist).

Because the lookup is fenced on sessions.end_reason, the generation that must
rotate is already durable: session_reset (/new), session_switch, idle, daily,
suspended and resume_pending_expired all return None, so a new conversation
gets a new id and a cold affinity scope, and a retired generation can never be
resolved again. No counter, no new persisted field, and no new precedence rule
in the cache-scope resolver -- /branch, delegate and tool children keep the
isolation of #79161/#79017 byte for byte.

Precedence is unchanged where it already worked: an explicit body session_id
and the previous_response_id chain both still outrank the declared key, and a
request that declares nothing keeps its per-request id. Recording is opt-in
(bind_declared_conversation), so no other _run_agent caller's rows change.

Refs #96811

(cherry picked from commit e7c83dddf36784d1012bf483240ebc7f6b2ef9aa)
2026-09-01 02:14:35 -07:00
joaomarcos d7995bffaf fix(cache): qualify the declared key with the conversation generation
The declared key is a per-CHAT identifier and outlives the conversation it
names: reset_session() mints a fresh physical id on /new but keeps the key, and
the idle/daily/suspended policy resets do the same. Hashing the key alone
therefore mapped the conversation before a reset and the one after it onto one
gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and
@kshitijk4poor reproduced on #97709.

No counter is introduced. The generation that must rotate is already durable:
every one of those boundaries closes the outgoing row with an
_RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads
the most recent one and declared_conversation_scope hashes 'key|generation'.

That makes the carrier stable across a host's per-response physical ids -- a
host that never resets writes no boundary, so every reply hashes the same value
-- while rotating on every conversation replacement, /new and the policy
auto-resets alike. ended_at only moves forward, so a retired generation can
never be reused: no ABA.

It also cannot drift from the rest of the codebase's notion of a conversation
boundary, because find_latest_gateway_session_for_peer fences on the same set.

The read is on the memoized resolution path, not per API call, and both lookups
fail closed: an unqualified key would span a /new, so a DB error degrades to the
physical-id scope. A SessionDB without the lookup keeps the previous behaviour.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 65672e3a93 fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).

Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.

Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.

- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
  into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
  rotation AND across per-response ids). Hashed because, unlike a session id,
  the key embeds platform/chat/user identifiers and leaves the process
  verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
  when a host declared one. The providers read the attribution id when it is
  unset, so delegate trees keep sharing their parent's sticky key and every
  host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
  rules that keep /branch children, delegate subagents and tool children off
  their parent's chat key. Background-review forks clone the live runtime, so
  _persist_disabled excludes them for the same reason (#79161).

Refs #96570
Fixes #96811
2026-09-01 02:14:35 -07:00
kshitijk4poor 29d4c0ebfd refactor(compression): extract preflight seed predicate onto ContextCompressor
Address review follow-ups on the seed fix:

- Move the 'seed only from the 0 state' guard from an inline block in
  build_turn_context() into
  ContextCompressor.maybe_seed_preflight_display_tokens(), co-locating
  the predicate with the rest of the speculative-seed lifecycle
  (snapshot_preflight_display_tokens /
  rollback_interrupted_preflight_display_tokens). Callers now use the
  method via a getattr guard so test doubles and external context
  engines without it are unaffected.
- Rewrite the TestPreflightSentinelGuard docstring, which still
  described the old >=0 guard ('treats any negative value as no real
  usage yet'); the ==0 policy protects ALL non-zero readings.
- Drop the _seed mirror-helper: the tests now call the real production
  method on the compressor fixture, eliminating mirror-drift risk (the
  helper comment had already drifted once).
- Note the accepted trade-off (partial-usage providers pin the meter
  low until their next report) in the method docstring.
2026-09-01 14:33:23 +05:30
Turgut Kural 80b836fa9e fix(compression): preflight display-seed must not overwrite real provider usage
The preflight rough-estimate seed used 'last >= 0' semantics, so any
provider that reports real prompt_tokens got its reading replaced by
the schema/reasoning-inflated rough estimate whenever the estimate was
larger. last_prompt_tokens feeds both the CLI context meter (cli.py)
and the post-response compression gate (conversation_loop.py), so one
seed made the bar jump to an inflated number AND could push the
real-usage gate over threshold on estimator noise.

Observed in a production CLI session on a 1M-token window with a
reasoning-heavy history: the status bar showed ~492K real provider
prompt tokens; the next turn's preflight estimated ~685K (rough
estimates can inflate 1.4-2.5x on reasoning-heavy sessions, #81481)
and seeded it into last_prompt_tokens; compression then fired at ~69%
of the window while real usage was ~49%.

Policy change: a real provider reading (>0) always wins. The seed now
only fills the 0 state ('no reading yet'), keeping the status bar live
for usage-less providers (#34282's motivation); -1 remains protected
as the post-compression sentinel (#36718).

Updates TestPreflightSentinelGuard to encode the new policy and adds a
regression test with the measured 492K/685K numbers.
2026-09-01 14:33:23 +05:30
VJ Pixel 3571118218 fix(agent): run init-time fallback for ANY exhausted primary provider
The #17929 init-time fallback block was nested inside the
'_explicit not in {auto, openrouter, custom}' guard, so an exhausted
openrouter credential pool skipped fallback_providers entirely and
AIAgent.__init__ raised 'No LLM provider configured' — surfacing on
Telegram only as the generic 'unexpected error' message.

2026-08-23 outage (~22:09-23:50, 60 occurrences incl. a cron job):
single-entry openrouter pool hit daily free-tier quota; chain had a
healthy local Ollama entry that never got tried because provider was
the default 'openrouter'.

Un-nest the block so any primary without usable credentials walks the
chain before failing; providers explicitly chosen by name keep their
dedicated missing-key diagnostic when both primary AND chain fail.
Regression tests cover the openrouter-exhausted path both with and
without a usable chain entry.
2026-09-01 02:00:05 -07:00
ehz0ah 7fef5c7898 fix(openviking): clarify remember submission status 2026-09-01 14:11:55 +05:30
ehz0ah 6446e19cb5 fix(openviking): route remember through session extraction 2026-09-01 14:11:55 +05:30
Teknium a1c25d393a feat(desktop): built-in optional-skills catalog in Capabilities → Skills with one-click install
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.

- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
  scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
  install-state spinners off $hubActions, and an OfficialSkillDetail pane
  (hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
2026-09-01 01:39:59 -07:00
liuhao1024 46ac31e84c fix(desktop): sanitized deferral evidence for ledgered manual serve blockers
The Desktop venv-blocker scan (since #99724) defers ledger-verified
serve/dashboard holders to the CLI updater's stop+relaunch rungs, but the
scan output only carried an opaque deferred_backends count — nothing
explained WHICH holders the deferral consumed or why they vanished from
processes.

Add sanitized decision evidence (#98350): deferred_backend_evidence lists
structured ledger identity only (pid, purpose, recorded port) — never the
command line, which can carry tokens or private endpoints. Adds a desktop
parser contract fixture proving the consumer tolerates the diagnostics
while keeping blocked/processes authoritative.

Salvaged from PR #98350; the exemption half of that PR was independently
consolidated on main via #99724 (_is_updater_owned_backend).
2026-09-01 01:30:09 -07:00
kshitijk4poor 801466765b fix: harden temp fallback against predictable-path attacks; neutral error text; docs
Final-diff review findings (/simplify-code on the full 3-commit stack):

- Temp-dir fallback uses a per-uid name (hermes-profile-exports-<uid>) and
  get_profile_export_path refuses a pre-existing symlink or a directory
  owned by another user — a fixed /tmp/hermes-profile-exports is a
  predictable shared path a local attacker could pre-create to receive the
  secret-bearing archive. Regression test mutation-checked.
- Fail-closed message reworded interface-neutrally (the web API surfaces it
  as HTTP 400 detail where '-o' alone made no sense).
- Docs now cover the temp fallback and the fail-closed refusal.
2026-09-01 01:00:23 -07:00
kshitijk4poor 525dd12da7 fix: fail closed when no safe export destination exists; polish salvage edges
- _profile_export_directory(): when the managed store, the home-sibling
  store, AND the temp dir all resolve inside Git checkouts, raise a clear
  ValueError instead of warning and proceeding — a stderr warning would not
  stop a scripted export from staging a secret-bearing archive in a source
  tree, which is the exact #92457 incident class. All three callers already
  surface ValueError cleanly (CLI/TUI print Error: + exit, API returns 400).
- .dockerignore: drop the /default.tar.gz line made redundant by the global
  *.tar.gz pattern this PR adds.
- hermes profile export -o help text: stop advertising the old
  <name>.tar.gz cwd default.
- Tests: cwd-in-unrelated-checkout topology (the second production shape
  from the blocking review) and the fail-closed path. Mutation-checked:
  both fail on the pre-fix helper.
2026-09-01 01:00:23 -07:00
kshitijk4poor 26fb8f60e6 fix: anchor checkout detection to the export path, not cwd
Follow-up to the salvaged #92689:

- _profile_export_directory() now proves safety on the export dir's OWN
  ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
  heuristic missed the checkout whenever HERMES_HOME sat inside one but
  the process ran from elsewhere (cron, service manager) — the export
  landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
  violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
  and catch OSError too — a bad profile name or read-only home printed a
  raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
  pollution in the tests/hermes_cli sweep can't divorce monkeypatches
  from the code under test; add regression tests for the cwd-independent
  topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
2026-09-01 01:00:23 -07:00
joaomarcos c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
Brooklyn Nicholson b293157f7d feat(tools): withdraw tip and tour when the user switches them off
Off should mean the model is never told the tool exists. A switch that
only made the call fail leaves Hermes offering walkthroughs it cannot
give and promising to point at things it cannot point at, which reads as
a broken agent rather than a respected preference.

Both gate on the switch through a shared desktop_ui.user_enabled helper,
which is the reactions check_fn generalized — same config read, same
reason it has to be config rather than an env var: the switch belongs to
the session's client, and the client may be on another machine.
2026-08-31 21:11:11 -05:00
Brooklyn Nicholson d0e41cda69 fix(gateway): let config.set write the desktop's Appearance switches
config.set matches an explicit key list and answers 4002 for anything
else, so a renderer mirroring an unlisted key wrote nothing at all. The
reactions toggle shipped that way: every write was rejected into a
swallowed .catch(), and react_to_message stayed dark no matter what the
user picked.

Adds the display booleans as a recognized group so the toggle reaches
the config of whichever gateway the app is actually talking to, which is
the only place a check_fn can read it.
2026-08-31 21:11:11 -05:00
Brooklyn Nicholson 1b47277067 test(tui_gateway): cover stop-vs-marker recovery races
Pin that session.interrupt retires the original and rotated marker keys
on ACK, and that a Stop arriving before the disk write cannot leave a
resume-able marker behind.

Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
2026-08-31 20:50:39 -05:00
Brooklyn Nicholson adb23c13cb fix(bot-mode): keep Bot Chat resume on a proven compression tip
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-08-31 20:47:53 -05:00
Gille 4abf5c0790 test(bot-mode): scope title lookup regressions 2026-08-31 20:47:53 -05:00
Gille bb28056efd fix(bot-mode): keep canonical lookup on compression lineage 2026-08-31 20:47:53 -05:00
Ben Barclay 56916841b5 refactor(dashboard-auth): replace PKCE cookie payload with base64url(JSON) codec (#99210)
The PKCE cookie's payload has now needed three serialization fixes at
the same spot: the original flat 'k=v;k=v' string tripped http.cookies'
\073 quoted form (dropped whole by strict cookie parsers like Go's
net/http — #83832 field case), and #99176 URL-encoded the whole flat
payload to stay inside the RFC 6265 cookie-octet set. The stacked
layers (single-encoded next=, ';' joins, whole-payload encoding, legacy
discriminator) were the recurring defect source.

Kill the bug class instead of patching it again: the payload is a dict
end-to-end and goes on the wire as base64url(JSON) — the urlsafe
alphabet is a strict subset of cookie-octets, and JSON framing means no
segment value can ever collide with a delimiter. parse_pkce_payload
keeps a three-rung compatibility ladder (base64url(JSON) -> oldest flat
form split-as-is -> #99176 unquote-then-split) for in-flight cookies
during a rolling upgrade (10-minute TTL); a new cookie hitting an old
server fails the OAuth state check and the user just retries.

The 'next' segment is stored as its plain validated path — no extra
encoding layer, so the post-login redirect Location is byte-for-byte
the original target.

Refs #99176, #84065.
2026-09-01 11:47:38 +10:00
xxxigm c99a3919a8 fix(dashboard): drop systemd-only advice from the model-picker code-skew 503 (#97046)
Desktop-owned serve and macOS hosts have no hermes-dashboard unit, so the
restart hint now follows HERMES_SERVE_HEADLESS instead of hardcoding systemctl.
2026-08-31 20:43:36 -05:00
Brooklyn Nicholson 5c6dbe22c3 fix(compression): take reasoning_content when the summarizer leaves content empty
Local and thinking backends (DeepSeek, Qwen, Kimi) often return a usable
summary in reasoning fields. Treat that as the summary instead of burning
another 100s+ empty-content retry. Leave the wire max_tokens omit intact.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
Brooklyn Nicholson 02458e67ad fix(auxiliary): accept dict and object messages in extract_content_or_reasoning
Compression and some OpenAI-compatible proxies hand us a dict-shaped
response or a bare message, not a ChatCompletion. Reuse the existing
helper instead of a second extractor, and bound an optional reasoning
fallback so a chain-of-thought dump cannot become the summary.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
Brooklyn Nicholson 0e7eebc266 fix(desktop): scope remote model catalog and primary-label REST to the focused profile
A shared dashboard's launch HERMES_HOME is not the selected profile. model.options now runs under @_profile_scoped, and global-remote REST keeps ?profile= even for the primary label.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-08-31 20:40:09 -05:00
kshitijk4poor cc0931d235 test(agent): loosen brittle error/final_response equality to substring
result['error'] and result['final_response'] are independently settable
keys that only coincidentally share _COMPRESSION_TIMEOUT_FINAL_RESPONSE
today; assert the actionable substring instead so a benign prefix or
rewording does not break the terminal-contract test.
2026-09-01 03:48:40 +05:30
kshitijk4poor 1e8f6a0491 fix(agent): surface preflight compression timeout as typed result, not generic error
When the turn-start fail-closed boundary (#98424) raises
PreflightCompressionTimedOut, the exception escaped run_conversation to
the surfaces' generic exception handlers. The gateway deliberately never
exposes raw exception text, so users saw 'Sorry, I encountered an
unexpected error... Try again or use /reset' instead of the boundary's
actionable guidance, and the compression_exhausted clean-session
recovery contract (#9893/#35809) never engaged.

Catch it at the build_turn_context callsite and convert it into the
same typed recovery dict the in-loop timeout consumers return
(salvaged #98741 / PR #99710): failed=True, partial=True,
compression_exhausted=True, turn_exit_reason=context_compression_timeout,
with the actionable message in final_response and error.

Regression test proves the exception no longer escapes and the typed
contract fields survive to the caller (mutation-checked: test fails on
main without the handler).
2026-09-01 03:48:40 +05:30
Teknium 5a407a0c52 test(gateway): pin routing-index harness to the tmp HERMES_HOME store
The #66887 fix pins the routing index to HERMES_HOME state.db; the
fast-path harness still read entries through the ambient store. Point
get_hermes_home at the test tmp so both are the same file, matching
the new single-store contract.
2026-08-31 14:54:18 -07:00
caya8205-2 8d74cb52da fix(gateway): give the routing index one store instead of the ambient one
Second half of #66887. _entries is a single flat dict holding every
profile's keys, so the index it persists to has to be a single file — but it
was read and written through _db, which resolves whichever profile scope is
active. A whole-index rewrite during one profile's turn copied every other
profile's routing rows into that profile's store, and startup, which runs
unscoped, then loaded a different copy than the last writer produced.

That is why the startup recovery pass never sees a secondary profile's crash
marker, which is the half this issue's title names. mark_turn_active()
persists through the single-entry fast path (state.db only, no sessions.json
mirror), so a marker written during a profile's turn landed in that
profile's store and _recover_unclean_sessions(), running with no scope, read
a store that had never heard of it. The turn was silently never promoted to
resume_pending.

Capture the gateway's own home at construction — the store is built at
startup before any profile scope exists — and route the index through it:
_ensure_loaded_locked, _reconcile_recovered_routing_locked,
_persist_routing_data and _save_entry now use _routing_db. A pinned handle
still wins, so suites that install a fake or disable the DB are unaffected.

_prune_stale_sessions_locked is the mixed case and is split accordingly: it
now asks _db_for_key(key) whether each session ended, because that is a
per-session question, while the index write stays on the single store. One
ambient handle previously answered it for every profile at once, which could
prune a live secondary-profile route on the strength of the root store's
copy of that session.

Regression as requested on the issue: mark a turn active under a secondary
profile's scope, then build a fresh store with no scope and run
recover_interrupted_turns(). It promotes exactly one turn to resume_pending
here and promotes zero against the previous behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 a60e04a32c fix(gateway): prove the compression child's owner before writing to it
Review P1 #2. _append_to_transcript_serialized() writes the compression
continuation to child_id BEFORE publishing either _transcript_reroutes or
the _entries update — that ordering is load-bearing for backlog order, so it
must not move. At that moment nothing in the routing index points at the
child, so _db_for_session_id(child_id) missed its scan and fell through to
_db_for_key(None), i.e. the ambient store. The fail-closed guard did not fire
because root is a live handle.

The row therefore targeted root rather than the already-proven parent owner.
With no child row there the append is rejected by the FOREIGN KEY constraint,
the pending queue never drains and the reroute cannot advance; against a
split-brain root the message would instead be written cross-profile.

Record ownership before the mutation instead of moving the publication: a
private _session_owner_hints map carries session_id -> owning key for ids
whose owner is proven but not yet published, consulted by the new
_owner_key_for_session_id() after the index scan misses, and dropped as soon
as routing publishes. Signatures are unchanged, so the existing suites that
stub _append_transcript_message keep working untouched; the map is read
through getattr for stores built via object.__new__.

The regression is physical rather than mocked: an ended compression parent
and a live child that exist only in profiles/fitness/state.db, no active
profile scope, append to the parent, then assert all four effects — the row
lands on the child in the profile store, the pending queue drains, the
reroute and the routing entry advance, and root state.db stays untouched.
Without the hint it fails exactly as the review predicted, on
"FOREIGN KEY constraint failed" against root.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00