8523402db00a9f477dd2ee28cfaf974aec3ed555
6 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6b9b3e0145 |
chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and listed five cleanups. All five are here. 1. scratch/repro_96811.py is deleted. It would have landed on main as a tracked file: scratch/ is not gitignored and has never existed on main, so this PR was creating the directory. Nothing referenced the probe, and TestConversationGenerationRotates / TestGenerationSurvivesPruning / TestPeerIdentityIsSourceQualified already carry all four of its stages, so it is dropped rather than parked under tests/. 2. Upgrade notes are written into this commit body (below) and the PR body. There is no committed changelog to add them to: scripts/release.py generates .release_notes.md from commit SUBJECTS at release time, and .gitignore keeps that file out of the tree. 3. declared_conversation_scope() now reads the sessions row ONCE. The fork verdict and the source the peer queries match on both live on that row, and asking for them separately read it twice per resolution. The new SessionDB.declared_scope_identity() returns the pair and keeps the marker rules beside is_explicit_fork_child() instead of re-implementing them in the caller. A SessionDB that does not expose the combined view keeps the original two-call path, so nothing that predates it changes behaviour -- including the three doubles that certify the fail-closed contract, which are untouched. TestOneIdentityReadPerResolution pins the single read, the two-call fallback, the fail-closed degrade and the fork refusal; removing the fold turns the first of those red. The third read stays: the generation lives in conversation_generations, a different table, and cannot be folded into a sessions lookup. 4. _declared_conversation_session() documents the concurrent first-turn race. Two simultaneous first requests on one declared key can each miss the lookup, mint a row and both bind, because each row is unkeyed at bind time and the mismatch guard does not fire. That converges rather than crossing: both rows carry the same key under the same source, so the lookup returns the later one for every subsequent reply and the earlier row is an abandoned transcript, never another conversation's identity. The same docstring still claimed the generation was durable in sessions.end_reason and that "nothing here needs a counter". That stopped being true in 09004753c9, which moved the generation into conversation_generations precisely because deriving it from prunable session rows was ABA. Corrected, along with the same stale sentence on TestConversationBoundariesRotate. 5. conversation_generations rows are now documented as deliberately never collected, rather than merely uncollected. Dropping one resets that peer to "no generation", so its next boundary writes 1 again and re-issues a gwk_ scope a retired conversation already used -- the exact ABA the table exists to close. Worth stating because the repo already carries both patterns a maintainer would extend: delete_session() cascades to messages, and gateway_hygiene_state is already swept by session_key. Upgrade notes, one-time on merge: - One cold prompt-cache bucket per keyed conversation. Every gateway platform declares gateway_session_key, so each keyed conversation's affinity scope moves once from its compression-lineage root session id to the gwk_ hash. One cache miss per live conversation, on its next turn only. - hermes status counts more sessions. A declared API conversation is now recorded as a keyed row and appears in "Active: N session(s)" where it was invisible. Those sessions already existed; only their visibility changes. - A database upgraded mid-conversation starts with no generation and takes its first from the next boundary written, so a conversation that reset before the upgrade shares its predecessor's scope once. One warm bucket, never a crossed identity. Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new), 33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py, 25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in test_cross_process_turn_lease.py, and 526 across test_hermes_state.py + tests/hermes_state/ + tests/state/. ruff clean. Found in review by @teknium1. Refs #96811 |
||
|
|
832d68aba4 |
fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The first two are defects I introduced in 99f2d4394f by replacing the wrong occurrence of an identical call site. 1. _run_agent raised NameError on every opted-in declared bind. Its worker finally evaluated `if _declared_selected:`, a local of _handle_responses / _handle_runs that is neither a parameter nor an enclosing binding here, so the successful declared-key paths failed at settlement after the agent run. bind_declared_conversation already IS the gate; the inner name is gone. 2. /v1/runs never received the gate at all -- it landed on _run_agent instead. _run_sync bound unconditionally, so an explicit body session_id that existed with an empty session_key was adopted by the header key even though the header lost precedence. It now carries the same gate. 3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse. delete_session() deletes the selected row and bulk prune selects ended rows, so the aggregate can return a pair it already emitted: (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation a retired affinity identity. The backwards-clock shape needs no pruning at all. The generation now lives in a conversation_generations table keyed by (source, session_key), advanced by _bump_conversation_generation inside the same transaction that writes each boundary -- outside prunable session history, wall-clock-free, and increment-only. end_session() and promote_to_session_reset() both advance it, and only when they actually wrote a boundary, so a repeated end cannot double-count. 4. The carrier could be memoized under the wrong source. _agent_source() fell back to agent.platform before the row landed while persistence uses _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a declared scope is non-None immediately, resolve_prompt_cache_scope memoizes it and never re-resolves once the authoritative row appears, so under an override both sides of a /new read the platform domain and hashed the same scope. The pre-row path now uses the persistence resolver itself. Coverage answers the review's specific objection that mocked tests proved the mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit row and waits for the worker to retire before asserting. Both were verified by mutation: reinstating the inner name fails two of them, and removing the /v1/runs gate fails the explicit-session one. The first version of that test passed with the gate removed -- it asserted before settlement -- and would have been the same empty proof the review called out. TestGenerationSurvivesPruning covers deleting the newest boundary, deleting every boundary, the backwards-clock-then-prune shape, compression and accidental ends not advancing it, repeated ends not double-counting, promotion advancing it, unkeyed rows advancing nothing, and peer scoping. TestSourceOverrideDomain covers the override across a reset. Found in review by @andrexibiza, whose analysis located each of these defects and specified what a correct fix had to prove. Refs #96811 Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com> |
||
|
|
d63e5d8a10 |
fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee. 1. The generation lookup was not in the same identity domain as recovery. latest_conversation_boundary() selected on session_key alone, while _declared_conversation_session() is qualified by (source, session_key). X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an API conversation may legally carry the same key as a Telegram row in one database -- a /new over there rotated this conversation's gwk_ generation while recovery correctly refused to cross the same line, moving the affinity identity out from under a physical identity that had not moved. The boundary read now takes (session_key, source), and the carrier is 'source|key|generation' rather than 'key|generation' -- keying on the string alone would also collapse two same-key conversations from different sources onto one routing key, since this value leaves the process verbatim as OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes from the agent's own session row, falling back to the platform the row will be created with before it lands. 2. The declared key's stated lower precedence did not survive settlement. Both handlers let stored_session_id / an explicit body session_id win, then called _bind_declared_conversation() unconditionally. record_gateway_session_peer() does SET session_key = ? across compression ancestors, so a request carrying conversation A's chain plus header key B silently rebound A to B: A could no longer be recovered by its own key, and B recovered A's session. Recording is now gated on the declared key having actually selected or minted the session, on both paths. Behind that gate the bind itself refuses to overwrite a row already bound to a different key, so a future caller cannot reintroduce the same defect by opting in wrongly. test_declaration_outranks_the_lineage_root asserted the pre-qualification contract by comparing a DB-backed agent against a DB-less one; it now makes the stronger statement it was written for -- one declared conversation reached through two different physical ids on the same peer. Refs #96811 Found in review by @andrexibiza, whose analysis located each of these defects and specified what a correct fix had to prove. Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com> |
||
|
|
e5bce4df4b |
fix(cache): make the conversation generation survive a backwards clock
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an NTP correction between two resets writes a SMALLER boundary, MAX keeps returning the older one, and the next conversation silently reuses the previous generation -- two conversations on one routing key, which is the defect this PR exists to remove. latest_conversation_boundary now returns (count, latest_ended_at) and the marker is 'count:ended_at'. The two halves fail under different conditions -- a backwards clock defeats the timestamp, retention pruning of an old ended row decrements the count -- so a generation repeats only if both happen at once. The pair is deliberately biased toward changing: a spurious change costs one cold prompt-cache bucket, a repeat would merge two conversations. Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites the second boundary to land before the first and asserts three conversations still resolve to three distinct scopes. Refs #96811 |
||
|
|
d7995bffaf |
fix(cache): qualify the declared key with the conversation generation
The declared key is a per-CHAT identifier and outlives the conversation it names: reset_session() mints a fresh physical id on /new but keeps the key, and the idle/daily/suspended policy resets do the same. Hashing the key alone therefore mapped the conversation before a reset and the one after it onto one gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and @kshitijk4poor reproduced on #97709. No counter is introduced. The generation that must rotate is already durable: every one of those boundaries closes the outgoing row with an _RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads the most recent one and declared_conversation_scope hashes 'key|generation'. That makes the carrier stable across a host's per-response physical ids -- a host that never resets writes no boundary, so every reply hashes the same value -- while rotating on every conversation replacement, /new and the policy auto-resets alike. ended_at only moves forward, so a retired generation can never be reused: no ABA. It also cannot drift from the rest of the codebase's notion of a conversation boundary, because find_latest_gateway_session_for_peer fences on the same set. The read is on the memoized resolution path, not per API call, and both lookups fail closed: an unqualified key would span a /new, so a DB error degrades to the physical-id scope. A SessionDB without the lookup keeps the previous behaviour. Refs #96811 |
||
|
|
65672e3a93 |
fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints one physical session per RESPONSE re-keys all four on every reply, so the conversation never lands back on the routing bucket it just warmed (#96811). Two hosts do exactly that. Hermes Studio's group chat mints gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after, and POST /v1/responses with client-managed history mints str(uuid4()) per request — while parsing X-Hermes-Session-Key one screen earlier and handing it to the agent. Hermes must not infer the logical conversation from the id's syntax: that rule merges independent client-supplied ids and Studio members truncated past its 96-character boundary (the #79017 failure class). It does not have to. gateway_session_key is already the "stable per-chat key" built by gateway.session.build_session_key from that header, and branching deliberately does not key off it. The affinity path simply never consulted it. - agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across rotation AND across per-response ids). Hashed because, unlike a session id, the key embeds platform/chat/user identifiers and leaves the process verbatim as a sticky id and as x-grok-conv-id. - agent/portal_tags.py: a separate ambient scope for ROUTING, published only when a host declared one. The providers read the attribution id when it is unset, so delegate trees keep sharing their parent's sticky key and every host that keeps one id per conversation is byte-identical to before. - hermes_state.py: is_explicit_fork_child() — the public view of the marker rules that keep /branch children, delegate subagents and tool children off their parent's chat key. Background-review forks clone the live runtime, so _persist_disabled excludes them for the same reason (#79161). Refs #96570 Fixes #96811 |