Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:
- `_request_owns_run` no longer admits run state that exists without an
owner stamp. Under gateway.multiplex_profiles every served profile holds
a valid key, so the "backward compatibility" branch made the boundary
allow-all whenever provenance was missing. Unstamped state now fails
closed; only an in-memory owner match or a durable idempotency record
under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
inside the request's profile scope, so its run is confined to the
creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
(`_release_run_owner_if_forgotten`) and runs at every retirement point
(task finally, SSE stream close, both sweep loops, chat-stream finally),
not only the terminal-status sweep — no stranded entries, no stateful id
ever left unowned.
Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).
Fixes#93689Fixes#90415
Supersedes #93747, #93704, #92822
Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
Two compounding bugs that cause WebUI to discard or misrender agent
responses when using GLM models on Ollama Cloud:
1. _is_ollama_glm_backend() matched "ollama" in base URL, which
included Ollama Cloud (ollama.com). The hosted service correctly
reports finish_reason and is not affected by the local Ollama
stop-reason bug. Exclude "ollama.com" before the substring check.
2. _handle_session_chat_stream() hardcoded "partial": False in the
assistant.completed SSE event instead of reading result.get("partial").
The WebUI could not detect truncation and rendered partial responses
incorrectly (showing only the continuation instead of the full text).
Read the partial flag from the agent result, matching the pattern
used by other SSE paths in the same file.
Fixes#72316
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.
1. scratch/repro_96811.py is deleted. It would have landed on main as a
tracked file: scratch/ is not gitignored and has never existed on main, so
this PR was creating the directory. Nothing referenced the probe, and
TestConversationGenerationRotates / TestGenerationSurvivesPruning /
TestPeerIdentityIsSourceQualified already carry all four of its stages, so
it is dropped rather than parked under tests/.
2. Upgrade notes are written into this commit body (below) and the PR body.
There is no committed changelog to add them to: scripts/release.py
generates .release_notes.md from commit SUBJECTS at release time, and
.gitignore keeps that file out of the tree.
3. declared_conversation_scope() now reads the sessions row ONCE. The fork
verdict and the source the peer queries match on both live on that row, and
asking for them separately read it twice per resolution. The new
SessionDB.declared_scope_identity() returns the pair and keeps the marker
rules beside is_explicit_fork_child() instead of re-implementing them in the
caller. A SessionDB that does not expose the combined view keeps the
original two-call path, so nothing that predates it changes behaviour --
including the three doubles that certify the fail-closed contract, which are
untouched. TestOneIdentityReadPerResolution pins the single read, the
two-call fallback, the fail-closed degrade and the fork refusal; removing
the fold turns the first of those red.
The third read stays: the generation lives in conversation_generations, a
different table, and cannot be folded into a sessions lookup.
4. _declared_conversation_session() documents the concurrent first-turn race.
Two simultaneous first requests on one declared key can each miss the
lookup, mint a row and both bind, because each row is unkeyed at bind time
and the mismatch guard does not fire. That converges rather than crossing:
both rows carry the same key under the same source, so the lookup returns
the later one for every subsequent reply and the earlier row is an abandoned
transcript, never another conversation's identity.
The same docstring still claimed the generation was durable in
sessions.end_reason and that "nothing here needs a counter". That stopped
being true in 09004753c9, which moved the generation into
conversation_generations precisely because deriving it from prunable session
rows was ABA. Corrected, along with the same stale sentence on
TestConversationBoundariesRotate.
5. conversation_generations rows are now documented as deliberately never
collected, rather than merely uncollected. Dropping one resets that peer to
"no generation", so its next boundary writes 1 again and re-issues a gwk_
scope a retired conversation already used -- the exact ABA the table exists
to close. Worth stating because the repo already carries both patterns a
maintainer would extend: delete_session() cascades to messages, and
gateway_hygiene_state is already swept by session_key.
Upgrade notes, one-time on merge:
- One cold prompt-cache bucket per keyed conversation. Every gateway platform
declares gateway_session_key, so each keyed conversation's affinity scope
moves once from its compression-lineage root session id to the gwk_ hash.
One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
recorded as a keyed row and appears in "Active: N session(s)" where it was
invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
first from the next boundary written, so a conversation that reset before the
upgrade shares its predecessor's scope once. One warm bucket, never a crossed
identity.
Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.
Found in review by @teknium1.
Refs #96811
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.
1. _run_agent raised NameError on every opted-in declared bind. Its worker
finally evaluated `if _declared_selected:`, a local of _handle_responses /
_handle_runs that is neither a parameter nor an enclosing binding here, so
the successful declared-key paths failed at settlement after the agent run.
bind_declared_conversation already IS the gate; the inner name is gone.
2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
_run_sync bound unconditionally, so an explicit body session_id that existed
with an empty session_key was adopted by the header key even though the
header lost precedence. It now carries the same gate.
3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
delete_session() deletes the selected row and bulk prune selects ended rows,
so the aggregate can return a pair it already emitted:
(1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
a retired affinity identity. The backwards-clock shape needs no pruning at
all. The generation now lives in a conversation_generations table keyed by
(source, session_key), advanced by _bump_conversation_generation inside the
same transaction that writes each boundary -- outside prunable session
history, wall-clock-free, and increment-only. end_session() and
promote_to_session_reset() both advance it, and only when they actually
wrote a boundary, so a repeated end cannot double-count.
4. The carrier could be memoized under the wrong source. _agent_source() fell
back to agent.platform before the row landed while persistence uses
_session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
it and never re-resolves once the authoritative row appears, so under an
override both sides of a /new read the platform domain and hashed the same
scope. The pre-row path now uses the persistence resolver itself.
Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.
TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Refs #96811
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
Both blockers from @andrexibiza's review of 28a2d7f0ee.
1. The generation lookup was not in the same identity domain as recovery.
latest_conversation_boundary() selected on session_key alone, while
_declared_conversation_session() is qualified by (source, session_key).
X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
API conversation may legally carry the same key as a Telegram row in one
database -- a /new over there rotated this conversation's gwk_ generation
while recovery correctly refused to cross the same line, moving the affinity
identity out from under a physical identity that had not moved.
The boundary read now takes (session_key, source), and the carrier is
'source|key|generation' rather than 'key|generation' -- keying on the string
alone would also collapse two same-key conversations from different sources
onto one routing key, since this value leaves the process verbatim as
OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
from the agent's own session row, falling back to the platform the row will
be created with before it lands.
2. The declared key's stated lower precedence did not survive settlement. Both
handlers let stored_session_id / an explicit body session_id win, then called
_bind_declared_conversation() unconditionally. record_gateway_session_peer()
does SET session_key = ? across compression ancestors, so a request carrying
conversation A's chain plus header key B silently rebound A to B: A could no
longer be recovered by its own key, and B recovered A's session.
Recording is now gated on the declared key having actually selected or
minted the session, on both paths. Behind that gate the bind itself refuses
to overwrite a row already bound to a different key, so a future caller
cannot reintroduce the same defect by opting in wrongly.
test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.
Refs #96811
Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.
Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
POST /v1/responses and POST /v1/runs parse and authenticate the client's
X-Hermes-Session-Key, pass it downstream for memory scoping, and then mint a
throwaway physical session id anyway whenever the client manages its own
history (no previous_response_id chain to carry one forward).
Every conversation-affinity hint Hermes sends is derived from that physical
id, so all four re-keyed on every single reply: prompt_cache_key on both
OpenAI-wire transports, the OpenRouter and Nous sticky session_id, and xAI's
x-grok-conv-id. The conversation never landed back on a warm prefix.
Fix the identity rather than the four consumers. The declared key resolves to
its live session through find_latest_gateway_session_for_peer -- the same
reset-fenced recovery every native gateway platform already uses -- and the
turn records the row it ended on through record_gateway_session_peer, which
AIAgent._ensure_db_session never did (it knows the key and writes the row
unkeyed, so the mapping the next reply needs did not exist).
Because the lookup is fenced on sessions.end_reason, the generation that must
rotate is already durable: session_reset (/new), session_switch, idle, daily,
suspended and resume_pending_expired all return None, so a new conversation
gets a new id and a cold affinity scope, and a retired generation can never be
resolved again. No counter, no new persisted field, and no new precedence rule
in the cache-scope resolver -- /branch, delegate and tool children keep the
isolation of #79161/#79017 byte for byte.
Precedence is unchanged where it already worked: an explicit body session_id
and the previous_response_id chain both still outrank the declared key, and a
request that declares nothing keeps its per-request id. Recording is opt-in
(bind_declared_conversation), so no other _run_agent caller's rows change.
Refs #96811
(cherry picked from commit e7c83dddf36784d1012bf483240ebc7f6b2ef9aa)
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
- Clear the accidental end stamp on resurrection (at the lineage tip):
a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
archive auto-resurrect on the next lookup — the user could never retire
the canonical chat. Test pins the resurrect -> deliberate-archive ->
stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
compressed lineage carries end_reason='compression', so tip-stamped
accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
dm resolution) filtered archived rows out via list_sessions_rich and
still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.
Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.
Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.
Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.
Fixes#91583 (defect 2). Repro and live validation by @kubaboski.
The prefix is an address: the caller is naming WHICH agent the request is
for. With gateway.multiplex_profiles off, _resolve_request_profile ignored
the prefix entirely — "don't 404 a would-be valid route" — so a request
explicitly addressed to one agent was silently answered by a different one.
Observed live (Aug 2026): `hermes peer dm mini/researcher` was answered by
the mini's DEFAULT agent with no error on either side, because that host
runs one LaunchDaemon per profile and only the default daemon hosted an
api_server. A wrong-agent answer is strictly worse than an error: the
sender believes the addressee got the message.
With multiplexing off the process serves exactly one profile, so the prefix
is honored when it names that profile (peers address single-profile daemons
this way without knowing the host's topology — get_active_profile_name() is
the same identity the file already uses for model resolution) and rejected
otherwise through the existing _PROFILE_REJECTED path (404). A process that
cannot resolve its own identity rejects too: if it cannot prove who it is,
it must not answer as anyone.
Unprefixed requests are untouched, and multiplexed hosts are untouched —
the change is confined to the prefix-present, multiplexing-off branch that
previously discarded the caller's addressing.
Bot Mode always hides canonical 'Bot Chat' sessions, but _find_bot_chat's
GET /api/sessions listing used the default include_hidden=False path, so
the existing hidden row was invisible, _ensure_bot_chat tried to create a
duplicate, and the peer DB's UNIQUE(title) guard rejected it — DM failed.
- api_server: GET /api/sessions now accepts an exact-title lookup
(?title=...) and honors include_hidden=1 ONLY alongside a title filter,
so canonical hidden rows resolve without exposing a blanket hidden
listing on the client surface. The title needle is pushed into SQL
(search_query) so old hidden rows outside the recency window are found.
- peer dm client: _find_bot_chat sends title + include_hidden=1; older
peers ignore the unknown params and degrade to today's behavior.
- Clear diagnosable error on the older-peer duplicate-create rejection,
naming the hidden canonical chat and the PATCH hidden:false workaround.
- Unit tests (hidden resolution, no duplicate create, older-peer error,
older-peer visible fallback) + real-gateway E2E over a real state.db.
Root-cause analysis and regression recipe by @kubaboski in #91583.
Fixes#91583
_turn_transcript_messages pre-classified every message with
_is_compressed_summary_message (full content flatten + prefix scan), then
_message_response re-ran the same classifier inside its projection --
2x per non-summary row, 3x per summary row on every run.completed emit.
The outer guard was redundant: _message_response already yields
display_kind hidden for pure handoffs. One projection call per row now.
Surfaced by the post-merge simplify re-review of #91517/#91535.
The frame handler returns reply dicts (heartbeat/detach acks) that the WS
reader loop sends back; the -> None annotation was the only new ty
diagnostic vs origin/main.
Addresses both merge blockers from @andrexibiza's review of #85351:
1. HTTP-uploaded artifacts could never be consumed by broker dispatch:
artifact_scope_key hashed (principal, session, family), the HTTP routes
store with an EMPTY session (API-key auth has no server session) while
broker validation carries a session-bearing ControllerScope — every
real upload->dispatch journey died with ArtifactScopeMismatch
(reproduced before fixing). Canonical ownership is now
principal/transport-family (documented in the scope-key docstring);
ids stay unguessable server-minted 32-hex and downloads one-shot.
New composition regression: HTTP-shape upload -> registered controller
scope -> broker artifact dispatch, mutation-checked (re-adding session
to the key makes it fail).
2. The 'profile-scoped' artifact store was first-profile-wins process
state: one adapter-level singleton pinned profile B to profile A's
physical root on multiplex listeners (same frozen-handle class as
#88734). Stores are now cached by resolved profile, and the broker
selects the store from the controller scope's profile_id (default-slot
fallback preserves single-profile/test behaviour). New A/B multiplex
regression proves distinct physical roots regardless of touch order.
Also documents the advertised ticket_expires_at as best-effort wall clock
(broker enforces expiry monotonically) per review feedback.
- Rename the broker's TicketInvalid to ControllerTicketInvalid: the same
exception name already exists in hermes_cli/dashboard_auth/ws_tickets.py
and BOTH are caught in the same WS auth flow this feature touches — two
unrelated same-named exception types in one blast radius invited a wrong
except clause.
- Import the 'server-internal' sentinel identity from its canonical
definition (ws_tickets.INTERNAL_USER_ID/INTERNAL_PROVIDER) instead of
re-declaring the strings; drift would have silently broken the
internal-peer exclusion in _is_authenticated_identity.
Surfaced during review of PR #85351.
attach/disconnect/detach acquire a per-controller threading.Lock that a
worker-thread dispatch can hold for up to 10s while blocking on the event
loop to transmit its command frame (run_coroutine_threadsafe +
result(timeout=10)). Acquiring that lock synchronously from loop context
(controller WS finally, frame handler, gateway WS teardown) could park the
ENTIRE gateway event loop behind the send bridge — a deterministic
multi-second global stall whenever controller teardown raced an in-flight
command. All loop-context broker calls now go through asyncio.to_thread,
matching the existing offload pattern for _close_sessions_for_transport.
Surfaced during review of PR #85351.
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.
Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.
Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
The 7-key internal-fields tuple was inlined twice (agent/compaction_display.py
and _project_client_message); a drift between the copies would silently leak
one internal field class through the API projection. Surfaced during review
of PR #85442.
Project client-visible session messages through the canonical compaction classifier. Hide standalone handoffs, unwrap merged carriers to their authentic prior-tail content, strip inherited internal fields, and keep model-facing recovery history unchanged.
_request_reasoning_config() whitelisted none..xhigh, so a client sending
max or ultra (valid /reasoning + config.yaml levels) fell through to the
default effort with no error. The server now accepts the full internal
ladder (hermes_constants.VALID_REASONING_EFFORTS); per-provider wire
clamping happens downstream via agent.reasoning_effort, same as every
other entry surface. Salvages the api_server hunk of #78216 (credit
@snowzlmbot); the un-clamping half of that PR was rejected separately.
The #35314 empty-dispatch recovery cache accepts any previously
resolved model string, including the advertised virtual model
(hermes-agent by default) — which is never a dispatchable model.
Guard both directions: never store the alias in _last_resolved_model,
and never serve a cached value equal to it. The alias now survives at
most one turn and cannot propagate across turns through the cache.
Class bug behind #79101; complements the session-row symptom fixes
(#72739, #79102).
Salvaged from #72739 (net diff onto current main; the read-side legacy-row
guard now flows through the shared _stored_session_model() helper that
landed in #88751, keeping one resolver for both chat sites).
POST /api/sessions persisted the advertised virtual alias (hermes-agent)
whenever the request omitted model or echoed the alias back; later turns
replayed it upstream as a real model id and every turn on the session
failed. Null the alias in _session_runtime_request_from_body() — shared by
session create, chat, chat-stream, and model-lock — and stop the
create-handler's raw-body fallback from bypassing that normalization.
Two bugs found by a REAL two-gateway live test (two isolated HERMES_HOMEs,
bravo running the api_server platform, alpha's agent autonomously running
`hermes peer dm` from its Bot Chat protocol; reply relayed correctly and
persisted in bravo's canonical Bot Chat):
1. peer dm parsed the session-create response flat, but api_server wraps
the row: {"object": "hermes.session", "session": {...}} — every first DM
to a fresh peer failed with "Peer did not return a session id" (and the
orphaned Bot Chat then 400'd retries with duplicate-title). Parse the
wrapped shape; the test fake now mirrors the real response shape so this
class can't pass green again.
2. api_server: a session created with no model persists the advertised
virtual model ("hermes-agent") on the row; session chat then replayed it
as a REAL model id and the provider 400'd ("hermes-agent is not a valid
model ID"). _request_agent_overrides already filters the virtual model
for per-request bodies — apply the same filter to the stored session
model at both chat sites (sync + stream), so it means "gateway default"
exactly like the request-body path.
Live E2E transcript (bravo's Bot Chat, via /api/sessions/{id}/messages):
user: Message from 🤖 alpha (@alpha): What is your callsign?
assistant: CALLSIGN-BRAVO-7
peer cmd unit suite 10/10 with the corrected fake.
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)
Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.
- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
incl. the compression-lineage recursive CTE); list_sessions_rich gains
include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
from every listing path (and the REST sidebar endpoints inherit it with no
change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
set_session_hidden; _session_response exposes it.
Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).
* fix: teach lost-and-found recovery about the 55-column sessions layout
Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
The session chat stream registered its wrapper task in _active_run_tasks,
but that turn is already counted by active_agent_work_count() via
_inflight_agent_runs (_run_agent) — the drain saw 2 for one turn
(test_session_chat_sse_turn_is_interrupted). Keep only the agent-ref
registration; run-scoped steer control doesn't need the task entry.
Adds POST /v1/runs/{run_id}/steer and bridges Browser-Extension/WebUI
session chat streams into the active run registry so live runs on those
surfaces are steerable too.
- steer accepted only while run status is exactly 'running'; stop/stopping/
terminal states return 409 run_not_accepting_steer even while cooperative
shutdown retains the agent reference
- session SSE disconnect/cancellation interrupts and drains the executor-
backed run instead of cancelling only the async wrapper; control refs stay
registered until the turn actually exits
- undelivered steer text (accepted after the final response) is preserved as
pending_steer on the terminal run.completed event/status so clients can
replay it as the next user turn instead of losing it
- docs for the endpoint, next-tool-boundary delivery, acceptance-vs-delivery
semantics
Salvaged from PR #54466 by @abundantbeing.
* fix(gateway): pass live adapters to cron fire webhook's fire_due
The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).
Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.
Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).
* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)
The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.
Restore the invariant that the GATEWAY owns cron execution:
- Dashboard route: after verifying the NAS JWT and resolving the job's
profile, FORWARD the fire to the gateway api_server's own
/api/cron/fire on loopback, NAS bearer preserved (the gateway
re-verifies the JWT — defense in depth, no new trust link), and pass
the gateway's response through. Gateway unreachable → 503 so NAS
retries per the Chronos contract (non-2xx = retryable; the store CAS
de-dupes the eventual double fire). Deliberately NO local-execution
fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
per target profile (config.yaml extra.port → API_SERVER_PORT from
process env or the profile's .env → 8642), with /p/<profile>/ prefix
routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
first boot when absent (never overwrites an operator value), so the
loopback api_server passes its startup guard on hosted images. The
fire route itself is NAS-JWT-authed; the key gates the rest of the
api_server surface. The listener binds 127.0.0.1 by default and the
Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
topology and the 503-retry semantics.
Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.
* fix(cron): read the profile api_server port via the canonical config loader
CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).
Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.
* fix(gateway): only messaging platforms count for the scale-to-zero arm gate
The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).
The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
e81d18dfb collapsed six per-surface copies of reasoning resolution onto
resolve_reasoning_config() and, in its own words, "fixes the gateway
resolving reasoning against config model.default instead of the session's
effective model". It did not touch gateway/platforms/api_server.py, which
kept that defect.
_create_agent() called GatewayRunner._load_reasoning_config() with no
model on its first line — before the model precedence chain (browser lock
-> session /model -> session row -> route -> per-request -> defaults) has
run. Per-model agent.reasoning_overrides therefore keyed off model.default
on the one surface where every request names its own model: a request for
a model with an override silently got the global effort instead.
Resolve after the chain settles, so the override follows the model the
request actually runs. An explicit per-request reasoning parameter still
takes precedence over config.
The existing test stub for _load_reasoning_config took no arguments (it
mirrored the old call); it now matches the real signature, as the sibling
stub in the same file already did.
The non-streaming /v1/responses path built function_call and
function_call_output output items with no status field (and no item id),
while the SSE streaming path correctly emits status in_progress ->
completed. Spec-strict OpenAI clients reading the non-streaming output
array could interpret the status-less function_call items as pending
calls the CLIENT must execute — but these tools were already executed
server-side by the Hermes agent and are replayed for structured tool UI
only. Reported by a community user whose GPT-5.6 client concluded 'a
server should not tell an OpenAI client to execute a tool the server
already executed itself'.
- _extract_output_items now stamps status: completed and spec-shaped
item ids (fc_/fco_) on replayed items, matching the streaming path
- test updated to pin status + id shape
- docs example updated + explicit note that output tool calls are
replayed, never pending
The shutdown drain ACCOUNTS for API-server work but never INTERRUPTS it.
`_drain_active_agents()` folds `_active_api_run_count()` into both its wait
loop and its `timed_out` verdict, while `_interrupt_running_agents()` iterates
`self._running_agents` only -- a dict no API turn ever enters, because the
API server owns its own agent lifecycle. `gateway/run.py` states the gap
against itself: "API-server / desk sessions have the same structural gap
(#63529)."
The user-visible result is that every gateway restart with a live API or
desktop turn burns the full drain timeout and then runs
`_kill_tool_subprocesses("post-interrupt")`, which amputates the turn's tool
subprocesses with no cooperative interrupt and no resume marker.
There are seven API agent-entry points. Six funnel through `_run_agent()`
(both session-chat routes, and `/v1/chat/completions` + `/v1/responses` in
streaming and non-streaming form) and are counted by `_inflight_agent_runs`;
the seventh, `/v1/runs`, runs its own lifecycle and is counted through
`_active_run_tasks`. None of the six has a run_id, so the run_id-keyed
`_active_run_agents` cannot reach them, and only two pass `agent_ref` -- which
lands in a caller-local list, not a registry.
So register once at the single unconditional creation site inside
`_run_agent`, beside the existing `_publish_turn_process_ownership()` call,
and unregister in the same `finally` that already clears it. That one
symmetric pair covers all six callers. The registry is adapter-owned and
keyed by object identity, kept separate from `_active_run_agents` because
that dict is run_id-keyed and scoped to the public `/v1/runs` stop API.
`interrupt_active_runs()` then walks both registries, deduped by identity, so
the interrupt set matches the set the drain waits on. The settle window after
the interrupt now polls API work as well: the interrupt is cooperative, and
without this the window closes the instant `_running_agents` is empty -- which
it always is for API turns -- and the tool kill lands on a turn that was asked
to stop microseconds earlier.
Pinned was capped at half the viewport by its own nested scroller, so past
roughly a dozen pins the rest were reachable only by scrolling inside a
scroller — a pin you have to go hunting for isn't doing its job.
Drop the cap and let the section grow into the sidebar's existing scroll,
and stop virtualizing Pinned: virtualization needs a bounded viewport to
measure against, which is exactly what's being removed. No count badge, no
"show more" — pin as many as you want and they all render.
Also back-fill pins on the API-server list route, which was the one list
path still windowing purely on recency.
PATCH /api/sessions/{id} only accepted title and end_reason, so the
`pinned` flag the desktop sends was rejected as an unsupported field —
and the client swallows that error. Pins lived in one app's localStorage
and never reached state.db, which also meant the server-side auto-archive
sweep was free to hide the chats a pin exists to keep.
Accept pinned and archived as booleans, route them to the SessionDB
setters that already existed, and include both in the serialized session
so clients can reconcile against server truth.
The session event stream (api_server.py:~2236) was the one genuinely
unicode-distinct SSE writer — json.dumps(payload, ensure_ascii=False) +
.encode('utf-8'). Every other writer uses plain json.dumps. Route it
through _sse_frame(..., ensure_ascii=False) so _sse_frame is now the single
source of truth for ALL SSE frame serialization in the module (chat-
completion, responses._write_event, /v1/runs, and the session stream).
Byte-identical for non-ASCII payloads: verified against the historical
inline encoder (raw bytes preserved). The ensure_ascii=False path is now
exercised by test_sse_frame_ensure_ascii_false_reproduces_session_event_stream.
_extend _sse_frame with an explicit ensure_ascii param (default True,
byte-identical to a bare json.dumps) and route the two sibling writers
through it: _write_sse_responses._write_event and the /v1/runs event
stream. This completes the dedup PR #65009 — previously only the five
_write_sse_chat_completion sites used the helper, leaving the other two
writers on inline json.dumps with no shared shape.
No behavior change: every writer's emitted bytes are unchanged (verified
byte-for-byte, including non-ASCII payloads where the default
ensure_ascii=True matches the original inline encoders). The ensure_ascii
option is exposed so a future writer can opt into raw non-ASCII bytes
without fractalizing the format again.
Adds tests/gateway/test_sse_frame.py asserting the byte-contract
invariant between _sse_frame and the historical inline encoders.