Commit Graph

235 Commits

Author SHA1 Message Date
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium 2031c819fe simplify(compat): gateway — re-land 92d0bd0d73 (reverted by stale-index commit b818085298)
Re-applies the gateway compat removal byte-for-byte; see 92d0bd0d73 for the
full inventory (30 re-exports/aliases + 2 shim modules dropped, 3 shim-only
names re-removed, 24 callers + 34 test files repointed). No new changes.
2026-09-03 13:12:50 -07:00
Teknium b818085298 simplify(compat): doctor/status — drop 13 re-exports + the doctor_* globals() facade (97 names), repoint 6 callers / 13 tests 2026-09-03 13:10:52 -07:00
Teknium 92d0bd0d73 simplify(compat): gateway — drop 30 re-exports/aliases + 2 shim modules, re-remove 3 shim-only names, repoint 24 callers + 34 test files
Per COMPAT_REMOVAL.md (internal import paths are not a stable API):

Re-exports removed
- gateway/run.py: atomic_json_write, load_dotenv, resolve_delivery_transport,
  TurnRunner, merge_pending_message_event, _arm_loop_floor_timer,
  start_loop_liveness_watchdog, DEFAULT_GATEWAY_POST_INTERRUPT_GRACE_TIMEOUT,
  _UNSET (9) — run_* mixins and tests now import from the defining module
  (gateway.delivery / gateway.run_turn_runner / gateway.platforms.base /
  gateway.shutdown_watchdog / gateway.restart / utils).
- gateway/session.py: SessionResetPolicy, normalize_whatsapp_identifier,
  TranscriptReadError, auto_continue_freshness_window (+ "_now & co." noqa
  facade) — gateway/__init__ takes SessionResetPolicy from .config; callers
  take TranscriptReadError from gateway.session_transcript.
- gateway/kanban_watchers.py: _wake_scope_id + "tests import via origin" noqa
  facade; tests import from kanban_watchers_common / _notifier.
- gateway/slash_commands.py: _model_switch_skew_guard, HISTORY_UNREADABLE.
- gateway/stream_consumer.py: escape_code_fences_for_display.
- gateway/platforms/api_server.py: "re-exported" RunIdempotencyStore comment;
  tui_gateway + tests import gateway.platforms.api_server_run_idempotency.
- gateway/platforms/__init__.py: PEP 562 __getattr__/__dir__ lazy QQAdapter /
  YuanbaoAdapter facade (no in-tree importer).
- gateway/startup_watchdog.py: whole re-export shim module deleted; the three
  in-tree callers import hermes_startup_watchdog directly.

Aliases removed
- gateway/platforms/signal.py: SignalAdapter._markdown_to_signal.
- gateway/shutdown_forensics.py: _parse_systemd_duration_to_us.
- gateway/platforms/yuanbao.py: OutboundManager.start_slow_notifier /
  cancel_slow_notifier / get_chat_lock / _chat_locks / CHAT_DICT_MAX_SIZE
  delegates; module-level get_active_adapter / send_yuanbao_direct;
  MarkdownProcessor has_unclosed_fence / ends_with_table_row /
  split_at_paragraph_boundary static pass-throughs (chunk_markdown_text stays —
  it carries yuanbao's chunking policy). tools/send_message_senders +
  tools/yuanbao_tools call YuanbaoAdapter.get_active() / sender.send_direct().

Shim-only names re-removed (earlier review-fix round 96c104c903)
- CapabilityDescriptor.from_platform_entry (+ tests/gateway/relay/test_descriptor_from_entry.py)
- SessionTurnLeaseRegistry.__len__ (+ its test)
- is_relay_media_url KEPT: download() uses it, real internal helper.

Tests that pinned a shim (startup_watchdog re-export identity, api_server
RunIdempotencyStore identity, signal wrapper parity) are dropped; tests that
pinned live behavior are repointed at the implementation.

Note: the gateway/run.py hunk of this change was swept into 00a3cfe5c9 by a
concurrent commit on the shared worktree; this commit carries the rest.

Verified in an isolated worktree (HEAD + this change only): ruff clean,
import-smoke of all 33 touched modules under a fresh HERMES_HOME,
605 passed / 1 skipped across the 39 covering test files.
2026-09-03 13:10:49 -07:00
Teknium 474eed838e review-fix(suppress-audit): patch_parser/skills_sync/update_cmd_fleet/api_server — restore BASE exception semantics 2026-09-03 09:56:39 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 58e54136a9 refactor(gateway/platforms): api_server — shared _call_verifier (sync/async off-loop), default-then-suppress for best-effort reads 2026-09-02 23:06:21 -07:00
Teknium 46bd3e9031 refactor(gateway/platforms): api_server — shared _session_headers builder, tuple unpacks in session chat handlers 2026-09-02 23:02:05 -07:00
Teknium 7035c2c82c refactor(gateway/platforms): api_server — reuse _shared.coerce_port, _requested_ids/_apply_provider_runtime helpers, inline single-use cron checks 2026-09-02 22:52:56 -07:00
Teknium 1f93b1a84f refactor(gateway/platforms): api_server — compact docstrings/comments (WHY kept), generated run/room-grant delegators 2026-09-02 22:41:52 -07:00
Teknium e1b3c757d9 refactor(gateway/platforms): api_server — single AIOHTTP_AVAILABLE middleware block, packed JSON literals, collapsed guard ladders 2026-09-02 22:30:27 -07:00
Teknium 73a91014df refactor(gateway/platforms): api_server second pass — contextlib.suppress for swallow-only excepts, shared cron _job_response tail, ternary/dict-comp collapses 2026-09-02 21:53:12 -07:00
Teknium 118581167b refactor(gateway/platforms): api_server import/constant packing, one-expression predicates, direct cron callables 2026-09-02 19:51:23 -07:00
Teknium 5e28d524dc refactor(gateway/platforms): extract _SessionEventQueue from _handle_session_chat_stream (148 -> 124 LOC) 2026-09-02 19:47:30 -07:00
Teknium b2e55ef3fd refactor(gateway/platforms): AST-neutral closer/bracket hug pass on api_server (-18 LOC) 2026-09-02 19:43:15 -07:00
Teknium d84dd007af refactor(gateway/platforms): api_server _require_auth decorator (14 handlers), reservation release via one helper 2026-09-02 19:42:25 -07:00
Teknium 957387cc9a refactor(gateway/platforms): api_server sessions/chat-stream/cron/_run_agent — shared run_kwargs, _finish_turn_result, _track_background_task, _session_db_unavailable (4388 -> 4216 LOC) 2026-09-02 19:36:13 -07:00
Teknium 836327e3b1 refactor(gateway/platforms): api_server _create_agent/health/models/browser-control/artifact handlers compacted (4592 -> 4388 LOC) 2026-09-02 19:21:15 -07:00
Teknium fc3d4894c7 refactor(gateway/platforms): api_server runtime selection — lift _resolve_provider_runtime/_recover_or_record_model, _auth_failed_response, compact infra docstrings (4826 -> 4592 LOC) 2026-09-02 19:13:49 -07:00
Teknium 4a3b4b691e refactor(gateway/platforms): api_server module helpers — shared text/list caps, _normalize_image_part, compact docstrings (5055 -> 4826 LOC) 2026-09-02 18:57:49 -07:00
Teknium bfc48f8089 refactor(gateway/platforms): openai routes — _ResponsesStream state object, shared finish/route/spawn helpers (1409 -> 1081 LOC) 2026-09-02 18:41:00 -07:00
Teknium a109980a3e refactor(gateway/platforms): AST-neutral layout compaction of api_server + openai routes (7019 -> 6459 LOC) 2026-09-02 18:25:26 -07:00
Teknium 581d97e545 refactor(adapters/api_server): 10589->8871; extract OpenAI-compatible routes into api_server_openai_routes.py, dedupe SSE/response builders, drop dead idempotency acknowledge path 2026-09-02 14:06:32 -07:00
Teknium 7a86397a46 fix(api_server): fail closed on unstamped runs; claim session-chat-stream run owner (#93689)
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:

- `_request_owns_run` no longer admits run state that exists without an
  owner stamp. Under gateway.multiplex_profiles every served profile holds
  a valid key, so the "backward compatibility" branch made the boundary
  allow-all whenever provenance was missing. Unstamped state now fails
  closed; only an in-memory owner match or a durable idempotency record
  under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
  inside the request's profile scope, so its run is confined to the
  creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
  (`_release_run_owner_if_forgotten`) and runs at every retirement point
  (task finally, SSE stream close, both sweep loops, chat-stream finally),
  not only the terminal-status sweep — no stranded entries, no stateful id
  ever left unowned.

Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).

Fixes #93689
Fixes #90415
Supersedes #93747, #93704, #92822

Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
JonthanaHanh 5360886f54 fix(gateway): exclude Ollama Cloud from GLM truncation detection; propagate partial flag (#72316)
Two compounding bugs that cause WebUI to discard or misrender agent
responses when using GLM models on Ollama Cloud:

1. _is_ollama_glm_backend() matched "ollama" in base URL, which
   included Ollama Cloud (ollama.com). The hosted service correctly
   reports finish_reason and is not affected by the local Ollama
   stop-reason bug.  Exclude "ollama.com" before the substring check.

2. _handle_session_chat_stream() hardcoded "partial": False in the
   assistant.completed SSE event instead of reading result.get("partial").
   The WebUI could not detect truncation and rendered partial responses
   incorrectly (showing only the continuation instead of the full text).
   Read the partial flag from the agent result, matching the pattern
   used by other SSE paths in the same file.

Fixes #72316
2026-09-01 23:27:10 -07:00
joaomarcos 6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos d63e5d8a10 fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee.

1. The generation lookup was not in the same identity domain as recovery.
   latest_conversation_boundary() selected on session_key alone, while
   _declared_conversation_session() is qualified by (source, session_key).
   X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
   API conversation may legally carry the same key as a Telegram row in one
   database -- a /new over there rotated this conversation's gwk_ generation
   while recovery correctly refused to cross the same line, moving the affinity
   identity out from under a physical identity that had not moved.

   The boundary read now takes (session_key, source), and the carrier is
   'source|key|generation' rather than 'key|generation' -- keying on the string
   alone would also collapse two same-key conversations from different sources
   onto one routing key, since this value leaves the process verbatim as
   OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
   from the agent's own session row, falling back to the platform the row will
   be created with before it lands.

2. The declared key's stated lower precedence did not survive settlement. Both
   handlers let stored_session_id / an explicit body session_id win, then called
   _bind_declared_conversation() unconditionally. record_gateway_session_peer()
   does SET session_key = ? across compression ancestors, so a request carrying
   conversation A's chain plus header key B silently rebound A to B: A could no
   longer be recovered by its own key, and B recovered A's session.

   Recording is now gated on the declared key having actually selected or
   minted the session, on both paths. Behind that gate the bind itself refuses
   to overwrite a row already bound to a different key, so a future caller
   cannot reintroduce the same defect by opting in wrongly.

test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.

Refs #96811

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos 3739cf3b86 fix(api): resolve the declared conversation instead of minting a session per request
POST /v1/responses and POST /v1/runs parse and authenticate the client's
X-Hermes-Session-Key, pass it downstream for memory scoping, and then mint a
throwaway physical session id anyway whenever the client manages its own
history (no previous_response_id chain to carry one forward).

Every conversation-affinity hint Hermes sends is derived from that physical
id, so all four re-keyed on every single reply: prompt_cache_key on both
OpenAI-wire transports, the OpenRouter and Nous sticky session_id, and xAI's
x-grok-conv-id. The conversation never landed back on a warm prefix.

Fix the identity rather than the four consumers. The declared key resolves to
its live session through find_latest_gateway_session_for_peer -- the same
reset-fenced recovery every native gateway platform already uses -- and the
turn records the row it ended on through record_gateway_session_peer, which
AIAgent._ensure_db_session never did (it knows the key and writes the row
unkeyed, so the mapping the next reply needs did not exist).

Because the lookup is fenced on sessions.end_reason, the generation that must
rotate is already durable: session_reset (/new), session_switch, idle, daily,
suspended and resume_pending_expired all return None, so a new conversation
gets a new id and a cold affinity scope, and a retired generation can never be
resolved again. No counter, no new persisted field, and no new precedence rule
in the cache-scope resolver -- /branch, delegate and tool children keep the
isolation of #79161/#79017 byte for byte.

Precedence is unchanged where it already worked: an explicit body session_id
and the previous_response_id chain both still outrank the declared key, and a
request that declares nothing keeps its per-request id. Recording is opt-in
(bind_declared_conversation), so no other _run_agent caller's rows change.

Refs #96811

(cherry picked from commit e7c83dddf36784d1012bf483240ebc7f6b2ef9aa)
2026-09-01 02:14:35 -07:00
David Dudok de Wit e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
Teknium 31579f781e fix(cron): transient run prompt survives the relay-fronted gateway forward
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
2026-08-27 20:53:02 -07:00
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
loulanyue 4e8419dadb fix(gateway): honor approval scope capabilities 2026-08-23 17:45:47 -05:00
kshitij 4865194772 fix(bot-mode): review follow-ups for recoverable-archive resurrection
- Clear the accidental end stamp on resurrection (at the lineage tip):
  a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
  archive auto-resurrect on the next lookup — the user could never retire
  the canonical chat. Test pins the resurrect -> deliberate-archive ->
  stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
  compressed lineage carries end_reason='compression', so tip-stamped
  accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
  dm resolution) filtered archived rows out via list_sessions_rich and
  still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
  hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
  into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
  by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
2026-08-24 03:27:11 +05:30
Teknium 6cb1085d3d fix(gateway): /p/<profile>/ on a non-multiplex gateway fails closed instead of serving the owner profile
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.

Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.

Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.

Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.

Fixes #91583 (defect 2). Repro and live validation by @kubaboski.
2026-08-23 03:57:47 -07:00
John Paul Soliva 265bdcac82 fix(api-server): a /p/<profile> prefix on a non-multiplexed gateway fails closed instead of misdelivering
The prefix is an address: the caller is naming WHICH agent the request is
for. With gateway.multiplex_profiles off, _resolve_request_profile ignored
the prefix entirely — "don't 404 a would-be valid route" — so a request
explicitly addressed to one agent was silently answered by a different one.
Observed live (Aug 2026): `hermes peer dm mini/researcher` was answered by
the mini's DEFAULT agent with no error on either side, because that host
runs one LaunchDaemon per profile and only the default daemon hosted an
api_server. A wrong-agent answer is strictly worse than an error: the
sender believes the addressee got the message.

With multiplexing off the process serves exactly one profile, so the prefix
is honored when it names that profile (peers address single-profile daemons
this way without knowing the host's topology — get_active_profile_name() is
the same identity the file already uses for model resolution) and rejected
otherwise through the existing _PROFILE_REJECTED path (404). A process that
cannot resolve its own identity rejects too: if it cannot prove who it is,
it must not answer as anyone.

Unprefixed requests are untouched, and multiplexed hosts are untouched —
the change is confined to the prefix-present, multiplexing-off branch that
previously discarded the caller's addressing.
2026-08-23 03:57:47 -07:00
Teknium 231e613d3d fix(peer): resolve hidden canonical Bot Chats in hermes peer dm
Bot Mode always hides canonical 'Bot Chat' sessions, but _find_bot_chat's
GET /api/sessions listing used the default include_hidden=False path, so
the existing hidden row was invisible, _ensure_bot_chat tried to create a
duplicate, and the peer DB's UNIQUE(title) guard rejected it — DM failed.

- api_server: GET /api/sessions now accepts an exact-title lookup
  (?title=...) and honors include_hidden=1 ONLY alongside a title filter,
  so canonical hidden rows resolve without exposing a blanket hidden
  listing on the client surface. The title needle is pushed into SQL
  (search_query) so old hidden rows outside the recency window are found.
- peer dm client: _find_bot_chat sends title + include_hidden=1; older
  peers ignore the unknown params and degrade to today's behavior.
- Clear diagnosable error on the older-peer duplicate-create rejection,
  naming the hidden canonical chat and the PATCH hidden:false workaround.
- Unit tests (hidden resolution, no duplicate create, older-peer error,
  older-peer visible fallback) + real-gateway E2E over a real state.db.

Root-cause analysis and regression recipe by @kubaboski in #91583.

Fixes #91583
2026-08-23 03:56:28 -07:00
kshitijk4poor 0b8a848754 perf(api): classify compaction rows once per message in run.completed transcript
_turn_transcript_messages pre-classified every message with
_is_compressed_summary_message (full content flatten + prefix scan), then
_message_response re-ran the same classifier inside its projection --
2x per non-summary row, 3x per summary row on every run.completed emit.
The outer guard was redundant: _message_response already yields
display_kind hidden for pure handoffs. One projection call per row now.
Surfaced by the post-merge simplify re-review of #91517/#91535.
2026-08-21 22:54:13 +05:30
kshitijk4poor 23a64a97ec fix(api): correct _handle_browser_control_frame return annotation
The frame handler returns reply dicts (heartbeat/detach acks) that the WS
reader loop sends back; the -> None annotation was the only new ty
diagnostic vs origin/main.
2026-08-21 22:33:45 +05:30
kshitijk4poor 847289864d fix(browser): make the artifact boundary compose end-to-end and scope stores per profile
Addresses both merge blockers from @andrexibiza's review of #85351:

1. HTTP-uploaded artifacts could never be consumed by broker dispatch:
   artifact_scope_key hashed (principal, session, family), the HTTP routes
   store with an EMPTY session (API-key auth has no server session) while
   broker validation carries a session-bearing ControllerScope — every
   real upload->dispatch journey died with ArtifactScopeMismatch
   (reproduced before fixing). Canonical ownership is now
   principal/transport-family (documented in the scope-key docstring);
   ids stay unguessable server-minted 32-hex and downloads one-shot.
   New composition regression: HTTP-shape upload -> registered controller
   scope -> broker artifact dispatch, mutation-checked (re-adding session
   to the key makes it fail).

2. The 'profile-scoped' artifact store was first-profile-wins process
   state: one adapter-level singleton pinned profile B to profile A's
   physical root on multiplex listeners (same frozen-handle class as
   #88734). Stores are now cached by resolved profile, and the broker
   selects the store from the controller scope's profile_id (default-slot
   fallback preserves single-profile/test behaviour). New A/B multiplex
   regression proves distinct physical roots regardless of touch order.

Also documents the advertised ticket_expires_at as best-effort wall clock
(broker enforces expiry monotonically) per review feedback.
2026-08-21 22:33:45 +05:30
kshitijk4poor 13f209d4fd refactor(browser): dedupe auth-flow names and sentinel identity
- Rename the broker's TicketInvalid to ControllerTicketInvalid: the same
  exception name already exists in hermes_cli/dashboard_auth/ws_tickets.py
  and BOTH are caught in the same WS auth flow this feature touches — two
  unrelated same-named exception types in one blast radius invited a wrong
  except clause.
- Import the 'server-internal' sentinel identity from its canonical
  definition (ws_tickets.INTERNAL_USER_ID/INTERNAL_PROVIDER) instead of
  re-declaring the strings; drift would have silently broken the
  internal-peer exclusion in _is_authenticated_identity.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
kshitijk4poor 45078eb941 fix(browser): offload broker lock acquisition off the event loop
attach/disconnect/detach acquire a per-controller threading.Lock that a
worker-thread dispatch can hold for up to 10s while blocking on the event
loop to transmit its command frame (run_coroutine_threadsafe +
result(timeout=10)). Acquiring that lock synchronously from loop context
(controller WS finally, frame handler, gateway WS teardown) could park the
ENTIRE gateway event loop behind the send bridge — a deterministic
multi-second global stall whenever controller teardown raced an in-flight
command. All loop-context broker calls now go through asyncio.to_thread,
matching the existing offload pattern for _close_sessions_for_transport.

Surfaced during review of PR #85351.
2026-08-21 22:33:45 +05:30
abundantbeing 1977c3d2eb feat(browser): add scoped artifact endpoints, broker permission gates, and companion journal 2026-08-21 22:33:45 +05:30
abundantbeing 095a1d078c fix(browser): preserve controller work across reconnects
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.

Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
2026-08-21 22:33:45 +05:30
abundantbeing d524cc9a16 fix(browser): harden extension controller routing
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.

Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
2026-08-21 22:33:45 +05:30
abundantbeing c9fd5223f6 feat(browser): enable extension controller actions 2026-08-21 22:33:45 +05:30
abundantbeing 5df1d0e113 feat(browser): add authenticated control broker 2026-08-21 22:33:45 +05:30
kshitijk4poor b2c4f1f376 refactor(api): reuse _COMPACTION_INTERNAL_FIELDS from compaction_display
The 7-key internal-fields tuple was inlined twice (agent/compaction_display.py
and _project_client_message); a drift between the copies would silently leak
one internal field class through the API projection. Surfaced during review
of PR #85442.
2026-08-21 21:53:11 +05:30
abundantbeing a2a23a8f7e fix(clients): hide compaction carriers across surfaces 2026-08-21 21:53:11 +05:30
abundantbeing 97e32d49ac fix(api): hide compaction scaffolding from clients
Project client-visible session messages through the canonical compaction classifier. Hide standalone handoffs, unwrap merged carriers to their authentic prior-tail content, strip inherited internal fields, and keep model-facing recovery history unchanged.
2026-08-21 21:53:11 +05:30