Commit Graph

249 Commits

Author SHA1 Message Date
xielevi 4ba717df12 fix(profiles): migrate session/routing identity on profile rename
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.

The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:

- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
  tables, matching the agent:<name>: namespace by exact prefix (substr, not
  LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
  the profile inside routing/origin JSON, and REFUSING on a target collision
  (routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
  (keys + origin.profile) then persist — the half a DB write cannot reach.
  Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
  params only to handlers that declare them, bare handlers unchanged) so a
  live gateway rekeys its in-memory copy AND both durable stores (routing home
  + the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
  does NOT fall back to a racing CLI-side write: it prints a warning telling
  the operator to restart the gateway and retry. With no live gateway it
  performs the durable rewrite itself (safe: nothing else holds the store
  open).

Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.

Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
2026-09-16 00:32:15 -07:00
teknium1 ebe11403c6 fix(gateway): a profile named 'main' gets its own session namespace
`main` is a valid profile name (only hermes/default/test/tmp/root/sudo are
reserved), but _session_key_namespace mapped it to `agent:main` — the default
profile's namespace. Both profiles then built byte-identical keys: one routing
entry, one cached agent, and, since 75ae2859b9 pinned default-namespace
keys to the launch store, profiles/main's scoped sessions were written into
the ROOT state.db instead of profiles/main/state.db.

Key the `main` profile as `agent:main~` (`~` is outside the profile-id
alphabet, so the marked form cannot be any other profile's id) and give the
namespace slot one inverse, profile_from_session_key_namespace, used by the
store's key parser, _parse_session_key, the update-marker profile reader and
the profile-delete eviction prefix. Default keys stay byte-identical.
2026-09-13 15:41:01 -07:00
Siddharth Balyan cbcf7b72f7 feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop

* feat(gateway): /signin signs the free tier into a Nous account from a DM

* feat(cli): chat surfaces name /signin as the sign-in verb

* fix(auth): review follow-ups for the shared sign-in flow and /signin

* fix(i18n): carry the /status free-tier line in every locale catalog

* refactor(cli): the chat sign-in command is /login

* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
2026-09-11 03:45:33 +05:30
Teknium 7d50f99fbb fix(gateway): honor privacy policy in busy message origins
Reuse the effective gateway config and shared session platform policy before hashing model-facing metadata. Preserve original routing state and cover enabled/disabled redaction across all busy injection routes.
2026-09-07 07:12:06 -07:00
Teknium 34e512ae58 fix(sessions): remove remaining timer migration and guidance 2026-09-07 06:10:54 -07:00
Teknium 1d5d059410 fix(gateway): stop time-triggered conversation rotation 2026-09-07 06:10:54 -07:00
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium 2031c819fe simplify(compat): gateway — re-land 92d0bd0d73 (reverted by stale-index commit b818085298)
Re-applies the gateway compat removal byte-for-byte; see 92d0bd0d73 for the
full inventory (30 re-exports/aliases + 2 shim modules dropped, 3 shim-only
names re-removed, 24 callers + 34 test files repointed). No new changes.
2026-09-03 13:12:50 -07:00
Teknium b818085298 simplify(compat): doctor/status — drop 13 re-exports + the doctor_* globals() facade (97 names), repoint 6 callers / 13 tests 2026-09-03 13:10:52 -07:00
Teknium 92d0bd0d73 simplify(compat): gateway — drop 30 re-exports/aliases + 2 shim modules, re-remove 3 shim-only names, repoint 24 callers + 34 test files
Per COMPAT_REMOVAL.md (internal import paths are not a stable API):

Re-exports removed
- gateway/run.py: atomic_json_write, load_dotenv, resolve_delivery_transport,
  TurnRunner, merge_pending_message_event, _arm_loop_floor_timer,
  start_loop_liveness_watchdog, DEFAULT_GATEWAY_POST_INTERRUPT_GRACE_TIMEOUT,
  _UNSET (9) — run_* mixins and tests now import from the defining module
  (gateway.delivery / gateway.run_turn_runner / gateway.platforms.base /
  gateway.shutdown_watchdog / gateway.restart / utils).
- gateway/session.py: SessionResetPolicy, normalize_whatsapp_identifier,
  TranscriptReadError, auto_continue_freshness_window (+ "_now & co." noqa
  facade) — gateway/__init__ takes SessionResetPolicy from .config; callers
  take TranscriptReadError from gateway.session_transcript.
- gateway/kanban_watchers.py: _wake_scope_id + "tests import via origin" noqa
  facade; tests import from kanban_watchers_common / _notifier.
- gateway/slash_commands.py: _model_switch_skew_guard, HISTORY_UNREADABLE.
- gateway/stream_consumer.py: escape_code_fences_for_display.
- gateway/platforms/api_server.py: "re-exported" RunIdempotencyStore comment;
  tui_gateway + tests import gateway.platforms.api_server_run_idempotency.
- gateway/platforms/__init__.py: PEP 562 __getattr__/__dir__ lazy QQAdapter /
  YuanbaoAdapter facade (no in-tree importer).
- gateway/startup_watchdog.py: whole re-export shim module deleted; the three
  in-tree callers import hermes_startup_watchdog directly.

Aliases removed
- gateway/platforms/signal.py: SignalAdapter._markdown_to_signal.
- gateway/shutdown_forensics.py: _parse_systemd_duration_to_us.
- gateway/platforms/yuanbao.py: OutboundManager.start_slow_notifier /
  cancel_slow_notifier / get_chat_lock / _chat_locks / CHAT_DICT_MAX_SIZE
  delegates; module-level get_active_adapter / send_yuanbao_direct;
  MarkdownProcessor has_unclosed_fence / ends_with_table_row /
  split_at_paragraph_boundary static pass-throughs (chunk_markdown_text stays —
  it carries yuanbao's chunking policy). tools/send_message_senders +
  tools/yuanbao_tools call YuanbaoAdapter.get_active() / sender.send_direct().

Shim-only names re-removed (earlier review-fix round 96c104c903)
- CapabilityDescriptor.from_platform_entry (+ tests/gateway/relay/test_descriptor_from_entry.py)
- SessionTurnLeaseRegistry.__len__ (+ its test)
- is_relay_media_url KEPT: download() uses it, real internal helper.

Tests that pinned a shim (startup_watchdog re-export identity, api_server
RunIdempotencyStore identity, signal wrapper parity) are dropped; tests that
pinned live behavior are repointed at the implementation.

Note: the gateway/run.py hunk of this change was swept into 00a3cfe5c9 by a
concurrent commit on the shared worktree; this commit carries the rest.

Verified in an isolated worktree (HEAD + this change only): ruff clean,
import-smoke of all 33 touched modules under a fresh HERMES_HOME,
605 passed / 1 skipped across the 39 covering test files.
2026-09-03 13:10:49 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 28bbdc4bb8 refactor(gateway): session.py — hoist Slack notes constants, fold key/context builders (prompt golden byte-identical) 2026-09-02 23:42:48 -07:00
Teknium e8d734565c refactor(gateway): status.py — compact docstrings by hand, fold small guards 2026-09-02 23:35:26 -07:00
Teknium 2622c2b06a refactor(gateway): session — simplify pii-safe resolution, from_dict ordering, hash prefix 2026-09-02 21:29:40 -07:00
Teknium 873c805495 refactor(gateway): session — read participant id lazily in build_session_key (duck-typed sources) 2026-09-02 21:10:12 -07:00
Teknium 1239f383ba refactor(gateway): session — _RouteDecision.schedule_reset, inline route checks, compact store/state/context comments 2026-09-02 21:04:36 -07:00
Teknium c4655a638c refactor(gateway): session — tighten module header and path-guard docstring 2026-09-02 20:10:00 -07:00
Teknium 527521fced refactor(gateway): session — repack prompt-text literals to 100 cols 2026-09-02 20:07:00 -07:00
Teknium d6d2cad44c refactor(gateway): session — unify build_session_key branches, compact dataclass comments/docstrings 2026-09-02 20:00:41 -07:00
Teknium 46258d89c5 refactor(gateway): session — rewrap comment/docstring paragraphs to 100 cols 2026-09-02 19:38:43 -07:00
Teknium 75fdd85316 refactor(gateway): session — pack exploded argument lists 2026-09-02 19:36:48 -07:00
Teknium 4f943a1cdd refactor(gateway): session — hug short parameter/argument lists 2026-09-02 19:33:00 -07:00
Teknium e890c57a2b refactor(gateway): session — fold short multi-line signatures/calls 2026-09-02 19:29:01 -07:00
Teknium ef4f0f4f61 refactor(gateway): session — checkpoint: phase helpers for get_or_create_session, shared clock/id helpers in lifecycle, context/state/stall folding 2026-09-02 19:21:38 -07:00
Teknium d7bdf2788d refactor(gateway/session): split SessionStore into persistence/recovery/lifecycle/transcript mixins by call-graph cohesion; compact wire helpers 2026-09-02 16:15:42 -07:00
Teknium c89819bec2 refactor(gateway/session): split the two >200-LOC methods into collaborators, unify entry mutators, table-drive platform prompt notes 2026-09-02 15:46:47 -07:00
Teknium 659c6b9fff refactor(gateway): simplify session, status, stream_consumer, kanban_watchers, relay adapter, hosted-room driver/discussion/replicas, shutdown/lifecycle, pairing, channel_directory, control_socket and small modules 2026-09-02 13:31:54 -07:00
Teknium aed6720dab refactor(gateway/run, slash_commands): dispatch tables, helper unification and hand-reviewed comment compaction
run.py:
- built-in adapter creation: 9-branch if/elif -> _BUILTIN_ADAPTERS table
- idle slash-command routing: 35 `if canonical == ...` branches -> _gateway_idle_command_handlers()
- shared helpers: _send_command_ack (4 sites), _command_origin_for_source (2), _session_entry_for_manager
  (goal/heartbeat), _toggle_adapter_auto_tts_set (2), _load_env_or_agent_cfg_timeout (2), _float_env reuse (2),
  _resolve_session_key_or_none (3), _running_agent_ids (4), _schedule_rename_from_title_thread (2),
  _write_runtime_status_quiet (5), _AUTO_RESET_CONTEXT_NOTES/_auto_reset_reason_text
- ruff SIM102/SIM103/SIM105/SIM108/SIM118 + F401 across gateway/ (semantics re-reviewed; sqlite Row
  `.keys()` and side-effecting assignments kept)
- two hand-reviewed comment/docstring compaction passes (AST-identical, rationale kept)

slash_commands.py:
- /model: typed path and picker callback shared one 200-line commit block -> _perform_model_switch +
  _commit_model_switch
- comment/docstring compaction (AST-identical)
2026-09-02 13:31:53 -07:00
Teknium 7463cd1202 fix(session): fence multiplex peer-fallback recovery and profile inheritance by owner (#74285, #88381)
The per-profile store partition (17ba992108, 5ffaed6e45, 5cc3da6827)
already keeps fresh rows apart, but legacy rows written to root state.db
before the partition still sat where the default profile's peer-tuple
fallback could adopt them: a Telegram DM's tuple (chat_id == user_id, no
thread) is identical for every bot. Three residual holes, closed with the
smallest predicate that fits main's design:

- hermes_state.find_latest_gateway_session_for_peer: the fallback query
  now requires COALESCE(s.profile_name, <store owner>) = <store owner>
  (owner via SessionDB._own_profile_name). Handles NULL legacy rows and
  the single→multiplex migration case a key-namespace fence would break;
  stores outside the profile tree (no derivable owner) are unchanged.
- gateway/session._recovered_row_allowed_for_active_profile: under
  multiplexing no longer `return True` — the recovered row's agent:<ns>:
  must match the REQUESTED key's namespace (the active profile is
  meaningless when several profiles serve concurrently). Single-profile
  behavior unchanged; keyless/unnamespaced rows stay adoptable.
- hermes_state create_session parent COALESCE: profile_name inherits only
  when parent and child agree on agent:<ns>: (or either is keyless), so a
  default child forked from a sibling row is not durably mislabelled.

Co-authored-by: pcaruba <31041167+pcaruba@users.noreply.github.com>
Co-authored-by: jiangtaoliu-source <308256854+jiangtaoliu-source@users.noreply.github.com>
Co-authored-by: 69k4xmdfm2-blip <275826864+69k4xmdfm2-blip@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
leomcamilo bcc2e65818 fix(state): quarantine SessionDB handle after structural corruption
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.

Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".

Refs #90837, #90950, #97940, #89332, #45383

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
2026-09-02 16:57:21 +05:30
fangliquanflq 2d37a3a23c fix(gateway): fail closed when transcript reads fail 2026-09-01 23:58:01 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium cf21e28d39 fix(gateway): _routing_db tolerates bare test instances (object.__new__)
Bare SessionStore instances built without __init__ lack _db_pinned,
_routing_home, and the handle cache behind the _db property. Restore
main's old getattr contract for them: report no DB and fall through to
the sessions.json path instead of raising AttributeError.
2026-08-31 14:54:18 -07:00
caya8205-2 8d74cb52da fix(gateway): give the routing index one store instead of the ambient one
Second half of #66887. _entries is a single flat dict holding every
profile's keys, so the index it persists to has to be a single file — but it
was read and written through _db, which resolves whichever profile scope is
active. A whole-index rewrite during one profile's turn copied every other
profile's routing rows into that profile's store, and startup, which runs
unscoped, then loaded a different copy than the last writer produced.

That is why the startup recovery pass never sees a secondary profile's crash
marker, which is the half this issue's title names. mark_turn_active()
persists through the single-entry fast path (state.db only, no sessions.json
mirror), so a marker written during a profile's turn landed in that
profile's store and _recover_unclean_sessions(), running with no scope, read
a store that had never heard of it. The turn was silently never promoted to
resume_pending.

Capture the gateway's own home at construction — the store is built at
startup before any profile scope exists — and route the index through it:
_ensure_loaded_locked, _reconcile_recovered_routing_locked,
_persist_routing_data and _save_entry now use _routing_db. A pinned handle
still wins, so suites that install a fake or disable the DB are unaffected.

_prune_stale_sessions_locked is the mixed case and is split accordingly: it
now asks _db_for_key(key) whether each session ended, because that is a
per-session question, while the index write stays on the single store. One
ambient handle previously answered it for every profile at once, which could
prune a live secondary-profile route on the strength of the root store's
copy of that session.

Regression as requested on the issue: mark a turn active under a secondary
profile's scope, then build a fresh store with no scope and run
recover_interrupted_turns(). It promotes exactly one turn to resume_pending
here and promotes zero against the previous behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 a60e04a32c fix(gateway): prove the compression child's owner before writing to it
Review P1 #2. _append_to_transcript_serialized() writes the compression
continuation to child_id BEFORE publishing either _transcript_reroutes or
the _entries update — that ordering is load-bearing for backlog order, so it
must not move. At that moment nothing in the routing index points at the
child, so _db_for_session_id(child_id) missed its scan and fell through to
_db_for_key(None), i.e. the ambient store. The fail-closed guard did not fire
because root is a live handle.

The row therefore targeted root rather than the already-proven parent owner.
With no child row there the append is rejected by the FOREIGN KEY constraint,
the pending queue never drains and the reroute cannot advance; against a
split-brain root the message would instead be written cross-profile.

Record ownership before the mutation instead of moving the publication: a
private _session_owner_hints map carries session_id -> owning key for ids
whose owner is proven but not yet published, consulted by the new
_owner_key_for_session_id() after the index scan misses, and dropped as soon
as routing publishes. Signatures are unchanged, so the existing suites that
stub _append_transcript_message keep working untouched; the map is read
through getattr for stores built via object.__new__.

The regression is physical rather than mocked: an ended compression parent
and a live child that exist only in profiles/fitness/state.db, no active
profile scope, append to the parent, then assert all four effects — the row
lands on the child in the profile store, the pending queue drains, the
reroute and the routing entry advance, and root state.db stays untouched.
Without the hint it fails exactly as the review predicted, on
"FOREIGN KEY constraint failed" against root.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 49fb58523f fix(gateway): fail closed when a named profile has no resolvable home
Review P1. _profile_home_for_key() returned the same None for three
different states — multiplexing off / legacy agent:main namespace, a named
profile whose directory does not exist yet, and a resolution error — and
_db_for_key() collapsed all of them to the ambient store.

That recreated the very split this change removes. The enrollment bridge
provisions profiles/<name>/ at runtime, so a key such as
agent:fitness:telegram:dm:1 can legitimately be seen first: the first lookup
landed in root state.db, and the next one, after provisioning, in
profiles/fitness/state.db. One qualified session identity, two physical
stores. The resolver-exception path fell open the same way.

Ownership is now tri-state:
  - no named owner            -> ambient DB (single-profile behavior intact)
  - named owner + home        -> that profile's DB
  - named owner, unresolvable -> None, and a warning; never root

Callers already treat a missing DB as "skip the mutation", which is the
defer-don't-misroute behavior wanted here. _append_transcript_message is the
one path reached with an id the entry-point guard did not check (the
compression-child id), so it now raises explicitly and lets the caller's
retry queue hold the row instead of relying on an AttributeError.

Tests exercise the effect boundary rather than cache state: a named key
before its profile exists leaves root untouched and lands only in the
profile store once provisioned, and a resolver exception fails closed too.
Both fail against the previous two-state behavior by returning a live
SessionDB where None is required.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 58dcc10049 fix(gateway): memoize only resolved profile homes, never the miss
Review feedback on the memo introduced with _profile_home_for_key. The
problem is sharper than "no invalidation on profile deletion": caching the
miss pinned a profile that appears AFTER the gateway started to the ambient
store for the life of the process, which is the exact failure this helper
exists to prevent.

That is not hypothetical — an enrollment bridge can provision
profiles/<name>/ at runtime, so a key is legitimately seen before its
directory exists.

Memoize hits only. A miss costs one profile_exists() stat and recurs only
for profiles that genuinely do not exist, so the hot path for real profiles
is still a dict hit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
caya8205-2 5ffaed6e45 fix(gateway): resolve session storage from the key's profile, not ambient scope
#88734 made SessionStore._db follow the ambient HERMES_HOME so a multiplexed
profile's rows reach its own state.db. That is correct for the inbound message
path, which installs the scope via _profile_runtime_scope. Nothing else does.

_session_expiry_watcher (gateway/run.py) walks the single process-wide
_entries dict — every profile's keys — and finalizes expired sessions with no
scope installed, so _db resolved the ROOT store for rows that live under
profiles/<name>/state.db. The scoped inbound path and the unscoped background
path then maintained two copies of the same logical session whose end_reason
drifted apart independently. Once they disagreed, the #54878 stale-routing
guard read one copy while the routing index pointed at the other, and a live
conversation was dropped and recreated — silently, since that branch only sets
was_auto_reset when a reset policy also fired.

Field evidence from a live two-profile install: session 20260814_234313 was
end_reason=None in the root store but agent_close in the profile store, while
20260822_225807 was inverted. Both directions, which rules out a single
mis-scoped writer.

The owning profile is already encoded in the session key, so derive the store
from it: _profile_home_for_key / _db_for_key, plus _db_for_session_id for the
entry points addressed by session id. 40 self._db uses across 14 methods now
resolve that way. No signature changed and no existing test was modified.

_profile_home_for_key returns None when multiplexing is off, when the key
carries the legacy agent:main namespace, or when the profile has no live
directory, so single-profile installs resolve exactly where they always did.
The explicit-path branch still goes through SessionDB.__init__ ->
_ensure_test_isolation, keeping the live-DB guard over per-profile paths.

Part of #66887. The routing-index half — _routing_scope() and the sessions.json
mirror still pinned to one frozen sessions_dir while the handle moves — is left
for a follow-up rather than mixed in here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 14:54:18 -07:00
rainbowgits 71256dfd01 fix(state): fail loudly when state.db is replaced under a live process
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:02:55 -07:00
Teknium f680a4dc7d fix(gateway): require FTS provenance before transcript rebuild-and-retry
Widen #96038's fail-closed classifier to the gateway transcript retry
path: SessionStore._is_fts_corruption_error no longer treats a generic
'database disk image is malformed' as FTS-only damage. It now delegates
to SessionDB._is_fts_write_corruption_error (SQLITE_CORRUPT_VTAB result
code or explicit fts5 corrupt-structure text) and only keeps the
messages_fts-named cases. Structural corruption falls through to the
bounded retry/backoff path instead of rebuilding FTS and retrying writes
against a damaged database.

Sibling site spotted in PR #98090 by @fangliquanflq.
2026-08-31 11:42:23 -07:00
kshitijk4poor 726f0ce1b5 fix: keep original entry object on same-id recovery — preserve live state (review follow-up)
Review found _create_entry_from_recovered_row builds a minimal entry:
replacing the live object would silently drop model_override, token/cost
counters, resume_pending/queued-work markers, and metadata. Keep the
original entry (routing is unchanged, so no sessions.json rewrite either)
and log at INFO — this is a success path, not a corrective action.
Regression test now asserts state preservation, reopen_session call, and
no save.
2026-08-27 20:06:16 +05:30
Jackal991 ff3f25e041 fix(gateway): keep sessions.json entry when startup recovery succeeds with same session id
Closes #95957
2026-08-27 20:06:16 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
fangliquanflq 80cec2785d fix(gateway): preserve routing state across recovery 2026-08-23 18:25:12 -07:00
fangliquanflq 4b659f0e33 fix(gateway): retry failed session database opens 2026-08-23 18:25:12 -07:00
poisdahl fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
kshitijk4poor c45e2b19c3 fix(state): guard gateway FTS rebuild + comment early flag-set
Add the foreign-holder guard to gateway/session.py::_rebuild_fts_once(),
the third FTS rebuild path that was not covered by the original fix.
Also add a comment explaining why _fts_runtime_rebuild_attempted is set
before the foreign-holder check: the fail-open path that follows
persists FTS_STALE_KEY so the next startup retries via _recover_stale_fts.
2026-08-22 03:56:13 +05:30