Symptom (Ghostty, display.pet on, display.bell_on_complete on): at the end of a turn the
input line fills with several rows of base64 and a stale copy of the status bar + pet
stays above the response panel.
Root cause: _ring_bell runs on the agent thread and terminal_notify wrote the OSC 9 /
OSC 777 sequence through its own open("/dev/tty") (or sys.stdout). The prompt_toolkit
loop thread may at that moment be mid-write of a 12 KB kitty APC pet frame, which the tty
drains ~1 KB at a time. The second writer splices into the frame; the foreign ESC aborts
the APC and the terminal paints the remainder of the payload as text at the input cursor.
The wrapped garbage scrolls the screen, so the panel that follows is printed against a
stale cursor position and the old chrome survives above it.
Change: when the CLI's Application is running, _ring_bell hands "\a" + the notification
sequence to the app loop (_run_on_app_loop -> _write_terminal_sequence), serializing it
behind the renderer and the after_render frame writer. terminal_notify.notify() keeps the
/dev/tty path for callers without a running app; the sequence builder is split out as
notification_sequence().
Verification: pty A/B with the real Application + after_render frame writer, 400 rings
vs 91 frames — base: 3 leaks (4,324 base64 chars painted); fixed: 0 leaks, 400/400
notifications delivered, 0 aborted frames.
The dashboard endpoint GET /api/skills/hub/search passes its user-supplied
`source` straight into parallel_search_sources and never applied the merged
provider cut, so ?source=nvidia returned a mixed set. That was the fourth
caller of the walker; the cut was copy-pasted at three of them and missing
at the fourth.
parallel_search_sources already computes the normalized provider filter, so
the cut now lives there — applied per source before results are counted and
merged. Every caller (CLI search via unified_search, CLI browse, TUI-gateway
browse, dashboard router) sees the same rule with no provider logic of its
own, source_counts stop reporting rows that are then dropped, and the three
duplicated call-site cuts are deleted. do_browse keeps its provider-specific
"No skills found for provider" message.
Also:
- HermesIndexSource.search now treats a whitespace-only provider_filter as
"no filter", matching GitHubSource.search (the two adapters previously
disagreed on the same keyword argument; unreachable through the walker,
which pre-normalizes).
- The regression-test fixture seeds tap caches by github_provider_for label
instead of case-sensitive repo literals, and serializes metas through
_skill_meta_to_dict, so a DEFAULT_TAPS casing change can no longer silently
unseed the fixture.
Validation: 121 targeted tests green; disabling the walker cut fails the
pre-existing test_unified_search_provider_filter_keeps_index_source with the
expected clawhub leak; 4/4 regression cases still red on unpatched main.
Follow-up polish on the provider-filter-before-limit fix:
- GitHubSource.search now skips taps whose repo maps to a different provider
instead of enumerating every tap and filtering afterwards. A tap's repo fixes
the provider of every result it yields (github_provider_for is the only source
of extra.provider in this adapter), so the skip is lossless and avoids up to 23
useless tap enumerations per provider-filtered search — real GitHub API calls
against the 60/hr unauthenticated budget and the 30s overall timeout when the
index is unavailable. The now-redundant post-loop filter is dropped.
- _provider_filter_of() is the single owner of "does --source name a provider";
it replaces the four inline copies of the strip/lower/membership idiom in
_select_active_sources, parallel_search_sources, unified_search and do_browse.
- _tap_cache_key() is shared by _list_skills_in_repo and the regression test so
the seeded tap cache can never drift from the production key format.
- _entry_provider() dedupes the raw-index provider extraction used by both the
pre-ranking filter and the scoring loop in HermesIndexSource.search.
- browse_skills (the TUI-gateway browse path) now applies the same merged
provider cut as do_browse; it accepted a provider value but returned
unfiltered results.
Validation: 121 targeted tests green; the regression tests go red on both
adapters when either the tap skip or the index pre-filter is neutralized, and
4/4 red on unpatched main; live CLI repro returns 0 results on main and 3/3
provider matches on this stack.
Slot arguments are always complete json.dumps output (Gemini re-sends full args), so the mid-stream JSON check could never fire; the id key needs no tool name; _new_call_id and the slot lookup now share _provider_call_id.
Two different calls to the same tool arriving in separate stream events
collided in one accumulator slot: Gemini 2.5 sends no call id and
part_index restarts at 0 per event, so the second call's arguments were
emitted as a delta on the first call's index and concatenated downstream
into unparseable JSON, dropping a call. Gemini 3 ids are now the slot
identity (part_index and thought signature drift across events of one
call); without an id, a call whose arguments are not a continuation or
resend of the slot's accumulated JSON opens its own slot, kept reachable
as key#N so a later resend lands on it.
Re-applied by hand onto the collapsed translate_stream_event on main from
#75528 (9371874010 + f4c8863cdc). #24676 by cdbartholomew (May 13) was
the first fix for this collision (value-based slot matching without the
id key) and is credited as co-author.
Co-authored-by: Chris Bartholomew <chris.bartholomew@vectorize.io>
The Electron listener now always emits iss (null when the server sent none)
and McpOauthCallbackParams is extra="forbid", so a new Desktop against a
backend without this change would fail every remote MCP OAuth login with a
4000 - including providers that never send iss. Send the key only when set.
Also drop the deliver_callback_flow test the RPC test subsumes.
The gateway loopback handler re-inlined tools.mcp_oauth._parse_redirect_query;
that copy is exactly how the gateway relay lost `iss` while the CLI path
kept it. Use the helper so the four callback keys have one owner, and
point the docstring at it instead of repeating the RFC 9207 rationale.
The oauth.callback handler parsed `iss` but never passed it to deliver_callback_flow, and McpOauthCallbackParams (extra="forbid") had no `iss` field, so the desktop renderer sending `iss: null` was rejected with 4000 "unknown key" — breaking every Desktop→remote-gateway MCP OAuth login. Add the field, forward it, and regenerate the OpenRPC/TS contract artifacts via scripts/gen_gateway_contracts.py.
Also update tests/hermes_cli/test_mcp_dashboard_oauth.py for the 3-tuple callback shape introduced by the cherry-picked commit (it was red on the stack).
mcp 2.x rejects an authorization response that omits the RFC 9207 `iss`
parameter when the authorization server advertised
`authorization_response_iss_parameter_supported`. Cloudflare advertises it
AND sends it; the CLI loopback handler has always forwarded it, but every
other callback producer parsed only code/state/error, so the SDK raised:
OAuthFlowError: Authorization response missing iss parameter
advertised by the authorization server
and the server parked. Same machine, same config, `hermes mcp login <name>`
from a terminal succeeded — the failure is specific to the non-CLI relays.
Forward `iss` on every producer, matching `_make_callback_handler()`:
- tools/mcp_dashboard_oauth.py: `deliver_callback()` accepts `iss`;
`wait_for_callback()` returns `(code, state, iss)`. The bridge in
tools/mcp_oauth.py already splats that tuple into
`_authorization_code_result(code, state, iss)`, so it needs no change.
- tui_gateway/mcp_oauth_sessions.py: the gateway-hosted loopback listener
parses `iss`, and `deliver_callback_flow()` forwards it.
- tui_gateway/methods_tools.py: the `oauth.callback` RPC passes `iss`.
- hermes_cli/web_routers/mcp.py: the dashboard callback route accepts it.
- apps/desktop/electron/mcp-oauth-callback-ipc.ts: the one-shot listener
reads `iss` off the redirect (the renderer already spreads the whole
callback object into the RPC, so it flows through unchanged).
Providers that omit `iss` round-trip as `None`/`null` rather than being
dropped, so servers that do not advertise RFC 9207 keep working.
Verified live on Windows against mcp.cloudflare.com, whose metadata sets
`authorization_response_iss_parameter_supported: true`: the server that
previously parked on the missing-iss error now reports
`Authenticated — 3452 tool(s) available` and `hermes mcp test cloudflare`
connects. State-mismatch and replay rejection are unchanged.
Tests (each fails on base, passes with the fix):
- test_dashboard_flow_preserves_rfc9207_iss
- test_deliver_callback_forwards_iss (client-redirect relay)
- test_loopback_listener_forwards_iss (real HTTP redirect)
- two vitest cases on the Electron listener, incl. the iss-absent case
Refs #92758, #99984. PR #92765 fixes the dashboard route and the loopback
listener but not the client-redirect relay
(`deliver_callback_flow` / `oauth.callback` / the Electron listener), which
is the path Desktop drives against a remote backend.
hermes_state_common pulls in agent.* at import, so the URI builder moves to
hermes_state_holders (errno/os/sqlite3/pathlib only) where the gateway
readiness probe and backup can adopt it in a follow-up sweep. The doctor
structural-damage branch is one helper instead of two copies, the holder
scan goes through hermes_state_repair._live_writer_holds_db, the migration
hint uses _schema_not_built (the startswith("no such ") check also matched
"no such module: fts5"), and the hermes_state import is hoisted so an import
failure cannot mask itself as UnboundLocalError.
read_only_db_uri() replaces four inline mode=ro URI sites (two of which
still used the raw f-string that truncates on ?/# in the home path:
state_db_has_structural_damage and collect_state_db_stats). The doctor
write probe now applies the live-holder gate in both modes: a quiet store
is probed in place as on main, a held store is probed through a read-only
snapshot, and a held store over 1 GB is skipped with an info line unless
--fix is given (the unconditional copy cost one full DB write per plain
doctor run). Connect/backup failures propagate to the existing
classification instead of being reported as FTS write-health failures.
Observational sessions commands print a migration hint instead of a raw
traceback when a read-only opener meets an older schema.
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
- Keep (a) list/stats/pinned open SessionDB(read_only=True) — one parametrized test — and (b) a missing store prints empty results and is never created.
- Drop the insights read-only test (already covered on main), the status test, the mutating-action/live-writer/doctor-isolation tests, and the two doctor factory tests.
- Replace _EmptyObservationalStore + error-string sniffing with a plain '_default_db_path() does not exist' branch printing each action's empty output; the fake-store tests in test_sessions_pin keep working because the branch only runs when the open fails.
- _session_count: back to main's raw sqlite mode=ro COUNT(*) via as_uri() — routing it through SessionDB(read_only=True) both re-introduced the raw f-string URI ('?'/'#' in the home path truncate it) and queries columns (s.archived) an unmigrated store lacks, so doctor would report a healthy DB as broken.
- _write_health_reason: snapshot source URI built with as_uri() for the same reason; the --fix live probe (_db_opens_cleanly runs BEGIN IMMEDIATE) now falls back to the snapshot unless live_writer_holds_db proves the store quiet, matching _state_db_wal — hermes doctor --fix never becomes a second writer against a gateway's state.db (#103339).
- SessionDB._connect_read_only: same as_uri() form so every read-only opener is safe in a home containing '?' or '#'.
- test_sessions_export_output_dir: fixture accepts the read_only kwarg the PR introduced.
- Drop the two doctor tests that pinned the SessionDB factory kwargs; main's URI-reserved-chars test covers _session_count.
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
## What does this PR do?
Makes observational CLI commands open `state.db` in read-only mode, so they can inspect a live Hermes installation without participating in writable WAL lifecycle handling.
### Symptom
Running `hermes status`, `hermes doctor` without `--fix`, `hermes sessions list`, `hermes sessions stats`, or `hermes insights` while a gateway owns the store could open another writable session handle. The live turn could then lose its WAL generation and stop.
### Impact
Users inspecting status or session history during an active turn could lose that in-flight turn and leave the gateway halted until recovery.
### Bug Cause
**Trigger:** observational CLI helpers constructed `SessionDB()` with its writable default.
**Causal chain:**
1. A live gateway holds the `state.db` WAL generation.
2. A nested observational CLI command opens a second writable handle.
3. Writable-handle close behavior can participate in WAL lifecycle work and retire the generation used by the live writer.
**Why it is wrong:** these commands only query state and should not have writer privileges.
**Working sibling / contrast:** repair and mutating session commands still use writable access intentionally.
**Ruled out:** no state schema, migration, or WAL checkpoint implementation changes are included.
### Fix
Routes status, non-fixing doctor state inspection, sessions list/stats, and both insights entrypoints through `SessionDB(read_only=True)`. Repair and mutating paths remain writable, and regression tests cover WAL preservation with a live writer.
## Related Issue
Fixes#110173
## Type of Change
- ✅ Bug fix (non-breaking change that fixes an issue)
## Changes Made
- `hermes_cli/status.py`, `hermes_cli/doctor_state.py`, and insights helpers — open observational state readers read-only.
- `hermes_cli/sessions_cmd.py` — make only `list` and `stats` read-only; retain writable access for mutations.
- `tests/hermes_cli/test_observational_sessiondb_modes.py` — verify access modes and a live writer's WAL remains usable.
## How to Test
- ✅ `scripts/run_tests.sh tests/hermes_cli/test_observational_sessiondb_modes.py tests/hermes_cli/test_cli_insights_command.py` — 9 passed.
- ✅ `scripts/run_tests.sh tests/hermes_cli/test_doctor.py tests/hermes_cli/test_doctor_structural_corruption.py tests/hermes_cli/test_sessions_error_exit_codes.py` — 75 passed; two sandbox-only failures came from blocked host process/symlink operations.
- ✅ A live `SessionDB` writer remains able to create and retrieve a session after `sessions stats` reads the store.
## Checklist
### Code
- ✅ I've read the Contributing Guide
- ✅ My commit messages follow Conventional Commits
- ✅ I searched for existing PRs to make sure this isn't a duplicate
- ✅ My PR contains only changes related to this fix
- ✅ I've run relevant tests locally (see How to Test)
- ✅ I've added tests for my changes
- ✅ I've tested on my platform: macOS
### Documentation & Housekeeping
- ✅ Documentation update: N/A
- ✅ `cli-config.yaml.example`: N/A
- ✅ `CONTRIBUTING.md` or `AGENTS.md`: N/A
- ✅ Cross-platform impact considered
- ✅ Tool descriptions/schemas: N/A
The stale-attempt socket shutdown re-implemented two blocks that already
live in agent_runtime_helpers: the settimeout(0)+shutdown(SHUT_RDWR)
body of force_close_tcp_sockets (now _shutdown_socket) and the
candidate->socket lookup of _iter_pool_sockets (now _socket_from_candidate).
The hand-unrolled _httpcore_stream unwrapping is dead since
_connection_candidates walks _stream/_httpcore_stream itself, so the helper
starts from the network_stream extension and the response stream only.
Also add ReadError to _TRANSIENT_TRANSPORT_ERRORS (the third classifier of
the same abort-induced read; the other two were already updated), drop the
incidental gettimeout() assertion from the shape test, and cut the E2E
from ~3 s to ~1.5 s (stale budget 1 s, serve_forever poll 50 ms).
Keep the real httpx 0.28 wrapper-shape test (proves the shutdown reaches
the socket through BoundSyncStream/ResponseStream/PoolByteStream) and the
loopback E2E (a parked reader unwinds within its stale budget and the
retry lands). The other four were narrower restatements of the same paths.
Also treat httpx.ReadError as a transport error in codex_runtime: it is the
same abort-induced-read class the streaming retry loop now recovers from.
Capping ``read`` to ``stale`` silently overrode an explicit
HERMES_STREAM_READ_TIMEOUT and the local-endpoint ``read = base`` branch,
and let the stale-kill E2E pass via ReadTimeout alone rather than through
the socket-shutdown unwedge this change is about. The shutdown path is
sufficient on its own: the E2E still passes with the cap removed.
The stale-stream monitor aborted a wedged provider stream only via
force_close_tcp_sockets() -> shutdown(SHUT_RDWR). That is best-effort: a
parked body read is not unblocked on every platform (Windows keeps the
pending recv parked) and the sweep can miss the socket. The worker then
stayed blocked in the provider read, so the retry loop never retried; the
monitor re-killed every stale interval and the call only ended at the
byte-read timeout, far past the stale budget - the reported
"No response from provider for 180-240s ... Reconnecting" loop ending in
"The model server is not responding".
- _kill_stale_stream now also closes the killed attempt's own provider
response (identity-guarded self._attempt_stream_response), which is what
actually unblocks a parked reader; a racing retry's fresh response is
never touched.
- an abort-induced httpx.ReadError counts as a transient connection error,
so the aborted attempt reconnects instead of ending the turn.
Reproduced with a local SSE server: before, the worker stayed parked and no
second request was issued (recovery only at the byte-read timeout); after,
the kill unblocks the reader at the stale budget and the retry lands.
Move _non_exportable_entries next to its first caller, fold the .pyc/.pyo
suffixes into it (the clone-all closure kept its own copy), and route the
last un-ignored profile copytree (the skills/ copy in _bootstrap_profile_dir)
through it. Cut the three repeated "sockets abort copytree" comments down to
the helper docstring. Split the clone-all special-file case into its own
POSIX-marked test so the cron-jobs assertion keeps its Windows coverage, and
use monkeypatch.chdir in the socket-binding test helper.
Route the --clone-all copytree ignore through _non_exportable_entries so a
live source profile holding a gateway or agent-browser socket (or a FIFO)
no longer aborts the clone with [Errno 6] No such device or address.
.pyc/.pyo and the root exclude sets keep their existing handling.
hermes_cli/profile_distribution.py:_copy_dist_payload is left alone: it
copies from a freshly extracted distribution archive (a staged temp tree),
never from a live profile, so it cannot meet a socket.
Extends test_clone_all_does_not_copy_cron_jobs to cover the clone path.
Named-profile export staged its copy with an ignore callable that only
excluded credential files, so any Unix socket in the profile (e.g. a stale
agent-browser control socket under home/.agent-browser/) made
shutil.copytree collect "[Errno 6] No such device or address" and raise
shutil.Error, failing the entire export. The default-profile export
already excluded *.sock by suffix, but a socket without that suffix (or a
FIFO, or a device node) failed it the same way.
Extract the universal exclusions into _non_exportable_entries(), which
keeps the __pycache__/*.sock/*.tmp name rules and additionally drops any
entry that is not a regular file, directory, or symlink (os.lstat mode
check), and use it in both the default and named export branches.
The internal_notification rule lived in three stringly-typed places
(_prepare_turn, the queued follow-up, MACHINERY_DISPLAY_KINDS); hoist it to
response_filters.display_kind_for_event(). should_swallow_silence() re-ran
the silence predicate its callers had already evaluated, so reduce it to
is_machinery_display_kind() and drop the try/except staticmethod wrapper
(response_filters has no gateway imports, so a module-level import is safe).
Drop the predicate-only unit test (the integration tests bind the contract)
and sentence-case the fallback like its sibling warnings.
"internal_notification" is the only persist_user_display_kind run_turn.py ever
sets; "model_switch"/"auto_continue" had no producer. Drop the
agent_result["already_sent"] = False reset in the shaper: the stream consumer
already retracts previews and clears its turn-final flags on a bare marker
(_suppress_silence_marker), so _run_agent_mark_streamed_delivery never sets
already_sent for one and the reset was dead.
The recursive _run_agent for a queued (/queue) follow-up passed no
persist_user_display_kind, so an internal (self-injected) follow-up ending
in a bare silence marker got the visible "silence marker rejected" fallback
meant for humans, and a human follow-up behind an internal opener could be
swallowed. Pass the same rule the top-level turn uses, stamp the terminal
turn's kind next to queued_terminal_inbound_id, and let the shaper prefer
it over the opener's kind.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The room-level test is the reported defect: two composer sends, two
threads, and neither backend transcript may contain the other thread's
prompt. Proven red on the unfixed code — both threads resolved to
sid-research-1.
Also pins same-thread continuity, the pre-thread adoption (first thread
continues the old conversation, later threads do not), and that the
member half of the key stays source-qualified so a remote `research` and
a local `research` never share a session.
Co-authored-by: Wenfengcheng <30426178+Wenfengcheng@users.noreply.github.com>
The co-keyed readers of `room.sessions` had to follow session identity or
they would split from it: the stop interrupt targeted a member's only
session regardless of which thread issued the stop, and the stranded-reply
harvest resumed by bare member key even though its marker already carries
`{before, thread}`.
The clarify mirror moves too — with per-thread sessions one member can be
blocked in two threads at once, and a room-and-member key let thread B's
question silently replace thread A's card. The sweep's owner lookup stays
a MEMBER question, so it reads the member half of the key and keeps the
`::` source-qualifier check on that half alone.
Co-authored-by: Wenfengcheng <30426178+Wenfengcheng@users.noreply.github.com>
A group member's hidden plumbing session was keyed by member alone, so
every thread in a room collapsed onto one backend transcript: thread B
resumed thread A's session and answered carrying A's context.
Everything else in the room engine is already thread-aware — the delta is
filtered by `groupThreadOf(e) === thread` and the watermark is
`${thread}::${memberKey}` — so only session identity was missing. Key and
title it per thread, and adopt a room's pre-thread pointer into the first
thread that speaks so an upgraded room's history is continued rather than
orphaned.
Co-authored-by: Wenfengcheng <30426178+Wenfengcheng@users.noreply.github.com>
The website key-features bullet (EN + zh-Hans) omitted /reset from the
retry points while the README lists it; also drops a local import in the
test file shadowed by the module-level one.
- none-returning client counts as success (sentinel contract)
- pending buffer drops oldest past the turn cap and the byte cap
- the shutdown interleaving test now pre-seeds a pending turn so the
no-resend assert discriminates on content, not timing: with the lock
removed it fails on assert 2 == 1 (duplicate send), verified by mutation
Two review findings on the salvage stack:
- _write_turns' _quietly consolidation keyed failure on a None result,
implicitly assuming add_memory never legitimately returns None. A None-
returning stub (the most common mock idiom) would mark every successful
write failed and re-append the batch forever. Module-level _FAILED
sentinel: only a raised exception re-queues.
- _pending_turns had no bound: a persistently failing service accumulated
one entry per turn for the process lifetime (gateway runs never
re-initialize), and every retry re-sent the whole accumulated payload —
O(n^2) upload bytes, and a size-rejected batch could never shrink. Cap
the buffer (50 turns / 256 KiB, drop oldest, warn once per trim).
Answers the open review ask on #109359 with the corrected contract: a hung
remote add_memory is bounded by the SDK timeout (empirically 1.03s at
timeout=1.0, max_retries=0), so shutdown's flush waiting on the capture
lock is bounded, not forever. The guard: while the worker owns the write
(blocked inside add_memory), a concurrent shutdown must wait and must not
re-send the same pending batch. Mutation-checked: lock -> nullcontext
makes this test and the session-switch interleaving test fail.
Eight tests compare a recorded custom_id against a freshly computed
_capture_custom_id; near a 4h-bucket edge the provider's write and the
test's expectation read now() in different buckets and the equality
fails. The frozen_capture_clock fixture pins the module datetime so both
reads are identical by construction.
The per-turn capture rewrite removed the raw urllib /v4/conversations
ingest, but stale references survived outside the diff hunks:
- README "Behavior" still carried the "written once via the conversations
endpoint" paragraph contradicted by the new bullets right above it.
- website memory-providers (EN + zh-Hans) still listed full-session ingest,
session-end /v4/conversations ingest, and ingest in the base-url and
api_timeout rows — the PR had updated one line per file but missed the
rest of the section.
Now every surface describes per-turn documents.add capture with retry.
Two cleanups on the salvaged per-turn capture path (review of #109359):
- _write_turns hand-rolled the try/except/log that the file's own _quietly
helper already provides (same exception class, level, exc_info shape);
the deleted _ingest path used _quietly for the identical call. Consolidate:
add_memory always returns a dict, so a None result is the failure sentinel.
- _capture_custom_id's `or 'hermes'` never fires: _sanitize_tag already
returns _DEFAULT_CONTAINER_TAG on empty input, and the literal duplicated
the constant the helper owns. Verified behavior-preserving (tests green
with the fallback artificially restored).
sync_turn runs on the MemoryManager worker, but on_session_switch and
shutdown run on the caller thread, so two _write_turns() calls could
snapshot the same pending batch and each replace the whole list (duplicate
append or lost pending turn). A capture lock now covers the
snapshot/write/replace sequence; _write_turns() reads pending turns inside
the lock instead of taking a caller-built list.
Retries are at-least-once: the documents API appends on a shared custom_id
and does not dedupe by content, so a write it accepted but whose response
was lost is appended again. Stated in the docstring and README instead of
implied.
Addresses review on #109359.
A failed flush at on_session_switch() restored the pending buffer and
then cleared it on the next line, so an unavailable service at the switch
boundary still lost every pending turn. Pending turns now carry their own
session_id; _write_turns() batches per session, so a later retry (next
turn, session end, shutdown) writes old-session turns under the old
session's custom_id even after the switch.
Addresses review on #109359.
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).
Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.
Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.
Metadata stays type/session_id/timestamp plus the existing sm_source.
Answers the P1 review on #111187: build_profile_secret_scope() held only
<profile>/.env plus that profile's external-source snapshot, never the
administrator-managed .env. The launch process applies that file LAST with
override (_apply_managed_env), so a managed key beats the user's own value in
os.environ. Under multiplex semantics get_secret() stops falling back to
os.environ on a scope miss, so inside a routed cron fire (and equally inside a
real multiplex gateway turn, which builds its scope through the same function
via gateway/run.py::_load_profile_secret_scope) a managed-only credential
resolved as absent and a managed-vs-user collision resolved to the USER value:
reversed precedence.
Fix at the source: build_profile_secret_scope() overlays load_managed_env()
last, after the profile .env and external sources, skipping process-global
names exactly as it does for the other two layers. Every multiplex-authoritative
scope (gateway turn, routed desktop fire, external worker env build) is built
here, so managed authority is composed once instead of restored per consumer.
No generic ambient-env fallback is reintroduced: only the managed file's own
keys enter the scope, and only with the managed file's values.
Regression (parametrized, two invariants): inside a routed fire a managed-only
key resolves through get_secret(); a managed-vs-user collision yields the
managed value.