Symptom (Ghostty, display.pet on, display.bell_on_complete on): at the end of a turn the
input line fills with several rows of base64 and a stale copy of the status bar + pet
stays above the response panel.
Root cause: _ring_bell runs on the agent thread and terminal_notify wrote the OSC 9 /
OSC 777 sequence through its own open("/dev/tty") (or sys.stdout). The prompt_toolkit
loop thread may at that moment be mid-write of a 12 KB kitty APC pet frame, which the tty
drains ~1 KB at a time. The second writer splices into the frame; the foreign ESC aborts
the APC and the terminal paints the remainder of the payload as text at the input cursor.
The wrapped garbage scrolls the screen, so the panel that follows is printed against a
stale cursor position and the old chrome survives above it.
Change: when the CLI's Application is running, _ring_bell hands "\a" + the notification
sequence to the app loop (_run_on_app_loop -> _write_terminal_sequence), serializing it
behind the renderer and the after_render frame writer. terminal_notify.notify() keeps the
/dev/tty path for callers without a running app; the sequence builder is split out as
notification_sequence().
Verification: pty A/B with the real Application + after_render frame writer, 400 rings
vs 91 frames — base: 3 leaks (4,324 base64 chars painted); fixed: 0 leaks, 400/400
notifications delivered, 0 aborted frames.
The dashboard endpoint GET /api/skills/hub/search passes its user-supplied
`source` straight into parallel_search_sources and never applied the merged
provider cut, so ?source=nvidia returned a mixed set. That was the fourth
caller of the walker; the cut was copy-pasted at three of them and missing
at the fourth.
parallel_search_sources already computes the normalized provider filter, so
the cut now lives there — applied per source before results are counted and
merged. Every caller (CLI search via unified_search, CLI browse, TUI-gateway
browse, dashboard router) sees the same rule with no provider logic of its
own, source_counts stop reporting rows that are then dropped, and the three
duplicated call-site cuts are deleted. do_browse keeps its provider-specific
"No skills found for provider" message.
Also:
- HermesIndexSource.search now treats a whitespace-only provider_filter as
"no filter", matching GitHubSource.search (the two adapters previously
disagreed on the same keyword argument; unreachable through the walker,
which pre-normalizes).
- The regression-test fixture seeds tap caches by github_provider_for label
instead of case-sensitive repo literals, and serializes metas through
_skill_meta_to_dict, so a DEFAULT_TAPS casing change can no longer silently
unseed the fixture.
Validation: 121 targeted tests green; disabling the walker cut fails the
pre-existing test_unified_search_provider_filter_keeps_index_source with the
expected clawhub leak; 4/4 regression cases still red on unpatched main.
Follow-up polish on the provider-filter-before-limit fix:
- GitHubSource.search now skips taps whose repo maps to a different provider
instead of enumerating every tap and filtering afterwards. A tap's repo fixes
the provider of every result it yields (github_provider_for is the only source
of extra.provider in this adapter), so the skip is lossless and avoids up to 23
useless tap enumerations per provider-filtered search — real GitHub API calls
against the 60/hr unauthenticated budget and the 30s overall timeout when the
index is unavailable. The now-redundant post-loop filter is dropped.
- _provider_filter_of() is the single owner of "does --source name a provider";
it replaces the four inline copies of the strip/lower/membership idiom in
_select_active_sources, parallel_search_sources, unified_search and do_browse.
- _tap_cache_key() is shared by _list_skills_in_repo and the regression test so
the seeded tap cache can never drift from the production key format.
- _entry_provider() dedupes the raw-index provider extraction used by both the
pre-ranking filter and the scoring loop in HermesIndexSource.search.
- browse_skills (the TUI-gateway browse path) now applies the same merged
provider cut as do_browse; it accepted a provider value but returned
unfiltered results.
Validation: 121 targeted tests green; the regression tests go red on both
adapters when either the tap skip or the index pre-filter is neutralized, and
4/4 red on unpatched main; live CLI repro returns 0 results on main and 3/3
provider matches on this stack.
Two different calls to the same tool arriving in separate stream events
collided in one accumulator slot: Gemini 2.5 sends no call id and
part_index restarts at 0 per event, so the second call's arguments were
emitted as a delta on the first call's index and concatenated downstream
into unparseable JSON, dropping a call. Gemini 3 ids are now the slot
identity (part_index and thought signature drift across events of one
call); without an id, a call whose arguments are not a continuation or
resend of the slot's accumulated JSON opens its own slot, kept reachable
as key#N so a later resend lands on it.
Re-applied by hand onto the collapsed translate_stream_event on main from
#75528 (9371874010 + f4c8863cdc). #24676 by cdbartholomew (May 13) was
the first fix for this collision (value-based slot matching without the
id key) and is credited as co-author.
Co-authored-by: Chris Bartholomew <chris.bartholomew@vectorize.io>
The Electron listener now always emits iss (null when the server sent none)
and McpOauthCallbackParams is extra="forbid", so a new Desktop against a
backend without this change would fail every remote MCP OAuth login with a
4000 - including providers that never send iss. Send the key only when set.
Also drop the deliver_callback_flow test the RPC test subsumes.
The oauth.callback handler parsed `iss` but never passed it to deliver_callback_flow, and McpOauthCallbackParams (extra="forbid") had no `iss` field, so the desktop renderer sending `iss: null` was rejected with 4000 "unknown key" — breaking every Desktop→remote-gateway MCP OAuth login. Add the field, forward it, and regenerate the OpenRPC/TS contract artifacts via scripts/gen_gateway_contracts.py.
Also update tests/hermes_cli/test_mcp_dashboard_oauth.py for the 3-tuple callback shape introduced by the cherry-picked commit (it was red on the stack).
mcp 2.x rejects an authorization response that omits the RFC 9207 `iss`
parameter when the authorization server advertised
`authorization_response_iss_parameter_supported`. Cloudflare advertises it
AND sends it; the CLI loopback handler has always forwarded it, but every
other callback producer parsed only code/state/error, so the SDK raised:
OAuthFlowError: Authorization response missing iss parameter
advertised by the authorization server
and the server parked. Same machine, same config, `hermes mcp login <name>`
from a terminal succeeded — the failure is specific to the non-CLI relays.
Forward `iss` on every producer, matching `_make_callback_handler()`:
- tools/mcp_dashboard_oauth.py: `deliver_callback()` accepts `iss`;
`wait_for_callback()` returns `(code, state, iss)`. The bridge in
tools/mcp_oauth.py already splats that tuple into
`_authorization_code_result(code, state, iss)`, so it needs no change.
- tui_gateway/mcp_oauth_sessions.py: the gateway-hosted loopback listener
parses `iss`, and `deliver_callback_flow()` forwards it.
- tui_gateway/methods_tools.py: the `oauth.callback` RPC passes `iss`.
- hermes_cli/web_routers/mcp.py: the dashboard callback route accepts it.
- apps/desktop/electron/mcp-oauth-callback-ipc.ts: the one-shot listener
reads `iss` off the redirect (the renderer already spreads the whole
callback object into the RPC, so it flows through unchanged).
Providers that omit `iss` round-trip as `None`/`null` rather than being
dropped, so servers that do not advertise RFC 9207 keep working.
Verified live on Windows against mcp.cloudflare.com, whose metadata sets
`authorization_response_iss_parameter_supported: true`: the server that
previously parked on the missing-iss error now reports
`Authenticated — 3452 tool(s) available` and `hermes mcp test cloudflare`
connects. State-mismatch and replay rejection are unchanged.
Tests (each fails on base, passes with the fix):
- test_dashboard_flow_preserves_rfc9207_iss
- test_deliver_callback_forwards_iss (client-redirect relay)
- test_loopback_listener_forwards_iss (real HTTP redirect)
- two vitest cases on the Electron listener, incl. the iss-absent case
Refs #92758, #99984. PR #92765 fixes the dashboard route and the loopback
listener but not the client-redirect relay
(`deliver_callback_flow` / `oauth.callback` / the Electron listener), which
is the path Desktop drives against a remote backend.
read_only_db_uri() replaces four inline mode=ro URI sites (two of which
still used the raw f-string that truncates on ?/# in the home path:
state_db_has_structural_damage and collect_state_db_stats). The doctor
write probe now applies the live-holder gate in both modes: a quiet store
is probed in place as on main, a held store is probed through a read-only
snapshot, and a held store over 1 GB is skipped with an info line unless
--fix is given (the unconditional copy cost one full DB write per plain
doctor run). Connect/backup failures propagate to the existing
classification instead of being reported as FTS write-health failures.
Observational sessions commands print a migration hint instead of a raw
traceback when a read-only opener meets an older schema.
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
- Keep (a) list/stats/pinned open SessionDB(read_only=True) — one parametrized test — and (b) a missing store prints empty results and is never created.
- Drop the insights read-only test (already covered on main), the status test, the mutating-action/live-writer/doctor-isolation tests, and the two doctor factory tests.
- Replace _EmptyObservationalStore + error-string sniffing with a plain '_default_db_path() does not exist' branch printing each action's empty output; the fake-store tests in test_sessions_pin keep working because the branch only runs when the open fails.
- _session_count: back to main's raw sqlite mode=ro COUNT(*) via as_uri() — routing it through SessionDB(read_only=True) both re-introduced the raw f-string URI ('?'/'#' in the home path truncate it) and queries columns (s.archived) an unmigrated store lacks, so doctor would report a healthy DB as broken.
- _write_health_reason: snapshot source URI built with as_uri() for the same reason; the --fix live probe (_db_opens_cleanly runs BEGIN IMMEDIATE) now falls back to the snapshot unless live_writer_holds_db proves the store quiet, matching _state_db_wal — hermes doctor --fix never becomes a second writer against a gateway's state.db (#103339).
- SessionDB._connect_read_only: same as_uri() form so every read-only opener is safe in a home containing '?' or '#'.
- test_sessions_export_output_dir: fixture accepts the read_only kwarg the PR introduced.
- Drop the two doctor tests that pinned the SessionDB factory kwargs; main's URI-reserved-chars test covers _session_count.
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
## What does this PR do?
Makes observational CLI commands open `state.db` in read-only mode, so they can inspect a live Hermes installation without participating in writable WAL lifecycle handling.
### Symptom
Running `hermes status`, `hermes doctor` without `--fix`, `hermes sessions list`, `hermes sessions stats`, or `hermes insights` while a gateway owns the store could open another writable session handle. The live turn could then lose its WAL generation and stop.
### Impact
Users inspecting status or session history during an active turn could lose that in-flight turn and leave the gateway halted until recovery.
### Bug Cause
**Trigger:** observational CLI helpers constructed `SessionDB()` with its writable default.
**Causal chain:**
1. A live gateway holds the `state.db` WAL generation.
2. A nested observational CLI command opens a second writable handle.
3. Writable-handle close behavior can participate in WAL lifecycle work and retire the generation used by the live writer.
**Why it is wrong:** these commands only query state and should not have writer privileges.
**Working sibling / contrast:** repair and mutating session commands still use writable access intentionally.
**Ruled out:** no state schema, migration, or WAL checkpoint implementation changes are included.
### Fix
Routes status, non-fixing doctor state inspection, sessions list/stats, and both insights entrypoints through `SessionDB(read_only=True)`. Repair and mutating paths remain writable, and regression tests cover WAL preservation with a live writer.
## Related Issue
Fixes#110173
## Type of Change
- ✅ Bug fix (non-breaking change that fixes an issue)
## Changes Made
- `hermes_cli/status.py`, `hermes_cli/doctor_state.py`, and insights helpers — open observational state readers read-only.
- `hermes_cli/sessions_cmd.py` — make only `list` and `stats` read-only; retain writable access for mutations.
- `tests/hermes_cli/test_observational_sessiondb_modes.py` — verify access modes and a live writer's WAL remains usable.
## How to Test
- ✅ `scripts/run_tests.sh tests/hermes_cli/test_observational_sessiondb_modes.py tests/hermes_cli/test_cli_insights_command.py` — 9 passed.
- ✅ `scripts/run_tests.sh tests/hermes_cli/test_doctor.py tests/hermes_cli/test_doctor_structural_corruption.py tests/hermes_cli/test_sessions_error_exit_codes.py` — 75 passed; two sandbox-only failures came from blocked host process/symlink operations.
- ✅ A live `SessionDB` writer remains able to create and retrieve a session after `sessions stats` reads the store.
## Checklist
### Code
- ✅ I've read the Contributing Guide
- ✅ My commit messages follow Conventional Commits
- ✅ I searched for existing PRs to make sure this isn't a duplicate
- ✅ My PR contains only changes related to this fix
- ✅ I've run relevant tests locally (see How to Test)
- ✅ I've added tests for my changes
- ✅ I've tested on my platform: macOS
### Documentation & Housekeeping
- ✅ Documentation update: N/A
- ✅ `cli-config.yaml.example`: N/A
- ✅ `CONTRIBUTING.md` or `AGENTS.md`: N/A
- ✅ Cross-platform impact considered
- ✅ Tool descriptions/schemas: N/A
The stale-attempt socket shutdown re-implemented two blocks that already
live in agent_runtime_helpers: the settimeout(0)+shutdown(SHUT_RDWR)
body of force_close_tcp_sockets (now _shutdown_socket) and the
candidate->socket lookup of _iter_pool_sockets (now _socket_from_candidate).
The hand-unrolled _httpcore_stream unwrapping is dead since
_connection_candidates walks _stream/_httpcore_stream itself, so the helper
starts from the network_stream extension and the response stream only.
Also add ReadError to _TRANSIENT_TRANSPORT_ERRORS (the third classifier of
the same abort-induced read; the other two were already updated), drop the
incidental gettimeout() assertion from the shape test, and cut the E2E
from ~3 s to ~1.5 s (stale budget 1 s, serve_forever poll 50 ms).
Keep the real httpx 0.28 wrapper-shape test (proves the shutdown reaches
the socket through BoundSyncStream/ResponseStream/PoolByteStream) and the
loopback E2E (a parked reader unwinds within its stale budget and the
retry lands). The other four were narrower restatements of the same paths.
Also treat httpx.ReadError as a transport error in codex_runtime: it is the
same abort-induced-read class the streaming retry loop now recovers from.
The stale-stream monitor aborted a wedged provider stream only via
force_close_tcp_sockets() -> shutdown(SHUT_RDWR). That is best-effort: a
parked body read is not unblocked on every platform (Windows keeps the
pending recv parked) and the sweep can miss the socket. The worker then
stayed blocked in the provider read, so the retry loop never retried; the
monitor re-killed every stale interval and the call only ended at the
byte-read timeout, far past the stale budget - the reported
"No response from provider for 180-240s ... Reconnecting" loop ending in
"The model server is not responding".
- _kill_stale_stream now also closes the killed attempt's own provider
response (identity-guarded self._attempt_stream_response), which is what
actually unblocks a parked reader; a racing retry's fresh response is
never touched.
- an abort-induced httpx.ReadError counts as a transient connection error,
so the aborted attempt reconnects instead of ending the turn.
Reproduced with a local SSE server: before, the worker stayed parked and no
second request was issued (recovery only at the byte-read timeout); after,
the kill unblocks the reader at the stale budget and the retry lands.
Move _non_exportable_entries next to its first caller, fold the .pyc/.pyo
suffixes into it (the clone-all closure kept its own copy), and route the
last un-ignored profile copytree (the skills/ copy in _bootstrap_profile_dir)
through it. Cut the three repeated "sockets abort copytree" comments down to
the helper docstring. Split the clone-all special-file case into its own
POSIX-marked test so the cron-jobs assertion keeps its Windows coverage, and
use monkeypatch.chdir in the socket-binding test helper.
Route the --clone-all copytree ignore through _non_exportable_entries so a
live source profile holding a gateway or agent-browser socket (or a FIFO)
no longer aborts the clone with [Errno 6] No such device or address.
.pyc/.pyo and the root exclude sets keep their existing handling.
hermes_cli/profile_distribution.py:_copy_dist_payload is left alone: it
copies from a freshly extracted distribution archive (a staged temp tree),
never from a live profile, so it cannot meet a socket.
Extends test_clone_all_does_not_copy_cron_jobs to cover the clone path.
Named-profile export staged its copy with an ignore callable that only
excluded credential files, so any Unix socket in the profile (e.g. a stale
agent-browser control socket under home/.agent-browser/) made
shutil.copytree collect "[Errno 6] No such device or address" and raise
shutil.Error, failing the entire export. The default-profile export
already excluded *.sock by suffix, but a socket without that suffix (or a
FIFO, or a device node) failed it the same way.
Extract the universal exclusions into _non_exportable_entries(), which
keeps the __pycache__/*.sock/*.tmp name rules and additionally drops any
entry that is not a regular file, directory, or symlink (os.lstat mode
check), and use it in both the default and named export branches.
The internal_notification rule lived in three stringly-typed places
(_prepare_turn, the queued follow-up, MACHINERY_DISPLAY_KINDS); hoist it to
response_filters.display_kind_for_event(). should_swallow_silence() re-ran
the silence predicate its callers had already evaluated, so reduce it to
is_machinery_display_kind() and drop the try/except staticmethod wrapper
(response_filters has no gateway imports, so a module-level import is safe).
Drop the predicate-only unit test (the integration tests bind the contract)
and sentence-case the fallback like its sibling warnings.
"internal_notification" is the only persist_user_display_kind run_turn.py ever
sets; "model_switch"/"auto_continue" had no producer. Drop the
agent_result["already_sent"] = False reset in the shaper: the stream consumer
already retracts previews and clears its turn-final flags on a bare marker
(_suppress_silence_marker), so _run_agent_mark_streamed_delivery never sets
already_sent for one and the reset was dead.
The recursive _run_agent for a queued (/queue) follow-up passed no
persist_user_display_kind, so an internal (self-injected) follow-up ending
in a bare silence marker got the visible "silence marker rejected" fallback
meant for humans, and a human follow-up behind an internal opener could be
swallowed. Pass the same rule the top-level turn uses, stamp the terminal
turn's kind next to queued_terminal_inbound_id, and let the shaper prefer
it over the opener's kind.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
- none-returning client counts as success (sentinel contract)
- pending buffer drops oldest past the turn cap and the byte cap
- the shutdown interleaving test now pre-seeds a pending turn so the
no-resend assert discriminates on content, not timing: with the lock
removed it fails on assert 2 == 1 (duplicate send), verified by mutation
Answers the open review ask on #109359 with the corrected contract: a hung
remote add_memory is bounded by the SDK timeout (empirically 1.03s at
timeout=1.0, max_retries=0), so shutdown's flush waiting on the capture
lock is bounded, not forever. The guard: while the worker owns the write
(blocked inside add_memory), a concurrent shutdown must wait and must not
re-send the same pending batch. Mutation-checked: lock -> nullcontext
makes this test and the session-switch interleaving test fail.
Eight tests compare a recorded custom_id against a freshly computed
_capture_custom_id; near a 4h-bucket edge the provider's write and the
test's expectation read now() in different buckets and the equality
fails. The frozen_capture_clock fixture pins the module datetime so both
reads are identical by construction.
sync_turn runs on the MemoryManager worker, but on_session_switch and
shutdown run on the caller thread, so two _write_turns() calls could
snapshot the same pending batch and each replace the whole list (duplicate
append or lost pending turn). A capture lock now covers the
snapshot/write/replace sequence; _write_turns() reads pending turns inside
the lock instead of taking a caller-built list.
Retries are at-least-once: the documents API appends on a shared custom_id
and does not dedupe by content, so a write it accepted but whose response
was lost is appended again. Stated in the docstring and README instead of
implied.
Addresses review on #109359.
A failed flush at on_session_switch() restored the pending buffer and
then cleared it on the next line, so an unavailable service at the switch
boundary still lost every pending turn. Pending turns now carry their own
session_id; _write_turns() batches per session, so a later retry (next
turn, session end, shutdown) writes old-session turns under the old
session's custom_id even after the switch.
Addresses review on #109359.
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).
Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.
Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.
Metadata stays type/session_id/timestamp plus the existing sm_source.
Answers the P1 review on #111187: build_profile_secret_scope() held only
<profile>/.env plus that profile's external-source snapshot, never the
administrator-managed .env. The launch process applies that file LAST with
override (_apply_managed_env), so a managed key beats the user's own value in
os.environ. Under multiplex semantics get_secret() stops falling back to
os.environ on a scope miss, so inside a routed cron fire (and equally inside a
real multiplex gateway turn, which builds its scope through the same function
via gateway/run.py::_load_profile_secret_scope) a managed-only credential
resolved as absent and a managed-vs-user collision resolved to the USER value:
reversed precedence.
Fix at the source: build_profile_secret_scope() overlays load_managed_env()
last, after the profile .env and external sources, skipping process-global
names exactly as it does for the other two layers. Every multiplex-authoritative
scope (gateway turn, routed desktop fire, external worker env build) is built
here, so managed authority is composed once instead of restored per consumer.
No generic ambient-env fallback is reintroduced: only the managed file's own
keys enter the scope, and only with the managed file's values.
Regression (parametrized, two invariants): inside a routed fire a managed-only
key resolves through get_secret(); a managed-vs-user collision yields the
managed value.
The earlier trim dropped the only test of the `routed_profile_fire() and
not is_multiplex_active()` branch in the external-worker handoff: the
process flag is OFF (a desktop tick is not a multiplexer) yet a fire routed
to a sibling profile must serialize `multiplex_active=True` and hand the
worker an env without the launch profile's residue, and the context must
not outlive the handoff span. Without it that branch could regress to the
pre-fix behaviour unnoticed.
The PR shipped nineteen tests across six files, most of them variations of
one boundary. Keep the three that pin distinct behaviour:
- a routed desktop-ticker fire runs under multiplex semantics for exactly
its scope: a scope miss returns None instead of the launch credential and
the parent os.environ is byte-identical afterwards;
- a routed no_agent child never sees a launch-only name, whether the launch
.env defined it or a launch external source supplied it (applied or lost
to a pre-existing process value), while its own values come through;
- administrator-managed keys keep policy precedence over the routed
profile's own value.
Everything else was either a positive control of the same seam, a
set-membership check on a module-level constant, or a re-statement through
a different entry point.
Review findings on f5f88d5058. Three are defects the previous round introduced.
Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.
Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.
Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.
Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.
Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.
(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
Review findings on d8c467f223, each reproduced through its production path.
Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.
Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.
Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).
Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.
(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
The source-name strip added alongside the external-source fix ran
unconditionally. Outside multiplexing there is no other profile to leak from --
os.environ IS this profile's environment -- so popping those names relied on the
routed scope overlay putting each one back, which in turn relies on the
per-home snapshot recorded at boot. Correct today, but it made a single-profile
child's credentials depend on bookkeeping that has nothing to do with isolation.
Guard it the way strip_launch_profile_env guards itself: no multiplexing, no
strip. A single-profile no_agent child now keeps a byte-identical env even if a
source's snapshot were ever missing. Pinned by a regression that runs a real
child with a source-owned name in os.environ and no multiplex context; making
the strip unconditional fails it.
(cherry picked from commit afa429b30a1c9c9b4c011f7d2426a9f099911596)
Two credential-isolation gaps found in review of the previous head.
1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.
2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.
Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.
(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
The routed no_agent child env started from all of os.environ and only overwrote the
names present in the installed scope. A name defined only by the LAUNCH profile's .env
and absent from the routed profile therefore reached the routed child with the launch
value instead of unset -- the secret scrub only knows classified names, so a custom or
unclassified secret crossed the profile boundary (review finding on the first head).
strip_launch_profile_env (main, 284d220ba4) is the primitive the external-worker path
already uses for exactly this: it drops the launch profile's dotenv-owned keys and the
bridged TERMINAL_* settings, and is a no-op outside multiplex or when the target IS the
launch profile. Apply it to the base BEFORE the scope overlay (so a shared name keeps its
routed value) and BEFORE the sanitizer (so routed values still pass the scrub and
passthrough rules). Pinned by a child-process negative control: the launch-only name
arrives <unset>, the shared name arrives routed, and the parent process is unchanged.
(cherry picked from commit 69349b527b5144c1f1540a4c10035d5ec798c1db)
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).
Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).
Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
`_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.
Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.
(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
`/*.db-wal` / `/*.db-shm` only match at the checkout root, so on a flat
install `git stash push --include-untracked` still swept
`cron/executions.db-wal` / `-shm` while the scheduler held the WAL-mode
database open (cron.executions._connect opens it via open_db in WAL mode).
The base file stayed put but its WAL vanished under a live writer, so the
next `_connect()` failed with `disk I/O error` (review finding on #111175).
Use `/cron/executions.db*` like the gateway recovery db rule already does,
covering -wal/-shm/-journal and retired-WAL dirs. Every other non-root db
rule in the block already uses the glob. The regression tuple gains the two
sidecars, and the stash test gets the reviewer's repro against the real
`_stash_local_changes_if_needed`: an open WAL connection with one committed
row must still be readable from a fresh connection afterwards.
The flat-install block listed state.db and kanban.db sidecars one by one,
so any other root-level SQLite store (response_store.db, a future ledger)
and its -wal/-shm/-journal sidecars would still be swept by the updater's
`git stash push --include-untracked` and unlinked under the running
gateway (#110648). Replace the per-file lines with root-anchored globs
(`/*.db`, `/*.db-wal`, ...). No tracked root-level *.db exists, and
`git ls-files -ci --exclude-standard` is unchanged before/after, so the
globs newly ignore nothing that is committed.
Also add the rest of the flat-install runtime roots the previous fold
missed: the credential siblings from
gateway/platforms/base.py::_ROOT_CREDENTIAL_PATHS (.anthropic_oauth.json,
google_token.json, google_oauth_pending.json, auth/,
webhook_subscriptions.json), the active pairing location platforms/
(gateway/pairing.py), kanban/, gateway_state.json, processes.json,
cron.pid, the channel directory/alias and feishu pairing stores,
pending_messages/, checkpoints/, plugin-data/, hooks/, and the Discord
message-recovery db under gateway/ (gateway/ itself is tracked, so only
that file pattern is ignored). `/.credentials/` had no producer -- the
real dir is `credentials/` (web_routers/files.py, _ROOT_CREDENTIAL_PATHS)
-- so it is replaced. `/state-snapshots/` is dropped: the existing
unanchored `*-snapshots/` rule already matches it.
The test tuple now carries one representative per ignored class and its
comment no longer claims _ROOT_CREDENTIAL_PATHS enumerates the sidecar
set (that is `_sqlite_files`).
The same `git stash push --include-untracked` sweep that took state.db
on a flat install (#110648) also takes every other untracked file at the
$HERMES_HOME root: config.yaml, auth.json/auth.lock, memories/,
profiles/, .credentials/, mcp-tokens/ and pairing/. Losing those on a
declined or failed restore strands the user's credentials and profile
config just as badly as losing the session store.
Extend the root-anchored block with those paths (none are tracked or
already ignored on main) and append them to the test's
FLAT_INSTALL_RUNTIME_STATE list so the existing stash invariant covers
them without a new test.
The per-profile job store lives at HERMES_HOME/cron/jobs.json, so a
flat install keeps it beside executions.db inside the checkout-root
stash domain. Without an ignore rule the untracked autostash of
hermes update sweeps it away with the rest of the runtime state.
Absorb the path (noted in #110670) and its regression assertion into
the runtime-state carrier.
(cherry picked from commit 66282e6dc3ac5f276cdee5b92a22856f146c9a45)
On a flat install (checkout root == $HERMES_HOME) the untracked autostash of
`hermes update` sweeps the live state.db/-wal, snapshots, cron ledger and
lock/pid files into the stash and unlinks them under the running gateway; the
restart recreates an empty store at the same path and the declined restore
leaves the profile with no transcripts (#110648).
Root-anchor the runtime state set in .gitignore, mirroring the
.hermes-bootstrap-complete (#38529) and /.install_method (#66189) precedent
and the $HERMES_HOME-root enumeration in the platforms base module, so the
stash step is never entered for runtime state alone. Regression test runs the
exact stash command against a real repo carrying the tracked .gitignore.
(cherry picked from commit 6c74c24e6131af189801e54f1e11b260a529b74a)
Fold the transient/permanent sync-loop tests into one parametrized table
and add the two classifier branches that had no loop-level coverage: a
401 whose body was rewritten to HTML by a reverse proxy (errcode dropped,
so only http_status can stop the loop) and a 429 M_LIMIT_EXCEEDED that
must be retried because neither errcode nor status is an auth signal.
Deleting the http_status fallback in _is_permanent_matrix_auth_error now
fails the 401-html case.
Hoist _sync_error to module level so parametrize can call it directly
instead of the staticmethod.__func__ workaround. Trim the
test_ws_auth_retry docstring, which still described a Matrix test class
that moved to test_matrix.py.
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.
With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).
The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.
A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.
Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.
Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.
(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.
I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.
I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.
While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.
Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.
On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.
(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.
I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.
Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.
(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)