Fixes#88654.
After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:
restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
if restart_mode is None:
continue
A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:
cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'
with no operator signal of any kind.
Two changes.
_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.
When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.
Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.
Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.
Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.
Fixes#91583 (defect 2). Repro and live validation by @kubaboski.
The prefix is an address: the caller is naming WHICH agent the request is
for. With gateway.multiplex_profiles off, _resolve_request_profile ignored
the prefix entirely — "don't 404 a would-be valid route" — so a request
explicitly addressed to one agent was silently answered by a different one.
Observed live (Aug 2026): `hermes peer dm mini/researcher` was answered by
the mini's DEFAULT agent with no error on either side, because that host
runs one LaunchDaemon per profile and only the default daemon hosted an
api_server. A wrong-agent answer is strictly worse than an error: the
sender believes the addressee got the message.
With multiplexing off the process serves exactly one profile, so the prefix
is honored when it names that profile (peers address single-profile daemons
this way without knowing the host's topology — get_active_profile_name() is
the same identity the file already uses for model resolution) and rejected
otherwise through the existing _PROFILE_REJECTED path (404). A process that
cannot resolve its own identity rejects too: if it cannot prove who it is,
it must not answer as anyone.
Unprefixed requests are untouched, and multiplexed hosts are untouched —
the change is confined to the prefix-present, multiplexing-off branch that
previously discarded the caller's addressing.
Widen the DM tempfile-leak fix (#91902/#92407) to the sibling sites
PR #92784 introduced:
- tools/bot_relay.py: expose the 6h stale sweep as
cleanup_bot_relay_artifacts() (cleanup_*_cache contract) and wire it
into gateway housekeeping — previously it ran only when the Desktop
drained the outbox, so plaintext envelopes/replies queued while the
Desktop was away could sit on disk forever.
- tui_gateway/methods_bot_relay.py: move the payload write inside the
try/finally so a failed write no longer leaks hermes-relay-dm-*.txt.
- tools/bot_mode_dm.py: _spawn_delivery takes dm_file=None for relay
waiter deliveries, which have no plaintext DM tempfile to reclaim.
The in-band sweep in _write_dm_file only runs when another DM is
written — a gateway that never sends one keeps orphans forever. Expose
the sweep as cleanup_bot_dm_cache() with the same contract as the other
cleanup_*_cache helpers (returns files removed) and wire it into the
gateway housekeeping loop on the hourly media-cache cadence. Also sweeps
legacy hermes-dm-*.txt and hermes-relay-dm-*.txt orphans in the OS temp
root.
Folded in from #92407 (mehmetkr-31), adapted to the runner-owned
cleanup design salvaged from #91902.
The Electron main has routed a profile with a `profiles.<name>` remote
entry in connection.json to its own pooled backend since
profileRemoteOverride() landed — but the only way to WRITE that entry was
hand-editing connection.json (#91349, design intent from #90223 /
6170f844: this belongs on the profile rail, not the machine-level
Gateways page).
- New "Connect to a remote host…" action on the profile-rail square's
context menu, opening a URL + token dialog that writes the exact
`profiles.<name>` shape through the existing typed
getConnectionConfig/applyConnectionConfig bridge (renderer never
touches connection.json; tokens ride the existing safeStorage
encryption path, with the allowPlainTextToken opt-in surfaced on
keyring-less machines).
- First-time connect shows a one-time confirmation with a plain-language
risk note; editing an existing override (token rotation) skips it.
- Overridden profile squares carry a "remote" globe badge, and the
tooltip/aria label names the host. A "Remove remote connection" button
clears the override (mode: local via the existing coerce path).
- Registry name collision: the dialog warns when the profile name
matches a v2 connections-registry id/label.
- Token rotation: when switching to an overridden profile fails with an
auth-shaped error (401/forbidden/invalid token), a re-enter-token
toast opens the same dialog instead of leaving a silently dead
profile. Connectivity failures stay generic.
- All five locales updated.
Closes#91349
Design-intent analysis credit: @otfnfn
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).
Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.
- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
— bounded, interrupt-safe wait over pending notify_on_complete
sessions; reconciles orphaned-pipe exits (#17327) each pass so a
wedged reader cannot burn the full bound; KeyboardInterrupt aborts
the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
session flush / cleanup (covers -q and -Q, i.e. the DM recipient
shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.
Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.
Fixes#90879
A Desktop per-profile alias (e.g. moxie with a Cloud override) routes to a
remote backend's root profile: route { connectionId, profile: 'moxie',
targetProfile: 'default' }. Once the hosted backend answers the roster
itself, the row's identity is (connection, 'default') — a different key
than the alias meta — so the friendly name regressed to the raw Cloud
hostname after activation, and Cloud-only rosters showed generic 'Hermes'
instead of the configured alias (#89131).
Add a connection-exact alias index built from the credential-free route
inventory, keyed by (connectionId, targetProfile). displayName,
botRosterMeta, and botFriendlyNames consult it so the claimed backend row
reads as the alias (and its title/meta), while:
- same-named defaults on OTHER connections never borrow the identity
- two aliases claiming one backend row fail closed
- the local default and un-aliased remote defaults keep existing behavior
Evidence: @TheAirick's controlled candidate testing on #89131.
Bot Mode always hides canonical 'Bot Chat' sessions, but _find_bot_chat's
GET /api/sessions listing used the default include_hidden=False path, so
the existing hidden row was invisible, _ensure_bot_chat tried to create a
duplicate, and the peer DB's UNIQUE(title) guard rejected it — DM failed.
- api_server: GET /api/sessions now accepts an exact-title lookup
(?title=...) and honors include_hidden=1 ONLY alongside a title filter,
so canonical hidden rows resolve without exposing a blanket hidden
listing on the client surface. The title needle is pushed into SQL
(search_query) so old hidden rows outside the recency window are found.
- peer dm client: _find_bot_chat sends title + include_hidden=1; older
peers ignore the unknown params and degrade to today's behavior.
- Clear diagnosable error on the older-peer duplicate-create rejection,
naming the hidden canonical chat and the PATCH hidden:false workaround.
- Unit tests (hidden resolution, no duplicate create, older-peer error,
older-peer visible fallback) + real-gateway E2E over a real state.db.
Root-cause analysis and regression recipe by @kubaboski in #91583.
Fixes#91583
The salvaged commit called managed_python_env() at the git-path sync
without an in-scope import (UnboundLocalError on every git update — CI
red). The repair test pinned the raw {**os.environ, VIRTUAL_ENV} dict, a
change-detector on exactly the construction #83914 replaces; it now
asserts the managed-env contract.
The salvaged fix covered the git-path sync; the same raw-os.environ
construction existed at the main update path and the interrupted-install
recovery path. All three now build their uv env via managed_python_env()
(#83914 class — same bug, all sites).
A/B-proven with real uv: poisoned UV_PYTHON/UV_SYSTEM_PYTHON steers the
merge-base construction into the hijacker's interpreter (VERDICT:
HIJACKED); the managed construction installs into the install's venv
(VERDICT: ISOLATED). Compose-checked with #92824's stale-VIRTUAL_ENV pin:
isolation + pin together install into the running interpreter on the
site-packages shape.
Address review feedback:
- Add two unit tests asserting the update's uv_env contract: third-party
UV_PYTHON_INSTALL_DIR is dropped, managed pins (UV_MANAGED_PYTHON=1,
UV_NO_CONFIG=1) are set, VIRTUAL_ENV points at this install's venv, and
the managed store stays under .hermes-runtime.
- Drop the inline dated comment in favor of intent description.
uv respects UV_PYTHON_INSTALL_DIR from the process environment. When a
third-party app (e.g. WorkBuddy) sets a User-level UV_PYTHON_INSTALL_DIR,
the update's uv pip install can target the wrong interpreter and fail
installing extras, leaving the venv entry-point shims missing. Use the
official managed_python_env() isolation (drops VIRTUAL_ENV/PYTHONPATH/
UV_PYTHON, forces UV_PYTHON_INSTALL_DIR to .hermes-runtime/python,
UV_NO_CONFIG=1) and then point VIRTUAL_ENV at this install's venv.
CI has no faster-whisper and lazy installs disabled, so the 'local'
provider resolves through a different unavailability branch than a dev
box — the reason string differs but the relay verdict (and no key in the
payload) is the actual contract.
The speak-stream WebSocket resolved its URL through the bare v1
getConnection/getGatewayWsUrl pair, which answers for the PRIMARY backend.
When a registry remote connection rides over a machine that also has a
local Hermes install (the common case — the installer always installs the
full agent), spoken replies dialed the LOCAL backend and hit its
unconfigured TTS, while chat (connectionScoped REST) correctly went remote.
Users saw 'configure STT/TTS' although their remote gateway had voice
fully configured.
Resolve the PCM socket through the same (connectionId, profile) bridges
store/gateway's openSecondary uses, and never overwrite a backend-namespace
profile the registry mint already wrote into the URL (SSH remoteProfile
aliasing, sharedRemote scoping).
Every REST audio call already carried connectionScoped(); this was the one
remaining self-built audio URL. Contract pinned by
voice-playback.routing.test.ts (sabotage-verified: 2/4 fail on the old
resolver).
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.
Backend:
- tools/voice_client_config.py: single resolver; per-provider client
wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
elevenlabs-tts). Server-host-only providers (local whisper, edge,
command/plugin) and missing credentials resolve to {mode: relay}.
xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).
Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
profile) with 60s TTL, provider-direct STT + TTS calls, sentence
cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
below it. Barge-in via the same stopVoicePlayback sequence bump.
Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.
Docs: voice-mode.md client-direct section ships in this PR.
The 85% compaction autoraise exists to stop wasting the small advertised
272K Codex window. -900k large-context picker variants (#92797) run at
~900K, where the global compression.threshold (default 50%, ~450K) is the
right behavior — autoraising them to 85% (~765K) would delay compaction
far past what the user configured.
- _is_codex_gpt54_or_gpt55() excludes valid -900k variants, so both the
85% override and the one-time autoraise notice skip those sessions.
- Base slugs are unchanged: 272K window + 85% autoraise.
- Tests: variant/base threshold pairs incl. namespaced ids; docs note in
the -900k section.
Review pass findings on the force_jpeg change:
- Broaden the JPEG mode guard from {RGBA, P} to 'not in {RGB, L}':
force_jpeg newly routes PNG inputs to the JPEG encoder, and an
LA-mode PNG (grayscale+alpha) would crash img.save() with
'cannot write mode LA as JPEG'.
- browser_use_cli's _native_screenshot_result is the THIRD native
history-embed site: it baked the data URL into a _multimodal tool
result with the 5 MB one-shot default and no dimension cap. Apply
the same 256KB/1568px/force_jpeg history-reuse policy as the two
sites already migrated.
PNG has no quality ladder, so a text-dense screenshot over the 256KB
history-embed cap (#92699 / #92783) could only shrink by halving
dimensions — 1568px dropped to ~784px and on-screen text became
unreadable, the exact fidelity screenshot QA depends on.
Add force_jpeg to _resize_image_for_vision: the two history-embed call
sites (vision_analyze native, browser_vision native) re-encode
resize-needing screenshots as JPEG so the quality ladder (85/70/50)
absorbs the byte pressure and the readable resolution survives.
Under-cap images are untouched and stay PNG; one-shot/reactive paths
keep their existing format behavior.
Flagged during the #92783 salvage review.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
The salvaged _interpreter_scripts_dir hand-rolled the Scripts/bin layout,
which the AST lint-test in test_update_zip_two_phase forbids — route it
through the canonical hermes_constants.venv_bin_dir instead, with the
interpreter's own dir as fallback for non-venv layouts.
Sibling-test blast radius from #92617: the salvaged fixture's fake
_run_quarantined_install predates the strict_quarantine kwarg the
update sync now passes.
- _is_uv_command: detect 'python -m uv'/'python -m uvx' and launcher
wrappers, not just a uv basename (review: naive check missed module form)
- _insert_python_pin: never duplicate a caller-supplied --python (review:
last-wins ambiguity)
- _interpreter_scripts_dir: when pinning to sys.executable on Windows with
no project venv, quarantine the running interpreter's Scripts dir so the
hermes.exe shims uv rewrites are actually unlocked (review: quarantine
path diverged from pinned interpreter)
- tests: rewritten to repo English convention; added python -m uv,
--python-guard and Windows quarantine-target cases (5 total)
When Hermes is installed via pip / site-packages (e.g. the Windows
installer), PROJECT_ROOT is the interpreter's site-packages directory and
PROJECT_ROOT/venv is never created. The update and interrupted-install
recovery paths still set VIRTUAL_ENV=PROJECT_ROOT/venv, so uv fails with
'Failed to inspect Python interpreter from active virtual environment'
before installing anything — leaving the install partially updated.
Detect the nonexistent VIRTUAL_ENV in the shared dependency-install helper
and pin uv to the running interpreter (uv pip install --python
sys.executable) instead, matching the fix already applied to lazy-deps
(#83335) and the ZIP update path (#71510).
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.
Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
Follow-up on the #91828 salvage: regression test for the subtle
resolveDiskPluginEntry branch that rejects a DIRECTORY literally named
plugin.js, and lint fix. The truncation-fix semantics from #92809 are
preserved: an oversize source keeps its error inventory row (returns
true — the file exists and was fully probed), while a vanished file
returns false so the scanner reconciles the ghost.
On Windows the updater renames the live `hermes*.exe` shims aside
(`hermes.exe.old.<unix-ms>`) so uv can write replacements. When that quarantine
succeeds but the install then fails, the recovery path could leave the install
with no `hermes` on PATH at all — unrecoverable in place, because the command
that would repair it IS `hermes update` (#75584).
Restoring a quarantined shim happens at three sites: the updater, the
early-recovery installer, and the startup sweep's orphan rescue. Each was a
single un-retried rename whose OSError was swallowed in silence, while the
OUTBOUND quarantine rename already retried a lock. That is backwards — a failed
quarantine merely aborts an update, a failed restore removes `hermes` from
PATH — and the two sites that had messages had already drifted apart.
- `_early_recovery.restore_quarantined_shims()` is now the single
implementation: retry ladder, one recovery message, returns the pairs it
could not restore. It lives in the stdlib-only module that both `main` and
`_install_repair` already import, so the layers cannot drift again. A pair is
not a failure when the original reappeared or the quarantine file vanished —
two processes sweeping the same orphan must not produce a spurious error.
- `_cleanup_quarantined_exes` unlinked every `*.exe.old.*` on each invocation.
When the original shim was already missing, that .old file was the ONLY
surviving copy — deleting it converted a one-rename recovery into a full
reinstall. It now rescues the orphan through the shared helper instead, and
leaves anything inside a 15-minute grace window alone so it cannot destroy a
concurrent update's in-flight quarantine.
- Ordering is by the PARSED `.old.<unix-ms>` stamp, not the raw filename.
Lexicographic ordering only tracks recency while every stamp shares a digit
width; a stray `.old.999` sorts above a 13-digit epoch-ms stamp and would be
the copy rescued onto the live shim name.
- Names whose suffix does not parse as int-ms are not ours: never rescued,
never deleted. The sweep should not destroy files whose provenance it cannot
establish, and they are not produced by the quarantiner.
The stamp is read from the filename rather than st_mtime because `rename`
preserves the original shim's mtime, which records when uv wrote the shim —
days earlier, in general — not when it was quarantined. A regression test pins
that distinction.
Messages go to stderr: the sweep runs on EVERY hermes invocation and
`hermes acp` speaks JSON-RPC on stdout.
Scope note: `_quarantine_running_hermes_exe` is deliberately byte-identical to
main here. Why the outbound rename fails in the first place (the launcher
holding its own image without FILE_SHARE_DELETE) is #88121's subject; this is
the net underneath, covering the case where quarantine SUCCEEDS and the install
dies afterwards. The two touch disjoint functions and can merge in either order.
Reproduced and verified on Windows 11 (26200), Python 3.11.15: stranded the
shims, confirmed a normal `hermes` invocation now rescues the orphan instead of
deleting it, and confirmed an exhausted rescue prints the recovery command.
16 new tests; 35 pass across the four quarantine suites.
Runtime desktop plugins read their source through hermes:readFileText,
the preview IPC that silently truncates at TEXT_PREVIEW_MAX_BYTES
(512 KiB) — a larger plugin.js evaluated as a partial file (cryptic
syntax error, or worse, a half-module that parses).
- electron/main.ts: dedicated hermes:readPluginSource handler — full
read, 16 MiB cap enforced as a hard EFBIG via resolveReadableFileForIpc
(same path hardening + sensitive-file blocking), never truncation.
- preload.ts / global.d.ts: bridge + types (optional — older shells).
- runtime-loader.ts: loadDiskPlugin reads via readPluginSource; on older
shells falls back to readFileText but FAILS LOUDLY on truncated:true
(error toast + error inventory row) instead of evaluating a partial
file. Existence probe keeps the preview read (metadata is enough).
- tests: full-read path, loud old-shell truncation failure (sabotage-
verified), small-plugin fallback.
CPython interns identifier-like string literals, so 'is' cannot
distinguish an import alias from a copy-pasted literal (verified:
two exec'd namespaces each defining the literal share one object).
The equality assertions three lines above are the full honest guard.
Also reword a comment: raw == is marker-SENSITIVE, not asymmetric.
Reviewer point on #92539: nothing documented that an in-place content
mutation of a stamped dict must pop _DB_PERSISTED_MARKER (and invalidate
the bounded flush-scan prefix) or the DB silently goes stale. Both
existing mutators (turn_finalizer fill-empty-tail, context_compressor
micro-compaction defrag) already follow the contract; this states it at
the constant so the next one does too.