DPR-sized canvas backing store separated from CSS footprint, tracking
zoom/display changes live. Rendering-fix subset of PR #75307; the
overlay-placement feature portion is out of scope here.
Covers the devicePixelRatio half of #83216.
All three artifact timestamp sources (message.timestamp,
session.last_active, session.started_at) are epoch SECONDS — the
transcript reader and session-date-groups both multiply by 1000 — but
the collector passed them straight to new Date() (ms), so every
artifact rendered as 1970-01-21. Normalize seconds to ms once at
collection; the Date.now() fallback stays ms.
Local file artifacts (e.g. D:\ComfyUI\output\*.png) fell through to
mediaExternalUrl() which yields a file:// URL the renderer cannot
load. Route through the desktop fs bridge whenever it exists —
readDesktopFileDataUrl already dispatches remote REST vs local
Electron internally (#83380).
Require explicit provenance for tool-result artifacts while preserving assistant links, MEDIA deliveries, generated outputs, file mutations, and browser screenshots. Normalize persisted Unix-second timestamps at collection time and retain millisecond fallbacks.
Consolidates current-main-compatible work from #41156 and #48577.
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Co-authored-by: tt-a1i <53142663+tt-a1i@users.noreply.github.com>
The Artifacts view reads message.timestamp, session.last_active, and
session.started_at from the SQLite database, which stores all timestamps
as Unix epoch seconds (REAL). These values were passed directly to
JavaScript's Date() constructor, which expects milliseconds — causing
every artifact timestamp to display as January 1970 dates.
Fix by multiplying the database value by 1000 to convert seconds to
milliseconds at the storage point, using nullish coalescing (??) instead
of logical OR (||) so that valid zero timestamps are not skipped.
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)
Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.
- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
incl. the compression-lineage recursive CTE); list_sessions_rich gains
include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
from every listing path (and the REST sidebar endpoints inherit it with no
change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
set_session_hidden; _session_response exposes it.
Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).
* fix: teach lost-and-found recovery about the 55-column sessions layout
Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
cron.manage resolved its jobs store from the process HERMES_HOME, so a profile
whose cron lives in ~/.hermes/profiles/<name>/cron/ was invisible to the default
gateway (and any bot/plugin querying per-profile routines saw 'no cron jobs').
Add an optional 'profile' param that scopes the whole action via
set_hermes_home_override, exactly mirroring the adjacent skills.manage handler:
resolve get_profile_dir(profile), 404 (err 4064) if missing, override in a
try/finally that always reset_hermes_home_override. Omitted/None keeps the
launch-profile behavior, so existing callers are unaffected. cronjob() itself is
unchanged (it already keys off HERMES_HOME).
Enables the Hermes-Bot-Mode plugin to show a bot's real routines
(NousResearch/Hermes-Bot-Mode#37). Needs a SERVE-backend gateway restart to take
effect live. 2/2 in the new focused test.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
Sweeper follow-up: the browser-screenshot and video kwargs captures now
also assert max_tokens is absent, protecting the central auxiliary
no-cap policy against refactors that would restore the hardcoded caps.
Covers the max-tokens-knob contract: vision call_kwargs omit max_tokens
entirely (configured values, defaults, and even an explicit
auxiliary.vision.max_tokens config entry must never be forwarded), so
providers use their full output budget.
The vision tools' call_kwargs hardcode max_tokens caps (2000 for
vision_analyze/browser_vision, 4000 for video analysis), truncating
descriptions of complex images at the cap. The centralized aux client
already omits max_tokens by default (#34845) so providers use their
model max output; these three call sites were the leftovers that
bypassed that policy.
Remove the hardcoded caps entirely — the aux client handles the
mandatory-max_tokens Anthropic wire via _resolve_anthropic_messages_max_tokens
(model output ceiling) and Gemini native omits maxOutputTokens (65K ceiling),
so no wire needs an explicit cap.
delegation.max_concurrent_children caps how many delegated children run in
parallel per batch (and concurrent background delegation units). The old default
of 3 needlessly serialized independent fan-outs (e.g. reviewing/investigating N
PRs or issues at once), so large batches ran in slow chunks of 3.
Raise the shipped default to 10, which sits at/below the existing high-cost
advisory threshold (>10), so the default never trips the warning. Each child
still consumes API tokens independently, so this is a throughput/latency win the
user pays for in parallel token spend — the floor stays 1 and there is no
ceiling, so anyone can tune it down or up.
- config_defaults.py: default 3 -> 10; _config_version 36 -> 37.
- delegate_tool.py: _DEFAULT_MAX_CONCURRENT_CHILDREN 3 -> 10 (+ docstring).
- config_migrations.py: _migrate_to_37 lifts configs pinned at exactly the old
default 3 to 10 (deliberate non-3 overrides preserved; unset inherits 10).
- cli-config.yaml.example: documented default updated.
Verified: default/fallback read 10, version 37, and the migration lifts 3->10,
preserves an explicit 5, and leaves unset untouched.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
Self-review follow-up: cover call_converse_stream's max_tokens=None path
(same builder, previously unpinned) and document why the shim reads the
caller cap with truthiness rather than 'is None' (parity with the
Anthropic shim's reading).
The Bedrock Converse shim hardcoded 'else 4096' when the caller passed no
max_tokens, so auxiliary vision descriptions stayed capped at 4096 tokens
on the Bedrock wire even after #75253 removed the vision call sites' own
caps (#10809 was only partially fixed there).
Converse's inferenceConfig.maxTokens is optional; when omitted, Bedrock
defaults to the model's maximum allowed output. Thread an explicit
max_tokens=None through build_converse_kwargs/call_converse to omit the
field, and drop an all-empty inferenceConfig from the wire request
entirely. The 4096 default is unchanged for every existing caller (main
transport passes params.get('max_tokens', 4096) explicitly), so only
no-cap aux calls opt in.
Surfaced during review of #75253.
On Windows _get_service_pids() is empty (no systemd/launchd query), so a
Scheduled-Task-supervised gateway whose gateway.pid record is missing or
stale is invisible to both the service-PID and recorded-PID exclusions the
reaper already applies (#86658) — and gets SIGTERM'd on every desktop open
(#86098 class, pidfile-less path).
Add a Windows-only backstop: any reaper candidate whose parent chain
reaches services.exe (the Task Scheduler launches tasks under the services
tree) is spared even with no pidfile.
The backstop is deliberately inert on POSIX: every process there has PID 1
(launchd/init/systemd) in its ancestry — and a genuine orphan is reparented
directly to PID 1 — so supervisor-name ancestry carries zero supervision
signal and would disable the reaper entirely on macOS/WSL (#51325, #75936).
POSIX supervised gateways are already covered pidfile-independently by the
_get_service_pids() exclusion.
Known limitation (fail-open, documented): if the Task-launched bootstrap
parent has already exited, Windows does not reparent the gateway, the chain
breaks before services.exe, and the gateway is treated as an orphan.
Salvaged from #86702 by @EvanProgramming (authorship preserved); reduced to
the genuinely-new Windows backstop — the PR's other two hunks were already
merged on main via #86658 (one in a strictly stronger full-parent-chain
form) and its POSIX ancestry checks were dropped as unsound (verified
empirically: a true double-fork orphan's psutil parent IS launchd).
Review fixes from #86679 comments (trevorgordon981, helix4u, kshitijk4poor):
- Edit inheritance: mergeConnectionInput preserves fields the editor does
not carry (cloud org, ssh remoteHermesPath/remoteProfile) so a rename no
longer wipes them. When the payload carries an ssh host string, stored
user/port are NOT inherited — the composite host field is authoritative,
fixing the stale user/port resurrection on edit.
- Token hygiene: tokens only persist on token-auth remotes; switching an
entry to oauth (or cloud) clears the stale envelope.
- Plain-text opt-in: the panel now surfaces the same consent dialog as
Settings -> Gateway on keyring-less machines (registry list exposes
secureTokenStorage; save retries with allowPlainTextToken after consent).
- Registry test isolation: hermes:connections:test builds the probe directly
from the registry entry instead of coercing against v1 connection.json —
no more inheriting the v1 global token for a different host, and the local
entry now probes the app-managed backend (never v1 remote/ssh state, so
the test button can no longer trigger a v1 file write).
- 'local' id reserved at the validation boundary: a crafted IPC payload can
no longer replace the local entry via upsert.
- Cloud creation hidden in the editor (a dialable cloud entry comes from the
Cloud sign-in/discovery flow); migrated cloud entries stay editable.
- First-run migration write is guarded: a failed write keeps the migrated
registry in memory instead of hard-failing every connections IPC call.
- uniqueLabel(): single label-dedup helper — counts up instead of "X 2 2",
clamps 253-char migrated URL-host labels under LABEL_MAX; used by
normalizeRegistry and both migration paths.
- UI copy: staged-rollout note replaces the "side by side" claim; test
failure toast leads with the failure wording; dropped unused i18n keys.
Tests: +9 pure-module cases (reserved id, token-drop rules, merge
inheritance, ssh host precedence, uniqueLabel); electron+settings suites
1355 passed.
First slice of multi-source agent support: the desktop can now persist ANY
number of named backends (local runtime, remote gateways, Hermes Cloud
instances, SSH hosts) side by side instead of one global connection plus
per-profile overrides.
- electron/connection-registry.ts: pure v2 registry module — required
case-insensitively-unique labels (device names), @name-device handle rule
for duplicate profile names across sources (agentHandle), defensive
normalizeRegistry for corrupt files, one-time v1→v2 migration that imports
the global block + per-profile overrides (deduped by URL/host) and leaves
connection.json untouched for older builds.
- main.ts: connections.json storage beside connection.json (same secret
posture: safeStorage-encrypted tokens, 0600, tighten-before-parse, mtime
cache) + hermes:connections:* IPC (list/save/remove/set-primary/test).
Test maps registry entries onto the existing testDesktopConnectionConfig
probe stack — no new probe code.
- Settings → Connections: manage the registry (add/edit/remove/test/make
primary) with forced naming; local entry is non-removable; removing the
primary retargets to local. en + zh locales.
Storage-level only by design: routing/pool generalization to composite
(connection, profile) keys, the multi-source roster, plugin SDK surface, and
fan-out updates land as follow-up PRs.
/simplify-code finding: turn_context evaluated resolve_prompt_cache_scope()
inside set_runtime_main's argument list under the umbrella try/except — a
resolution failure would silently skip the ENTIRE runtime binding
(provider/model/base_url/api_key/session_id for all aux calls that turn),
not just the cache scope.
- prompt_cache_scope: add resolve_prompt_cache_scope_safe() (never raises,
returns None on failure/empty).
- turn_context: resolve the scope into a local via the safe variant BEFORE
the set_runtime_main call, so a failure can only lose the scope.
- chat_completion_helpers: _prompt_cache_scope_for_agent delegates to the
shared safe variant (guarded import retained).
- tests: +1 (hostile-property agent -> None; normal/empty passthrough).
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
_session_db re-resolves instead of staying pinned to the physical id);
_persist_disabled agents (background-review forks that never get a DB row)
memoize the fallback instead of re-querying the lineage per API call;
module docstring cross-references get_conversation_root and why the two
lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
_prompt_cache_scope_for_agent(agent) call to a single local above the
OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
Legacy compaction mode (compression.in_place: false) rotates the physical
session_id mid-conversation. The prompt-cache scope introduced in #79161 was
derived from that physical id, so every rotation moved the same conversation
into a fresh cache bucket - the prompt cache went cold at every rotation
boundary (#79017).
Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT
of the current session (SessionDB.get_compression_lineage, fork-aware
post-#79193) - once per turn, memoized per transcript segment, and prefer it
over the physical session_id at every prompt_cache_key derivation site:
- agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) -
lineage-root walk with per-segment memo; falls back to the physical id
when no DB is attached or the walk fails, degrading to pre-fix behavior.
- transports/codex.py: build_kwargs accepts cache_scope_id and prefers it
for the body prompt_cache_key, the xAI x-grok-conv-id header, and the
Codex x-client-request-id routing header. The Codex session_id header
keeps the raw physical id (transcript identity, #57012 contract).
- transports/chat_completions.py: _add_prompt_cache_key accepts
cache_scope_id with the same precedence.
- chat_completion_helpers.py: build_api_kwargs threads the resolved scope
into all three build_kwargs call sites (codex, profile, legacy).
- auxiliary_client.py: set_runtime_main carries cache_scope; the aux
Responses cache-key site prefers it over the physical session_id.
- turn_context.py: resolves the scope once per turn and threads it through
set_runtime_main (no DB walk on the per-API-call hot path).
Scope semantics preserved from #79161: /new starts a fresh scope (new
lineage), /branch children, delegate subagents, and tool children stay
isolated (explicit-fork exclusion in get_compression_lineage), unrelated
sessions keep distinct buckets, and cron per-fire timestamps still
normalize via _cache_scope_from_session_id.
Default installs compact in place (session_id never rotates), so they hit
the memo and produce byte-identical keys to before.
Fixes#79017
Follow-up to the salvaged #76252 addressing both review gaps:
- New TestMinorLineFallForward class with a direct test of the
explicit-patch fallback branch: bare '3.12' resolves to a VULNERABLE
build while an explicit 3.12.x patch is fixed, so recovery must go
through _list_available_patches on the next minor line. Asserts the
exact `uv python install` request sequence.
- New all-minors-exhausted test: everything vulnerable on 3.11-3.13
returns None with per-line attempts bounded by _MAX_PATCH_RETRIES and
no requests beyond 3.13 (requires-python is <3.14).
- test_retry_is_bounded_by_max_retries_constant now actually uses its
counting wrapper and asserts the collected install calls (the
previous version collected them into a dead variable).
Also dedupes the fallback loop the same way the same-minor loop does:
_attempt_install_generation can now record the probed candidate version
into a caller-supplied tried_versions set, so the explicit-patch pass
skips the version the bare-minor request already resolved to and
rejected -- previously that wasted a full download+install+probe+delete
cycle per minor line re-trying a known-vulnerable build.
When every patch on the current minor line (e.g. 3.11) still links a
vulnerable SQLite (e.g. 3.50.4 on Windows), the provisioner now tries
the next supported minor line (3.12, then 3.13) before giving up.
Previously, _install_safe_python_generation only tried patches within
the same minor line. On Windows, where python-build-standalone may not
publish a fixed build for the installed patch, users were stuck with a
repeated warning on every `hermes update` with no path forward.
The requires-python constraint (>=3.11,<3.14) and the downstream
import smoke test already gate compatibility, so the minor-line
upgrade is safe.
Adds allow_minor_upgrade parameter to _attempt_install_generation to
relax the same-minor-line version guard when called from the fallback
path.
Fixes#76106
The guard's "use a separate worktree or temporary clone" advice sent
agents to /tmp by default. /tmp is RAM-backed tmpfs on most distros, and
parallel salvage clones each running npm ci (~1.6GB per clone) filled a
32GB tmpfs to 97% during a 15-subagent campaign, ENOSPC-ing sibling test
runs. The message now recommends `git clone --shared <root> ~/.hermes/scratch/<task>`
(honoring HERMES_HOME), warns that dependency installs belong on real
disk, and tells the agent to delete the clone once the branch is pushed.
Install-Uv piped the astral installer's entire output to Out-Null, so any
real failure (proxy block, AV quarantine, permissions) surfaced only as the
generic "uv installed but not found" message, and astral.sh was the sole
install source even though corporate proxies commonly block it while the
byte-identical GitHub releases installer downloads fine.
Three-rung ladder, all inside Install-Uv:
1. astral.sh installer with output captured via Tee-Object.
2. GitHub releases installer mirror (same UV_INSTALL_DIR).
3. Salvage an existing uv.exe (Get-Command uv, or the astral default
%USERPROFILE%\.local\bin\uv.exe) by copying it into $HermesHome\bin so
the managed-first invariant holds.
On total failure, print the last 15 lines of captured installer output plus
the existing manual-install pointer.
Reported by @BitBernd; proxy diagnosis by @gakugaku; Out-Null suppression
first identified by @webtecnica in #69366.
Closes#69216
The composer's "Use skill: X" pill only checked the draft text and the
workspace-name collision - it happily re-offered a skill the session had
already loaded (skill_view), edited (skill_manage), or that the user had
invoked via its /name command. Clicking it would re-inject the full
SKILL.md into a context that already carries it.
The draft provider now scans the session transcript (per-runtime
$sessionStates mirror, $messages for the active session) for
skill_view/skill_manage tool calls naming the skill - exact or qualified
(category/name, plugin:name), from parsed args or hydrated argsText -
and for user turns starting with the skill's slash command, and stands
down on a hit. The scan only runs when the draft actually matched a
skill, so ordinary typing pays nothing.
Three fixes for generated/displayed images in the desktop chat:
- Shell fallback context menu no longer swallows right-clicks on images:
the guard now yields to Electron's native image menu (Copy Image, Copy
Image Address, Save Image As...) for img/picture/video/canvas targets,
matching the existing editable/selection carve-outs.
- Save Image As / download button: generated-image URLs (fal.media etc.)
end in an extensionless content hash, so saves produced an unopenable
"All Files" blob. The main-process save dialog and the renderer anchor
fallback both now append a MIME-derived extension, add image type
filters, and default to the user's Downloads directory instead of the
process cwd (win-unpacked on packaged Windows installs).
- New will-download handler routes any Chromium-initiated download
through the same Downloads-dir + guaranteed-extension policy.
Validation: new unit tests for the filename derivation (6 passing);
npm run check:lint green (tsc x3 + eslint, 0 errors).
`connect()` caches every path it has initialized in the process-local
`_INITIALIZED_PATHS` set and then skips all first-open work for it —
header validation, integrity probe, `SCHEMA_SQL`, additive migrations.
That cache is keyed on a path, but the schema it stands for lives in a
file, and the two can drift apart: delete or replace `kanban.db` under a
live gateway/dispatcher/dashboard process and the next `connect()` takes
the fast path, lets SQLite create a fresh empty database, and hands back
a connection with no tables in it.
Nothing notices. Every query then fails with `no such table: tasks`,
`plugin_api._conn()` logs its init warning and carries on, and the board
renders empty. Because the cache entry survives, the process re-creates
the same schema-less ~4 KB file on every restart of the desktop app in
front of it — only killing the backing process clears it.
Verify the sentinel table on the fast path and self-heal when it is gone:
drop the stale cache entry and fall through to the existing init path,
which re-runs the probes and the schema script under the cross-process
init lock. The check is one `sqlite_master` lookup on the already-resident
page 1, so the steady-state path stays lock-free (#36644) and does no
schema work.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The unconditional 2s join before InterruptedError delayed interrupt
detection when Relay managed execution was not active (CI:
tests/run_agent/test_interrupt_propagation.py — detection took 2.34s
against a <1.0s budget, because the mocked worker sleeps 5s and there
is no Relay scope to unwind).
Extract the join into _join_worker_for_relay_teardown(), which no-ops
unless a Relay runtime exists AND managed execution consumers are
registered — the only case where an orphaned physical scope can corrupt
the LIFO stack (#81521). Applied at all three interrupt sites
(streaming, non-streaming, Bedrock streaming). The regression test now
simulates a live runtime so the join path stays covered.
The salvaged _close_scope_handle replaced direct scope.pop calls that
were bounded by _SCOPE_OP_TIMEOUT with an unbounded run_in_session
callback, regressing the bounded-finalization contract (CI:
tests/agent/test_relay_runtime_bounded_scope_ops.py — end_turn /
close_session / finish_logical_calls hung when the native pop wedged).
Pass timeout=_SCOPE_OP_TIMEOUT so the whole drain+close costs at most
one span and never blocks turn or session completion.
Follow-up to HexLab98's salvaged commits:
- Apply the same bounded worker join before raising InterruptedError at
the two sibling interrupt sites that share the raise-without-join
shape: the non-streaming API poll loop and the Bedrock streaming poll
loop. Both workers run Relay-managed physical attempts, so raising
immediately allowed turn teardown to race a still-open physical scope
exactly as in the streaming path.
- Address the #81601 review finding (egilewski): the pinned nemo-relay
binding's get_scope_stack() returns a native ScopeStack object which
scope.pop rejects with TypeError, so the orphan drain never drained
under the real binding. current_top() now prefers the version-correct
scope.get_handle() accessor and falls back to the old list-unwrap for
fake/legacy shapes. Handle comparisons go through same_handle(),
comparing by uuid, because native ScopeHandle instances do not
implement value equality.
- Add a real-binding regression test that reproduces the orphaned-scope
session close against the pinned native wheel (skips where the native
binding is unavailable), alongside the existing fake-based coverage.
Empty-stream stalls that trip interrupt were raising InterruptedError
before the stream worker closed its physical LLM scope, corrupting the
Relay LIFO stack and cascading into a CLI EIO redraw storm (#81521).
Replace raw substring checks ("azure.com" in full_url) with the existing
base_url_host_matches() helper at two sites in runtime_provider.py.
The substring approach misclassified URLs whose path (not hostname)
contained "azure.com" — e.g. https://example.invalid/proxy/azure.com/v1 —
causing the wrong credential (Azure key instead of explicit Anthropic token)
to be selected, and potentially leaking a more-privileged Azure key across
a trust boundary.
base_url_host_matches() parses the URL and validates only the hostname
against allowed Azure suffixes with proper boundary rules.
Fixes#74312
_to_openai_base_url() matched ZAI (open.bigmodel.cn, api.z.ai, bare
"bigmodel") and Kimi (api.kimi.com) via `substring in url`, so any custom
gateway whose base_url happened to contain one of those strings as a path
segment (e.g. a reverse-proxy prefix like /proxy/bigmodel-fallback/) was
silently misrouted to the wrong OpenAI-wire endpoint shape.
This is the same false-positive class 6f33f510e8 just fixed for the
MiniMax branch in the same function by switching to base_url_host_matches()
(hostname-anchored). Apply the same fix to the ZAI and Kimi branches, which
that commit didn't touch. Drops the bare "bigmodel" substring check since
open.bigmodel.cn is the only canonical bigmodel-family host referenced
anywhere else in the codebase (agent/model_metadata.py, hermes_cli/auth.py).
Added regression tests mirroring the MiniMax marker-in-path tests added in
the same commit.
The gateway auto-restart phase used to swallow every exception at debug
level, so tests driving cmd_update end-to-end never noticed it touching
real gateway discovery. With #78574 surfacing an aborted restart as a
failed update, an unmocked find_gateway_pids on a box with a live
gateway hits the conftest live-system guard and turns into a spurious
sys.exit(1).
Add an autouse fixture in test_cmd_update.py (discovery returns nothing,
systemd unsupported) and the same seams in test_update_head_moved_gate's
helper so the phase is a clean no-op for tests that do not assert on
gateway restarts.
Reviewer egilewski found the original defer was circular (#83590 comment):
the self-lock preflight wrote .update-incomplete and exited, but the next
launch only ran the full recovery AFTER main.py's third-party imports —
so a healthy venv's probes made the early pass a no-op, main.py imported
cryptography eagerly, the .pyd got mapped again, and the deferred install
re-hit the exact self-lock it was meant to escape.
Close the loop by making the marker guarantee the install runs BEFORE any
native extension can be imported:
- hermes_cli/_install_repair.py (new, stdlib-only): single source of truth
for the core .[all] reinstall — ensurepip bootstrap, uv-pip/pip
resolution with VIRTUAL_ENV, Termux env stripping, Windows hermes*.exe
quarantine, per-extra fallback ladder, and fd1→fd2 routing for acp
safety. Deliberately free of managed_uv/hermes_constants imports so it
stays importable in the corrupted-venv state it exists to repair.
- hermes_cli/_early_recovery.py: recover_if_needed now completes a pending
.update-incomplete install BEFORE the import probes, on every launch
that sees the marker (unless argv is update). Success clears the
marker; failure bumps an attempts counter inside the marker body and
keeps it. A 3-attempt ceiling stops a persistently-failing install
from reinstall-hammering every launch (hermes acp included) — past the
ceiling the late post-import recovery takes over with its manual
recovery instructions. Single-flight lock shared with the late path.
- hermes_cli/main.py: _recover_core_update_marker_locked delegates the
install to the shared executor (no duplicated logic); ensure_uv stays
in the late path so a venv whose uv vanished mid-update still
bootstraps it.
- tests: 7 new regressions — the reviewer's exact case (marker + healthy
venv → install runs while sys.modules has no cryptography), failure
keeps marker + increments attempts, retry ceiling, lazy marker does
not trigger core install (#58004 invariant), argv-update skip, and
corrupt/missing marker bodies. The key test was sabotage-verified:
removing the pre-import branch makes it fail with zero install calls,
while a lone-lazy-marker test still passes; restoring the branch makes
it pass again.
Refs #83569
Two gaps left every Windows git-checkout install unable to recover from
the exact failure state #83569 reports:
1. Self-lock detection. _detect_venv_python_processes() always excludes
the calling process by design — a CLI hermes update IS the venv python.
An updater that had already imported a native venv extension (the
canonical one being cryptography.hazmat.bindings._rust, mapped while
hermes_cli.main resolved external secret sources) passed every
preflight and then died mid-sync with os error 5 when uv tried to
rewrite the mapped .pyd, stranding the venv half-updated. A new
preflight now refuses the sync before touching the checkout, writes
the update-incomplete marker so the next fresh launch completes the
install, and exits 2. Verified on a live Windows 11 host: after
importing hermes_cli.main, tasklist /m _rust.pyd shows the .pyd mapped
in the caller, and a peer process cannot open it read-write
(Permission denied) — while a rename succeeds, matching how uv/pip
actually fail (truncate+write, not rename).
2. Early-recovery install path. _early_recovery._run_repair_install used
sys.executable -m pip unconditionally. Windows git checkouts install
on a uv-managed base interpreter (python-build-standalone), whose
EXTERNALLY-MANAGED marker makes plain pip abort with
externally-managed-environment — the repair no-oped and the venv
stayed broken. The repair now detects the PEP 668 marker, prefers
uv pip install with VIRTUAL_ENV pointed at the project venv, and
falls back to pip --break-system-packages when no uv binary exists.
Both fixes ship with subprocess/unit regressions (sabotage-verified):
the new tests fail on pre-fix code and pass with it. Complements #77517,
which keeps the updater from importing cryptography in the first place;
this PR is the defence-in-depth when any future path loads it anyway.
Fixes#83569
_cold_start_windows_gateway_after_update() printed the success line off a
successful Popen return alone, which only proves CreateProcess succeeded,
not that the child survived. On Windows, a job object denying
CREATE_BREAKAWAY_FROM_JOB hard-kills the child during updater teardown
before it logs anything, yet the updater still printed "Starting Windows
gateway after update (PID ...)" — leaving Telegram/Discord/etc. offline
with no indication anything failed (#84185).
Route the success report through gateway_windows._report_gateway_start(),
the same post-spawn liveness poll every other _spawn_detached() caller
already uses, so a dead child is reported as a failure with a
manual-recovery hint instead of a false success.
Review follow-up (#78574): the aborted-restart handler only flagged the fleet
stale when the post-failure survivor probe was None or non-empty. A positive
empty probe was treated as proof-of-safety — but `[]` is only safe when
nothing was running before the phase. If a gateway was discovered, stopped
(SIGTERM/drain), and its replacement never came back, the probe is empty at
exactly that unsafe moment and the update reported success — the fail-open
contract this fix exists to close.
Snapshot the pre-restart gateway PIDs before any stop/drain and route the
handler decision through a pure _restart_phase_failure_is_incomplete() helper
that fails closed on an empty survivor set whenever a gateway existed
pre-restart (or the pre-state could not be read). Add decision-level regression
tests covering the stopped-without-replacement gap, unknown pre-state, and the
truly-no-gateway positive control.