get_default_hermes_root() resolves HERMES_HOME against the platform
native home (~80us of path resolution) on EVERY call and is called at
31+ sites — every _load_global_auth_store() (per provider row in the
/model picker), kanban, backup, gateway, update. Its result depends
only on (HERMES_HOME, native home), so memoise it keyed on those two
inputs, compared for free on each call (freshness-correct even if a
test or plugin mutates HERMES_HOME mid-process).
_load_global_auth_store() re-read + re-parsed the global auth.json on
every call; read_credential_pool() -> load_pool() runs it once per
provider row in the /model picker even when the profile has entries and
the global fallback never fires. Memoise keyed on the global auth
file's path+mtime (same pattern as _nous_auth_status_cache); the store
only changes when a global-scope auth write touches the file.
Measured (profile mode, 30-provider global store): get_default_hermes_root
81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us ->
137us per call — ~2ms saved per /model picker render (20 provider rows).
Regression tests: hermes_constants memo pin (no path resolution on
repeat calls, HERMES_HOME change forces a fresh resolution); global-store
memo pins (store read once across repeats, mtime bump re-reads once,
absent store stays cheap).
(cherry picked from commit be348f32e5bd7479c26fabb652de549fb9c8a1e1)
The /tools and /personality completers run on every keystroke while the
user types those commands (complete_while_typing), and both re-read +
re-parse the full config on each keypress:
- _tools_completions called load_config() — the defensive deepcopy
(~340us/call on cache hit) even though it only reads toolset enable
state + MCP server names. Switched to load_config_readonly() (the
perf(agent) #74322 pattern; this per-keystroke site was missed).
- _personality_completions called load_cli_config() — a full YAML parse
+ deep merge of the built-in defaults (~110us) — on every keystroke.
Memoised keyed on the config file path+mtime (same pattern as load_env
/ _nous_auth_status_cache), so the parse runs once per config state.
Measured: /tools 357us -> 18us per keystroke; /personality parse drops
from 1-per-keystroke to 1-per-config-change (500 keystrokes -> 1 parse).
Regression tests: _tools_completions uses the readonly loader (deepcopy
loader never called); personality memo parses once across repeated
completions and re-parses once after a config mtime change.
(cherry picked from commit 2b3f897171f93dc6b099848ec5ab3763b7294870)
The main.py decomposition re-exported the sessions/update/dashboard command
surface with eager from-imports, so every hermes invocation (including
hermes --version) paid for update_cmd's dependency chain (jwt, click,
cryptography). Resolve the re-exports through the existing PEP 562 module
__getattr__ (same pattern as _PROVIDER_MODELS) so each module loads on
first actual use. Internal call sites go through a _self() helper because
bare-name lookups do not trigger __getattr__; _self() imports sys locally
since update tests patch hermes_cli.main.sys. The
_warn_stale_dashboard_processes back-compat alias moves into the lazy
surface, and the sessions argparse dispatch defers sessions_cmd to call
time. Monkeypatching hermes_cli.main.<name> keeps working: a patch sets a
real module attribute, which shadows __getattr__.
Measured (Windows 11, Python 3.11, median of 7 warm runs):
import hermes_cli.main 253ms -> 196ms (-22%).
(cherry picked from commit cad1083b71635f98815b698d69a18f7f58e15517)
Read the durable display transcript when creating a branch instead of copying the compacted model projection. Hydrate the Desktop branch boundary from persisted history, avoid stale whole-chat counts, and seed the new tile from the backend snapshot. Add regression coverage for compacted histories, visible-message counts, selected prefixes, and hydration races.
(cherry picked from commit c3d2d759ae104395adde215b2e293ccf8e895684)
CLI-origin sessions have no session_key; the Desktop history-truncation
path called replace_messages(session["session_key"], ...) with None,
whose reinsert violated the messages.session_id FK -> "FOREIGN KEY
constraint failed" -> "Restore failed" on resume. Key the persist off
the durable session id instead.
Extracted from PR #81904 (the scope=compacted API half was superseded by
include_compacted, #86595). The PR's companion change defaulting
session_key to the session id at insert time is deliberately NOT taken:
main treats a non-NULL session_key as "this is a gateway session"
(list_gateway_sessions, orphan gateway-session repair), so the default
would misclassify every CLI session.
(extracted from PR #81904, commit cef9b9b27d)
When the 1h provider_models_cache.json TTL lapses, the model picker
serially fetches /v1/models for each authenticated provider. With 10+
providers this stacks to 15-30s of blocking before the picker renders.
Add a parallel prefetch step before the serial picker build loops:
- _collect_authed_provider_slugs(): lightweight credential pre-scan
that mirrors sections 1/2/2b without fetching model lists
- _prefetch_provider_models_parallel(): ThreadPoolExecutor-based
concurrent fetch of stale/missing cache entries (max 8 workers)
- update_provider_cache_entry(): thread-safe single-entry cache writer
with threading.Lock to prevent concurrent write races
Guardrails:
- Skipped when <=3 authed providers (overhead not worth it)
- Skipped when refresh=True (serial path force-refreshes)
- Exception-isolated (falls back to serial path on any failure)
- No behavioral change (same model lists, same picker output)
Closes#80413
(cherry picked from commit 89dddd6cb5d53d73278e0518c375fb5b878e5c6b)
_create_request_anthropic_client() built a fresh anthropic.Anthropic
client (and httpx pool) on every single LLM call, and
_close_request_anthropic_client() always fully closed it right after
- unlike the OpenAI-wire path, which caches and reuses one warm
client across sequential calls via a single-slot cache keyed on the
effective client kwargs.
Add the same single-slot cache to the Anthropic-wire path: keyed on
credentials, base URL/Bedrock region, per-model timeout, and the
1M-beta flag; in_use guards concurrent calls from sharing one pool's
close/abort lifecycle; poisoned marks a cross-thread-aborted slot so
the owner-thread close discards it; reuse only on request_complete /
stream_request_complete (the same _REQUEST_CLIENT_REUSE_REASONS the
OpenAI path already uses). Wires a teardown hook into
release_clients()/close() mirroring _close_cached_request_openai_client.
Fixes #HPA-02
(cherry picked from commit 37f90df15593e6ded0390f827f8e5604bf0acc86)
Every search route in _search_messages_impl (FTS, CJK bigram, trigram,
LIKE fallback, rebuild-gap supplement) selected m.content, then the
result tail popped it unread. On DBs with multi-MB tool rows, each
search read and materialized up to `limit` full rows only to discard
them. Snippets come from snippet()/substr() in SQL and the context
window is re-fetched by id, so no code path ever read the column.
Drop the column from all six SELECT lists. Returned dicts are
unchanged: content was never part of the public result (the pop ran
before return), and tests/test_hermes_state.py already documents that
contract.
(cherry picked from commit d0c3af167e7dd4eb18e1bea29ba107a94911ea24)
Hermes Console registers `checkpoints prune`, `clear` and `clear-legacy` as
mutating, so it takes a console-level confirmation before dispatching any of
them. `_apply_confirmed_defaults` then exists to keep the CLI layer from
asking a second time — its docstring says so — but it only force-defaults
`clear` and `clear-legacy`. `prune` was left out, even though `cmd_prune`
gates its orphan preview on the identical `not args.force` shape.
`_capture_output` redirects stdout and stderr but never stdin, so the
unskipped `_confirm()` call hits `input()` with no terminal behind it:
`EOFError` propagates into `_confirm`, which returns False, and `cmd_prune`
prints "Aborted." and returns 1. The console turns that non-zero exit into a
ConsoleCommandError, so `checkpoints prune` fails outright for any user who
has at least one orphan checkpoint project — after that user already
confirmed. When the server does happen to inherit a foreground terminal, the
same call instead blocks a console worker thread and eats the operator's
keystrokes.
Forcing the flag is the documented behavior here rather than a weakening of
the recent orphan-allowlist hardening. `orphan_allowlist` binds a deletion to
the identities shown in the preview, guarding the window where a workdir
disappears while the command waits on `input()`. Under the console there is
no preview and no wait, which is exactly the `--force` case the comment on
`cmd_prune` describes as "no restriction".
Migrates 8 of 12 native dialog call sites in the kanban dashboard plugin
to the SDK's ConfirmDialog primitive (added in PR #50550):
- moveTask, moveSelected, applyBulk, deleteTask, deleteSelected,
archiveBoard, removeAttachment, doPatch
The 4 remaining carve-outs (window.prompt for completion summary,
window.alert for missing summary, cli_hint clipboard fallback) are
documented inline — the host's ConfirmDialog hardcodes onClick → unmount,
preventing the keep-open-across-validation behavior the completion-summary
form needs. Followup: upstream a `disabled` prop to ConfirmDialog and
rebuild the completion body using host Dialog components.
New architecture:
- useKanbanDialogs(t) — Promise-based dialog state machine at
KanbanPage scope. request({kind, ...}) returns {confirmed, summary?}.
- KanbanDialog component — renders ConfirmDialog from SDK for kind=confirm.
- performMoveTask(taskId, newStatus, count, summary) — extracted shared
dispatch path for single + bulk moves (optimistic UI + PATCH/POST +
error recovery).
- requestDialog prop threading — KanbanPage → BoardSwitcher,
TaskDrawer → TaskDetail → doPatch/AttachmentsSection. Every call
site has a defensive fallback to window.confirm if the prop is
missing (verified by test_dashboard_done_actions_prompt_for_completion_summary
counting the cancel guards + destructive:true markers in the bundle).
New host i18n keys (web/src/i18n/en.ts + types.ts):
- kanban.confirmDoneMany / confirmArchiveMany / confirmBlockedMany
- kanban.trash.confirmTitle / confirmManyTitle
Tests:
- Replaced bundle-string-only completion-summary test with behavioral
coverage: bundle cancel-guard count + destructive marker count, plus
backend tests that confirm cancel preserves old status and confirm
dispatches the expected PATCH/DELETE body.
- Removed the SDK_CONTRACT_VERSION snapshot test from
web/src/plugins/registry.test.ts (forbidden by AGENTS.md
"Don't write change-detector tests"; the two remaining tests in that
file already cover the new SDK surface behaviorally).
Closes#50547 (consumers of #50550).
Cross-vendor re-review: Gemini 3.5 Flash + GPT-OSS 120B (both SHOULD-FIX,
no remaining BLOCKERs after these fixes).
The salvaged #76716 adds git_metadata_generation to sessions (54 -> 55
columns). Update the synthetic-rebuild test's pinned widths and row
builders to the new current layout.
Creating a project whose resolved primary path already belongs to a
non-archived project now raises a clear ValueError naming the existing
project (create_project) — duplicated projects each seeded an identical
copy of the repo subtree, multiplying the duplicate-lane bug per copy.
The agent-facing project_create tool is idempotent instead: it re-activates
the existing project rather than erroring. allow_duplicate_path=True keeps
deliberate duplicates possible. Also updates the legacy non-git lane-id
expectation to the branch-style id introduced for #53329.
_build_repos emptied each lane's sessions array for the overview (hydrate=False)
payload BEFORE _sort_lanes ran. _lane_sort_key derives a lane's activity from
max(_session_time(s) for s in group["sessions"]), so with the rows already gone
every non-trunk lane scored activity=0.0 and the sort key
(is_trunk, is_kanban, -activity, label) collapsed to alphabetical-by-label. The
documented intent — branches and linked worktrees sort by most-recent activity,
then label — was silently defeated on the projects.tree RPC that feeds the
desktop sidebar overview, while the drill-in path (hydrate=True) kept the rows
and sorted correctly. Any repo with two or more non-trunk lanes showed a
different order in the overview than when opened.
Move the session-clearing to after _sort_lanes/_disambiguate_labels so the sort
reads real recency. Lane counts are still captured before clearing, so
sessionCount and the slim overview payload are unchanged — only the order is
fixed, and the overview now matches the drill-in.
_place_by_heuristic used the raw path as the lane key for non-git
project folders, while the desktop overlay independently computed
::branch::main for the same session (since git_branch was null).
The ID mismatch caused duplicate lanes — one from the backend with
the folder name, one from the overlay labeled 'main'.
Use _branch_lane_id(path, DEFAULT_BRANCH_LABEL) so the backend's
lane key matches the overlay's expected ::branch::main scheme,
eliminating the duplicate lane.
A gateway active-session lease is acquired against the root HERMES_HOME,
but release_active_session()/transfer_active_session() re-resolved the
registry path from the *current* HERMES_HOME. Under native multiplex a
routed turn runs agent cleanup inside _profile_runtime_scope, so the
release looked under the named profile while the root entry stayed
alive — after max_concurrent_sessions routed turns every new session was
rejected with 'Hermes is at the active session limit' (#85431).
Pin state/lock paths on the lease at acquisition time and prefer them on
release and transfer. Fixes#85431.
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
generate_base_drafts hardens every base draft to a transparent PNG with
_harden_transparency. When the provider returned a non-PNG file (webp,
jpg, or gif), the hardened PNG is saved under a new path and the original
draft is left in cache/images. Nothing prunes that directory outside the
gateway housekeeping loop, so a CLI or desktop draft round leaks one
original per non-PNG draft.
Remove the original after a successful hardening when the output path
differs from the input. PNG inputs (including mixed-case suffixes like
.PNG) are hardened in place so a case-insensitive filesystem cannot
treat with_suffix(".png") as a different file and unlink the output.
Each hatch generates one row strip per state into cache/images and
extracts the animation frames from it, but never removes the strip. The
only cleanup for that directory runs in the gateway housekeeping loop,
which a CLI, desktop, or cron hatch never starts, so the strips
accumulate for good.
Drop the strip after every attempt once its frames are decoded into
memory, including failed or retried attempts, so a hatch no longer grows
the image cache without bound.
The orphan reaper had two gaps that let agent-browser daemons accumulate
indefinitely inside a single long-lived hermes process:
1. `_reap_orphaned_browser_sessions()` ran exactly once, before the cleanup
loop started, so a leak appearing after boot could never be recovered.
2. `owner_alive is True` skipped unconditionally. In-memory session tracking
is lost on any exception path between spawn and registration, but the
owner PID stays up — so such a daemon was skipped forever.
The daemon-side `AGENT_BROWSER_IDLE_TIMEOUT_MS` is not a backstop for (2):
it does not fire when the daemon itself is wedged, e.g. after Chrome's
framework was replaced underneath it by an auto-update.
Observed on macOS: five agent-browser daemons (96 Chrome processes) built up
over 10 days inside an 18-day-uptime hermes process, holding roughly 5 CPU
cores busy and driving the load average past 100. Four of those processes
were still running a Chrome framework version that had since been replaced
on disk, spinning at ~85% CPU each.
Changes:
- Re-run the reaper every `BROWSER_ORPHAN_REAP_INTERVAL` (300s) from inside
the cleanup loop. Cycle 0 preserves the existing startup reap.
- When the owner is alive but the session is untracked, fall back to idle
age: reap past `BROWSER_ORPHAN_GRACE_SECONDS`, defined as
`max(1h, 20 x inactivity_timeout)`. Unknown age fails safe.
- Add `_socket_dir_idle_seconds()` — the newest mtime under a session's
socket dir. Every browser command writes `_stdout_<cmd>` / `_stderr_<cmd>`
there, making it a last-activity marker that survives hermes restarts and
does not depend on in-memory bookkeeping surviving an exception path. It
scans directory entries rather than reading the directory mtime alone:
command names repeat, and rewriting an existing `_stdout_click` updates
that file's mtime but not the directory's, so a dir-mtime-only check would
report a busy session as idle and reap it.
Sessions still present in `_active_sessions` are never touched at any age,
and the new path still goes through `_verify_reapable_browser_daemon`, so
the anti-spoof / anti-PID-recycle guarantees from #14073 are unchanged.
Adds 9 tests: idle-age unit tests (including the dir-mtime regression),
spared/reaped/fail-safe cases for a live owner, the identity-guard gate on
the new path, and a periodic-reap test asserting more than one reap per
cleanup-thread lifetime.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
pet.info accepts knownRevision; when it matches the active sheet's
revision the multi-MB spritesheetBase64 is elided and
spritesheetUnchanged=true is returned. The desktop floating pet passes
the revision it already holds and keeps its cached bytes, so backstop
refreshes no longer resend ~3.2MB frames over the WS (write-loop stalls,
disconnect storms). Legacy callers omitting knownRevision get the full
payload unchanged.
Review follow-up (final round PASS with a repeated suggestion): the pets
CLI regression test only exercised real bools, leaving the exact bug this
fix shipped untested. Add a quoted-'false' test driving _has_active_pet
(now False) and toggle_pet_display (now takes the ENABLE branch,
distinguished by the 'no pets installed' error instead of err=None from
the old wrong-way disable). Verified RED on the pre-fix pets.py and GREEN
on the fix.
Three bare bool() reads of display.pet.enabled (deep-merged config, so a
hand-edited quoted YAML value lands as the string 'false'): the pet.cells
gate, the pet.gallery enabled echo, and the shared pet-state helper.
bool('false') is True, so a quoted value kept the mascot enabled against
the operator's explicit intent.
All three now go through utils.is_truthy_value (default False, matching
DEFAULT_CONFIG). Regression test drives pet.gallery with a quoted 'false'
config and asserts enabled=False; verified RED on the old code.
* feat(sessions): generic 'hidden' session flag (sidebar-hide, still resumable)
Adds a source-orthogonal, archive-orthogonal 'hidden' session flag meaning
'don't show in the global Sessions sidebar, but stay fully resumable by the
surface that owns it'. Mirrors the existing archived/pinned capability end to
end, so it's a generic widening (any plugin that owns its own session lifecycle
- kanban, Bot Mode, future plugins - can keep its sessions out of the shared
recents list) rather than a per-plugin special-case.
- Schema: hidden INTEGER NOT NULL DEFAULT 0 on sessions (additive; lands on
existing DBs via the declarative _reconcile_columns ADD COLUMN path, same as
archived/pinned - no version-gated migration).
- DB: SessionDB.set_session_hidden(session_id, hidden) (clones set_session_pinned
incl. the compression-lineage recursive CTE); list_sessions_rich gains
include_hidden=False, appending 's.hidden = 0' by default so hidden rows drop
from every listing path (and the REST sidebar endpoints inherit it with no
change).
- Gateway: session.set_hidden RPC (mirrors session.title); session.create accepts
hidden=true, deferred via pending_hidden and applied in _ensure_session_db_row
when the row is lazily created (mirrors pending_title).
- REST parity: PATCH /api/sessions/{id} accepts+bool-validates 'hidden' ->
set_session_hidden; _session_response exposes it.
Enables Hermes-Bot-Mode to hide canonical 'Bot Chat' sessions from the sidebar
(NousResearch/Hermes-Bot-Mode#46) WITHOUT retagging source (which would mis-set
the agent platform). Bot Chats keep source=desktop. Gateway RPC needs a
SERVE-backend restart to take effect live. 1 focused test (default-exclude /
include_hidden / unhide round-trip).
* fix: teach lost-and-found recovery about the 55-column sessions layout
Adding the 'hidden' column makes the current sessions table 55 columns. The
SQLite lost-and-found recovery classifier keys off the physical field count
(SESSIONS_LAYOUT_NFIELDS) to identify a salvaged sessions row, so a recovered
current-layout row (nfield=55) would otherwise be unrecognized and dropped.
Add 55 to the frozenset (54/52 stay as historical prefixes) and update the
column-count assertions + synthetic current-layout insert in the recovery test.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
cron.manage resolved its jobs store from the process HERMES_HOME, so a profile
whose cron lives in ~/.hermes/profiles/<name>/cron/ was invisible to the default
gateway (and any bot/plugin querying per-profile routines saw 'no cron jobs').
Add an optional 'profile' param that scopes the whole action via
set_hermes_home_override, exactly mirroring the adjacent skills.manage handler:
resolve get_profile_dir(profile), 404 (err 4064) if missing, override in a
try/finally that always reset_hermes_home_override. Omitted/None keeps the
launch-profile behavior, so existing callers are unaffected. cronjob() itself is
unchanged (it already keys off HERMES_HOME).
Enables the Hermes-Bot-Mode plugin to show a bot's real routines
(NousResearch/Hermes-Bot-Mode#37). Needs a SERVE-backend gateway restart to take
effect live. 2/2 in the new focused test.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
Sweeper follow-up: the browser-screenshot and video kwargs captures now
also assert max_tokens is absent, protecting the central auxiliary
no-cap policy against refactors that would restore the hardcoded caps.
Covers the max-tokens-knob contract: vision call_kwargs omit max_tokens
entirely (configured values, defaults, and even an explicit
auxiliary.vision.max_tokens config entry must never be forwarded), so
providers use their full output budget.
Self-review follow-up: cover call_converse_stream's max_tokens=None path
(same builder, previously unpinned) and document why the shim reads the
caller cap with truthiness rather than 'is None' (parity with the
Anthropic shim's reading).
The Bedrock Converse shim hardcoded 'else 4096' when the caller passed no
max_tokens, so auxiliary vision descriptions stayed capped at 4096 tokens
on the Bedrock wire even after #75253 removed the vision call sites' own
caps (#10809 was only partially fixed there).
Converse's inferenceConfig.maxTokens is optional; when omitted, Bedrock
defaults to the model's maximum allowed output. Thread an explicit
max_tokens=None through build_converse_kwargs/call_converse to omit the
field, and drop an all-empty inferenceConfig from the wire request
entirely. The 4096 default is unchanged for every existing caller (main
transport passes params.get('max_tokens', 4096) explicitly), so only
no-cap aux calls opt in.
Surfaced during review of #75253.
On Windows _get_service_pids() is empty (no systemd/launchd query), so a
Scheduled-Task-supervised gateway whose gateway.pid record is missing or
stale is invisible to both the service-PID and recorded-PID exclusions the
reaper already applies (#86658) — and gets SIGTERM'd on every desktop open
(#86098 class, pidfile-less path).
Add a Windows-only backstop: any reaper candidate whose parent chain
reaches services.exe (the Task Scheduler launches tasks under the services
tree) is spared even with no pidfile.
The backstop is deliberately inert on POSIX: every process there has PID 1
(launchd/init/systemd) in its ancestry — and a genuine orphan is reparented
directly to PID 1 — so supervisor-name ancestry carries zero supervision
signal and would disable the reaper entirely on macOS/WSL (#51325, #75936).
POSIX supervised gateways are already covered pidfile-independently by the
_get_service_pids() exclusion.
Known limitation (fail-open, documented): if the Task-launched bootstrap
parent has already exited, Windows does not reparent the gateway, the chain
breaks before services.exe, and the gateway is treated as an orphan.
Salvaged from #86702 by @EvanProgramming (authorship preserved); reduced to
the genuinely-new Windows backstop — the PR's other two hunks were already
merged on main via #86658 (one in a strictly stronger full-parent-chain
form) and its POSIX ancestry checks were dropped as unsound (verified
empirically: a true double-fork orphan's psutil parent IS launchd).
/simplify-code finding: turn_context evaluated resolve_prompt_cache_scope()
inside set_runtime_main's argument list under the umbrella try/except — a
resolution failure would silently skip the ENTIRE runtime binding
(provider/model/base_url/api_key/session_id for all aux calls that turn),
not just the cache scope.
- prompt_cache_scope: add resolve_prompt_cache_scope_safe() (never raises,
returns None on failure/empty).
- turn_context: resolve the scope into a local via the safe variant BEFORE
the set_runtime_main call, so a failure can only lose the scope.
- chat_completion_helpers: _prompt_cache_scope_for_agent delegates to the
shared safe variant (guarded import retained).
- tests: +1 (hostile-property agent -> None; normal/empty passthrough).
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
_session_db re-resolves instead of staying pinned to the physical id);
_persist_disabled agents (background-review forks that never get a DB row)
memoize the fallback instead of re-querying the lineage per API call;
module docstring cross-references get_conversation_root and why the two
lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
_prompt_cache_scope_for_agent(agent) call to a single local above the
OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
Legacy compaction mode (compression.in_place: false) rotates the physical
session_id mid-conversation. The prompt-cache scope introduced in #79161 was
derived from that physical id, so every rotation moved the same conversation
into a fresh cache bucket - the prompt cache went cold at every rotation
boundary (#79017).
Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT
of the current session (SessionDB.get_compression_lineage, fork-aware
post-#79193) - once per turn, memoized per transcript segment, and prefer it
over the physical session_id at every prompt_cache_key derivation site:
- agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) -
lineage-root walk with per-segment memo; falls back to the physical id
when no DB is attached or the walk fails, degrading to pre-fix behavior.
- transports/codex.py: build_kwargs accepts cache_scope_id and prefers it
for the body prompt_cache_key, the xAI x-grok-conv-id header, and the
Codex x-client-request-id routing header. The Codex session_id header
keeps the raw physical id (transcript identity, #57012 contract).
- transports/chat_completions.py: _add_prompt_cache_key accepts
cache_scope_id with the same precedence.
- chat_completion_helpers.py: build_api_kwargs threads the resolved scope
into all three build_kwargs call sites (codex, profile, legacy).
- auxiliary_client.py: set_runtime_main carries cache_scope; the aux
Responses cache-key site prefers it over the physical session_id.
- turn_context.py: resolves the scope once per turn and threads it through
set_runtime_main (no DB walk on the per-API-call hot path).
Scope semantics preserved from #79161: /new starts a fresh scope (new
lineage), /branch children, delegate subagents, and tool children stay
isolated (explicit-fork exclusion in get_compression_lineage), unrelated
sessions keep distinct buckets, and cron per-fire timestamps still
normalize via _cache_scope_from_session_id.
Default installs compact in place (session_id never rotates), so they hit
the memo and produce byte-identical keys to before.
Fixes#79017
Follow-up to the salvaged #76252 addressing both review gaps:
- New TestMinorLineFallForward class with a direct test of the
explicit-patch fallback branch: bare '3.12' resolves to a VULNERABLE
build while an explicit 3.12.x patch is fixed, so recovery must go
through _list_available_patches on the next minor line. Asserts the
exact `uv python install` request sequence.
- New all-minors-exhausted test: everything vulnerable on 3.11-3.13
returns None with per-line attempts bounded by _MAX_PATCH_RETRIES and
no requests beyond 3.13 (requires-python is <3.14).
- test_retry_is_bounded_by_max_retries_constant now actually uses its
counting wrapper and asserts the collected install calls (the
previous version collected them into a dead variable).
Also dedupes the fallback loop the same way the same-minor loop does:
_attempt_install_generation can now record the probed candidate version
into a caller-supplied tried_versions set, so the explicit-patch pass
skips the version the bare-minor request already resolved to and
rejected -- previously that wasted a full download+install+probe+delete
cycle per minor line re-trying a known-vulnerable build.
When every patch on the current minor line (e.g. 3.11) still links a
vulnerable SQLite (e.g. 3.50.4 on Windows), the provisioner now tries
the next supported minor line (3.12, then 3.13) before giving up.
Previously, _install_safe_python_generation only tried patches within
the same minor line. On Windows, where python-build-standalone may not
publish a fixed build for the installed patch, users were stuck with a
repeated warning on every `hermes update` with no path forward.
The requires-python constraint (>=3.11,<3.14) and the downstream
import smoke test already gate compatibility, so the minor-line
upgrade is safe.
Adds allow_minor_upgrade parameter to _attempt_install_generation to
relax the same-minor-line version guard when called from the fallback
path.
Fixes#76106
The guard's "use a separate worktree or temporary clone" advice sent
agents to /tmp by default. /tmp is RAM-backed tmpfs on most distros, and
parallel salvage clones each running npm ci (~1.6GB per clone) filled a
32GB tmpfs to 97% during a 15-subagent campaign, ENOSPC-ing sibling test
runs. The message now recommends `git clone --shared <root> ~/.hermes/scratch/<task>`
(honoring HERMES_HOME), warns that dependency installs belong on real
disk, and tells the agent to delete the clone once the branch is pushed.
Install-Uv piped the astral installer's entire output to Out-Null, so any
real failure (proxy block, AV quarantine, permissions) surfaced only as the
generic "uv installed but not found" message, and astral.sh was the sole
install source even though corporate proxies commonly block it while the
byte-identical GitHub releases installer downloads fine.
Three-rung ladder, all inside Install-Uv:
1. astral.sh installer with output captured via Tee-Object.
2. GitHub releases installer mirror (same UV_INSTALL_DIR).
3. Salvage an existing uv.exe (Get-Command uv, or the astral default
%USERPROFILE%\.local\bin\uv.exe) by copying it into $HermesHome\bin so
the managed-first invariant holds.
On total failure, print the last 15 lines of captured installer output plus
the existing manual-install pointer.
Reported by @BitBernd; proxy diagnosis by @gakugaku; Out-Null suppression
first identified by @webtecnica in #69366.
Closes#69216