- tests/agent/test_session_activity.py asserts against
ACTIVITY_DESCRIPTION_MAX instead of the literal 120.
- The session-stall WARNING log line names its config knob
(agent.session_stall_timeout) so operators can find the setting.
- hermes_state.py: collapse the triple blank line near line 191.
- hermes_cli/status.py no longer imports the private
hermes_cli.main._relative_time: the helper moved to a public home
(hermes_cli.timefmt.relative_time); main._relative_time stays as a
thin back-compat wrapper (sessions_cmd and external patchers keep
working).
- Config docs now describe session_stall_timeout precisely: a RECOVERY
notifier for an in-process AIAgent with an adapter-queued follow-up —
not a general gateway/session stall detector — with a per-AIAgent scan
cadence (not globally coordinated per durable session).
- import_sessions documents the deliberate export-includes /
import-resets asymmetry for the activity fields (no resurrected
'working' labels on machines where no agent runs), with a regression
pinning both halves.
- Strip trailing whitespace in contributors/emails/fangliquan@qq.com
(git diff --check housekeeping).
PR #76354 review, scope/contract items + housekeeping.
The post-begin_commit() waiter previously called unbounded future.result(),
so the advertised compression.context_total_ceiling_seconds was silently
unenforced for commit-phase hangs. The commit still must complete (abandoning
an in-flight SessionDB mutation would diverge live messages from durable
state), but the wait is now bounded in increments against the remaining
ceiling: on ceiling breach the overrun is logged (WARNING escalating to
ERROR), surfaced once through the user-visible warning channel via the new
on_commit_overrun callback (wired to _emit_warning in run_agent.py), and the
host keeps waiting in bounded slices until the commit finishes.
Documented guarantee (config comment + docs, en/zh): summary phase bounded
by the ceiling; commit phase logged + surfaced if it exceeds it — never
silently hung, never abandoned mid-commit.
Test updated to assert the surfacing fires (previously accepted a silent
over-ceiling wait); adds coverage that a raising overrun callback cannot
break the commit wait.
Once begin_commit() wins, SessionDB mutation cannot be fence-cancelled;
document that context_total_ceiling_seconds covers the summary phase only
and pin the hang-wait contract in tests.
Three code-reuse fixes applied during salvage:
1. Reuse _relative_time from hermes_cli/main.py instead of duplicating
the relative-time formatting logic in hermes_cli/status.py.
2. Extract _stamp_hygiene_compression_provenance helper in gateway/run.py
to deduplicate the two nearly-identical try/except blocks that stamp
compression timeout/abort provenance in the hygiene path.
3. Add ContextCompressor.record_timeout_failure() method and use it from
the in-agent compress_context timeout callback instead of re-implementing
the (60, 300, 900) cooldown ladder inline. The existing summary-LLM
exception handler already has this ladder — now both paths share one
method.
Three mechanisms to detect and notify when gateway sessions stall silently:
1. Mid-turn activity heartbeats stamped to SessionDB so hermes sessions list
and hermes status show progress during long turns without new message rows.
2. Stall watchdog: when a busy session has pending inbound and the shared
activity clock is idle past agent.session_stall_timeout (default 300),
log a WARNING and notify the user once to try /new. Notify-only; does
not kill the turn.
3. Compaction timeout: fenceless compress_context callers get a progress-aware
host budget (compression.context_timeout_seconds default 120 idle,
compression.context_total_ceiling_seconds default 600 ceiling). On timeout,
cancel via commit fence, skip compaction without dropping messages, and
continue the turn.
Closes#72016 (slices 1-3; slice 4 cumulative SSE stream-retry deadline
remains a follow-up).
Cherry-picked from PR #72424 by @fangliquanflq.
The a2a client tools are registered unconditionally by the plugin, but a
newly-registered plugin toolset defaults to ENABLED for every platform until
the user has seen it in 'hermes tools'. That force-injected 'a2a' into every
agent's enabled_toolsets, leaking 3 tool schemas to all users and breaking
tests that assert exact toolset membership
(test_api_server_toolset::test_create_agent_respects_config_override).
Add 'a2a' to _DEFAULT_OFF_TOOLSETS so it stays opt-in (user enables via
'hermes tools'), matching the spotify precedent. The inbound platform
adapter is already opt-in (only instantiated when the a2a platform is
enabled); this aligns the outbound client tools with the same posture.
Builds on the adapter list_channels() hook (cherry-picked from #43545 by
@Guoen0):
- plugins/platforms/simplex: implement list_channels() — enumerates
contacts (/contacts) and groups (/groups) over the live daemon
WebSocket into the channel directory. Returns None when the WS is
down so the directory falls back to session discovery instead of
wiping known targets.
- hermes send --list: merge configured-but-undiscovered platforms into
the listing. Previously a platform configured only via env (e.g. a
fresh SimpleX setup used for outbound sends) was silently omitted,
leaving users guessing at platform names.
- format_directory_for_display(): accept an explicit platforms view and
render empty platforms with a targeting hint instead of hiding them.
- docs: simplex hermes-send section.
Reported by Fedpostoffice on Discord (simplex missing from
hermes send --list; guessed platform names simplex-chat/simplex-relay).
The inverse of the inbound webhook platform: hooks.outbound in
config.yaml lists HTTP targets + the plugin-hook events they subscribe
to (on_session_end, subagent_stop, post_tool_call, ...). Each firing
POSTs a JSON payload (same top-level shape as shell hooks' stdin wire)
signed GitHub-style with HMAC-SHA256 (X-Hermes-Signature-256).
Rides the existing hook bus — notify-only callbacks registered on the
plugin manager at the same CLI/gateway/main entry points as shell
hooks. Delivery is fire-and-forget via a bounded queue + single daemon
worker thread, so a dead endpoint can never stall a tool call. Bounded
retries (5xx/conn errors once; 4xx never). secret_env preferred over
inline secret. HERMES_SAFE_MODE skips registration. hermes hooks list
shows outbound targets with signed/UNSIGNED status.
Zero new model tools, zero new subsystems.
_format_price_per_mtok collapsed any per-Mtok price below one cent to
"$0.00" (and readers treated near-zero as free). Nous Portal's DeepSeek
V4 Flash 0731 promo prices cache hits at $0.0018/Mtok, which rendered as
free in the /model picker and hermes model listings.
Prices under $0.01/Mtok now widen precision to the first significant
digit plus one, with trailing zeros trimmed: 0.0000000018/tok →
$0.0018. Standard prices keep the aligned two-decimal format; exact
zero still renders as "free".
- Pass extra_exclude={pid} to _reap_unsupervised_gateway_orphans so the
killed PID isn't double-killed during the sweep (#75936).
- Add extra_exclude param to _reap_unsupervised_gateway_orphans signature.
- Replace bare except:pass with logger.debug for diagnosability.
- Fix existing test (mock _reap_unsupervised_gateway_orphans so it
doesn't scan real processes and trigger conftest live-system guard).
- Add regression test asserting the killed PID is excluded from the sweep.
stop_profile_gateway() only killed the single PID recorded in the pid
file. On repeated restarts, the new process overwrites the pid file
before the old one exits, making older gateway instances invisible
to subsequent stops. Each restart stacked another orphan — observed
as N threads and N replies per Discord message.
After killing the recorded PID, also call _reap_unsupervised_gateway_orphans()
to sweep any remaining gateway processes for this profile. The reap function
already handles the no-systemd guard and SIGTERM+SIGKILL escalation.
Fixes#75936
/api/status is the desktop's boot liveness probe (polled ~1/s) but since
#60537 every call ran a full topology scan — per-profile yaml.safe_load
(pure-Python loader), psutil process probes, realpath walks — in the
default executor. On multi-profile installs concurrent polls pile up and
hold the GIL 14-16s, starving the event loop: the WS sidecar cannot
flush gateway.ready, the desktop times out into the next stall, and boot
escalates to the 'Hermes couldn't start' overlay (#60800).
Memoize the scan behind a 10s TTL with a collapse lock so concurrent
polls share one scan. Topology only changes on gateway start/stop, so a
<=10s stale badge is an acceptable trade for not starving the loop. The
cache also keys on the collector's identity: tests monkeypatch
_collect_profile_gateway_topology per case, and the identity check keeps
them hermetic (a swapped collector is a miss) without a reset hook.
py-spy captures during a failing boot land in _profile_platform_ports ->
yaml.safe_load on executor threads (7 profiles, Windows). After: one
cold-start scan, zero recurring stalls, desktop boots.
Efficiency-pass follow-up on the #66355 salvage: force=True bypassed the
cooldown entirely, and AIAgent.close() fires a forced trim for EVERY
in-process child subagent close (delegate_tool child.close(), parent
close step 5). A delegate batch of N children closing back-to-back in
the gateway process stacked N+1 uncooled full gc.collect()+malloc_trim
passes (50-500ms each with a large live heap). Forced trims now honor a
5s floor — bursts coalesce, the parent's final close-trim still fires.
Guard test mutation-checked (floor zeroed -> test fails).
Simplify-pass follow-up on the #66355 salvage:
1. _config_settings runs on EVERY trim attempt (before the cooldown
check) and only reads — swap load_config for load_config_readonly.
Deep-copying the whole config per attempt generates exactly the
allocator garbage this module exists to release. Tests re-seamed.
2. Trim-failure logs demoted warning->debug at all 3 periodic sites
(gateway housekeeping, idle reaper, slash worker): sibling failure
branches in the same loops log at debug, and a persistent failure
(e.g. broken import after a partial update) would otherwise warn
every 60s forever.
3. The frame-inspection test now asserts the expected locals exist
before reading them — a rename in _run_prompt_submit fails the test
loudly instead of vacuously passing on None.
Add config-driven glibc malloc_trim for long-lived Hermes processes:
- hermes_cli/mem_trim.py: trim_memory() with configurable cooldown,
RSS snapshot telemetry, and forced-trim INFO logging
- gateway/run.py: periodic trim in gateway housekeeping loop
- tui_gateway/server.py: trim in idle reaper (~every 5 min)
- tui_gateway/slash_worker.py: trim on turn boundary
- run_agent.py: force trim on agent close
- hermes_cli/config.py: context.memory_trim config section
(enabled, cooldown_seconds, log_every_n, info_log_min_delta_mb)
CSA tier-4 reviewed (4 rounds, 0 HIGH/MEDIUM/CRITICAL remaining).
Supersedes PR #63708 + #64591 with enhanced telemetry and gateway/slash_worker coverage.
NOTION_API_KEY/LINEAR_API_KEY defaults, DEEPINFRA_API_KEY model gate, FAL_KEY via the file's own _scoped_credential helper, XAI/VERCEL/DAYTONA presence checks, GITHUB_TOKEN/GH_TOKEN + GitHub App creds in tirith/skills_hub, and the masked HERMES_API_KEY display in tui_gateway config.show all route through get_secret.
The interactive /model picker probes the current custom endpoint live via
fetch_api_models(), which defaults to a 5s timeout. A slow or flaky custom
endpoint blocks the picker for up to 5s on open. The lmstudio picker probe
already uses a 1.5s timeout for exactly this reason; apply the same fast-fail
bound to the three custom-provider probe sites, gated on for_picker so the
non-picker (5s) path is unchanged.
Follow-ups on the startup-burst memo:
- Populate the memo on the valid-token fast path as well. The startup
burst usually finds a VALID token, and each check_fn call still paid
two cross-process file locks + state reads to reach that return; the
original memo only engaged after a refresh. The token has at least
refresh_skew_seconds (>=120s) of life at that return, so a 5s memo can
never serve an expired token.
- Clear the module-level memo in test_nous_portal_staging_allowlist's
refresh-capture helper: with the fast-path populate, a token memoized
by an earlier test would otherwise short-circuit the refresh these
tests assert on (3 tests failed without this).
- Add dedicated memo behavior tests (TTL hit, TTL expiry, insecure
bypass) — the original PR shipped none. Mutation-checked: all 3 fail
against main's un-memoized function, pass on this branch.
check_tool_availability runs once per managed-tool check_fn (browser,
image_gen, etc.) during banner render. Each one independently triggers a
~15s blocking Nous Portal token-refresh network call when the stored token
is expired. On a slow/constrained host (e.g. a small monitoring CT) that
serial burst stretched startup to many minutes, appearing 'stalled'.
Add a per-process memo (5s TTL) so the burst collapses into a single network
round-trip. Only successful, non-forced resolutions are cached; force_fresh
and insecure/ca_bundle callers bypass and don't populate the cache, so
normal refresh semantics are unchanged.
Verified: 3 rapid resolve_nous_access_token() calls -> 1 underlying refresh.
Architecture fix for the bug class behind the Termux --version NameError
(live on main since eb4040242): version-printing kept being reimplemented
as *_fast() copies at the top of hermes_cli/main.py, each duplicating
canonical logic (project-root resolution, container detection, profile
detection). The copies drift silently — eb4040242 edited the canonical
output and referenced the PROJECT_ROOT module constant inside the fast
function, which doesn't exist yet at the fast exit point.
- hermes_cli/_startup_fast.py: THE implementations, stdlib-only. main.py's
*_fast() names become thin delegates (kept for test/back-compat), and
PROJECT_ROOT itself derives from the same helper — the constant and the
fast path can no longer disagree.
- Fast output now includes the .install_method stamp (one cheap file read)
and a 'Run hermes version for update status' pointer, so globalizing the
fast path doesn't silently drop slow-path info.
- Guard tests: (1) import-weight — subprocess-imports _startup_fast and
fails if any heavy module (config/yaml/argparse/cli/run_agent/httpx)
lands in sys.modules; (2) subprocess parity on+off Termux — the test
that would have caught eb4040242 the day it landed; (3) install-method
stamp surfacing.
hermes --version: ~3.8s cold / 0.2-0.4s warm -> 0.01-0.02s everywhere.
- _print_fast_version_info referenced the PROJECT_ROOT module constant,
which is defined AFTER the ultrafast exit point. On current main this
is a LIVE latent bug: the Termux fast path NameErrors on --version
(eb4040242 changed the print to use PROJECT_ROOT without noticing the
constant doesn't exist yet on that path). Compute the root locally.
- PR tests asserted the old 'Project:' label; main renamed it to
'Install directory:' (eb4040242). Expectations updated.
- agent_import.dump_yaml_file now calls utils.atomic_yaml_write instead
of hand-rolling safe_dump + atomic_write_text — same temp+fsync+atomic
rename and symlink preservation, plus mode/owner preservation a
0600-secured config.yaml needs
- openclaw script: the EXDEV/EBUSY copy fallback gains copystat + target
fsync so the docstring's 'mirrors utils.atomic_replace' durability
claim is true on cross-device deployments
- trim load_yaml_file's docstring to the behavior contract
agent_import.py carries a private load_yaml_file/dump_yaml_file pair that
returned {} for an absent file AND for a present file it could not read or
parse. Three importers -- import_permission_allowlist, import_permission_denylist
and import_mcp_servers -- read config.yaml through it, merge one section into
the result, and write the whole mapping straight back. So a YAML syntax error,
a permission problem or a broken mount meant the importer replaced every
setting the user had with only the one to three keys it merged, and still
reported the item as "imported". The write was a bare path.write_text(), so an
interrupted import truncated the file instead.
Distinguish the two cases at the read. Absent, or present but empty, still
yields {} so first-time creation works. Present but unreadable, unparseable, or
not a mapping raises ConfigReadError; the three sites funnel through a new
load_target_config() that records the refusal as a per-item error and leaves the
file byte-identical. Dry-run refuses too, rather than previewing an "imported"
that would destroy the config. dump_yaml_file now writes through
utils.atomic_write_text, which the module already imports and uses for the
memory store.
This is the invariant hermes_cli/config.py already enforces for its own writers
via require_readable_config_before_write / atomic_config_write, whose docstring
names this exact root cause and calls itself "the single chokepoint every
config-update path should use". agent_import.py has its own helper pair and so
was never covered; it was the last config.yaml writer without the guard.
The identical helper pair lives in openclaw_to_hermes.py, the script this module
was ported from, where twelve config.yaml read-modify-write sites share the same
defect; fixed there too. Its refusal is recorded at the run_if_selected dispatch
point, which flips the existing _config_apply_blocked flag so the remaining
config-mutating options short-circuit instead of each rediscovering the same
unreadable file. The atomic write is inlined with tempfile + os.replace because
that script runs standalone with only the stdlib on its path.
/simplify-code review found _poll_for_token has a second caller:
web_server._nous_poller (dashboard/desktop device login), which surfaces
str(e) as the UI error_message — so wrapping only in
_nous_device_code_login left the dashboard showing the bare timeout.
Move the enrichment into _poll_for_token's deadline raise so every
caller inherits the guidance, and drop the now-redundant try/except
wrap in the CLI login. Add a source-level regression test driving the
real poll loop (authorization_pending stub client) to the deadline.
A bare 'Timed out waiting for device authorization' gives the user
nothing to act on. The most common cause is Portal sign-in failing in
the opened browser tab (including the server-side CAPTCHA loop from
issue #20605), so point at the Portal login page and the hermes portal
retry command.
Salvaged from PR #75290 by @HexLab98 (timeout-guidance kernel only).
The URL-rewrite portion of that PR was dropped: the live Portal has no
/device route (verified 404 with a real user_code), so rewriting the
manage-subscription verification URL would break login entirely.
Guidance text reworded to reference only real URLs.
/simplify-code findings on the salvage stack:
- extract _category_skill_dirs() as the single category detector; the
install guard and hermes_cli._existing_categories() now share it
(third copy of the heuristic eliminated)
- filter rglob hits through is_excluded_skill_path so vendored /
support-dir SKILL.md files (node_modules, references/pkg) no longer
misclassify a plain directory as a category and block install
- fix inaccurate WHY comment (lock-file check only guards hub-installed
skills, not hand-authored dirs), drop underscore prefixes on locals,
fold the file-collision guard under the single exists() check
The salvaged hydrate_profile_secret_sources (#74549) seeded its
profile-local env from <home>/.env only, but the documented 1Password
bootstrap flow puts OP_SERVICE_ACCOUNT_TOKEN in the gitignored
<home>/.op.env (mirrored from load_hermes_dotenv). A cold profile using
that flow still failed 1Password hydration — the one unaddressed item
from the sweeper review on #74549. Seed .op.env via setdefault so .env
values win; never touches os.environ. Two regression tests.
`hermes update` aborted its managed-Python runtime repair with an error
that reads like a contradiction:
⚠ Managed Python runtime repair skipped: cannot import name
'venv_python_path' from 'hermes_constants'
(/home/teknium/.hermes/hermes-agent/hermes_constants.py)
The named file does contain the symbol. The module in memory does not.
main.py imports hermes_constants from the OLD checkout, `git pull` then
replaces that file on disk, and the freshly-pulled managed_uv runs its
lazy `from hermes_constants import venv_python_path` against the module
object Python already cached in sys.modules — the pre-upgrade one. The
ImportError reports the path of the new file, so it looks like the symbol
is missing from a file that plainly has it.
Same update-boundary class already documented on `_UvResult` for the
ensure_uv() arity skew. It fires on the first update from any release
older than 83314ca38, which introduced the symbol.
Both lazy call sites now reload the module from disk on ImportError:
- hermes_cli/managed_uv.py::_venv_python — the reported path
- hermes_cli/gateway.py::get_python_path — same flaw, same fix; a gateway
restarted mid-update hits it identically
Reload rather than a local fallback on purpose. Recomputing the layout
inline would hand-roll `Scripts`/`bin` a second time — exactly what #76105
deduped into venv_bin_dir()/venv_python_path(), and what
test_no_open_coded_venv_layout_remains_in_hermes_cli bans. Reloading fixes
the actual problem (a stale module) and keeps one owner for the layout.
hermes_cli/update_cmd.py imports the symbol at module scope, which is a
different failure mode (the module fails to import at all) and is already
covered by the installer's retry-once for the update-boundary crash.
Tests reproduce the stale-module state by deleting the attribute from the
imported hermes_constants: recovery resolves through the reloaded shared
helper (asserted via a sentinel, so an open-coded copy cannot pass), and
the normal path never reloads. Against origin/main's managed_uv they fail
with the exact reported ImportError.
#76499 landed the npm floor as >=11.17.0 on a Node >=20 baseline. This
branch takes the other half of the same constraint: the vendored Node 26
tree now installs npm 12 into itself, so the toolchain satisfies the
stricter floor rather than the manifest relaxing to meet the tarball.
Resolved package.json + package-lock.json to node >=26.0.0 / npm >=12.0.0
and refreshed npm_engine.py's illustrative range to match. website/'s
mirror keeps #76499's >=11.17.0 — it is not a root workspace and builds
on its own Node.
`_append_node_dir_for_service()` had two bugs, both caught by
`test_system_unit_uses_target_user_home_not_calling_user`:
1. It crashed. `iter_hermes_node_dirs()` defaults to the *calling* user's
Hermes home, so under sudo it stats `/root/.hermes/node/bin` — which
raises `PermissionError` for a non-root caller rather than returning
False. An unreadable candidate dir means "skip this rung", not "kill
the generator", so the probe now swallows OSError.
2. Worse than the crash: had the stat succeeded, a `--system` unit
targeting alice would have baked *root's* managed Node into alice's
PATH. The generator now skips the managed-Node rung on the system
path and re-runs it after `_hermes_home_for_target_user()` resolves,
passing that home explicitly. Entries are prepended so the managed
Node still outranks the remapped shell-PATH entries, matching the
user-unit ordering.
The launchd generator is unaffected — it has no target-user remapping,
so the default home is already correct there.
Hermes installs runtimes for itself — `uv` at `$HERMES_HOME/bin/uv`, Node
at `$HERMES_HOME/node` — and neither directory is on an arbitrary
process's PATH. Every `shutil.which("node"/"npm"/"npx"/"uv")` in Hermes's
own code therefore has two failure modes: the managed runtime is invisible,
so the caller reports "not installed" or degrades to a slower tier on a
machine that has exactly what it needed; and when a system copy also
exists, the one Hermes does not own wins.
Routed the Hermes-owned call sites through managed-aware resolvers:
- `agent/lsp/install.py`, `hermes_cli/dep_ensure.py`, `hermes_cli/main.py`
(`_make_tui_argv`), `hermes_cli/tools_config.py` (`_run_post_setup`) now
use `find_node_executable()`.
- `hermes_cli/tools_config.py::_pip_install` and `hermes_cli/setup.py`'s
vercel install use `ensure_uv()` (installing uv is in scope during setup,
and the Windows installer's `uv venv` does not seed pip, so the fallback
tier is "No module named pip"). `tools/lazy_deps.py` uses `resolve_uv()`
— a lookup, not a bootstrap, because it runs mid-turn for an optional
dependency and downloading a runtime as a side effect exceeds what the
caller asked for.
- `hermes_cli/gateway.py`: extracted `_append_node_dir_for_service()`,
shared by the systemd unit and launchd plist generators, which appends
the managed dirs before the PATH-resolved one. A service definition is
written once and survives reboots, so resolving a system Node that
happens to lead the installing shell's PATH bakes the wrong interpreter
in permanently. Managed dirs are profile-scoped, so each profile's unit
still names its own Node; the existing symlink-parent rule (don't
`.resolve()`) is preserved verbatim.
- `tools/environments/local.py`: the terminal tool's subshell PATH gains
the managed dirs, appended alongside the sane entries rather than
prepended — a tool the user deliberately put on their own PATH still
wins, and the managed one only fills a gap. This is also what makes the
bare `which("uv")` in `tools/env_probe.py` correct: that probe reports
the environment the *model* sees, and the model can only run what is on
that subshell's PATH.
`scripts/install.ps1`: the persisted User PATH update becomes
`Set-ManagedNodeFirstOnUserPath`, a move-to-front rather than an
add-if-missing. Installs made by an older install.ps1 already have the
managed dir in User PATH — at the tail, behind a system Node — and an
add-if-missing check sees it present and leaves that ordering in place
forever, so the users the bug hurt would never be repaired. Unrelated
entries keep their relative order (empty segments included; a trailing
`;` is legal and the installer's other PATH code preserves them),
duplicates collapse, and it writes only when the string actually changes.
Tests:
- `tests/test_managed_runtime_resolution.py` — AST guard that fails any
new bare `which()` for a managed runtime, with a short justified
allow-list and a companion test that fails when an allow-list entry goes
stale. Reading source is banned by AGENTS.md and this is the documented
exception: the property is "no call site anywhere spells it this way",
which no runtime seam can observe.
- `scripts/ci/test_install_ps1_path_migration.ps1` — behavioral, not a
source regex: it lifts the real `Set-ManagedNodeFirstOnUserPath` out of
install.ps1's AST and rewrites only the two registry calls into an
in-memory store, so the shipped split/dedupe/prepend/change-detection
logic executes for real. Not in the default lane (Linux runners have no
PowerShell host); runs under `pwsh`. 13/13 assertions pass.
The SQLite runtime repair staged its replacement environment with
`uv sync --extra all --locked --no-config`, and managed_python_env also
exports UV_NO_CONFIG=1. Both drop `[tool.uv]` from pyproject.toml —
including `exclude-newer = "14 days"`, which uv.lock was generated with.
uv 0.12 treats the missing setting as a resolver change, re-resolves, and
then refuses to write under `--locked`:
Resolving despite existing lockfile due to removal of global exclude newer
error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.
So every repair attempt failed at the dependency-sync gate and reported
"replacement environment did not pass dependency and import smoke tests",
leaving vulnerable-SQLite installs stuck on journal_mode=DELETE with a
guaranteed-failure warning on each `hermes update`.
Drop `--no-config` from the sync argv and pop UV_NO_CONFIG from its env.
Interpreter provisioning keeps both: only the sync has to agree with the
lockfile the project shipped.
Salvaged premise from #67065 (@webtecnica, issue #67027), reimplemented:
get_env_value() read os.environ first with no secret-scope check, so a
multiplexed profile turn could serve another profile's credential. Its
siblings get_env_value_prefer_dotenv and gateway.config._getenv were
already scope-aware.
Reimplementation note: the original diff called get_secret() but fell
through to os.environ on a scoped miss — re-opening the exact leak it
targeted (flagged by the sweeper review). This version delegates policy
fully to agent.secret_scope.get_secret (global vars pass through; scope
authoritative under multiplexing; legacy environ behavior when off;
UnscopedSecretError propagates fail-closed), then falls back to .env.
6 regression tests incl. the #67027 repro (envless profile + multiplexed
turn -> None, not the other profile's key); sabotage-verified RED on the
old implementation.
The salvaged cleanup (#75197) scrubbed every known Hermes key absent from
the profile .env — deleting user-shell-exported credentials
(export OPENAI_API_KEY=...) on every hermes invocation, a documented flow
the author's own failing test_dump_flags_shell_only_key_not_in_dotenv
confirmed. A child process cannot distinguish shell exports from
parent-process leakage, so the scrub now covers ONLY
_PROFILE_MANAGED_ENV_KEYS (ACP routing keys: HERMES_ACP_*,
HERMES_COPILOT_ACP_*, COPILOT_CLI_PATH, COPILOT_ACP_BASE_URL) —
the vector from #75141. Cross-profile credential isolation is owned at
read time by agent.secret_scope.get_secret.
Adds shell-export survival regression + a scope-invariant test that fails
if the scrub set is ever widened toward credential-shaped keys.
Align load_hermes_dotenv() with reload_env() so known Hermes env vars
absent from the active profile .env are removed from os.environ instead
of leaking from a parent process / other profile.
Register ACP-related keys (HERMES_ACP_AUTH_METHOD, HERMES_COPILOT_ACP_*,
COPILOT_CLI_PATH, COPILOT_ACP_BASE_URL) in _EXTRA_ENV_KEYS so they
participate in known-key cleanup.
This is the same isolation gap class as #68367 / #66930, but:
- Not Desktop-only spawn scrub — CLI/gateway restart inheritance
- Not Matrix/messaging auto-enable only — copilot-acp provider/ACP config
- Startup dotenv clear so *any* inheritance path is covered
Example: HERMES_ACP_AUTH_METHOD=cursor_login leaking into a Claude Code
ACP profile caused authenticate -> Internal error -> Discord
'model provider failed after retries'.
The npm 12 requirement (f88ed6c717) strands every system-Node install:
no shipping Node bundles npm >=12, engine-strict makes EBADENGINE fatal,
and the recovery in npm_engine.py refuses to touch a system npm — so
'hermes update' leaves the install in a mixed state (updated code, stale
Node deps, no TUI/web/desktop rebuild) with only a manual-fix hint.
Instead of modifying the user's toolchain (still never done), the
EBADENGINE recovery now provisions Hermes' own managed Node tree under
$HERMES_HOME/node — the same pinned-nodejs.org path install.sh and
install.ps1 use — upgrades THAT npm into the required range, and hands
the caller the managed npm for its single retry.
- hermes_constants.bootstrap_hermes_managed_node(): cross-platform
provisioning (POSIX via node-bootstrap.sh _nb_install_bundled_node,
Windows via the existing portable-zip download); reuses a healthy tree.
- node-bootstrap.sh: HERMES_NODE_SKIP_LINKS=1 skips the ~/.local/bin
node/npm/npx symlinks so the private tree never shadows the user's
own toolchain on PATH.
- maybe_repair_npm_engine() now returns the npm path to retry with
(managed-in-place upgrade or freshly provisioned runtime); both call
sites retry with the returned path and put the managed tree first on
PATH so npm lifecycle scripts resolve the managed node.
- Node-only mismatches on a foreign npm are now also recoverable (the
managed tree ships a supported Node); on a managed npm they still
correctly decline.
E2E (real download, temp HERMES_HOME): provisioned node v22.23.2,
upgraded bundled npm 10.9.4 -> 12.0.2, system npm byte-identical after,
no ~/.local/bin links re-pointed, healthy-tree reuse in 0.05s.
Cancelling the API-key prompt mid-wizard (Enter → 'Cancelled.') let the
wizard continue through Terminal/Gateway/Tools and finish 'successfully'
with no model configured — the user exits believing they're set up, then
hits a broken chat.
_print_setup_summary() (called by every setup path: full, quick,
blank-slate, portal) now probes resolve_provider() and, when nothing is
configured, prints an unmissable warning with the two one-line fixes
(hermes model / hermes setup --portal).
Consumer-onboarding audit finding #7 (sev 4), Aug 2026.
A completely unconfigured install previously booted into a working-looking
chat (banner showed model 'unknown'), accepted a message, spun ~30s, then
failed with 'Set OPENROUTER_API_KEY' — a provider the user never chose —
and never offered setup.
- HermesCLI.run() now probes provider readiness at startup (TTY only) and
offers the shared provider picker (hermes model flow, which fronts Quick
Setup / Nous Portal OAuth) when nothing is configured. Decline is
respected; picker state re-syncs into the live CLI so the next turn works
without a restart.
- New silent probe _runtime_credentials_ready(): no printing, no state
mutation; handles keyless local endpoints and callable bearer providers.
- The empty-api-key error is provider-aware: names the actual resolved
provider and points at 'hermes model' / 'hermes setup' instead of
hardcoding OPENROUTER_API_KEY.
- Banner: unconfigured installs render 'no model configured — run /model'
in red instead of the silent 'unknown' model slug.
Consumer-onboarding audit finding #2 (sev 5), Aug 2026.