The checkout root is where the update just pulled into, so "root unreadable" cannot
happen after a successful pull — drop the OSError fallback and the static five-name
tuple it fell back to (the tuple was the drift that caused the bug). Drop the phantom
`hermes_cli.hermes_logging` protection entry: no such module exists; the real
`hermes_logging` is root-level and now protected by name. Keep two tests: the stale
root `utils` scenario (red on base) and the hermes_logging protection.
`_STALE_PURGE_PREFIXES` listed five package names, so every top-level module in
the checkout root survived the post-pull purge. `hermes update` then imported
new source against a cached pre-pull `utils`, `hermes_constants`, `plugins` or
`providers`.
Field failure (2026-09-12, macOS): updating 0.20.6 -> 0.21.2 crossed 3145986c20,
which added `base_url_origin` to `utils.py`. The restart phase's up-front
`from hermes_cli.gateway import ...` pulled `agent.auxiliary_client`, whose
`from utils import base_url_origin` hit the cached 0.20.6 `utils`:
Update incomplete - gateway auto-restart failed: cannot import name
'base_url_origin' from 'utils' (.../hermes-agent/utils.py)
The purge docstring already claims it evicts EVERY cached Hermes module; the
hardcoded tuple was the same "re-fixed per symptom" shape it replaced. Scan
`PROJECT_ROOT` instead: top-level `.py` files plus directories with an
`__init__.py`. 50 names here, and a newly added module can no longer drift out.
Two exclusions, both deliberate:
- `hermes_logging` joins `_STALE_PURGE_PROTECTED`. Its queue listener, handler
list and `_logging_initialized` flag are module globals, so a fresh copy
starts a second QueueListener over the same log files while the first runs.
- `tests` is never purged. pytest resolves fixtures through the identity of its
already-imported test modules.
Falls back to the old tuple when the root is unreadable.
`_config_mcp_servers` is `self.config.get("mcp_servers") or {}`; with the
test's own `{"mcp_servers": {}}` it is `{}` whether or not the import fix
is present, so the assertion discriminated nothing. One invariant per fix.
#111408 widened the MCP config watcher seed from mtime to
utils.file_signature but omitted the import in cli_tui_mixin.
The NameError only fires when config.yaml exists, so isolated-home
tests short-circuited past it and every real CLI launch crashed.
Signed-off-by: mr-r0b0t <adam.manning@gmail.com>
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
matches neither the content-policy nor the billing patterns, so a safety
refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
overwrites a record that has one (the loop racing the user's click).
Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
_welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
of freezing whole sentences.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Drive CLIChatTurnMixin._chat_render_turn through a minimal stub instead of a full
HermesCLI (prompt_toolkit stubs, cli reload): the invariants are that a (text, images)
payload reaches _pending_input intact and that several queued parts join their text
and merge their images. Both are red on origin/main.
_clui mixin bundles Enter submissions carrying images as (text, images)
tuples, and _tui_enter_while_busy puts those tuples on _interrupt_queue in
interrupt mode. _chat_render_turn then did "\n".join(all_parts) over the raw
payloads, raising TypeError on the tuple — swallowed by the outer chat()
handler, so the interrupt message was silently lost and no next turn started.
Unpack tuple payloads in the re-queue block: join the text parts and re-attach
all images as a single (text, images) tuple, matching the shape
_tui_process_one_input already accepts. Plain-string payloads are unchanged.
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; all changes reviewed and approved by the contributor
attach() now returns False both when the client dropped mid-replay (session
already detached, socket dead) and when the force-redraw write stalled
(socket still attached, worth closing with 1013). Only close the socket in
the second case; the registry detach stays as the guard for both.
Also trims the salvaged tests to two invariants, both red on origin/main:
a mid-replay drop leaves the session reap_idle()-reclaimable, and a drain
send failure detaches the current socket without touching a replacement
that attached during the send (#110849).
Three caches that read config.yaml still keyed change detection on
st_mtime (or st_mtime_ns + st_size), so a same-size replacement that
keeps the old timestamp (cp -p, rsync -t, a timestamp-pinning writer)
was never noticed:
- model_tools._tool_defs_cache_key: get_tool_definitions kept serving
stale dynamic tool schemas / mcp_servers for the process lifetime.
- CLI mcp_servers auto-reload watcher (cli_tui_mixin seed +
cli_info_mixin._check_config_mcp_changes): the replaced mcp_servers
section was never reloaded. The seed is now _config_sig.
- tui_gateway/server._load_cfg_raw / _save_cfg (_cfg_mtime -> _cfg_sig):
the raw-config cache served the stale document and the next _save_cfg
would write it back over the on-disk file.
All three now use utils.file_signature like the rest of the PR. Tests that
reset the renamed module/instance attributes follow the rename; one
pinned-mtime replacement test per cache, red on the previous head.
Also corrects two stale type/comment annotations in hermes_cli/config.py
(_env_cache key shape, _RAW_CONFIG_CACHE record shape).
Review finding: three sibling config.yaml caches (tool-defs memo, CLI mcp watcher, TUI-gateway raw cfg) still compared mtime/size only.
The cherry-picked commit extends the config/profile/MCP/managed-scope/completer/
OAuth/skills-manifest signatures. This commit finishes the class and trims it:
- `file_signature()` lives in `utils.py` next to the other stat/metadata helpers
instead of `hermes_cli.managed_scope` (gateway/ and agent/ callers no longer
reach into the managed-scope module for a generic stat helper).
- `hermes_cli/config_effective.py` was left comparing 2-/4-wide prefixes against
the widened `_RAW_CONFIG_CACHE` / `_load_config_cache_sig` records, so
`load_user_config_effective()` re-parsed on every call (3 parses for 3 calls on
an unchanged file, 1 before); index by the new widths.
- Sibling caches keyed on the same (mtime, size) shape and reading the SAME files
now use the helper: `load_env()` memo, `agent/skill_utils` raw-config and
external-dirs caches, `hermes_cli/model_switch` alias identity, `agent/moa_loop`
preset stamp, `hermes_cli/auth` global auth-store memo.
- Tests trimmed to one invariant each (pinned-mtime replacement invalidates; an
unchanged file still hits), both red on origin/main.
Left alone on purpose: `tools/registry.py`, `tools/skills_tool_dedup.py`,
`gateway/status.py`, `hermes_cli/banner.py`, `hermes_cli/main.py`,
`hermes_cli/session_recovery.py` — those fingerprint source files, PID/lock files
or write to persisted on-disk caches shared across processes, where an inode/ctime
key would churn on every checkout/copy rather than catch a replaced config.
Config caches, the profile re-scan watcher, the MCP reconciler, managed-scope
reads, the completer memo, OAuth heal marks, and the skill-manifest snapshot
keyed change detection on (st_mtime_ns, st_size) only, so a replacement that
preserved both (cp -p, rsync -t, timestamp-pinning scripts, sync clients) was
treated as unchanged and stale values were served until restart.
Add st_ino (fresh inode on atomic replace) and st_ctime_ns (cannot be
backdated via os.utime) to every signature via a shared
hermes_cli.managed_scope.file_signature() helper.
Fixes#111105
The flat-install block only ignored the bare cron/executions.db. All three
cron stores (executions, deliveries, notepad) are opened through
sqlite_util.open_db in WAL mode, so executions.db-wal/-shm exist whenever
the scheduler is live, and deliveries.db / notepad.db were not ignored at
all. `git stash push --include-untracked` therefore still unlinked the
live WAL/SHM and the whole deliveries/notepad stores under the running
gateway - the same mechanism this PR closes for the root state.db.
Switch the cron entries to the by-class shape already used at the root
(/cron/*.db plus -wal/-shm/-journal/retired-wal sidecars) and also ignore
the cron lock/heartbeat/output files and the per-launch root markers
(.update_check, gateway-starts.log, .clean_shutdown, active_profile,
.hermes_history, slack_tokens.json) that otherwise force every update
into the stash step. `git ls-files -i -c --exclude-standard` is unchanged
(no tracked file newly hidden). The test's FLAT_INSTALL_RUNTIME_STATE
gains one representative per new class; red before, green after.
Review finding: cron/executions.db-wal, cron/deliveries.db and cron/notepad.db were unignored and swept by the flat-install autostash.
`backups/` is where the pre-update snapshot the updater restores a swept
state.db FROM lives — leaving it unignored means the recovery copy is
swept together with the live database. `vault.key`/`vault.json.enc` are
the local secret vault.
Docs: website/docs/getting-started/updating.md explains that on a flat
install (checkout root == $HERMES_HOME) runtime state is git-ignored and
never enters the autostash.
create_task(parents=[open parent]) — the reporter's actual incident path —
parked the card in todo with only a `created` event, and kanban_create's
payload carried no `gated`, so the board still showed an unexplained todo
while only the link surface was fixed. create_task now appends the same
dependency_wait {reason: parent_not_done, parent} event and kanban_create
returns gated/gated_by, mirroring kanban_link.
link_tasks gated on `status != 'done'`, but _parents_satisfied and
recompute_ready treat `archived` as terminal: linking a ready child under an
archived parent demoted it to todo with a false parent_not_done event and the
next recompute promoted it straight back. Gate on not in ('done','archived').
Review finding: create_task(parents=...) emitted no dependency_wait/gated; link under an archived parent flapped ready->todo->ready with a false reason.
The dashboard's POST /links is the fourth writer of link_tasks (CLI, tool,
dashboard, plus the graph builder); return the same ``gated`` flag so every
surface that can create the deadlock can see it. Document the
``dependency_wait`` payload the link path emits and the delegation rule the
reporter derived (never link a support card under the card it unblocks).
Drops the CLI output test (a change-detector on prose); the two DB-level
invariants (event emitted on demotion / none for a done parent) stay.
A ready child linked under an unfinished parent drops to todo with no
event and no operator signal; the only trace used to be claim_rejected
after a forced promote. Record a dependency_wait event when the demotion
fires, return the gate from link_tasks, warn in the CLI link command,
report gated in the kanban_link tool, and document the gate.
The per-root merge in _copy_dist_payload treated the first level under
skills/ as the replace unit, but skills live at skills/<category>/<skill>.
A distribution shipping one skill inside a category rmtree'd the whole
category directory, wiping every skill the user (or hermes update) had
placed there — the exact scenario of issue #25120 that this PR claims to
fix. A shipped directory with no files of its own is now a container:
its children are merged one level down, and the replace unit becomes the
nearest directory that holds a file (a skill dir always holds SKILL.md).
A symlinked category (skills/devops -> shared dir) is a container too,
so the pre-write symlink check now covers it instead of silently
unlinking the link and dropping the shipped copy in its place.
Review finding: categorised skill payload rmtree'd skills/<category>/,
destroying sibling user and bundled skills; symlinked category was
unlinked rather than refused.
Fold the skills/cron/force-install cases into one per-root merge test and keep the
symlink-refusal test; the three separate tests asserted the same mechanism.
The symlink refusal lived in `_real_dir`, which runs per entry inside the copy
loop. By the time it fired for `skills/`, `SOUL.md`/`mcp.json`/`cron` had already
been replaced and `write_manifest` had not run yet, leaving a half-updated
profile that fails identically on every retry. Add a pre-flight pass over
`_owned_entries` that walks each target path (whole chain for directories,
parent chain for files) and raises before any mutation; `_real_dir` keeps its
check as the backstop.
The `.env.template` branch used a bare `shutil.copy2`, which writes through an
existing `.env.EXAMPLE` symlink into whatever it points at. Route it through
`_replace_entry` so the link is unlinked first, like every other owned entry.
The error text now tells the user what they can actually do (remove the link or
replace it with a real directory) instead of pointing at `distribution_owned`,
which is the author's manifest, not the user's.
The existing symlink test now changes every other shipped entry upstream and
asserts the profile is byte-for-byte untouched after the refusal.
A user who points `profiles/x/skills` at a shared directory did so on
purpose. The merge path replaced such a symlink with a real directory
silently, discarding that configuration, where the base rmtree at least
failed loudly. Raise DistributionError naming the link instead; the
file->directory replacement stays because a stray file there is never
deliberate configuration.
`_remove_existing` now removes anything that lexists but is not a real
directory via unlink, so fifos and sockets no longer slip through.
Reviewer P1 on the original PR (symlink container must survive intact).
Keep the tests that pin behaviour a regression would break: user-added and
stale skill roots survive an update while shipped roots are replaced, a
forced reinstall keeps a user skill, and a user-added cron file survives an
update. The allowlist and symlink-transition cases exercise the same code
path (`_real_dir` + `_replace_entry`) and were duplicates of those.
Refs #110933
The skills-only special case (a `preserve_skills` kwarg plus nested closures
keyed on the literal "skills") left cron/ and any future owned directory on
the old rmtree-then-copytree path, which deletes user-added files on update
and on --force reinstall. The rule is the same for all of them: a top-level
owned directory is a container of roots; replace only the roots the payload
ships, never the whole destination.
Apply that rule to every owned directory, drop the kwarg and the string
check, and lift the helpers to module level. `_remove_existing` keeps its
symlink-no-follow semantics and `_real_dir` replaces any symlink or file on
the destination path so the payload is never written outside the profile.
The fresh-install path had nothing to preserve, so it needs no flag.
Restore "skills" in DEFAULT_DIST_OWNED; removing it was never needed for the
fix. Add one invariant test: a user-added cron/ file survives an update.
Refs #110933
The websockets dials got proxy=None but the three HTTP /json/version dials
(CDP override discovery, is_browser_debug_ready used by Lightpanda and the
real-profile readiness check, and surviving-Chrome detection) still resolved
via getproxies(), so under a system/env proxy discovery fell back to the raw
http:// URL, readiness never fired and /browser connect reported not ready.
loopback_request_kwargs() sits next to loopback_connect_kwargs() and is used
at all three sites (ProxyHandler({}) opener for the urllib one).
add_loopback_no_proxy turned an operator NO_PROXY=* into '*,127.0.0.1,...',
which urllib/requests no longer treat as the wildcard, flipping bypass-all
configs into proxy-all. A wildcard in either casing now leaves env untouched.
is_loopback_host also accepts any loopback IP literal (127.x, ::ffff:127.0.0.1).
Review finding: HTTP /json/version dials still proxied loopback; NO_PROXY=* wildcard broken by append.
Follow-up to the salvaged commit from #111169 (@KoNit-K):
- Drop HERMES_KANBAN_TASK_TITLE. The worker already has HERMES_KANBAN_TASK
and HERMES_KANBAN_BOARD/HERMES_KANBAN_DB pinned in its env, so
maybe_auto_title reads the card title from the board itself (no new
HERMES_* env var for non-secret config; the dispatcher and the
delegation scrub list stay untouched).
- Unreadable or missing card: the session is named `Kanban task <id>`
with zero auxiliary calls (the fallback the issue asked for; the
#109743 seed left such workers untitled).
- The card title persists at `llm` authority via set_auto_title, so a
manual /title still wins and the upgrade thread never starts.
- Tests trimmed to two invariants against a real board + SessionDB
(card title, unreadable-card fallback), both red on origin/main.
A claim that expired without a worker ever spawning (worker_pid NULL) was
reclaimed and immediately re-claimed on every dispatcher tick, with
consecutive_failures stuck at 0 — nothing could trip the breaker. Route the
reclaim through _record_task_failure (own txn after the reclaim commit, same
shape as enforce_max_runtime) instead of the salvaged raw counter increment,
so per-task max_retries / kanban.failure_limit and the gave_up event apply
and last_failure_error carries the stale lock. reclaim_task (operator path)
still resets the counter; the live-worker extend path never reaches it.
Trims the salvaged tests to one invariant that walks the breaker to its trip.
release_stale_claims() reclaimed expired claims without counting the
failure, so a claim that was never spawned (worker_pid NULL) spun forever
with consecutive_failures stuck at 0 and the failure_threshold breaker
could never trip (#111306). The reclaim transaction now increments
consecutive_failures atomically. Extend-live-worker and manual operator
reclaim paths are untouched (extend is not a failure; reclaim_task
deliberately resets the counter).
Tests: 3 new regression tests in tests/hermes_cli/test_kanban_db.py —
reclaim without spawn increments (red on base), reclaim with dead worker
increments (red on base), live-worker extend does not. Related kanban
suites: 79 passed, 1 skipped.
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
Replace the CLI-subprocess parametrization (6 hermes processes per run) with a
direct normalizer invariant: non-finite cadence -> default cadence, non-finite
temperatures -> None while finite values (0, 0.75) survive.
Keep the two tests that were red on origin/main (probe failure keeps the clock and
unloads once telemetry recovers; confirmed busy still resets). Drop the is_idle() bool
contract test — it pins behaviour that did not change — and fold the probe-failure log
call onto two lines.
is_idle() treated any /slots or /metrics probe error as "busy", so sweep_idle()
popped the idle clock on every failed probe — one transient failure per sweep
and a resident model never reached IDLE_UNLOAD_S (21 GB pinned for hours with
zero requests).
Split the probe into a tri-state: _probe_idle() returns True (confirmed idle),
False (confirmed busy), None (probe failed). The sweep now resets the clock
only on a confirmed busy sighting; a failed probe keeps the existing clock and
logs at INFO, and unload still requires confirmed idleness past the threshold.
The public is_idle() bool contract is unchanged (probe failure reads as not
idle), so no caller outside the sweep ever unloads on a telemetry hiccup.
Closes#111154
llama.cpp b10964 (the pinned build) renamed --no-webui to --no-ui, so the
router refused to start with an unknown-argument error. The direct-I/O flag
is already selected per-build by _direct_io_args on main, so this reduces
to the one remaining rename and asserts it in the existing spawn test.
Refs #111323
_run_pending_fleet_restart has three supervisor branches; #110709 covers launchd, this covers
the Windows branch (gateway_windows.is_installed -> restart) so a Windows developer with an
installed gateway service does not restart it from the test suite.
test_run_pending_restart_true_when_no_gateways patches the PID scan and
the systemd unit listings but leaves the macOS launchd restart live, so
on a machine with a real hermes launchd fleet the restart phase drains
those units, reports the fleet restart incomplete, and the assertion
fails. CI never sees this because it runs Linux.
Patch _restart_macos_launchd_gateways to a no-op (the same pattern the
file's _patch_update_deps helper already uses) so the no-gateways state
covers the launchd supervisor scope too.
Fixes#110701
test_unavailable_directory_links_are_diagnosed_without_creating_targets already proves a link
with a missing target is refused, not materialized; keep the two new invariants (aliased
parent -> 0o700; operator link at the home boundary still owns its mode).
`_directory_links()` walks every parent up to `/`, and `initialize_home()` skipped
`_secure_dir()` whenever that list was non-empty, so any symlinked ancestor disabled
home hardening and left `~/.hermes` plus its `cron`/`sessions`/`logs`/`memories`
subdirectories at the mkdir default `0o755`.
macOS makes that the normal case: the default temp root is `/var/folders/...` and `/var`
is a symlink to `/private/var` (as are `/tmp` and `/etc`), so any home reached through
those prefixes was never secured. Reported on macOS arm64 as `493 != 448` (0o755 vs 0o700)
in tests/cron/test_file_permissions.py; the same numbers reproduce on Linux by aliasing a
parent directory.
Only links at or below the home boundary are operator-owned. The home itself, or a link
inside it, still transfers permission ownership to the operator — behaviour pinned by
test_initialization_preserves_external_directory_modes and unchanged here. Availability
diagnosis is untouched: a link above the home whose target is missing is still refused
instead of being materialized on the underlying filesystem.
get_profile_dir("default") re-resolves the platform-native root at call time; when pytest's
basetemp sits inside the operator's Hermes home the per-test sandbox counts as "under the
native home" and these two tests overwrote the live config.yaml / .env / MEMORY.md with their
stubs. Write through the profile_env fixture's home instead.
Taken from #111096 by @ruochu88s (the autouse native-home override and the session tripwire
from that PR are not included; the basetemp relocation in tests/conftest.py closes the class).
`_is_secret_config_key` matched only an exact set plus four suffixes, so
credentials that `hermes config set` routes to .env but end in bare `_KEY`
(FAL_KEY, VOICE_TOOLS_OPENAI_KEY, API_SERVER_KEY) or AWS_SECRET_ACCESS_KEY
printed raw with redaction on. Every `_is_env_config_key`-routed key is now a
credential unless its suffix names a non-secret shape (_URL/_HOST/_USER/_ID/
_DOMAIN/_SCHEME); `_key`/`_access_key` join the leaf suffixes. Header names are
folded `-`→`_` before matching so `mcp_servers.<s>.headers.X-API-Key` masks on
get and on the set echo. Bare `auth` leaves the exact set: `mcp_servers.<s>.auth`
is the documented `oauth` mode enum and rendered as `***`. Unresolved `${VAR}`
placeholders are printed as-is so the operator can see which env var the config
references.
Review finding: `config get` printed FAL_KEY / AWS_SECRET_ACCESS_KEY / X-API-Key raw, masked `auth: oauth`, and hid `${VAR}` placeholders.
`hermes config get providers`, `config get providers.<p>.api_key`, `config get
<PROVIDER>_API_KEY` (the .env-routed branch) and `config get mcp_servers.<s>.env.X_API_KEY`
all printed the full credential. The agent runs this command from sessions whose transcripts
persist and get forwarded (a Gemini key surfaced in a Discord DM log), so `print` output is a
leak path the logging redactor never sees.
`get_config_value` now applies the structural masker used by `config show` before printing,
honouring `security.redact_secrets` (default on), with a `--raw` flag for operators/scripts
that need the real value. `_is_secret_config_key` extends the exact-name set with the same
`*_API_KEY / *_TOKEN / *_SECRET / *_PASSWORD` suffixes `_is_env_config_key` already routes to
.env, so env-map leaves under `mcp_servers.*.env` mask too, and the `config set` echo uses the
same predicate.
Slim redo of #84153 by @webtecnica (same direction: mask in get_config_value; dropped the
redact_url_query_params re-export and the separate redaction-enabled reader in favour of
agent.redact._redact_enabled, which already resolves the profile-scoped policy).
Fixes#110758Fixes#84106
d8dcdfd620 (v2026.9.14) made apply_wal_with_fallback refuse to ENABLE WAL
on a fresh database whose directory is on a cross-VM bind mount (virtiofs/9p),
but a database that was already WAL on such a mount kept WAL — correctly, we
never live-downgrade under other openers — and emitted nothing. The operator
in #110848 ran exactly that shape (Podman applehv virtiofs bind mount) and got
"database disk image is malformed" within a minute with no prior signal.
- apply_wal_with_fallback: in the on-disk-WAL branch, log a once-per-process
ERROR ('cross_vm_fs_existing_wal') when the DB file is on a cross-VM
filesystem, naming the two remedies (offline PRAGMA journal_mode=DELETE
after stopping every process + database.journal_mode: delete, or move the
database to a native/named volume). The fresh-DB refusal is unchanged.
- hermes doctor: _report_database_journal_modes flags a WAL database on a
cross-VM filesystem with check_warn and the same remedy (ranked above the
WAL-reset exposure warning; the exposure bookkeeping is kept).
- docs: docker.md gains "Filesystem requirements for state.db in containers";
configuration.md's database comment no longer implies operators must set
delete by hand on virtiofs.
Detection stays /proc/self/mountinfo-based (runs inside the Linux container
on macOS/Windows hosts). locking_mode=EXCLUSIVE is deliberately not adopted:
gateway, cron and workers open state.db concurrently.
Fixes#110848