When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.
Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.
Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.
Fixes#87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
runuser/su/sudo -u from a root shell leaks XDG_RUNTIME_DIR=/run/user/0 into
the child. _user_systemd_socket_ready() stat-ed sockets under it with a bare
Path.exists(), which only suppresses ENOENT/ENOTDIR/EBADF/ELOOP — EACCES on
the 0700 root-owned dir escaped as a raw PermissionError traceback instead of
the documented UserSystemdUnavailableError remediation path.
- _path_exists_safe(): Path.exists() that treats EACCES as absent; used at
both the readiness and DBUS-detection call sites.
- _ensure_user_systemd_env(): drop an XDG_RUNTIME_DIR that is unset or owned
by another user in favour of our own /run/user/{uid}, so the restart
actually succeeds after su/sudo -u instead of only failing cleanly.
Regression tests cover the EACCES readiness probe, foreign-dir replacement,
and that preflight raises UserSystemdUnavailableError (not PermissionError).
hermes update decides restart ownership from the live grandchild argv,
not from an env marker. Newly generated plists now include the flag.
The stderr_timestamp wrapper upgrades only historical Hermes gateway
run shapes for stale plists and leaves arbitrary launchd children unmarked.
launchd only stamps XPC_SERVICE_NAME on its direct child. The timestamp
wrapper is that child, so the grandchild gateway sees XPC_SERVICE_NAME=0
and the supervised-conflict guard refuses the service's own spawn.
Forward HERMES_GATEWAY_EXTERNAL_SUPERVISOR=1 when the wrapper itself is
launchd-supervised. Interactive XPC_SERVICE_NAME=0 starts stay unmarked.
Fixes#86893
After the VBS/cmd launcher exits, Task Scheduler marks
Hermes_Gateway_* Ready while the detached gateway keeps running.
The orphan reaper only bailed on Running, then fail-opened the
parent-chain check and killed the live bot on desktop serve start.
Fixes#87001
Docker Desktop writes fpath=(~/.docker/completions ...) into .zshrc.
The referenced-script walk then opened that directory, saw a non-regular
file, and fail-closed — blocking source ~/.zshrc on every terminal
command. Directories are not scripts; devices stay fail-closed.
Fixes#86864.
Legacy custom_providers configs commonly used short/placeholder
api_keys ('123', 'm') for local no-auth services like Ollama --
harmless for the endpoint itself, since Ollama accepts any key or no
key. A stricter has_usable_secret(value, min_length=4) gate added
later now rejects these, but only the credential-POOL resolution path
lacked the same "no-key-required" exemption every OTHER resolution
path in this file already has for exactly this scenario:
- The config-based custom_providers fallback (non-pool path) already
ends with `api_key or "no-key-required"`.
- The "actual" provider's local-offline path already injects
ACTUAL_LOCAL_NOAUTH_PLACEHOLDER before the usable-secret gate for a
loopback base_url.
- _try_resolve_from_custom_pool() was the one gap: it returned the raw
short pool credential unchanged, which then failed the downstream
has_usable_secret() gate with a generic "No usable credentials found
for custom" error that contradicts setup.status ("configured
credentials" vs "runtime failed"), sending users hunting in the
wrong direction.
Fixed by substituting the same "no-key-required" placeholder when the
pool's stored credential fails has_usable_secret() AND the base_url
resolves to a loopback hostname (using the existing _loopback_hostname
helper, matching the exemption scope the issue itself requested:
localhost/127.0.0.1/::1 only, not arbitrary remote endpoints with a
genuinely-too-short key).
Added 4 regression tests extending the existing
test_runtime_provider_resolution.py file, following its established
credential-pool mocking pattern: the exact reported 3-char repro
('123'), a 1-char case, a non-loopback sanity check confirming the
exemption stays scoped (a short key for a remote endpoint is NOT
silently exempted), and a sanity check that a genuinely usable
loopback key passes through unmodified. Verified as a genuine
regression by reverting the fix and confirming 2 tests fail with the
exact raw short key leaking through unchanged.
59/59 pass in the extended test file; 14/14 across two more related
custom-provider test files (no regression).
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.
- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
and returns (block_message, modified_args); modify directives
shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
{"action": "modify", "args": {...}} and Claude Code-compatible
{"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).
Salvaged from PR #28953. Best fix for #18988.
Both quarantine wrappers (_run_quarantined_install in main.py and
_run_install_cmd in _install_repair.py) renamed live hermes*.exe shims
aside before invoking the installer, but only renamed them back on
FAILURE. A SUCCESSFUL install that never rewrites entry points — uv
audits an already-satisfied editable install as a no-op — left the
shims quarantined as hermes.exe.old.<ms> and `hermes` disappeared from
PATH after a green install (#75584; reproduced live on a Windows
install recovering from the #86735 self-lock deferral).
Switch both sites from except/re-raise to try/finally so restore runs
on every path. _restore_quarantined_exes already skips shims the
installer actually replaced, so fresh output is never clobbered and
failure behavior is unchanged.
Regression tests cover both wrappers x {no-op success, rewriting
success, failure}; the no-op cases fail on the previous code.
_select_plugin_image_gen_provider hardcoded image_gen.use_gateway = False.
The managed (Nous-subscription) flow writes use_gateway = True via
_write_provider_config, then this selector runs AFTER it — so picking FAL
through Nous Portal silently persisted provider: fal, use_gateway: false
and every generation billed the user's personal FAL_KEY instead of the
subscription (real incident: key drained to zero-balance lock while the
managed route sat unused).
Fix the class, not the site:
- _select_plugin_image_gen_provider gains the same use_gateway kwarg its
video twin (_select_plugin_video_gen_provider) already had; all four
call sites pass use_gateway=bool(managed_feature), matching the video
call sites, TTS, STT, browser, and web.
- Active-provider detection (the checkmark in `hermes tools`): the
image_gen_plugin_name branch now defers managed entries to the
managed_feature branch and requires use_gateway OFF for direct-key
entries — mirroring the video branch's existing guard, so a managed
FAL pick and a direct-key FAL pick no longer both report active.
Runtime side (prefers_gateway("image_gen")) was already correct; the bug
was purely the setup-time writer.
Tests: new tests/hermes_cli/test_imagegen_managed_gateway.py (3 cases:
managed flag survives, direct pick still clears, image/video selector
contract parity). Sabotage-verified: restoring the hardcoded False fails
2/3. Neighboring hermes_cli provider/managed suites: 180 passed.
Follow-up to the salvaged #86823: the guard queried a hardcoded
"HermesGateway" task, but `hermes gateway install` registers
Hermes_Gateway (Hermes_Gateway_<profile> for named profiles) via
gateway_windows.get_task_name(). Query that name so the supervisor
guard is active on standard installs; fall back to the default literal
if the module import fails. Test now asserts the profile-aware name is
what reaches the task-state query.
Also corrects the cherry-picked commit's placeholder author email to
the contributor's GitHub noreply address.
The orphan-reap sweep (_reap_unsupervised_gateway_orphans) must not kill a
gateway that Windows Task Scheduler is actively managing. The existing
services.exe parent-chain backstop fails open: when the Task-launched conhost
bootstrap has already exited, Windows does not reparent the gateway, the
chain breaks, and the supervised gateway is treated as an orphan. The reaper
then writes the planned-stop marker, the gateway exits cleanly with code 0,
and the scheduler never restarts it (RestartCount only fires on non-zero
exit) — silently killing A2A/messaging on every desktop-app launch.
Querying the task's own state is the authoritative signal and closes the gap
without depending on process ancestry: if HermesGateway is Running, skip the
reap entirely. Uses PowerShell Get-ScheduledTask (English State enum,
locale-stable) rather than schtasks (localized output + codepage mangling).
A gateway whose asyncio event loop is stalled (e.g. an in-loop
compression pass, #72707) cannot process SIGTERM/SIGUSR1 shutdown.
The updater's drain wait then burned the full 180s budget, warned
"Gateway PID X still running after 180.0s — restart may fail", and
`hermes update` could deadlock behind the wedged process — the user
cannot update their way out of the stall.
Fix: before any drain wait, read the loop-liveness heartbeat file the
gateway rewrites every 30s (#66892). Classification:
- alive (fresh heartbeat): busy-but-alive loop — take the normal
graceful drain, honoring the in-flight cron drain floor (#86684).
- wedged (heartbeat for this PID stale >90s = 3 missed beats): the
loop is provably dead; drain is pointless. Bounded escalation:
SIGTERM + 5s grace, then SIGKILL + 5s wait, then proceed (~10s
worst case, far under the 180s drain budget).
- unknown (missing/corrupt file, PID mismatch): never escalate on
ambiguity — full drain path.
Wired into launchd_restart, systemd_restart, and both updater
gateway-shutdown sites (systemd unit drain + manual profile
gateways). The probe is a local stat + JSON read (well inside the
10s query tier of the subprocess timeout tiering).
The cron drain floor from #86684 is bypassed ONLY when the loop is
provably dead — a merely busy gateway still refreshes its heartbeat
and keeps the full drain budget.
Root cause of the loop stall itself (compression blocking the loop)
is #72707 territory and deliberately out of scope here.
Fixes#81642
The #86687 self-lock preflight fired on every Windows `hermes update`:
bitwarden.py's module-level cryptography import (fixed in #86782 /
#86826-class change) meant cryptography._rust was ALWAYS mapped by the
time the preflight ran, so the update exited 2 before even fetching and
looped forever — including the Desktop in-app update (#86780).
Two structural fixes so the guard can never re-brick the flow it protects:
1. Version-gated detection: _detect_self_loaded_native_modules() now
consults _dependency_sync_would_rewrite(dist) — installed version vs
the on-disk pyproject pins (base deps + all extras, env markers
honored). A loaded module whose distribution the sync will not touch
is no lock risk and is not reported. Unknown → fail closed.
2. Relocated deferral: the check no longer runs pre-fetch. It runs via
_abort_dependency_sync_if_self_locked() immediately before each venv
rewrite (git-path dep sync, ZIP-path dep sync, current-checkout venv
repair) — AFTER the code swap. A deferral now leaves the user on NEW
code with only the dependency install pending (completed by the next
launch's marker recovery), instead of stranding them on the old
checkout in an exit-2 loop.
PyYAML's _yaml extension (loaded by every CLI process) joins the
registry — with version gating it is now safe to list.
Tests: version-gate unit coverage (no-change skip, stale pin, missing
dist, extras, markers, fail-closed None), deferral wiring (marker +
gateway resume + exit 2), placement guards (no detector call pre-fetch;
guard present at git/ZIP sync), and subprocess-verified import hygiene
(import hermes_cli.main and the update --check dispatch never load
cryptography._rust).
Follow-up to #86687 (Halldrix's #83590 salvage — the preflight's intent
stands as defence-in-depth; this makes it fire only when true).
Fixes#86735Fixes#86780Fixes#86781
Rework of salvaged PR #6372 (@ag9920) onto current main:
- /save promoted from CLI-only JSON snapshot to a cross-platform session
export: `/save [json|md|html] [filename] [redact]` on CLI and every
gateway platform (sent as a document via adapter.send_document).
- Rendering routes through the existing shared renderers
(hermes_cli/session_export.py + session_export_html.py) instead of the
PR's new hermes_state formatter — new helpers normalize_save_format /
render_session_for_save / default_save_filename are shared by both
surfaces.
- `redact` arg runs the export through the force-mode secret redaction
pass (session_export_md.redact_session_data) before writing.
- Gateway handler awaits AsyncSessionDB correctly, sanitizes user-supplied
filenames with basename, and lands in gateway/slash_commands.py (the
handlers moved out of gateway/run.py since the PR was authored).
- /export stays profile export (name collision resolved: session export
lives on /save).
- Slack 50-slash cap curation: /platform moves to the /hermes-only set to
free a native slot for /save (parity test updated rationale comment).
- Folds in PR #62268 (@briandevans): None title/model coalescing in the
single-session HTML export.
Closes#4249. Closes#51200.
Single-session HTML exports of an un-named session render
`<title>None</title>` and `Model: None`. An untitled session (title
`None`) is the default state until async title generation completes, so
this is the common case, not an edge case.
The browser-tab `<title>` (page_title) and the `Model:` meta line use
`dict.get(key, default)`, whose default only fires when the key is
absent — not when it is present with value `None`. `_escape_html(None)`
then stringifies to the literal "None". The on-page `<h1>` in the same
function already uses the None-safe `... or "Hermes Session"` idiom, so
the tab title and header were inconsistent for the same session.
Use `... or "<default>"` at both sites so the tab title and model meta
fall back consistently with the header.
The desktop app carried its own hardcoded list of 17 vendor MCP endpoints
(apps/desktop/src/lib/mcp-directory.ts) powering the composer suggestion
pills — a second PR-reviewed vendor list, overlapping and drifting from the
Nous-approved MCP catalog (optional-mcps/).
This makes the catalog the single source of truth:
- manifest schema: optional `suggest:` block (keywords + hosts), parsed,
validated, and normalized in mcp_catalog.py
- 15 new URL-only hosted-remote catalog entries (atlassian, sentry, datadog,
notion, stripe, vercel, supabase, netlify, hugging_face, asana, intercom,
airtable, webflow, paypal, square); figma + linear manifests gain suggest
blocks
- GET /api/mcp/catalog now serves the suggest metadata
- desktop suggestion provider builds its match index from the catalog;
the static directory remains only as a compatibility rung for older
backends without suggest metadata
- setup card source line prefers the catalog entry's transport URL
GitHub stays out of the catalog on purpose: its hosted MCP rejects generic
DCR and the bundled github/* skills (gh CLI) are the stronger integration.
New desktop `github` suggestion provider offers the github-auth skill
instead — gated on a new cached GET /api/git/gh-auth probe so already-
authenticated users never see the pill.
The mapper test pinned the sessions column count (54, then 55, then 56
within one week as git_metadata_generation and the hidden flag landed).
Every ordinary column addition broke it. Derive max_fields from
PRAGMA table_info at runtime with a >= floor so the test keeps asserting
the rebuild contract without change-detecting the schema width.
- The readonly-loader completer test now stubs
get_portable_mcp_server_names_nowait — real plugin discovery runs
load_config() during one-time process init, which is not the
per-keystroke read the test guards against.
- Cap _resolve_toolset_memo at 256 entries: generation-keyed entries
from stale generations are never hit again, so clear on overflow to
keep long sessions bounded.
get_default_hermes_root() resolves HERMES_HOME against the platform
native home (~80us of path resolution) on EVERY call and is called at
31+ sites — every _load_global_auth_store() (per provider row in the
/model picker), kanban, backup, gateway, update. Its result depends
only on (HERMES_HOME, native home), so memoise it keyed on those two
inputs, compared for free on each call (freshness-correct even if a
test or plugin mutates HERMES_HOME mid-process).
_load_global_auth_store() re-read + re-parsed the global auth.json on
every call; read_credential_pool() -> load_pool() runs it once per
provider row in the /model picker even when the profile has entries and
the global fallback never fires. Memoise keyed on the global auth
file's path+mtime (same pattern as _nous_auth_status_cache); the store
only changes when a global-scope auth write touches the file.
Measured (profile mode, 30-provider global store): get_default_hermes_root
81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us ->
137us per call — ~2ms saved per /model picker render (20 provider rows).
Regression tests: hermes_constants memo pin (no path resolution on
repeat calls, HERMES_HOME change forces a fresh resolution); global-store
memo pins (store read once across repeats, mtime bump re-reads once,
absent store stays cheap).
(cherry picked from commit be348f32e5bd7479c26fabb652de549fb9c8a1e1)
The /tools and /personality completers run on every keystroke while the
user types those commands (complete_while_typing), and both re-read +
re-parse the full config on each keypress:
- _tools_completions called load_config() — the defensive deepcopy
(~340us/call on cache hit) even though it only reads toolset enable
state + MCP server names. Switched to load_config_readonly() (the
perf(agent) #74322 pattern; this per-keystroke site was missed).
- _personality_completions called load_cli_config() — a full YAML parse
+ deep merge of the built-in defaults (~110us) — on every keystroke.
Memoised keyed on the config file path+mtime (same pattern as load_env
/ _nous_auth_status_cache), so the parse runs once per config state.
Measured: /tools 357us -> 18us per keystroke; /personality parse drops
from 1-per-keystroke to 1-per-config-change (500 keystrokes -> 1 parse).
Regression tests: _tools_completions uses the readonly loader (deepcopy
loader never called); personality memo parses once across repeated
completions and re-parses once after a config mtime change.
(cherry picked from commit 2b3f897171f93dc6b099848ec5ab3763b7294870)
The main.py decomposition re-exported the sessions/update/dashboard command
surface with eager from-imports, so every hermes invocation (including
hermes --version) paid for update_cmd's dependency chain (jwt, click,
cryptography). Resolve the re-exports through the existing PEP 562 module
__getattr__ (same pattern as _PROVIDER_MODELS) so each module loads on
first actual use. Internal call sites go through a _self() helper because
bare-name lookups do not trigger __getattr__; _self() imports sys locally
since update tests patch hermes_cli.main.sys. The
_warn_stale_dashboard_processes back-compat alias moves into the lazy
surface, and the sessions argparse dispatch defers sessions_cmd to call
time. Monkeypatching hermes_cli.main.<name> keeps working: a patch sets a
real module attribute, which shadows __getattr__.
Measured (Windows 11, Python 3.11, median of 7 warm runs):
import hermes_cli.main 253ms -> 196ms (-22%).
(cherry picked from commit cad1083b71635f98815b698d69a18f7f58e15517)
When the 1h provider_models_cache.json TTL lapses, the model picker
serially fetches /v1/models for each authenticated provider. With 10+
providers this stacks to 15-30s of blocking before the picker renders.
Add a parallel prefetch step before the serial picker build loops:
- _collect_authed_provider_slugs(): lightweight credential pre-scan
that mirrors sections 1/2/2b without fetching model lists
- _prefetch_provider_models_parallel(): ThreadPoolExecutor-based
concurrent fetch of stale/missing cache entries (max 8 workers)
- update_provider_cache_entry(): thread-safe single-entry cache writer
with threading.Lock to prevent concurrent write races
Guardrails:
- Skipped when <=3 authed providers (overhead not worth it)
- Skipped when refresh=True (serial path force-refreshes)
- Exception-isolated (falls back to serial path on any failure)
- No behavioral change (same model lists, same picker output)
Closes#80413
(cherry picked from commit 89dddd6cb5d53d73278e0518c375fb5b878e5c6b)
Hermes Console registers `checkpoints prune`, `clear` and `clear-legacy` as
mutating, so it takes a console-level confirmation before dispatching any of
them. `_apply_confirmed_defaults` then exists to keep the CLI layer from
asking a second time — its docstring says so — but it only force-defaults
`clear` and `clear-legacy`. `prune` was left out, even though `cmd_prune`
gates its orphan preview on the identical `not args.force` shape.
`_capture_output` redirects stdout and stderr but never stdin, so the
unskipped `_confirm()` call hits `input()` with no terminal behind it:
`EOFError` propagates into `_confirm`, which returns False, and `cmd_prune`
prints "Aborted." and returns 1. The console turns that non-zero exit into a
ConsoleCommandError, so `checkpoints prune` fails outright for any user who
has at least one orphan checkpoint project — after that user already
confirmed. When the server does happen to inherit a foreground terminal, the
same call instead blocks a console worker thread and eats the operator's
keystrokes.
Forcing the flag is the documented behavior here rather than a weakening of
the recent orphan-allowlist hardening. `orphan_allowlist` binds a deletion to
the identities shown in the preview, guarding the window where a workdir
disappears while the command waits on `input()`. Under the console there is
no preview and no wait, which is exactly the `--force` case the comment on
`cmd_prune` describes as "no restriction".
The salvaged #76716 adds git_metadata_generation to sessions (54 -> 55
columns). Update the synthetic-rebuild test's pinned widths and row
builders to the new current layout.
Creating a project whose resolved primary path already belongs to a
non-archived project now raises a clear ValueError naming the existing
project (create_project) — duplicated projects each seeded an identical
copy of the repo subtree, multiplying the duplicate-lane bug per copy.
The agent-facing project_create tool is idempotent instead: it re-activates
the existing project rather than erroring. allow_duplicate_path=True keeps
deliberate duplicates possible. Also updates the legacy non-git lane-id
expectation to the branch-style id introduced for #53329.