Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:
- Snapshot status/pid/claim INSIDE the archive txn so the kill only
happens when this caller wins the archive transition; a losing
concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
hold the SQLite write lock. Safe because archived is terminal — no
dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
-> no signal, no event), live E2E: worker survived archive on main,
terminated (<0.3s, clean SIGTERM) with the fix.
Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.
_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.
Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
The oauth router imported _nous_poller/_minimax_poller/_xai_device_poller
from web_server_oauth at module level, so tests patching the owning module
("hermes_cli.web_server_oauth._minimax_poller") patched a binding the
router never read. The REAL poller then ran on the leaked daemon thread,
called the live MiniMax token endpoint from CI, and the in-flight
getaddrinfo segfaulted the interpreter during a later test's fixture setup
(CI run 34323790818, tests/hermes_cli/test_web_oauth_dispatch.py flake).
Route the three pollers through the existing late() seam (web_deps), the
same mechanism every other monkeypatch-sensitive symbol in this router
already uses, so the patch wins at thread-spawn time. Regression test
proves the mock intercepts and the real poller body never runs; it fails
on the old module-level import (sabotage-verified).
`hermes mcp test` resolved Authorization headers and printed first4***last4
— still a reusable credential fragment — and probe exceptions that echoed
`Authorization: Bearer <value>` reached the CLI error line and the dashboard
`POST /api/mcp/servers/{name}/test` response verbatim.
Redact once at the `_probe_single_server` raise seam so every consumer
(`mcp add`, `mcp test`, `mcp login`, `mcp configure`, the dashboard probe,
`hermes doctor`, catalog probes) prints already-safe text. Recognized
credential header fields (Authorization/Proxy-Authorization plus
agent.redact._SECRET_HEADER_NAMES) have their complete value replaced with
***; bare Bearer/Basic/Token/Digest spans are covered; the generic redactor
runs force=True as a second pass. CLI header display fails closed: only pure
${ENV} template values print.
Salvaged from PR #97466 by @686f6c61 (base predated the mcp_config/web_routers
decomposition; re-applied onto current main, test seams repointed to the
defining modules tools.mcp_tool_loop / tools.mcp_tool_lifecycle).
Inspired by Claude Code 2.1.268: "Fixed /mcp and /plugin server details,
claude mcp list/get, and MCP login errors showing secrets resolved from
${VAR} placeholders in MCP configs."
Fixes#97460
Co-authored-by: 686f6c61 <github@00b.tech>
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.
`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).
Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.
`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.
The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).
Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.
Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
A user who picked `deepseek-v4.1-flash` on their own custom endpoint kept
landing on `deepseek-v4-flash-0731`. Three sites each "helped" by diffing
the pick against a catalog and moving it:
- hermes_cli/models_validate.py: the shared catalog matcher auto-corrected
any id within difflib ratio 0.9 of a listed one (`corrected_model`), and
model_switch applied it. Version bumps, dated snapshots and qualifiers
all sit inside 0.9 of a sibling, so a newer release the listing lacked
was swapped for the older one under the user's label. The matcher now
does exact membership -> suggestion text only; the id goes to the wire
verbatim and a genuine typo is refused with the listed siblings named.
Every branch that carried the correction (live listing, static catalog,
curated fallback, MiniMax, Anthropic, custom, OpenRouter preset base)
loses it in one place.
- hermes_cli/model_switch.py: a `providers.<key>` endpoint reached by its
bare key (the slug Desktop picker rows carry) validated as a built-in
and hit the hard-rejecting live-listing branch; the same endpoint as
`custom:<key>` soft-accepted. Both spellings now validate as the user's
custom endpoint.
- apps/desktop: `manualPickRemoved` (composer reseed) and
`reconcileSelectionAfterCatalogRefresh` (Refresh Models) retargeted a
sticky pick to the profile default / the row's first model whenever the
provider row did not list it. Rows are hints (discovered, curated,
capped); the gateway's switch result is the only authority on a pick.
Both helpers are removed; the pick stays put.
Tests: change-detectors pinning the swap are rewritten as invariants
(never `corrected_model`; unlisted id on a user endpoint is kept and
warned; typo is refused with a suggestion); proven red on origin/main.
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
them as `stopped`.
The check called the registry check_fn inside the updater's own process,
whose import caches predate the install just performed (and which may be
the outer Python entirely), so a healthy freshly installed SDK produced a
false "will fail to load" warning. Run the same registry check in the
target interpreter via the existing _venv_probe path used by the core
dependency verifier.
Found by independent review before merge.
When `.[all]` fails and the per-extra fallback also fails for e.g.
`feishu`, the update printed only "Skipped optional extras that still
failed" and finished green. The running gateway kept its already-imported
modules, so the loss surfaced hours later as "No adapter available for
feishu" on the next restart (#10651).
After the fallback, check every enabled+configured platform through its
registry `check_fn` (and MCP when `mcp_servers` is set) and print which
configured feature will fail to load, with its install hint. Unconfigured
extras stay a quiet skipped line.
Reworks PR #10733 (LeonSGP43) against the registry instead of a
hand-written platform->module table so plugin platforms are covered.
Fixes#10651
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Two regressions in the mirror support: (1) the canonical
https://openrouter.ai/api/v1 that `hermes setup` persists under
provider: openrouter was treated as a custom endpoint, dropping the
auth.json credential pool and returning an empty API key; (2) an
unrelated CUSTOM_BASE_URL (which outranks the config mirror) still
received OPENROUTER_API_KEY because key selection tested mirror
eligibility, not the endpoint actually selected. A config URL is a
mirror only when its host is not openrouter.ai, and the mirror key
branch fires only when base_url is the config URL.
Found by independent review before merge.
When config.yaml sets `model.provider: openrouter` together with a
`model.base_url` mirror/proxy, an explicit `--provider openrouter`
request ignored the mirror and sent traffic to the public OpenRouter
endpoint: the config base_url was only trusted for auto/custom, the
credential pool was still consulted (so a pooled key won over the
mirror), and even when the mirror URL was used its host failed the
openrouter.ai match so OPENROUTER_API_KEY was not selected for it.
Trust the config base_url for the explicit openrouter case, treat that
mirror as an OpenRouter context for key selection, and bypass the pool
like the other custom-endpoint cases already do.
Fixes#10622
Previously, setting HERMES_MANAGED=false (or 0, no, off) would be
interpreted as a literal managed-system name, causing is_managed()
to incorrectly return True and block update/config commands.
- Add _MANAGED_FALSE_VALUES tuple for canonical false strings
- Check false values before true values in get_managed_system()
- Add parametrized regression tests for all false variants
Fixes#12864
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
The runtime already registers nothing for an explicit empty include list
(fb1ec36a4b), but every CLI reader still coerced `[]` to "no filter":
`hermes mcp list` printed "all", `hermes mcp configure` and `hermes tools`
pre-checked every tool (so confirming the picker silently re-enabled all of
them), and a catalog reinstall pre-checked the manifest defaults over the
user's zero-tool choice.
`_tool_filters` now returns the list whenever the key holds a list; only an
absent/non-list key is None. The pickers and list output branch on `is not
None`, matching `tools/mcp_tool_registration.py`.
Fixes#12865. Builds on #13096 (@dingn42) and #52874 (@Bartok9).
`_append_entry` moved to `utils.atomic_json_write`, which dumps with
`ensure_ascii=False` through a utf-8 text handle. An argv token holding
surrogate-escaped bytes (a non-UTF-8 project path via os.fsdecode) makes
json.dump raise UnicodeEncodeError — a ValueError, so the `except OSError`
does not catch it and callers silently lose their registration.
Expose `ensure_ascii` on `atomic_json_write` (default unchanged) and pass
True at the ledger call site, restoring the previous json.dumps behaviour
while keeping mode=0o600. One round-trip test, red on base.
Follow-up to #109156.
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
The picked test bound an AF_UNIX socket at pytest's tmp_path, which overflows
the ~108-byte sun_path limit under scripts/run_tests.sh's deep temp root
("AF_UNIX path too long"). Bind by a relative name from inside the temp
HERMES_HOME instead; the walker still sees the same absolute entry.
Also list cache/ + runtime roots and non-regular entries in the `hermes backup`
"What's excluded" docs so the user-visible behaviour change is documented.
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).
POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.
Fixes#12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.
Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.
Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.
Fixes#9748
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes#9636
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.
A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.
Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.
Fixes#9516
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.
Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.
Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
Ports #63762 forward onto current main per teknium1's review.
refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).
1. refresh_launchd_plist_if_needed(): return False after retry
exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
of the success message.
Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.
Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.
5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.
Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.