Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
Slash-command handlers gained loop-safe awaiting in ca9a61ae38, but
`PluginManager.invoke_hook` still called `async def` hook callbacks directly:
the coroutine object was appended to the results (so `pre_llm_call` context
injection silently did nothing) and Python warned "coroutine was never
awaited". `_invoke_hook_callback` now routes every return through
`resolve_plugin_command_result`, which covers both the direct and the
timeout-bounded paths and is safe under the gateway's running loop.
Fixes#12449 (remaining hook half). Salvage of #63240 by @Bartok9, applied
one layer down so the bounded-worker path is covered too.
Co-authored-by: Bartok9 <Bartok9@users.noreply.github.com>
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).
POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.
Fixes#12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.
Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.
Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.
Fixes#9748
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.
`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.
Fixes#9636
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.
A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.
Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.
Fixes#9516
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.
Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.
Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
Ports #63762 forward onto current main per teknium1's review.
refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).
1. refresh_launchd_plist_if_needed(): return False after retry
exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
of the success message.
Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.
Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.
5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.
Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
The classic CLI froze for ~0.7-2s between the banner and the first prompt. py-spy +
strace on real PTY startups showed the main/REPL threads inside
refuse_deleted_wal_generation -> _iter_proc_fd_targets: a second full SessionDB open.
_init_session_store built a bare SessionDB(); a moment later the goal/loop/heartbeat
managers acquired the same state.db through hermes_state_registry from the REPL thread,
which is a different handle, so the whole open ran again — including the /proc-wide
deleted-WAL sidecar scan (~4.4k readlinks). Each readlink drops and re-takes the GIL
while the startup threads (plugin discovery, MCP, skill sync, banner git) are busy, so an
11ms scan stretched to 1.3s per pass, and the second pass landed exactly where the
prompt should have appeared.
Route the CLI's handle (init + the two re-open sites) through the registry so every
in-process consumer shares one writer. One scan per startup; live A/B on the same box,
interleaved x6: banner->prompt gap 0.37s mean -> 0.15s mean (plain), 1.77s -> 0.39s
under strace. The registry release path replaces close(), so /quit, /snapshot restore
and /handoff keep their semantics.
Leaving the context-length prompt blank in the custom-endpoint wizard said
"will auto-detect" and then went silent, so users could not tell whether their
endpoint runs on a detected window or the runtime's default fallback (which
shapes compression and prompt-cache behaviour). After the save prompt, run the
same resolver the runtime uses (with the endpoint's URL and key) and print
either "auto-detected N tokens" or "not detected — using the default N tokens".
Feedback only: the probe result is not persisted, and a failing probe never
blocks the save.
Fixes#2513. Approach from PR #2522 (@ygd58) and PR #85499 (@Luna161), both
written against the pre-decomposition wizard module.
Co-authored-by: Luna161 <268031236+Luna161@users.noreply.github.com>
A throttled GitHub fetch also yields index-metadata-without-bundle, so the new
stale-entry verdict would tell users a skill "no longer exists upstream" when
it does. Check the adapters' rate-limit flag first and keep the existing
rate-limit hint for that case (the keep_open review concern on #3261).
The per-search "results may be stale" note is dropped: it fires on every
skills.sh search whether or not anything is stale, and the install-time error
now names the condition precisely where it happens.
_resolve_source_meta_and_bundle already distinguishes index-hit-without-
files from unknown identifiers, but do_install printed the same generic
'Could not fetch' for both, sending users off to re-check spellings for
what is actually a stale skills.sh entry. Split the message, and add a
staleness caveat to do_search results from skills.sh.
Fixes#3259. Supersedes #3261 (stale since July — re-applied onto the
current _print_fetch_failure helper).
BotFather rejects setMyCommands descriptions containing em/en dashes
(U+2012-U+2015, U+2212). Fold them to ASCII hyphen at the two Telegram
sinks (telegram_bot_commands, telegram_menu_commands) so core, plugin,
and skill entries are all covered.
Fixes#2925.
Simplify-pass follow-up on the repair rework: the self-check rev-list walks
and the rollback restore now run inside the same _ShallowLock hold (rev-list
never takes shallow.lock), so lock contention can no longer defeat a failed
self-check's rollback and leave a broken .git/shallow in place. The tmp-write
+ os.replace sequence shared by repair and prune moves into _write_shallow.
Scope note added: fetch-by-SHA install tips (HEAD-reflog-only) are not repair
candidates; corruption of that shape is prevented by the prune's reflog
fail-safe. 68 focused tests green; two-cycle repair->prune E2E re-verified.
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:
- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
rev-list --all --reflog, so a boundary the repair just restored (one a
reflog-only commit still needs) is never dropped again; previously the
production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
message body is prose; _batch_missing_parents stops at the blank line
ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
not --batch-all-objects: unrelated object loss (a deleted parent of a
locally-created commit) is no longer re-labelled as shallow history;
fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
shallow.lock, so a depth-1 fetch between read and write fails fast
instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
(_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
test file footguns fixed (encoding=, as_uri()) and the missing
repair->prune end-to-end regression added, mutation-checked.
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.
Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.
This complements PR #108290, which owns the prevention half.
Refs #108286
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
hermes -p <name> gateway status, hermes gateway status and hermes status (under
Serves:) list the /p/<profile>/<path> URL per inbound-port platform the live
multiplexer serves, read from the <profile>:<platform> ingress_url in the default
home's gateway_state.json (hermes_cli/gateway_multiplex_served.py). The dashboard's
messaging payload carries the same ingress_url and the Channels page renders it.
The dashboard's 409 guard now covers only api_server/webhook (the mirrored pair):
enabling Twilio/LINE/Teams/... on a secondary is allowed because the gateway serves
it.
What `hermes update` does, blockers and fixes, URL change for
inbound-port profiles, the post-create restart reminder, rollback, and
the `gateway migrate` reference row.
GET /api/gateway/migrate/plan returns the CLI plan JSON; POST
/api/gateway/migrate spawns `hermes gateway migrate --multiplex --yes`
detached (action log gateway-migrate.log). The Gateway card shows the
button only for a multi-profile install that is not yet multiplexed, and
disables it while listing the blockers.
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
Moves a per-profile-gateway install onto one multiplexed default gateway:
table-driven preflight (duplicate credential via the gateway's own
fingerprint; secondary port-binders without a /p/<profile>/ ingress),
--dry-run, apply (stop + uninstall each secondary's service, record it in
<default>/gateway_migration.json, flip gateway.multiplex_profiles through
the config API, restart/install the default on the same service manager,
verify served_profiles), and --standalone rollback from the manifest.
Idempotent; refuses cleanly when already multiplexed or blocked.
`hermes profile create` points at `hermes gateway restart` when a live
multiplexer is detected (the served set is snapshotted at startup).
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.
Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.
Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
`_native_service_homes()` re-implemented the root+SUDO_USER -> `pw_dir/.hermes`
resolution that `hermes_cli.main._resolve_sudo_user_profile_env` already did.
Both now call `hermes_constants.sudo_invoker_default_home()`; main.py appends
`profiles/<name>` to it.
Also removes the local `from pathlib import Path as _Path` (Path is a module
import) and narrows the except to `KeyError`: `import pwd` cannot fail once
`os.geteuid` exists, and `getpwnam(str)` raises nothing else.
Review follow-ups to the unit-anchored service identity.
`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.
Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.
Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.
The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.
Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.
Refs #108674
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).
The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.
Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.
The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.
Refs #108674
_load_global_auth_store is memoised on (path, st_mtime_ns). A write-through
to the root (borrowed Codex cooldown clear, xAI/Anthropic root rotation)
followed by a fallback read in the same mtime tick — coarse-mtime
filesystems (exFAT, some network/overlay mounts) — kept serving the
pre-write store, so the resolve path could log "quota restored" and then
raise quota_exhausted from the stale memo. _save_auth_store(target_path=...)
now drops the memo.
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.
Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
`_read_codex_pool_entries` had become a pass-through whose only residual
behaviour was taking the active-store lock around a read every other
`read_credential_pool` caller does unlocked (and which never covered the
root file it fell back to). Both consumers now call the helper directly.
`clear_codex_pool_quota_cooldowns` decided "profile owns rows?" with an
unlocked pre-read, then re-read the same file under the lock — a TOCTOU
against a concurrent `hermes auth add` and a wasted parse. It now tries the
active store under its lock and falls back to the borrowed root only when
that store has no codex rows.