Commit Graph

4012 Commits

Author SHA1 Message Date
Teknium 053f8b1b17 fix(kanban): make archive-time worker termination race-safe and audited
Harden the cherry-picked fix (#42858, credit @PINKIIILQWQ; #100613 by
@moon2sun covers the same gap) per the sweeper review on #42858:

- Snapshot status/pid/claim INSIDE the archive txn so the kill only
  happens when this caller wins the archive transition; a losing
  concurrent archiver returns False without signalling anything.
- Signal only tasks that were actually running (never-claimed tasks
  skip the no-op helper call entirely).
- Kill runs post-commit: _poll_worker_exit can block ~5s and must not
  hold the SQLite write lock. Safe because archived is terminal — no
  dispatcher can respawn off the released claim.
- Termination outcome lands as its own archive_worker_termination
  event so the archived event stays atomic with the status flip.
- 2 invariant tests (running task -> signalled + audited; non-running
  -> no signal, no event), live E2E: worker survived archive on main,
  terminated (<0.3s, clean SIGTERM) with the fix.

Port trigger: lobehub PR scout; same bug class as lobehub#19220's
"failed verify cannot disarm the schedule" family (lifecycle actions
must reach the live process, not just the DB row).
2026-09-12 21:45:15 -07:00
Teknium 77f0c83ec3 fix(sessions): honor CLAUDE_CONFIG_DIR and CODEX_HOME in foreign session discovery
Port from cline/cline#13827: foreign-session discovery hardcoded
~/.claude/projects and ~/.codex/sessions, so Claude Code installs using
CLAUDE_CONFIG_DIR and Codex CLI installs using CODEX_HOME (both official
relocation vars the tools themselves honor, and which hermes_cli/auth_codex.py
already reads for credentials) silently found nothing to import.

_default_root() resolves each source's store from its env var, treating a
blank/whitespace value as unset so an empty override can never resolve to a
CWD-relative "projects" path. The _SOURCES tuple gained the env fields; the
browser sibling now reads the parser through the _parser() accessor instead
of a positional index that the wider tuple would have silently broken.

Live E2E: env-rooted Claude + Codex sessions discovered, imported, and
resumed; blank override falls back to ~; docs updated.
2026-09-12 21:39:10 -07:00
salch-cred ad0398eed8 fix(kanban): preserve sticky block on tasks created with initial_status=blocked (#107398) 2026-09-12 21:27:21 -07:00
Teknium 0f3199bd65 fix(dashboard): OAuth start routes resolve pollers late so test mocks intercept the spawned thread
The oauth router imported _nous_poller/_minimax_poller/_xai_device_poller
from web_server_oauth at module level, so tests patching the owning module
("hermes_cli.web_server_oauth._minimax_poller") patched a binding the
router never read. The REAL poller then ran on the leaked daemon thread,
called the live MiniMax token endpoint from CI, and the in-flight
getaddrinfo segfaulted the interpreter during a later test's fixture setup
(CI run 34323790818, tests/hermes_cli/test_web_oauth_dispatch.py flake).

Route the three pollers through the existing late() seam (web_deps), the
same mechanism every other monkeypatch-sensitive symbol in this router
already uses, so the patch wins at thread-spawn time. Regression test
proves the mock intercepts and the real poller body never runs; it fails
on the old module-level import (sabotage-verified).
2026-09-12 21:00:55 -07:00
686f6c61 2bc9ed9f23 fix(mcp): fully redact credential headers in MCP probe errors and test display (salvage #97466)
`hermes mcp test` resolved Authorization headers and printed first4***last4
— still a reusable credential fragment — and probe exceptions that echoed
`Authorization: Bearer <value>` reached the CLI error line and the dashboard
`POST /api/mcp/servers/{name}/test` response verbatim.

Redact once at the `_probe_single_server` raise seam so every consumer
(`mcp add`, `mcp test`, `mcp login`, `mcp configure`, the dashboard probe,
`hermes doctor`, catalog probes) prints already-safe text. Recognized
credential header fields (Authorization/Proxy-Authorization plus
agent.redact._SECRET_HEADER_NAMES) have their complete value replaced with
***; bare Bearer/Basic/Token/Digest spans are covered; the generic redactor
runs force=True as a second pass. CLI header display fails closed: only pure
${ENV} template values print.

Salvaged from PR #97466 by @686f6c61 (base predated the mcp_config/web_routers
decomposition; re-applied onto current main, test seams repointed to the
defining modules tools.mcp_tool_loop / tools.mcp_tool_lifecycle).

Inspired by Claude Code 2.1.268: "Fixed /mcp and /plugin server details,
claude mcp list/get, and MCP login errors showing secrets resolved from
${VAR} placeholders in MCP configs."

Fixes #97460

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-12 20:49:18 -07:00
Hermes fleet-fix 1916cb249d security: make state databases and snapshots owner-only 2026-09-12 20:43:42 -07:00
teknium1 205645ee42 fix(gateway): register gateway.multiplex_profiles; explicit migrate --multiplex flips it with no standalone secondary
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.

`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
2026-09-12 18:35:21 -07:00
teknium1 acbecf588a fix(profiles): --clone leaves messaging channels behind; --clone-channels opts in
A cloned profile carried the source's TELEGRAM_BOT_TOKEN, DISCORD_BOT_TOKEN,
allowlists, WHATSAPP_ENABLED, API_SERVER_KEY and the platforms:/telegram:/
discord: config sections byte-for-byte. Standalone, that made two gateways
fight over one bot's long-poll; under multiplex it blocked
`hermes gateway migrate --multiplex` with one duplicate-credential finding
per platform per clone (18 on a real 10-profile install).

Every clone entry point (CLI --clone/--clone-from/--clone-all, dashboard
POST /api/profiles, TUI/Desktop profiles.create incl. its mirror_credentials
.env copy) now strips channel settings after the copy. The key set is derived
from the adapters — Platform enum + plugin registry (required_env,
allowed_users_env, allow_all_env, cron_deliver_env_var), the gateway env table
(gateway.config_env._ENV_STEPS / _ENV_ENABLE_CREDENTIALS) and each platform's
env prefix — so a new adapter is covered without a hand list. --clone-all also
drops pairing/WhatsApp-session/gateway ledgers. Provider and tool keys, the
model block, memory, skills and SOUL.md are untouched.

`--clone-channels` (REST/RPC: clone_channels) keeps them; it is refused when a
live multiplexer already serves the source and otherwise warns which
platforms are now shared. `hermes profile list` prints the same warning for
existing clones whose bot credential is byte-identical to the default's.

The dashboard's per-platform env-prefix table moves into profile_channels so
Channels-page cards and the clone stripper share one definition.
2026-09-12 18:35:21 -07:00
teknium1 de2d6a1b93 fix(config): a fresh process recovers the last good config.yaml instead of running on defaults
CI / Detect affected areas (push) Has been cancelled
CI / OSV scan (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
Docker Build, Test, and Publish / Detect affected areas (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
Nix flake check / Detect affected areas (push) Has been cancelled
Build Skills Index / build-index (push) Has been cancelled
CI / Python tests (push) Has been cancelled
CI / OS-specific tests (push) Has been cancelled
CI / Python lints (push) Has been cancelled
CI / JS & TS checks (push) Has been cancelled
CI / Installer tests (push) Has been cancelled
CI / Rust tests (push) Has been cancelled
CI / Desktop E2E (push) Has been cancelled
CI / Docs Site (push) Has been cancelled
CI / Deny unrelated histories (push) Has been cancelled
CI / Check contributors (push) Has been cancelled
CI / Check uv.lock (push) Has been cancelled
CI / Check no committed infographics (push) Has been cancelled
CI / Profile artifact check (push) Has been cancelled
CI / Check no case-colliding filenames (push) Has been cancelled
CI / package-lock.json diff (push) Has been cancelled
CI / Lint Docker scripts (push) Has been cancelled
CI / Supply-chain scan (push) Has been cancelled
CI / Review label gate (push) Has been cancelled
CI / All required checks pass (push) Has been cancelled
CI / CI timing report (push) Has been cancelled
Docker Build, Test, and Publish / build (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest-32-core) (push) Has been cancelled
Docker Build, Test, and Publish / build (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-latest-32-arm-core) (push) Has been cancelled
Docker Build, Test, and Publish / publish (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest-32-core) (push) Has been cancelled
Docker Build, Test, and Publish / publish (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-latest-32-arm-core) (push) Has been cancelled
Docker Build, Test, and Publish / merge (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
Nix flake check / nix flake check (push) Has been cancelled
Build Skills Index / trigger-deploy (push) Has been cancelled
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).

Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.

Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
2026-09-12 16:17:04 -07:00
teknium1 d595e636c8 fix(model): a selected model id is never rewritten to a catalog neighbour
A user who picked `deepseek-v4.1-flash` on their own custom endpoint kept
landing on `deepseek-v4-flash-0731`. Three sites each "helped" by diffing
the pick against a catalog and moving it:

- hermes_cli/models_validate.py: the shared catalog matcher auto-corrected
  any id within difflib ratio 0.9 of a listed one (`corrected_model`), and
  model_switch applied it. Version bumps, dated snapshots and qualifiers
  all sit inside 0.9 of a sibling, so a newer release the listing lacked
  was swapped for the older one under the user's label. The matcher now
  does exact membership -> suggestion text only; the id goes to the wire
  verbatim and a genuine typo is refused with the listed siblings named.
  Every branch that carried the correction (live listing, static catalog,
  curated fallback, MiniMax, Anthropic, custom, OpenRouter preset base)
  loses it in one place.

- hermes_cli/model_switch.py: a `providers.<key>` endpoint reached by its
  bare key (the slug Desktop picker rows carry) validated as a built-in
  and hit the hard-rejecting live-listing branch; the same endpoint as
  `custom:<key>` soft-accepted. Both spellings now validate as the user's
  custom endpoint.

- apps/desktop: `manualPickRemoved` (composer reseed) and
  `reconcileSelectionAfterCatalogRefresh` (Refresh Models) retargeted a
  sticky pick to the profile default / the row's first model whenever the
  provider row did not list it. Rows are hints (discovered, curated,
  capped); the gateway's switch result is the only authority on a pick.
  Both helpers are removed; the pick stays put.

Tests: change-detectors pinning the swap are rewritten as invariants
(never `corrected_model`; unlisted id on a user endpoint is kept and
warned; typo is refused with a suggestion); proven red on origin/main.
2026-09-12 14:05:36 -07:00
teknium1 6a66a5d481 fix(desktop,dashboard): served profile's api_server/webhook read connected with their /p/<profile>/ URL; shared-gateway restart asks first
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.

- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
  webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
  (`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
  adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
  the real one, not a config guess. `hermes status` lists those URLs beside the other
  shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
  when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
  (statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
  shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
  Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
  start/stop on a served profile renders as an inline notice instead of a raw error toast.
2026-09-12 12:52:19 -07:00
Xipong 3d7f773bb4 fix(kanban): honor explicit platform tool opt-ins across configuration surfaces 2026-09-12 12:32:55 -07:00
teknium1 476d7223b5 test: utf-8 encodings in messaging profile tests 2026-09-12 08:49:16 -07:00
teknium1 2d121aa322 fix(gateway): hot-serve reaches pooled Desktop backends; deleted profiles leave no stale runtime entries
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
  ?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
  process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
  recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
  them as `stopped`.
2026-09-12 08:49:16 -07:00
teknium1 da451afb46 fix(update): probe configured-feature deps in the target venv, not the updater
The check called the registry check_fn inside the updater's own process,
whose import caches predate the install just performed (and which may be
the outer Python entirely), so a healthy freshly installed SDK produced a
false "will fail to load" warning. Run the same registry check in the
target interpreter via the existing _venv_probe path used by the core
dependency verifier.

Found by independent review before merge.
2026-09-12 08:47:21 -07:00
teknium1 2ef9c55a92 fix(update): name configured platforms whose extras failed to install
When `.[all]` fails and the per-extra fallback also fails for e.g.
`feishu`, the update printed only "Skipped optional extras that still
failed" and finished green. The running gateway kept its already-imported
modules, so the loss surfaced hours later as "No adapter available for
feishu" on the next restart (#10651).

After the fallback, check every enabled+configured platform through its
registry `check_fn` (and MCP when `mcp_servers` is set) and print which
configured feature will fail to load, with its install hint. Unconfigured
extras stay a quiet skipped line.

Reworks PR #10733 (LeonSGP43) against the registry instead of a
hand-written platform->module table so plugin platforms are covered.

Fixes #10651
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
2026-09-12 08:47:21 -07:00
teknium1 199544e054 fix(openrouter): canonical config URL keeps the pool; mirror key follows the selected endpoint
Two regressions in the mirror support: (1) the canonical
https://openrouter.ai/api/v1 that `hermes setup` persists under
provider: openrouter was treated as a custom endpoint, dropping the
auth.json credential pool and returning an empty API key; (2) an
unrelated CUSTOM_BASE_URL (which outranks the config mirror) still
received OPENROUTER_API_KEY because key selection tested mirror
eligibility, not the endpoint actually selected. A config URL is a
mirror only when its host is not openrouter.ai, and the mirror key
branch fires only when base_url is the config URL.

Found by independent review before merge.
2026-09-12 08:47:04 -07:00
JackJin 63f1016bea fix(cli): honor config base_url mirror for explicit openrouter provider
When config.yaml sets `model.provider: openrouter` together with a
`model.base_url` mirror/proxy, an explicit `--provider openrouter`
request ignored the mirror and sent traffic to the public OpenRouter
endpoint: the config base_url was only trusted for auto/custom, the
credential pool was still consulted (so a pooled key won over the
mirror), and even when the mirror URL was used its host failed the
openrouter.ai match so OPENROUTER_API_KEY was not selected for it.

Trust the config base_url for the explicit openrouter case, treat that
mirror as an OpenRouter context for key selection, and bypass the pool
like the other custom-endpoint cases already do.

Fixes #10622
2026-09-12 08:47:04 -07:00
周鹤0668001310 92a8398087 fix(config): treat explicit false values in HERMES_MANAGED as unmanaged
Previously, setting HERMES_MANAGED=false (or 0, no, off) would be
interpreted as a literal managed-system name, causing is_managed()
to incorrectly return True and block update/config commands.

- Add _MANAGED_FALSE_VALUES tuple for canonical false strings
- Check false values before true values in get_managed_system()
- Add parametrized regression tests for all false variants

Fixes #12864
2026-09-12 08:41:00 -07:00
teknium1 847369e6c6 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:34:43 -07:00
teknium1 18aae66a13 chore(tests): explicit utf-8 encoding in test_mcp_config (windows-footgun ratchet) 2026-09-12 08:34:43 -07:00
teknium1 b70e0f4603 fix(mcp): CLI readers no longer reopen include: [] as "all tools enabled"
The runtime already registers nothing for an explicit empty include list
(fb1ec36a4b), but every CLI reader still coerced `[]` to "no filter":
`hermes mcp list` printed "all", `hermes mcp configure` and `hermes tools`
pre-checked every tool (so confirming the picker silently re-enabled all of
them), and a catalog reinstall pre-checked the manifest defaults over the
user's zero-tool choice.

`_tool_filters` now returns the list whenever the key holds a list; only an
absent/non-list key is None. The pickers and list output branch on `is not
None`, matching `tools/mcp_tool_registration.py`.

Fixes #12865. Builds on #13096 (@dingn42) and #52874 (@Bartok9).
2026-09-12 08:34:43 -07:00
teknium1 9d0d886c43 fix(identity): spawn-ledger write survives surrogate-escaped argv
`_append_entry` moved to `utils.atomic_json_write`, which dumps with
`ensure_ascii=False` through a utf-8 text handle. An argv token holding
surrogate-escaped bytes (a non-UTF-8 project path via os.fsdecode) makes
json.dump raise UnicodeEncodeError — a ValueError, so the `except OSError`
does not catch it and callers silently lose their registration.

Expose `ensure_ascii` on `atomic_json_write` (default unchanged) and pass
True at the ledger call site, restoring the previous json.dumps behaviour
while keeping mode=0o600. One round-trip test, red on base.

Follow-up to #109156.
2026-09-12 08:27:53 -07:00
teknium1 12fe7684e1 fix(plugins): async-await helper thread runs under the caller's ContextVars
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
2026-09-12 08:26:48 -07:00
teknium1 1664e12fc7 chore(tests): explicit utf-8 encoding in test_plugins (windows-footgun ratchet) 2026-09-12 08:26:48 -07:00
teknium1 71c72e078a fix(cli): one unverified-accept message for custom endpoints without /models; tests trimmed
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
2026-09-12 08:26:23 -07:00
Victor Nogueira a3671787b0 fix(cli): clarify unverified custom model warning 2026-09-12 08:26:23 -07:00
Victor Nogueira eda1e6cb15 fix(cli): accept unverified custom models
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
2026-09-12 08:26:23 -07:00
bixycler 70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
teknium1 535fd88c70 fix(backup): keep durable cache/ artifacts (images, citation ledger) in full backups
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
2026-09-12 08:25:49 -07:00
teknium1 b97ae5c9b8 fix(backup): make the socket test runner-safe and document the new exclusions
The picked test bound an AF_UNIX socket at pytest's tmp_path, which overflows
the ~108-byte sun_path limit under scripts/run_tests.sh's deep temp root
("AF_UNIX path too long"). Bind by a relative name from inside the temp
HERMES_HOME instead; the walker still sees the same absolute entry.

Also list cache/ + runtime roots and non-regular entries in the `hermes backup`
"What's excluded" docs so the user-visible behaviour change is documented.
2026-09-12 08:25:49 -07:00
mrwanstudio 947e027f61 fix(backup): skip non-regular filesystem entries 2026-09-12 08:25:49 -07:00
Brad Estes 3b3f354933 fix(backup): exclude profile caches from full backups 2026-09-12 08:25:49 -07:00
teknium1 7817af2a16 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:25:36 -07:00
teknium1 2b685f08fa fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:24:52 -07:00
teknium1 07ead2249b chore(tests): explicit utf-8 encoding in test_claw (windows-footgun ratchet) 2026-09-12 08:24:52 -07:00
Drexuxux 45dd97a6f7 fix(claw): cleanup no longer mistakes any "openclaw" in argv for a running daemon
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).

POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.

Fixes #12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.

Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
2026-09-12 08:24:52 -07:00
teknium1 d267bc7f78 fix(model-switch): provider:model resolves like provider/model off aggregators too
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.

Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.

Fixes #9748
2026-09-12 08:24:42 -07:00
zhao c5cdd92254 fix(copilot): validate supported token prefixes 2026-09-12 08:24:29 -07:00
teknium1 fd15cd003e chore(tests): encoding="utf-8" on read_text/write_text in test_personality_single_owner.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:26 -07:00
teknium1 44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
teknium1 5685b76fde fix(doctor): image_gen hint covers a selected provider with a missing key or SDK
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
2026-09-12 08:24:11 -07:00
teknium1 6ea3b00e7f chore(tests): encoding="utf-8" on read_text/write_text in test_doctor.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:11 -07:00
skyc1e 7d25ae58a7 fix(doctor): image_gen reports "no provider configured", not "system dependency not met"
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.

A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.

Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.

Fixes #9516
2026-09-12 08:24:11 -07:00
teknium1 81cec06c77 test(cli): banner hero line is not center-padded
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
2026-09-12 08:23:58 -07:00
Season 6d20c321d6 fix(gateway): enable linger on systemd user-service repair path (#12863)
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.

Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.

Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
2026-09-12 08:23:44 -07:00
ygd58 5be0c4921c fix(gateway): surface launchctl bootstrap failures in refresh_launchd_plist_if_needed
Ports #63762 forward onto current main per teknium1's review.

refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).

1. refresh_launchd_plist_if_needed(): return False after retry
   exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
   of the success message.

Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.

Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.

5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
2026-09-12 08:23:44 -07:00
teknium1 f34933605a test(identity): make the 0600 ledger test prove the mode under a permissive umask
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
2026-09-12 08:02:16 -07:00
codeshipsingh 5a8e651183 fix(identity): enforce 0600 on spawn ledger writes 2026-09-12 08:02:16 -07:00
teddychenfeiyang-png a8ac7899b5 test(plugins): typo'd requires_hermes clause fails plugins validate admission
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.

Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.
2026-09-12 07:57:13 -07:00