Commit Graph

3979 Commits

Author SHA1 Message Date
teknium1 7817af2a16 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:25:36 -07:00
teknium1 2b685f08fa fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:24:52 -07:00
teknium1 07ead2249b chore(tests): explicit utf-8 encoding in test_claw (windows-footgun ratchet) 2026-09-12 08:24:52 -07:00
Drexuxux 45dd97a6f7 fix(claw): cleanup no longer mistakes any "openclaw" in argv for a running daemon
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).

POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.

Fixes #12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.

Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
2026-09-12 08:24:52 -07:00
teknium1 d267bc7f78 fix(model-switch): provider:model resolves like provider/model off aggregators too
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.

Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.

Fixes #9748
2026-09-12 08:24:42 -07:00
zhao c5cdd92254 fix(copilot): validate supported token prefixes 2026-09-12 08:24:29 -07:00
teknium1 fd15cd003e chore(tests): encoding="utf-8" on read_text/write_text in test_personality_single_owner.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:26 -07:00
teknium1 44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
teknium1 5685b76fde fix(doctor): image_gen hint covers a selected provider with a missing key or SDK
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
2026-09-12 08:24:11 -07:00
teknium1 6ea3b00e7f chore(tests): encoding="utf-8" on read_text/write_text in test_doctor.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:11 -07:00
skyc1e 7d25ae58a7 fix(doctor): image_gen reports "no provider configured", not "system dependency not met"
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.

A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.

Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.

Fixes #9516
2026-09-12 08:24:11 -07:00
teknium1 81cec06c77 test(cli): banner hero line is not center-padded
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
2026-09-12 08:23:58 -07:00
Season 6d20c321d6 fix(gateway): enable linger on systemd user-service repair path (#12863)
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.

Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.

Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
2026-09-12 08:23:44 -07:00
ygd58 5be0c4921c fix(gateway): surface launchctl bootstrap failures in refresh_launchd_plist_if_needed
Ports #63762 forward onto current main per teknium1's review.

refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).

1. refresh_launchd_plist_if_needed(): return False after retry
   exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
   of the success message.

Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.

Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.

5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
2026-09-12 08:23:44 -07:00
teknium1 f34933605a test(identity): make the 0600 ledger test prove the mode under a permissive umask
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
2026-09-12 08:02:16 -07:00
codeshipsingh 5a8e651183 fix(identity): enforce 0600 on spawn ledger writes 2026-09-12 08:02:16 -07:00
teddychenfeiyang-png a8ac7899b5 test(plugins): typo'd requires_hermes clause fails plugins validate admission
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.

Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.
2026-09-12 07:57:13 -07:00
jinli.yl c4c508579d fix(plugins): validate requires_hermes from manifest module 2026-09-12 07:57:13 -07:00
Teknium b0c383cdf7 fix(desktop): served profiles show running and route lifecycle to the multiplexer from a pooled local backend
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.

Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
2026-09-12 06:13:44 -07:00
Teknium be2f7e9c36 feat: curator prunes unused skills at 30 days (was 90), stale at 14
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
2026-09-12 05:56:15 -07:00
teknium1 e8016a18c1 fix(cli): report the auto-detected context length when a custom provider is saved without one
Leaving the context-length prompt blank in the custom-endpoint wizard said
"will auto-detect" and then went silent, so users could not tell whether their
endpoint runs on a detected window or the runtime's default fallback (which
shapes compression and prompt-cache behaviour). After the save prompt, run the
same resolver the runtime uses (with the endpoint's URL and key) and print
either "auto-detected N tokens" or "not detected — using the default N tokens".
Feedback only: the probe result is not persisted, and a failing probe never
blocks the save.

Fixes #2513. Approach from PR #2522 (@ygd58) and PR #85499 (@Luna161), both
written against the pre-decomposition wizard module.

Co-authored-by: Luna161 <268031236+Luna161@users.noreply.github.com>
2026-09-12 05:10:51 -07:00
teknium1 08bb58bd4a fix(skills): never call a rate-limited fetch a stale index entry; drop per-search caveat
A throttled GitHub fetch also yields index-metadata-without-bundle, so the new
stale-entry verdict would tell users a skill "no longer exists upstream" when
it does. Check the adapters' rate-limit flag first and keep the existing
rate-limit hint for that case (the keep_open review concern on #3261).

The per-search "results may be stale" note is dropped: it fires on every
skills.sh search whether or not anything is stale, and the install-time error
now names the condition precisely where it happens.
2026-09-12 05:10:31 -07:00
nikkoxgonzales 257a704d18 fix(skills): name stale index entries instead of generic fetch failure
_resolve_source_meta_and_bundle already distinguishes index-hit-without-
files from unknown identifiers, but do_install printed the same generic
'Could not fetch' for both, sending users off to re-check spellings for
what is actually a stale skills.sh entry. Split the message, and add a
staleness caveat to do_search results from skills.sh.

Fixes #3259. Supersedes #3261 (stale since July — re-applied onto the
current _print_fetch_failure helper).
2026-09-12 05:10:31 -07:00
nikkoxgonzales 75ade17617 fix(telegram): normalize unicode dashes in bot menu descriptions
BotFather rejects setMyCommands descriptions containing em/en dashes
(U+2012-U+2015, U+2212). Fold them to ASCII hyphen at the two Telegram
sinks (telegram_bot_commands, telegram_menu_commands) so core, plugin,
and skill entries are all covered.

Fixes #2925.
2026-09-12 05:10:11 -07:00
kshitijk4poor 966fb375fc fix(cli): make shallow-boundary repair survive the prune and the graph-safety review findings
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:

- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
  rev-list --all --reflog, so a boundary the repair just restored (one a
  reflog-only commit still needs) is never dropped again; previously the
  production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
  message body is prose; _batch_missing_parents stops at the blank line
  ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
  not --batch-all-objects: unrelated object loss (a deleted parent of a
  locally-created commit) is no longer re-labelled as shallow history;
  fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
  shallow.lock, so a depth-1 fetch between read and write fails fast
  instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
  rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
  (_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
  test file footguns fixed (encoding=, as_uri()) and the missing
  repair->prune end-to-end regression added, mutation-checked.
2026-09-12 15:00:37 +05:30
joaomarcos 2fb87f047f fix(cli): repair shallow boundaries already dropped by stale-graft prune
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR #108290, which owns the prevention half.

Refs #108286
2026-09-12 15:00:37 +05:30
Teknium d76856cc69 fix(migrate): shared-ingress adapters declare serves_profile_prefix so migration reports them as notices
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
2026-09-12 02:02:20 -07:00
Teknium 9ca7db8232 test(multiplex): shared-listener ingress invariants
- /p/<profile>/line/webhook is verified with the NAMED profile's channel secret under
  its runtime scope; another profile's secret is 401 at that URL; the bare path is
  untouched; unknown profile / profile without the adapter is 404.
- A shared-listener adapter binds no port and records its /p/<profile>/ ingress_url
  in runtime status; LINE media URLs use the shared prefix.
- Runner: a secondary's port-binders are constructed in shared-listener mode instead
  of refusing the whole profile; api_server/webhook are skipped as mirrors.
- Dashboard: only the mirrored pair is refused (409) on a secondary.
2026-09-12 01:53:15 -07:00
Teknium bcdb49ac7b feat(gateway): adapters declare serves_profile_prefix for /p/<profile>/ ingress
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
2026-09-12 01:49:28 -07:00
Teknium 5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
kshitijk4poor 443c2785fa docs(gateway): state the sudo mid-command rationale once, at the decision point
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
2026-09-12 12:15:07 +05:30
kshitijk4poor 1ecdfd18db fix(gateway): consult the installed unit only when root
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.

Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
2026-09-12 12:15:07 +05:30
kshitijk4poor b18c5752e3 test(gateway): keep one invariant test per bare-name owner
Trims the nine salvaged tests to four, one per contract: SUDO_USER's default
home keeps the bare name across the unit sync; a home the unit does not pin
keeps its suffix (the #105525 guard); a bare unit pinning a profiles/<name>
home beats the profile branch; the production sync itself preserves the name.
2026-09-12 12:15:07 +05:30
JoaoMarcos44 1ff56d9ed9 fix(gateway): keep the unit anchor Linux-only and pin the order it depends on
Review follow-ups to the unit-anchored service identity.

`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.

Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.

Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.

The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.

Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.

Refs #108674
2026-09-12 12:15:07 +05:30
JoaoMarcos44 a4a10eb520 fix(gateway): anchor system service identity on the installed unit
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).

The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.

Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.

The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.

Refs #108674
2026-09-12 12:15:07 +05:30
Hukla fdb9b50a15 fix(gateway): preserve sudo system service name 2026-09-12 12:15:07 +05:30
kshitijk4poor c71a411fe5 test(gateway): release-slot fakes accept the finalizer's run_generation kwarg
The turn finalizer now releases only its own generation
(`_release_running_agent_state(key, run_generation=...)`); two test doubles
were zero-kwarg lambdas and raised TypeError inside the finally.
2026-09-12 12:03:39 +05:30
kshitij 683afe45ba fix(auth): a global-root write-through invalidates the root store memo
_load_global_auth_store is memoised on (path, st_mtime_ns). A write-through
to the root (borrowed Codex cooldown clear, xAI/Anthropic root rotation)
followed by a fallback read in the same mtime tick — coarse-mtime
filesystems (exFAT, some network/overlay mounts) — kept serving the
pre-write store, so the resolve path could log "quota restored" and then
raise quota_exhausted from the stale memo. _save_auth_store(target_path=...)
now drops the memo.
2026-09-12 12:02:55 +05:30
kshitij ea42884e99 refactor(auth): codex cooldown clear reuses the pool ownership rule
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.

Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
2026-09-12 12:02:55 +05:30
kshitijk4poor 65c9d33b14 fix(auth): a borrowed Codex pool clears its quota cooldown in the root store
With reads now inheriting the global-root pool, a profile hitting a stale
root cooldown probes quota, sees it restored, and calls
clear_codex_pool_quota_cooldowns() — which only ever edited the (empty)
profile store, so the next resolve raised quota_exhausted again forever.
Pick the store the same way agent/credential_pool.py persists borrowed
rows (_profile_owns_pool_provider / _borrowed_single_use_pool_root) and
lock/save against that path. The fallback test now also binds
profile-wins precedence; one new test pins the root write.
2026-09-12 12:02:55 +05:30
nekwo 3767c2eacd fix(auth): inherit global Codex pool in profiles 2026-09-12 12:02:55 +05:30
Teknium fc71fb63e5 test(multiplex): invariants for the residue fixes; retire allowlist tests
- tests/gateway/test_multiplex_residue_parity.py: served profile reads its own
  sessions.*; per-turn bridge skips secondary scope; resolve_proxy_url reads the
  routed scope (and does not borrow the default's on a miss); a stale served
  turn never recreates an archived profile; MCP discovery runs per profile home.
- test_config.py: v43 migration removes multiplex_profile_allowlist.
- Allowlist tests deleted (feature removed) or rewritten to "served set = all
  live profiles"; the unserved cases now use a tombstone / missing dir.
- MCP discovery fixtures converted to the per-home set/dict slot.
2026-09-11 19:50:46 -07:00
BowmanStephen 4d47891b4a fix(dashboard): isolate the environment of named-profile actions
_spawn_hermes_action copied the dashboard's os.environ verbatim into every
detached `hermes ...` action. The dashboard runs inside the gateway and has
loaded its own profile's .env into the process environment, so
`hermes -p <other> gateway restart` (and every other dashboard-driven
profile action) started with the DEFAULT profile's platform credentials
and ports already present. load_hermes_dotenv does not override keys that
are already set, so the named profile's own .env could not displace them:
an A2A-only profile ended up claiming the default Discord bot token and
binding the default API server / BlueBubbles ports.

For actions carrying a profile selector (`-p X`, `--profile X`,
`--profile=X`), build the child env from the standard scrubbed subprocess
environment, drop _PROFILE_MANAGED_ENV_KEYS plus every key defined by the
dashboard/default profile's .env and its hydrated secret sources, and pin
HERMES_HOME to the target profile so the child's normal startup loads that
profile's .env. Only the leading selector is inspected: argv after the
subcommand may legitimately contain -p for a nested process. Actions
without a selector keep the historical environment byte-for-byte.

Test: the new case asserts the leaked keys are gone, benign keys survive,
and (by running the real dotenv loader in a fresh interpreter with the
captured env) that the target profile's values load without reviving any
default-profile value.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:50:46 -07:00
Teknium 446f6f79a8 fix(multiplex): dashboard and gateway stop see a served profile's gateway as the multiplexer
A profile served by the default multiplexer owns no gateway.pid / gateway_state.json,
so every surface that reads per-profile identity files called it stopped while the CLI
status surfaces (hermes -p X status / gateway status / cron status) said "running via the
default-profile multiplexer":

- `/api/status?profile=X` and `/api/messaging/platforms?profile=X` reported
  gateway_running=false / state=None / "gateway_stopped" in the same body that listed X
  under gateways[].served_profiles. The shared ladder `resolve_gateway_liveness` gains a
  fourth rung for a named profile_dir: the live default multiplexer that records X in
  served_profiles IS X's gateway (pid = multiplexer pid, runtime = its record, X's
  `<X>:<platform>` entries re-keyed to the standalone shape).
- `POST /api/gateway/stop?profile=X` spawned `hermes -p X gateway stop`, which printed
  "No gateway running for this profile" (exit 0) into the action log while the UI flipped
  to stopped and the multiplexer kept serving X; `/api/gateway/restart?profile=X` spawned
  a `-p X gateway restart` that only exits 78. start/stop now answer 409 with the
  multiplexer explanation (one helper shared with the existing start refusal) and restart
  targets the multiplexer, the process that actually serves X. A `--force`-started
  separate gateway for X (own pid file) keeps normal per-profile management.
- CLI `hermes -p X gateway stop` refuses with exit 78 like run/start/install/restart when
  X has no gateway of its own, instead of a contradictory exit-0 "not running".

Docs: multi-profile-gateways.md §1 and §5 describe stop + the dashboard behaviour.
2026-09-11 19:38:36 -07:00
Teknium fa425c942e test: stub the safe.directory pre-read by default; map privacydied's email
noninteractive_git_env() now spawns `git config --get-all safe.directory` before
building the env. Eight tests fake subprocess.run/Popen with a fixed sequence of
expected git calls (update check, plugin pull, MCP install, bounded probe) and the
extra spawn tripped them in CI. An autouse fixture stubs the read to "no entries";
the two carve-out invariant tests opt back in with @pytest.mark.real_safe_directory
(and were confirmed to still exercise the real read: the ordering test would fail
against the stub).

contributors/emails: pry@privacydied.net -> privacydied (check-attribution).
2026-09-11 19:01:47 -07:00
Teknium 66adfaee6b refactor(git): trim safe.directory tests to two invariants, memoise the config read
Tests: fold the ambient GIT_CONFIG_KEY_n negative and the isolation-still-in-force
assertions into the ordering test, and drop the two positive-only cases it subsumes.
What remains pins the injected sequence to git's own `config -z --get-all` output and
asserts the real trust decision (named repo usable, unrelated cross-owner repo refused).

Cost: noninteractive_git_env() runs on every internal git call, including the banner
startup probe, and the two `git config` children added ~10 ms per call against ~0.2 ms
before. Memoise per process on the inputs that select the config files plus the
system/global candidates' mtimes, so an edit to ~/.gitconfig is picked up without a
restart (verified live: 128 -> 0 within one process after appending safe.directory).
2026-09-11 19:01:47 -07:00
privacydied 02200f0b65 fix(git): preserve safe.directory ordering and reset markers when carrying it
Review of #107748 found the carry path serialized the user's trust policy
incorrectly, and the mistake could WIDEN trust rather than merely reformat it.

safe.directory is an ordered multi-valued protected setting: an empty value
resets every entry seen so far. That is the documented mechanism for revoking a
system-wide `safe.directory=*` and then naming only the repositories you
actually trust. The previous helper broke all three properties that make it work:

  * read `--global` before `--system` (git's precedence is system, then global)
  * dropped empty entries via `if entry:`, deleting the reset marker
  * de-duplicated values, though it is a sequence and not a set

Reproduced with real git (2.55.0 here, 2.47.3 by the reviewer). Given
system `safe.directory=*`, global `safe.directory=` then `/trusted/only`:

  real git, user's own policy                     -> rc=128 (dubious ownership)
  previous head's sequence `/trusted/only`, `*`   -> rc=0   (trust widened)
  this head's sequence `*`, ``, `/trusted/only`   -> rc=128 (matches real git)

So a user who had deliberately revoked a machine-wide wildcard silently got it
back inside Hermes's internal git calls. That contradicted the "read-only and
non-widening" claim the original change rested on.

Fix: read scopes lowest-precedence first (system, then global) and replay every
value verbatim -- no de-duplication, no dropping of empty reset markers. Read
with `git config -z` so a value containing whitespace or a newline stays the one
entry git reads it as, instead of being split into several bogus trust entries
by splitlines()/strip(). The trailing field of a -z stream is always empty and
is discarded; interior empty fields are real resets and survive.

Two invariant tests, both proven red on the previous implementation:

  * the injected sequence equals `git config -z --get-all safe.directory` for
    the same two config files, asserting git's documented reset shape
  * an E2E trust decision under GIT_TEST_ASSUME_DIFFERENT_OWNER=1: the named
    repo stays usable and an unrelated cross-owner repo is still refused,
    proving the reset still revokes the wildcard

13 passed in tests/hermes_cli/test_noninteractive_git.py, 69 in
tests/security/test_gitspawn_config_injection.py and
tests/tools/test_checkpoint_manager.py.
2026-09-11 19:01:47 -07:00
privacydied 01a3206e90 fix(git): carry user safe.directory past non-interactive config isolation
`noninteractive_git_env()` blanks GIT_CONFIG_GLOBAL/SYSTEM to /dev/null so a
user's config cannot hang Hermes's internal git plumbing with pagers, hooks or
credential prompts. Sound intent, but it also discards `safe.directory` — and
git honours that key ONLY from global/system config (it is rejected from
repo-level config by design, so a hostile repo cannot self-authorise).

Result: every internal git call fails on a repo whose st_uid != geteuid():

    $ hermes -w
    ✗ Failed to create worktree: fatal: detected dubious ownership in
      repository at '/mnt/nas/py/repo'

That hits NFS/CIFS mounts without idmapping, shared checkouts, and containers
with a remapped uid. The user's own `git config --global --add safe.directory`
is correctly set and their interactive git works — Hermes throws the setting
away before git reads it, so the error's own suggested remedy can never fix it.
There is no config or env escape hatch: the blanking is unconditional.

Reproducer (any repo where the checkout uid differs from the caller's):

    git rev-parse HEAD                                    # works
    GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null \
      GIT_CONFIG_NOSYSTEM=1 git rev-parse HEAD            # dubious ownership

Fix: read the user's real `safe.directory` values before the isolation is
applied, then re-inject them over the GIT_CONFIG_KEY_n channel, which survives
GIT_CONFIG_GLOBAL=/dev/null. Isolation is unchanged — global/system config stay
pointed at /dev/null and every hardening override still applies, since a later
key of the same name wins in git's config order.

Read-only and non-widening: only values already present in the user's own
config are carried, so this grants no trust they had not granted. Ambient
GIT_CONFIG_KEY_n injection is still stripped first, so a caller cannot launder
an attacker-controlled path in this way — covered by a regression test.

Tests: three cases in TestNoninteractiveGitEnv — entries carried past
isolation, no entries injected when the user configured none, and ambient
injection not trusted. All pin GIT_CONFIG_SYSTEM at an empty file, since a real
/etc/gitconfig on the test host can otherwise leak entries and mask the
assertions.

Verified on Arch Linux, git 2.x, repo on an NFSv4 mount (uid 1024 vs caller
1000): `git worktree add` under the patched env returns rc=0 where it
previously failed. 80 passed in tests/hermes_cli/test_noninteractive_git.py,
tests/security/test_gitspawn_config_injection.py and
tests/tools/test_checkpoint_manager.py, including the pre-existing assertions
that the isolation stays in force.
2026-09-11 19:01:47 -07:00
Teknium bdb5bf96f3 test(hooks): synthetic payload carries the new profile field like production 2026-09-11 15:44:00 -07:00
Teknium a6c5b7ada8 fix(serve): SSH-isolated idle-exit keeps the backend alive while a cron job runs
turn_in_flight read only the dashboard session table; an in-process cron run
never registers there, so the watchdog reported "no running turn" and exited
mid-job (tool calls then failed with "cannot schedule new futures after
interpreter shutdown", the execution was marked unknown, the slot lost). The
probe now also consults cron.scheduler.get_running_job_ids — the ledger the
gateway shutdown drain already uses.

Addresses #107485
2026-09-11 15:28:37 -07:00