Commit Graph

3989 Commits

Author SHA1 Message Date
teknium1 12fe7684e1 fix(plugins): async-await helper thread runs under the caller's ContextVars
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
2026-09-12 08:26:48 -07:00
teknium1 1664e12fc7 chore(tests): explicit utf-8 encoding in test_plugins (windows-footgun ratchet) 2026-09-12 08:26:48 -07:00
teknium1 71c72e078a fix(cli): one unverified-accept message for custom endpoints without /models; tests trimmed
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
2026-09-12 08:26:23 -07:00
Victor Nogueira a3671787b0 fix(cli): clarify unverified custom model warning 2026-09-12 08:26:23 -07:00
Victor Nogueira eda1e6cb15 fix(cli): accept unverified custom models
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
2026-09-12 08:26:23 -07:00
bixycler 70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
teknium1 535fd88c70 fix(backup): keep durable cache/ artifacts (images, citation ledger) in full backups
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
2026-09-12 08:25:49 -07:00
teknium1 b97ae5c9b8 fix(backup): make the socket test runner-safe and document the new exclusions
The picked test bound an AF_UNIX socket at pytest's tmp_path, which overflows
the ~108-byte sun_path limit under scripts/run_tests.sh's deep temp root
("AF_UNIX path too long"). Bind by a relative name from inside the temp
HERMES_HOME instead; the walker still sees the same absolute entry.

Also list cache/ + runtime roots and non-regular entries in the `hermes backup`
"What's excluded" docs so the user-visible behaviour change is documented.
2026-09-12 08:25:49 -07:00
mrwanstudio 947e027f61 fix(backup): skip non-regular filesystem entries 2026-09-12 08:25:49 -07:00
Brad Estes 3b3f354933 fix(backup): exclude profile caches from full backups 2026-09-12 08:25:49 -07:00
teknium1 7817af2a16 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:25:36 -07:00
teknium1 2b685f08fa fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:24:52 -07:00
teknium1 07ead2249b chore(tests): explicit utf-8 encoding in test_claw (windows-footgun ratchet) 2026-09-12 08:24:52 -07:00
Drexuxux 45dd97a6f7 fix(claw): cleanup no longer mistakes any "openclaw" in argv for a running daemon
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).

POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.

Fixes #12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.

Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
2026-09-12 08:24:52 -07:00
teknium1 d267bc7f78 fix(model-switch): provider:model resolves like provider/model off aggregators too
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.

Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.

Fixes #9748
2026-09-12 08:24:42 -07:00
zhao c5cdd92254 fix(copilot): validate supported token prefixes 2026-09-12 08:24:29 -07:00
teknium1 fd15cd003e chore(tests): encoding="utf-8" on read_text/write_text in test_personality_single_owner.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:26 -07:00
teknium1 44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
teknium1 5685b76fde fix(doctor): image_gen hint covers a selected provider with a missing key or SDK
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
2026-09-12 08:24:11 -07:00
teknium1 6ea3b00e7f chore(tests): encoding="utf-8" on read_text/write_text in test_doctor.py
Windows-footgun ratchet for the file touched by this fix (no behaviour change).
2026-09-12 08:24:11 -07:00
skyc1e 7d25ae58a7 fix(doctor): image_gen reports "no provider configured", not "system dependency not met"
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.

A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.

Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.

Fixes #9516
2026-09-12 08:24:11 -07:00
teknium1 81cec06c77 test(cli): banner hero line is not center-padded
Invariant test for #9879 (red on origin/main: Rich inserted centering spaces
before the braille-padded hero). Adapted from PR #9880's test to the current
banner internals.
2026-09-12 08:23:58 -07:00
Season 6d20c321d6 fix(gateway): enable linger on systemd user-service repair path (#12863)
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.

Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.

Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
2026-09-12 08:23:44 -07:00
ygd58 5be0c4921c fix(gateway): surface launchctl bootstrap failures in refresh_launchd_plist_if_needed
Ports #63762 forward onto current main per teknium1's review.

refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).

1. refresh_launchd_plist_if_needed(): return False after retry
   exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
   of the success message.

Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.

Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.

5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
2026-09-12 08:23:44 -07:00
teknium1 f34933605a test(identity): make the 0600 ledger test prove the mode under a permissive umask
Set umask 0o022 around the write so the assertion fails whenever the mode
comes from the environment instead of the writer, and gate it off Windows
like the sibling POSIX mode-bit tests (st_mode is synthesized there).
2026-09-12 08:02:16 -07:00
codeshipsingh 5a8e651183 fix(identity): enforce 0600 on spawn ledger writes 2026-09-12 08:02:16 -07:00
teddychenfeiyang-png a8ac7899b5 test(plugins): typo'd requires_hermes clause fails plugins validate admission
Companion to the import fix: a clause whose version segment does not parse
(`>=0.21.1,<0.x`) must fail admission rather than silently gate nothing.

Salvage note: the source hunk (same import fix) and the duplicate positive
test from PR #108842 were dropped in favour of the earlier #107553; only the
negative test is carried here.
2026-09-12 07:57:13 -07:00
jinli.yl c4c508579d fix(plugins): validate requires_hermes from manifest module 2026-09-12 07:57:13 -07:00
Teknium b0c383cdf7 fix(desktop): served profiles show running and route lifecycle to the multiplexer from a pooled local backend
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.

Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
2026-09-12 06:13:44 -07:00
Teknium be2f7e9c36 feat: curator prunes unused skills at 30 days (was 90), stale at 14
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
2026-09-12 05:56:15 -07:00
teknium1 e8016a18c1 fix(cli): report the auto-detected context length when a custom provider is saved without one
Leaving the context-length prompt blank in the custom-endpoint wizard said
"will auto-detect" and then went silent, so users could not tell whether their
endpoint runs on a detected window or the runtime's default fallback (which
shapes compression and prompt-cache behaviour). After the save prompt, run the
same resolver the runtime uses (with the endpoint's URL and key) and print
either "auto-detected N tokens" or "not detected — using the default N tokens".
Feedback only: the probe result is not persisted, and a failing probe never
blocks the save.

Fixes #2513. Approach from PR #2522 (@ygd58) and PR #85499 (@Luna161), both
written against the pre-decomposition wizard module.

Co-authored-by: Luna161 <268031236+Luna161@users.noreply.github.com>
2026-09-12 05:10:51 -07:00
teknium1 08bb58bd4a fix(skills): never call a rate-limited fetch a stale index entry; drop per-search caveat
A throttled GitHub fetch also yields index-metadata-without-bundle, so the new
stale-entry verdict would tell users a skill "no longer exists upstream" when
it does. Check the adapters' rate-limit flag first and keep the existing
rate-limit hint for that case (the keep_open review concern on #3261).

The per-search "results may be stale" note is dropped: it fires on every
skills.sh search whether or not anything is stale, and the install-time error
now names the condition precisely where it happens.
2026-09-12 05:10:31 -07:00
nikkoxgonzales 257a704d18 fix(skills): name stale index entries instead of generic fetch failure
_resolve_source_meta_and_bundle already distinguishes index-hit-without-
files from unknown identifiers, but do_install printed the same generic
'Could not fetch' for both, sending users off to re-check spellings for
what is actually a stale skills.sh entry. Split the message, and add a
staleness caveat to do_search results from skills.sh.

Fixes #3259. Supersedes #3261 (stale since July — re-applied onto the
current _print_fetch_failure helper).
2026-09-12 05:10:31 -07:00
nikkoxgonzales 75ade17617 fix(telegram): normalize unicode dashes in bot menu descriptions
BotFather rejects setMyCommands descriptions containing em/en dashes
(U+2012-U+2015, U+2212). Fold them to ASCII hyphen at the two Telegram
sinks (telegram_bot_commands, telegram_menu_commands) so core, plugin,
and skill entries are all covered.

Fixes #2925.
2026-09-12 05:10:11 -07:00
kshitijk4poor 966fb375fc fix(cli): make shallow-boundary repair survive the prune and the graph-safety review findings
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:

- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
  rev-list --all --reflog, so a boundary the repair just restored (one a
  reflog-only commit still needs) is never dropped again; previously the
  production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
  message body is prose; _batch_missing_parents stops at the blank line
  ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
  not --batch-all-objects: unrelated object loss (a deleted parent of a
  locally-created commit) is no longer re-labelled as shallow history;
  fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
  shallow.lock, so a depth-1 fetch between read and write fails fast
  instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
  rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
  (_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
  test file footguns fixed (encoding=, as_uri()) and the missing
  repair->prune end-to-end regression added, mutation-checked.
2026-09-12 15:00:37 +05:30
joaomarcos 2fb87f047f fix(cli): repair shallow boundaries already dropped by stale-graft prune
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR #108290, which owns the prevention half.

Refs #108286
2026-09-12 15:00:37 +05:30
Teknium d76856cc69 fix(migrate): shared-ingress adapters declare serves_profile_prefix so migration reports them as notices
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
2026-09-12 02:02:20 -07:00
Teknium 9ca7db8232 test(multiplex): shared-listener ingress invariants
- /p/<profile>/line/webhook is verified with the NAMED profile's channel secret under
  its runtime scope; another profile's secret is 401 at that URL; the bare path is
  untouched; unknown profile / profile without the adapter is 404.
- A shared-listener adapter binds no port and records its /p/<profile>/ ingress_url
  in runtime status; LINE media URLs use the shared prefix.
- Runner: a secondary's port-binders are constructed in shared-listener mode instead
  of refusing the whole profile; api_server/webhook are skipped as mirrors.
- Dashboard: only the mirrored pair is refused (409) on a secondary.
2026-09-12 01:53:15 -07:00
Teknium bcdb49ac7b feat(gateway): adapters declare serves_profile_prefix for /p/<profile>/ ingress
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
2026-09-12 01:49:28 -07:00
Teknium 5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
kshitijk4poor 443c2785fa docs(gateway): state the sudo mid-command rationale once, at the decision point
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
2026-09-12 12:15:07 +05:30
kshitijk4poor 1ecdfd18db fix(gateway): consult the installed unit only when root
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.

Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
2026-09-12 12:15:07 +05:30
kshitijk4poor b18c5752e3 test(gateway): keep one invariant test per bare-name owner
Trims the nine salvaged tests to four, one per contract: SUDO_USER's default
home keeps the bare name across the unit sync; a home the unit does not pin
keeps its suffix (the #105525 guard); a bare unit pinning a profiles/<name>
home beats the profile branch; the production sync itself preserves the name.
2026-09-12 12:15:07 +05:30
JoaoMarcos44 1ff56d9ed9 fix(gateway): keep the unit anchor Linux-only and pin the order it depends on
Review follow-ups to the unit-anchored service identity.

`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.

Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.

Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.

The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.

Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.

Refs #108674
2026-09-12 12:15:07 +05:30
JoaoMarcos44 a4a10eb520 fix(gateway): anchor system service identity on the installed unit
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).

The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.

Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.

The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.

Refs #108674
2026-09-12 12:15:07 +05:30
Hukla fdb9b50a15 fix(gateway): preserve sudo system service name 2026-09-12 12:15:07 +05:30
kshitijk4poor c71a411fe5 test(gateway): release-slot fakes accept the finalizer's run_generation kwarg
The turn finalizer now releases only its own generation
(`_release_running_agent_state(key, run_generation=...)`); two test doubles
were zero-kwarg lambdas and raised TypeError inside the finally.
2026-09-12 12:03:39 +05:30
kshitij 683afe45ba fix(auth): a global-root write-through invalidates the root store memo
_load_global_auth_store is memoised on (path, st_mtime_ns). A write-through
to the root (borrowed Codex cooldown clear, xAI/Anthropic root rotation)
followed by a fallback read in the same mtime tick — coarse-mtime
filesystems (exFAT, some network/overlay mounts) — kept serving the
pre-write store, so the resolve path could log "quota restored" and then
raise quota_exhausted from the stale memo. _save_auth_store(target_path=...)
now drops the memo.
2026-09-12 12:02:55 +05:30
kshitij ea42884e99 refactor(auth): codex cooldown clear reuses the pool ownership rule
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.

Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
2026-09-12 12:02:55 +05:30
kshitijk4poor 65c9d33b14 fix(auth): a borrowed Codex pool clears its quota cooldown in the root store
With reads now inheriting the global-root pool, a profile hitting a stale
root cooldown probes quota, sees it restored, and calls
clear_codex_pool_quota_cooldowns() — which only ever edited the (empty)
profile store, so the next resolve raised quota_exhausted again forever.
Pick the store the same way agent/credential_pool.py persists borrowed
rows (_profile_owns_pool_provider / _borrowed_single_use_pool_root) and
lock/save against that path. The fallback test now also binds
profile-wins precedence; one new test pins the root write.
2026-09-12 12:02:55 +05:30