Commit Graph

7379 Commits

Author SHA1 Message Date
teknium1 12fe7684e1 fix(plugins): async-await helper thread runs under the caller's ContextVars
Under a running loop `resolve_plugin_command_result` awaited the coroutine on
a raw thread, so an async hook saw the process-default HERMES_HOME and no
secret scope (get_secret -> UnscopedSecretError on a secondary profile).
Run the thread body through `contextvars.copy_context().run`, matching the
bounded hook worker. Also fixes async plugin slash commands the same way.
2026-09-12 08:26:48 -07:00
teknium1 24444e52eb fix(plugins): await async hook callbacks instead of collecting bare coroutines
Slash-command handlers gained loop-safe awaiting in ca9a61ae38, but
`PluginManager.invoke_hook` still called `async def` hook callbacks directly:
the coroutine object was appended to the results (so `pre_llm_call` context
injection silently did nothing) and Python warned "coroutine was never
awaited". `_invoke_hook_callback` now routes every return through
`resolve_plugin_command_result`, which covers both the direct and the
timeout-bounded paths and is safe under the gateway's running loop.

Fixes #12449 (remaining hook half). Salvage of #63240 by @Bartok9, applied
one layer down so the bounded-worker path is covered too.

Co-authored-by: Bartok9 <Bartok9@users.noreply.github.com>
2026-09-12 08:26:48 -07:00
teknium1 71c72e078a fix(cli): one unverified-accept message for custom endpoints without /models; tests trimmed
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
2026-09-12 08:26:23 -07:00
Victor Nogueira a3671787b0 fix(cli): clarify unverified custom model warning 2026-09-12 08:26:23 -07:00
Victor Nogueira eda1e6cb15 fix(cli): accept unverified custom models
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
2026-09-12 08:26:23 -07:00
bixycler 70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
teknium1 535fd88c70 fix(backup): keep durable cache/ artifacts (images, citation ledger) in full backups
Excluding cache/ wholesale at profile roots dropped media the gateway
delivered to or received from the user (cache/images, audio, videos,
documents, screenshots) and the grounded-citations evidence ledger
(cache/citations/ledger.json) — none of which can be regenerated.
Prune only the regenerable cache/<x> subtrees; keep those six.
2026-09-12 08:25:49 -07:00
mrwanstudio 947e027f61 fix(backup): skip non-regular filesystem entries 2026-09-12 08:25:49 -07:00
Brad Estes 3b3f354933 fix(backup): exclude profile caches from full backups 2026-09-12 08:25:49 -07:00
teknium1 7817af2a16 fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:25:36 -07:00
teknium1 2b685f08fa fix(claw): detect the gateway's process title openclaw-gateway
Upstream sets process.title="openclaw-gateway" in the gateway run loop, so
the real daemon has comm "openclaw-gatewa" (15-char truncation) and no
`node … openclaw` argv for the script probe to match. Add the exact comm
probe; substring matching stays out.
2026-09-12 08:24:52 -07:00
Drexuxux 45dd97a6f7 fix(claw): cleanup no longer mistakes any "openclaw" in argv for a running daemon
`_detect_openclaw_processes()` ran `pgrep -f openclaw`, which matches every
process whose command line contains the word: an editor open on
~/.openclaw/config.json, `tail -f openclaw.log`, even the checking shell.
`hermes claw cleanup` then warned "OpenClaw is still running" and aborted on
idle hosts (#12648).

POSIX detection now mirrors the Windows branch: exact binary names
(`pgrep -x openclaw`, `pgrep -x clawd`) plus node interpreters whose script
argv names openclaw/clawd (anchored ERE), deduplicated into one report.

Fixes #12648. Exact-name approach from #24121 by @Drexuxux, re-applied onto
the current `_posix_probe` helper.

Co-authored-by: Drexuxux <Drexuxux@users.noreply.github.com>
2026-09-12 08:24:52 -07:00
teknium1 d267bc7f78 fix(model-switch): provider:model resolves like provider/model off aggregators too
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.

Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.

Fixes #9748
2026-09-12 08:24:42 -07:00
zhao c5cdd92254 fix(copilot): validate supported token prefixes 2026-09-12 08:24:29 -07:00
teknium1 44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
teknium1 5685b76fde fix(doctor): image_gen hint covers a selected provider with a missing key or SDK
Review finding on #109136: "no provider configured" was wrong when a provider IS
selected but its SDK/key is absent. Word it as unavailable + where to look.
2026-09-12 08:24:11 -07:00
skyc1e 7d25ae58a7 fix(doctor): image_gen reports "no provider configured", not "system dependency not met"
image_gen has several setup paths (FAL_KEY, managed Nous image generation,
plugin providers) so it declares no single `requires_env`; doctor's generic
branch then labelled a missing credential a "system dependency not met" and
left it out of the "run hermes setup" summary.

A small per-toolset setup-hint table: image_gen gets an actionable line
pointing at `hermes tools`, counts toward the setup summary, and toolsets
with a genuine system dependency (homeassistant) keep the old wording.

Port of PR #9548 by @skyc1e onto `hermes_cli/doctor_tools.py`.

Fixes #9516
2026-09-12 08:24:11 -07:00
Guy Gascoigne-Piggford e38d63b7f0 fix(cli): left-align banner hero art 2026-09-12 08:23:58 -07:00
Season 6d20c321d6 fix(gateway): enable linger on systemd user-service repair path (#12863)
The repair branch in `systemd_install()` exits as soon as it rewrites an
outdated unit and re-runs `systemctl enable`, bypassing
`_ensure_linger_enabled()`. On headless Linux the command reports
success, but the repaired user service still stops at logout.

Call `_ensure_linger_enabled()` before the early return when the install
is user-scoped, mirroring what the fresh-install path already does.

Adds two regression tests in `tests/hermes_cli/test_gateway_linger.py`:
- repair path (user scope) calls the linger helper
- repair path (system scope) does not call it
2026-09-12 08:23:44 -07:00
ygd58 5be0c4921c fix(gateway): surface launchctl bootstrap failures in refresh_launchd_plist_if_needed
Ports #63762 forward onto current main per teknium1's review.

refresh_launchd_plist_if_needed() logged the retry failure but still
returned True and printed success. launchd_install() then
unconditionally printed '✓ Service definition updated' even when the
service was not registered with launchd (#12882).

1. refresh_launchd_plist_if_needed(): return False after retry
   exhaustion so callers can distinguish failure from success.
2. launchd_install(): check the bool; on False print a warning instead
   of the success message.

Per review: the warning now renders the reload-log location via
display_hermes_home() (the existing lazy-import convention used
elsewhere in this module for user-facing paths, e.g. the gateway.log
path prints a few lines away) instead of a hardcoded ~/.hermes path,
so named/custom Hermes home profiles show the correct location.

Existing _retry_launchctl_bootstrap_until_registered() retry/EIO/
timeout/verify logic unchanged.

5/5 tests pass (4 ported + 1 new for the display_hermes_home fix).
2026-09-12 08:23:44 -07:00
codeshipsingh 5a8e651183 fix(identity): enforce 0600 on spawn ledger writes 2026-09-12 08:02:16 -07:00
jinli.yl c4c508579d fix(plugins): validate requires_hermes from manifest module 2026-09-12 07:57:13 -07:00
Teknium b0c383cdf7 fix(desktop): served profiles show running and route lifecycle to the multiplexer from a pooled local backend
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.

Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
2026-09-12 06:13:44 -07:00
Teknium be2f7e9c36 feat: curator prunes unused skills at 30 days (was 90), stale at 14
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
2026-09-12 05:56:15 -07:00
Teknium 876e444e4e fix(cli): open the session store through the state.db registry, not a bare SessionDB()
The classic CLI froze for ~0.7-2s between the banner and the first prompt. py-spy +
strace on real PTY startups showed the main/REPL threads inside
refuse_deleted_wal_generation -> _iter_proc_fd_targets: a second full SessionDB open.

_init_session_store built a bare SessionDB(); a moment later the goal/loop/heartbeat
managers acquired the same state.db through hermes_state_registry from the REPL thread,
which is a different handle, so the whole open ran again — including the /proc-wide
deleted-WAL sidecar scan (~4.4k readlinks). Each readlink drops and re-takes the GIL
while the startup threads (plugin discovery, MCP, skill sync, banner git) are busy, so an
11ms scan stretched to 1.3s per pass, and the second pass landed exactly where the
prompt should have appeared.

Route the CLI's handle (init + the two re-open sites) through the registry so every
in-process consumer shares one writer. One scan per startup; live A/B on the same box,
interleaved x6: banner->prompt gap 0.37s mean -> 0.15s mean (plain), 1.77s -> 0.39s
under strace. The registry release path replaces close(), so /quit, /snapshot restore
and /handoff keep their semantics.
2026-09-12 05:13:50 -07:00
teknium1 e8016a18c1 fix(cli): report the auto-detected context length when a custom provider is saved without one
Leaving the context-length prompt blank in the custom-endpoint wizard said
"will auto-detect" and then went silent, so users could not tell whether their
endpoint runs on a detected window or the runtime's default fallback (which
shapes compression and prompt-cache behaviour). After the save prompt, run the
same resolver the runtime uses (with the endpoint's URL and key) and print
either "auto-detected N tokens" or "not detected — using the default N tokens".
Feedback only: the probe result is not persisted, and a failing probe never
blocks the save.

Fixes #2513. Approach from PR #2522 (@ygd58) and PR #85499 (@Luna161), both
written against the pre-decomposition wizard module.

Co-authored-by: Luna161 <268031236+Luna161@users.noreply.github.com>
2026-09-12 05:10:51 -07:00
teknium1 08bb58bd4a fix(skills): never call a rate-limited fetch a stale index entry; drop per-search caveat
A throttled GitHub fetch also yields index-metadata-without-bundle, so the new
stale-entry verdict would tell users a skill "no longer exists upstream" when
it does. Check the adapters' rate-limit flag first and keep the existing
rate-limit hint for that case (the keep_open review concern on #3261).

The per-search "results may be stale" note is dropped: it fires on every
skills.sh search whether or not anything is stale, and the install-time error
now names the condition precisely where it happens.
2026-09-12 05:10:31 -07:00
nikkoxgonzales 257a704d18 fix(skills): name stale index entries instead of generic fetch failure
_resolve_source_meta_and_bundle already distinguishes index-hit-without-
files from unknown identifiers, but do_install printed the same generic
'Could not fetch' for both, sending users off to re-check spellings for
what is actually a stale skills.sh entry. Split the message, and add a
staleness caveat to do_search results from skills.sh.

Fixes #3259. Supersedes #3261 (stale since July — re-applied onto the
current _print_fetch_failure helper).
2026-09-12 05:10:31 -07:00
nikkoxgonzales 75ade17617 fix(telegram): normalize unicode dashes in bot menu descriptions
BotFather rejects setMyCommands descriptions containing em/en dashes
(U+2012-U+2015, U+2212). Fold them to ASCII hyphen at the two Telegram
sinks (telegram_bot_commands, telegram_menu_commands) so core, plugin,
and skill entries are all covered.

Fixes #2925.
2026-09-12 05:10:11 -07:00
kshitijk4poor eec131b716 refactor(cli): hold shallow.lock across write-verify-restore; share the atomic write
Simplify-pass follow-up on the repair rework: the self-check rev-list walks
and the rollback restore now run inside the same _ShallowLock hold (rev-list
never takes shallow.lock), so lock contention can no longer defeat a failed
self-check's rollback and leave a broken .git/shallow in place. The tmp-write
+ os.replace sequence shared by repair and prune moves into _write_shallow.
Scope note added: fetch-by-SHA install tips (HEAD-reflog-only) are not repair
candidates; corruption of that shape is prevented by the prune's reflog
fail-safe. 68 focused tests green; two-cycle repair->prune E2E re-verified.
2026-09-12 15:00:37 +05:30
kshitijk4poor 966fb375fc fix(cli): make shallow-boundary repair survive the prune and the graph-safety review findings
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:

- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
  rev-list --all --reflog, so a boundary the repair just restored (one a
  reflog-only commit still needs) is never dropped again; previously the
  production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
  message body is prose; _batch_missing_parents stops at the blank line
  ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
  not --batch-all-objects: unrelated object loss (a deleted parent of a
  locally-created commit) is no longer re-labelled as shallow history;
  fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
  shallow.lock, so a depth-1 fetch between read and write fails fast
  instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
  rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
  (_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
  test file footguns fixed (encoding=, as_uri()) and the missing
  repair->prune end-to-end regression added, mutation-checked.
2026-09-12 15:00:37 +05:30
joaomarcos 2fb87f047f fix(cli): repair shallow boundaries already dropped by stale-graft prune
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR #108290, which owns the prevention half.

Refs #108286
2026-09-12 15:00:37 +05:30
Teknium d76856cc69 fix(migrate): shared-ingress adapters declare serves_profile_prefix so migration reports them as notices
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
2026-09-12 02:02:20 -07:00
Teknium 65fd3a2b9c feat(status): show each served profile's shared-listener callback URLs
hermes -p <name> gateway status, hermes gateway status and hermes status (under
Serves:) list the /p/<profile>/<path> URL per inbound-port platform the live
multiplexer serves, read from the <profile>:<platform> ingress_url in the default
home's gateway_state.json (hermes_cli/gateway_multiplex_served.py). The dashboard's
messaging payload carries the same ingress_url and the Channels page renders it.
The dashboard's 409 guard now covers only api_server/webhook (the mirrored pair):
enabling Twilio/LINE/Teams/... on a secondary is allowed because the gateway serves
it.
2026-09-12 01:53:15 -07:00
Teknium df51797e2e feat(cron): let planned downtime skip missed recurring runs 2026-09-12 01:51:59 -07:00
Teknium be82d52cb7 docs: migrating from per-profile gateways to the multiplexer
What `hermes update` does, blockers and fixes, URL change for
inbound-port profiles, the post-create restart reminder, rollback, and
the `gateway migrate` reference row.
2026-09-12 01:49:28 -07:00
Teknium df95f378a6 feat(dashboard): "Migrate to a single multiplexed gateway" on the System page
GET /api/gateway/migrate/plan returns the CLI plan JSON; POST
/api/gateway/migrate spawns `hermes gateway migrate --multiplex --yes`
detached (action log gateway-migrate.log). The Gateway card shows the
button only for a multi-profile install that is not yet multiplexed, and
disables it while listing the blockers.
2026-09-12 01:49:28 -07:00
Teknium 07a4ae016a feat(update): auto-migrate to one multiplexed gateway when unblocked
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
2026-09-12 01:49:28 -07:00
Teknium e2fc493427 feat(gateway): hermes gateway migrate --multiplex / --standalone
Moves a per-profile-gateway install onto one multiplexed default gateway:
table-driven preflight (duplicate credential via the gateway's own
fingerprint; secondary port-binders without a /p/<profile>/ ingress),
--dry-run, apply (stop + uninstall each secondary's service, record it in
<default>/gateway_migration.json, flip gateway.multiplex_profiles through
the config API, restart/install the default on the same service manager,
verify served_profiles), and --standalone rollback from the manifest.
Idempotent; refuses cleanly when already multiplexed or blocked.

`hermes profile create` points at `hermes gateway restart` when a live
multiplexer is detected (the served set is snapshotted at startup).
2026-09-12 01:49:28 -07:00
Teknium bcdb49ac7b feat(gateway): adapters declare serves_profile_prefix for /p/<profile>/ ingress
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
2026-09-12 01:49:28 -07:00
Teknium 5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
kshitijk4poor 443c2785fa docs(gateway): state the sudo mid-command rationale once, at the decision point
The three helpers each restated why sudo moves the naming basis; keep it in _profile_suffix and leave the helpers their unique reasons. Drops a footgun marker the scanner has no pattern for and a raising=False on an attribute that exists.
2026-09-12 12:15:07 +05:30
kshitijk4poor 1ecdfd18db fix(gateway): consult the installed unit only when root
`_bare_unit_pinned_home()` read the system unit for every caller, so an
unprivileged `hermes -p kimi gateway status` (user scope) resolved
`hermes-gateway` instead of `hermes-gateway-kimi` whenever the bare system
unit pinned that profile home — aliasing the profile onto the user's default
unit. Only an elevated process operates the system unit, so gate on root.

Also drops the unreachable `OSError` arm (non-strict resolve swallows it) and
routes the legacy-unit search through `_SYSTEM_UNIT_DIR`.
2026-09-12 12:15:07 +05:30
kshitijk4poor a32b8cd4a3 refactor(gateway): share the sudo-invoker home lookup, drop unreachable guards
`_native_service_homes()` re-implemented the root+SUDO_USER -> `pw_dir/.hermes`
resolution that `hermes_cli.main._resolve_sudo_user_profile_env` already did.
Both now call `hermes_constants.sudo_invoker_default_home()`; main.py appends
`profiles/<name>` to it.

Also removes the local `from pathlib import Path as _Path` (Path is a module
import) and narrows the except to `KeyError`: `import pwd` cannot fail once
`os.geteuid` exists, and `getpwnam(str)` raises nothing else.
2026-09-12 12:15:07 +05:30
JoaoMarcos44 1ff56d9ed9 fix(gateway): keep the unit anchor Linux-only and pin the order it depends on
Review follow-ups to the unit-anchored service identity.

`_bare_unit_pinned_home()` now returns early off Linux. `_profile_suffix()` is
shared by the launchd label/plist helpers, the Windows scheduled-task name and
the s6/multiplex `_current_profile_name()` fallback, and a systemd unit is not
an identity authority for any of them. The gate is `is_linux()` (a plain
`sys.platform` test) rather than `supports_systemd_services()`, which can shell
out to `systemctl is-system-running` on WSL and containers -- unacceptable in a
helper that runs on every name resolution.

Document why the unit-pinned check must precede the profile branch, and pin it
with a test: `sudo hermes gateway install --system` resolves the BARE name from
root's default home, then writes the invoking user's remapped home into the
unit, so the bare unit legitimately carries a `<root>/profiles/<name>` home. If
the profile branch ran first it would answer `hermes-gateway-kimi` for a unit
installed as `hermes-gateway`, which is the original bug class.

Three more regressions: the named-profile-pinned bare unit above; a run that
drives the real `_sync_hermes_home_from_systemd_unit()` instead of simulating
the adoption with `setenv`; and an unreadable unit, which must fall through to
the suffix branches rather than hand its bare name to an unrelated home.

The class is now `linux_only`, because the gate makes the behaviour genuinely
host-dependent -- so the tests belong on the host that has it, not behind a
faked platform.

Verified on a real Linux kernel (WSL2, Python 3.12.13), not by simulation:
7/7 pass on this branch; with `hermes_cli/gateway.py` restored from origin/main
and the tests kept, 4 fail with the reported symptom
(`'hermes-gateway-kimi' == 'hermes-gateway'`, `'hermes-gateway-54de6eee' ==
'hermes-gateway'`) and the 3 guard tests still pass. Whole file on Linux:
origin/main 4 failed/106 passed/1 skipped, this branch 4 failed/113 passed/1
skipped -- same four pre-existing failures, exactly seven new passes.

Refs #108674
2026-09-12 12:15:07 +05:30
JoaoMarcos44 a4a10eb520 fix(gateway): anchor system service identity on the installed unit
`sudo hermes gateway start|stop|restart|status|uninstall|install --system`
resolved the systemd unit as `hermes-gateway-<sha256[:8]>` while the installed
unit is `hermes-gateway.service`, failing with `Unit ... not found` (exit 5).

The service name was derived from the CURRENT PROCESS's HERMES_HOME, and under
sudo that value changes MID-COMMAND: sudo strips HERMES_HOME and sets
HOME=/root, so the `_require_service_installed()` pre-flight resolved the bare
name and passed; `_sync_hermes_home_from_systemd_unit()` then adopted the
unit's pinned `HERMES_HOME=/home/<user>/.hermes` into `os.environ` (deliberate,
for runtime-status/PID reads), and every later `get_service_name()` took the
hash branch. Regression from the #105525 fix, which correctly moved the
comparison basis to `_get_platform_default_hermes_home()` -- right for a
temp-dir/Docker home, but wrong for an elevated process whose `~/.hermes` is
not the home that owns the unit.

Read the naming basis from the unit instead of the process: the installed
`hermes-gateway.service` is the authority on which home owns the bare name.
`_bare_unit_pinned_home()` parses that unit's pinned HERMES_HOME, and
`_profile_suffix()` accepts it alongside the platform-native default. This is
stable for every elevated identity, including `sudo -i` and cron where
SUDO_USER is absent, and for a custom HERMES_HOME pinned in the unit.

The #105525 guard is untouched: with no installed bare unit, a temp-dir/Docker
/custom home still keeps its own hashed suffix and can never resolve to -- or
uninstall -- the operator's `hermes-gateway.service`. Only the single home that
unit pins is recognised; an unrelated home stays suffixed. Nothing is memoized,
because `hermes_cli/profiles.py::_cleanup_gateway_service` swaps HERMES_HOME
mid-process and depends on re-derivation. The native-default check stays first
so the common path short-circuits before any file I/O.

Refs #108674
2026-09-12 12:15:07 +05:30
Hukla fdb9b50a15 fix(gateway): preserve sudo system service name 2026-09-12 12:15:07 +05:30
kshitij 683afe45ba fix(auth): a global-root write-through invalidates the root store memo
_load_global_auth_store is memoised on (path, st_mtime_ns). A write-through
to the root (borrowed Codex cooldown clear, xAI/Anthropic root rotation)
followed by a fallback read in the same mtime tick — coarse-mtime
filesystems (exFAT, some network/overlay mounts) — kept serving the
pre-write store, so the resolve path could log "quota restored" and then
raise quota_exhausted from the stale memo. _save_auth_store(target_path=...)
now drops the memo.
2026-09-12 12:02:55 +05:30
kshitij ea42884e99 refactor(auth): codex cooldown clear reuses the pool ownership rule
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.

Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
2026-09-12 12:02:55 +05:30
kshitijk4poor fe6351000f refactor(auth): codex pool reads call read_credential_pool directly; cooldown clear decides ownership under the lock
`_read_codex_pool_entries` had become a pass-through whose only residual
behaviour was taking the active-store lock around a read every other
`read_credential_pool` caller does unlocked (and which never covered the
root file it fell back to). Both consumers now call the helper directly.

`clear_codex_pool_quota_cooldowns` decided "profile owns rows?" with an
unlocked pre-read, then re-read the same file under the lock — a TOCTOU
against a concurrent `hermes auth add` and a wasted parse. It now tries the
active store under its lock and falls back to the borrowed root only when
that store has no codex rows.
2026-09-12 12:02:55 +05:30