Commit Graph

3454 Commits

Author SHA1 Message Date
salch-cred 49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Teknium 71c4bcf7af fix(cron): surface missed-fire catch-up lateness in hermes cron list / status
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).

- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
  scheduled_at, dispatched_at, lateness_seconds, and kind
  (on_time / late / catch_up, classified against the ticker tolerance and
  the schedule's catch-up grace window). Manual triggers and one-shots are
  not stamped (no scheduled instant to be late against / retired beyond
  grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
  show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
  catch-up, in both the built-in ticker and external provider paths.

CLI surface only — no new tools, no policy engine.

Addresses the visibility half of #99879.
2026-09-01 08:31:37 -07:00
Justin Wilson fb18fedf29 fix(updater): rebuild desktop on Windows hand-off repair path
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.

Fixes #97343
2026-09-01 08:27:58 -07:00
JoaoMarcos44 86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
Teknium 35f4ababe1 test: the managed-dashboard restart now continues the serve scan
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
2026-09-01 07:00:54 -07:00
Teknium 5a677479e9 test: pin the no-authority contract with the serve probe disabled
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.
2026-09-01 07:00:54 -07:00
joaomarcos 27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
joaomarcos 2d70d67b15 fix(update): stop the managed dashboard restart from skipping serve backends
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.

The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
2026-09-01 07:00:54 -07:00
joaomarcos f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00
Teknium be597fc730 fix: extend corrupt-config fail-closed guard to gateway, serve, and cron surfaces
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
  construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
  on every surface
2026-09-01 07:00:22 -07:00
embwl0x 6f85df97fd fix(cli): keep quiet prompts interactive 2026-09-01 07:00:22 -07:00
embwl0x 55e7ecd260 test(cli): cover config guard edge cases 2026-09-01 07:00:22 -07:00
embwl0x c335dc734a fix(cli): reject corrupt config in noninteractive runs 2026-09-01 07:00:22 -07:00
lesseradmin 779482598f fix(desktop): pass --disable-setuid-sandbox on the userns launch path
When chrome-sandbox is present but not root-owned 4755, Chromium can still
abort via setuid_sandbox_host even though the namespace sandbox works.
After the userns probe skips sudo, append --disable-setuid-sandbox so
.desktop/no-TTY launches keep the namespace sandbox without a privilege
prompt. Does not add --no-sandbox.

Fixes #51327
2026-09-01 02:32:39 -07:00
4dlt 3a7f2234a6 fix(cli): use Chromium's namespace sandbox when userns is available on Linux
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).

On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.

Fixes #88032
Fixes #51327
Fixes #58593

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 02:32:39 -07:00
Teknium a1c25d393a feat(desktop): built-in optional-skills catalog in Capabilities → Skills with one-click install
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.

- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
  scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
  install-state spinners off $hubActions, and an OfficialSkillDetail pane
  (hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
2026-09-01 01:39:59 -07:00
liuhao1024 46ac31e84c fix(desktop): sanitized deferral evidence for ledgered manual serve blockers
The Desktop venv-blocker scan (since #99724) defers ledger-verified
serve/dashboard holders to the CLI updater's stop+relaunch rungs, but the
scan output only carried an opaque deferred_backends count — nothing
explained WHICH holders the deferral consumed or why they vanished from
processes.

Add sanitized decision evidence (#98350): deferred_backend_evidence lists
structured ledger identity only (pid, purpose, recorded port) — never the
command line, which can carry tokens or private endpoints. Adds a desktop
parser contract fixture proving the consumer tolerates the diagnostics
while keeping blocked/processes authoritative.

Salvaged from PR #98350; the exemption half of that PR was independently
consolidated on main via #99724 (_is_updater_owned_backend).
2026-09-01 01:30:09 -07:00
kshitijk4poor 801466765b fix: harden temp fallback against predictable-path attacks; neutral error text; docs
Final-diff review findings (/simplify-code on the full 3-commit stack):

- Temp-dir fallback uses a per-uid name (hermes-profile-exports-<uid>) and
  get_profile_export_path refuses a pre-existing symlink or a directory
  owned by another user — a fixed /tmp/hermes-profile-exports is a
  predictable shared path a local attacker could pre-create to receive the
  secret-bearing archive. Regression test mutation-checked.
- Fail-closed message reworded interface-neutrally (the web API surfaces it
  as HTTP 400 detail where '-o' alone made no sense).
- Docs now cover the temp fallback and the fail-closed refusal.
2026-09-01 01:00:23 -07:00
kshitijk4poor 525dd12da7 fix: fail closed when no safe export destination exists; polish salvage edges
- _profile_export_directory(): when the managed store, the home-sibling
  store, AND the temp dir all resolve inside Git checkouts, raise a clear
  ValueError instead of warning and proceeding — a stderr warning would not
  stop a scripted export from staging a secret-bearing archive in a source
  tree, which is the exact #92457 incident class. All three callers already
  surface ValueError cleanly (CLI/TUI print Error: + exit, API returns 400).
- .dockerignore: drop the /default.tar.gz line made redundant by the global
  *.tar.gz pattern this PR adds.
- hermes profile export -o help text: stop advertising the old
  <name>.tar.gz cwd default.
- Tests: cwd-in-unrelated-checkout topology (the second production shape
  from the blocking review) and the fail-closed path. Mutation-checked:
  both fail on the pre-fix helper.
2026-09-01 01:00:23 -07:00
kshitijk4poor 26fb8f60e6 fix: anchor checkout detection to the export path, not cwd
Follow-up to the salvaged #92689:

- _profile_export_directory() now proves safety on the export dir's OWN
  ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
  heuristic missed the checkout whenever HERMES_HOME sat inside one but
  the process ran from elsewhere (cron, service manager) — the export
  landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
  violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
  and catch OSError too — a bad profile name or read-only home printed a
  raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
  pollution in the tests/hermes_cli sweep can't divorce monkeypatches
  from the code under test; add regression tests for the cwd-independent
  topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
2026-09-01 01:00:23 -07:00
joaomarcos c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
Ben Barclay 56916841b5 refactor(dashboard-auth): replace PKCE cookie payload with base64url(JSON) codec (#99210)
The PKCE cookie's payload has now needed three serialization fixes at
the same spot: the original flat 'k=v;k=v' string tripped http.cookies'
\073 quoted form (dropped whole by strict cookie parsers like Go's
net/http — #83832 field case), and #99176 URL-encoded the whole flat
payload to stay inside the RFC 6265 cookie-octet set. The stacked
layers (single-encoded next=, ';' joins, whole-payload encoding, legacy
discriminator) were the recurring defect source.

Kill the bug class instead of patching it again: the payload is a dict
end-to-end and goes on the wire as base64url(JSON) — the urlsafe
alphabet is a strict subset of cookie-octets, and JSON framing means no
segment value can ever collide with a delimiter. parse_pkce_payload
keeps a three-rung compatibility ladder (base64url(JSON) -> oldest flat
form split-as-is -> #99176 unquote-then-split) for in-flight cookies
during a rolling upgrade (10-minute TTL); a new cookie hitting an old
server fails the OAuth state check and the user just retries.

The 'next' segment is stored as its plain validated path — no extra
encoding layer, so the post-login redirect Location is byte-for-byte
the original target.

Refs #99176, #84065.
2026-09-01 11:47:38 +10:00
Teknium d95682cd04 fix(gateway): announce declined 'all' scope in /sessions and /resume listings
A non-admin '/sessions all' or '/resume --all' silently downgraded to
chat-scoped listing with zero feedback, which reads as 'my session
vanished' (community Telegram report). Both surfaces now append a notice
that cross-chat listing requires a configured admin. Follows up the
salvaged current-session '(current)' marker (PR #68556, fixes #68547):
sibling tests updated to pin the new contract, new i18n key
gateway.resume.all_requires_admin added to all 17 locales, docs updated.
2026-08-31 14:40:09 -07:00
Xuezhao Lan dc65b75d73 fix(gateway): mark current session in listings 2026-08-31 14:40:09 -07:00
Teknium 3783fd9ffe fix(update): defer ledger-verified serve/dashboard holders to the updater instead of dead-ending the Desktop hand-off
On Windows, the Desktop update preflight scan (hermes_cli._scan_venv_blockers)
classified only python -m http.server as a stoppable blocker. A hermes serve
or dashboard backend that survived the Desktop teardown (or was launched
manually) either dead-ended the hand-off with venv-blocked, or slipped past
it and kept venv\Scripts\hermes.exe mapped, so the updater quarantine
failed with os error 32 (#98336).

The CLI updater downstream already owns exactly this holder class with
positive-identity rungs: _ledger_reapable_backend_pids reaps dead-spawner
orphans and _ledger_manual_serve_holders stops manual serves and relaunches
them on their recorded host/port. Mirror the existing pausable-gateway
exemption: defer ledger-verified serve/dashboard holders to those rungs
instead of reporting them as blockers.

Identity is positive-only, per the #99558 guard contract: token-parsed
subcommand (never substring, #90778) + live-verified (pid, create_time)
ledger entry with matching purpose + provable ownership (spawner dead,
unrecorded, or the hand-off Desktop itself — an ancestor of the scan,
verified by pid+create_time). A backend supervised by any other live
process keeps blocking; ledger unreadable fails closed.

Fixes #98336
2026-08-31 13:11:54 -07:00
fangliquanflq 9caf31394f fix(update): verify repaired current-checkout runtime 2026-08-31 12:58:14 -07:00
fangliquanflq 4f19808915 fix(update): fail current-checkout runtime verification 2026-08-31 12:58:14 -07:00
fangliquanflq 7e22725bb4 fix(update): verify SQLite runtime remediation 2026-08-31 12:58:14 -07:00
Finn763 57f653db90 test(managed_uv): make self-lock regression assert carry its message
The previous form was `mock_install.assert_not_called(), ("msg")` — a bare
tuple expression whose parenthetical never surfaces as a failure message.
Switch to `assert mock_install.call_count == 0, "msg"` so the diagnostic
actually appears when the guard regresses (review feedback on #93163).
2026-08-31 12:50:29 -07:00
Finn763 d0503b500f fix(update): defer Windows runtime repair when the updater holds the venv
On Windows, hermes update launched from the install's own venv can never
complete the managed-runtime repair: Windows keeps the image of the
updater's venv\Scripts\python.exe (and the waiting hermes.exe launcher
ancestor) mapped until exit, so the park rename in _cut_over_candidate
always fails with ERROR_ACCESS_DENIED. The pre-flight holder scan
deliberately excludes the calling process and its ancestors, so the guard
passes and the repair burns its retries on a structurally unwinnable
rename - forever, since the failure is non-fatal and the 'next update
will retry' message is misleading for this case.

Detect the self-lock (sys.executable or a launcher ancestor inside the
live venv) before provisioning and defer with actionable guidance instead
of walking into the doomed cutover. The deferral happens pre-provisioning,
so the incomplete generation-* leftovers the reporter observed are no
longer produced. No-op off Windows: POSIX renames work while the updater
maps the tree. Mirrors the existing _defer_update_for_self_lock pattern.

Regression tests prove the fix bites: neutralized guard -> repair proceeds
to provisioning (red); restored guard -> deferred before provisioning
(green). Verified on Windows 11 against a real venv.

Closes #93032
2026-08-31 12:50:29 -07:00
Teknium 5505042f40 fix(sessions): fail closed on ownership uncertainty and fence every turn source
Follow-ups on top of the #94595 cherry-pick, implementing the maintainer
review's two blockers:

Blocker 1 (turn-admission chokepoint): _run_prompt_submit itself now runs
the ownership admission, so synthesized turns that never pass through the
prompt.submit RPC handler (crash auto-continue from cold session.resume,
wake-ups) are fenced too. Auto-continue additionally checks ownership
BEFORE emitting message.start and leaves the marker in place, closing the
#94778 shape where backend B resumed a session backend A was actively
running and auto-continued A's fresh interrupted-turn marker into a
duplicate concurrent turn.

Blocker 2 (fail-closed registry semantics): try_acquire_active_session no
longer converts an unreadable/corrupt registry into an untracked go-ahead.
Ownership uncertainty is a distinct typed refusal —
SESSION_COORDINATION_UNAVAILABLE — because "could not prove ownership" must
never be collapsed into "no owner exists". The TUI gateway claim helper
fails closed on claim exceptions for every surface, not just desktop.

Also: empty session ids short-circuit to a no-op lease (nothing to fence,
and the strict registry schema rejects empty ids), and the existing
fail-open tests were updated to assert the new fail-closed contract.
2026-08-31 12:36:33 -07:00
Teknium 7acc65a399 test(windows): live E2E for the git trampoline self-heal on the wine2e lane
Real windows-latest coverage for the #88136 salvage: probes drive the
actual _git_is_trampoline/_locate_real_git/_ensure_non_trampoline_git
helpers against the runner's genuine Git-for-Windows install plus a real
fork-bomb-guard trampoline stand-in. Wired into the on-demand
windows-venv-e2e lane (wine2e/** pushes only).
2026-08-31 12:21:46 -07:00
SayHell0W0rld 6cccc2ef7e fix(update): locate PortableGit under the shared root, not profile home
Review feedback on #88136 (monerostar): a profile-scoped `hermes update`
sets HERMES_HOME to <root>/profiles/<name>, but the Hermes-managed
PortableGit tree lives under the SHARED root (<root>/git/...). The locator
checked get_hermes_home() only, so a broken trampoline during a
profile-scoped update was not swapped and fell through to ZIP.

Extract _portable_git_candidates() (shared root first, profile home as
fallback) and add a regression test for the profile layout.
2026-08-31 12:21:46 -07:00
SayHell0W0rld 0bab9ff9e7 fix(update): self-heal broken Git-for-Windows trampoline on Windows
A Git-for-Windows trampoline launcher (bin\git.exe / cmd\git.exe shim,
~46KB) that fails to re-exec the real git-core binary refuses every git
call with a "BUG (fork bomb)" guard instead of running it (#87876).

Detect the trampoline up front via `git --version`, locate a real git
binary (Git for Windows or Hermes-managed PortableGit locations), and
rebuild the git command with it so fetch/pull/checkout keep working with
a real git instead of degrading to the ZIP fallback. When no real binary
can be found, leave the command untouched so the existing fetch-failure
handler still falls back to the ZIP path on Windows (#88046).
2026-08-31 12:21:46 -07:00
Teknium a42aee9585 fix(config): greedy literal-key matching + loud phantom-sibling refusal for dotted key names
Builds on webtecnica's escape-aware _split_key_path (#84152, cherry-picked
with authorship preserved; earliest fix in the family was RelaxJonh's #80253
greedy-match approach — both behaviors now ship together):

- _greedy_literal_match: when navigating an EXISTING mapping, prefer an
  existing literal key equal to the dot-join of the next N path segments
  (longest match wins). Dotted model IDs are the norm, so the common
  unescaped command (config set providers.p.models.grok-4.6.supports_vision
  true) now hits the real key across set/get/unset instead of creating a
  phantom sibling. Plain dotted paths with no dotted-key collision split
  exactly as before.
- _phantom_sibling + ValueError in _set_nested: refuse to CREATE a new
  intermediate mapping that would shadow an existing dotted literal sibling
  (Soju06's fail-loudly suggestion on #84064); set_config_value surfaces it
  as a clean CLI error with the escaped spelling to use.
- utils.py::atomic_roundtrip_yaml_update (the second split site, #91607 —
  /model + TUI persistence) now uses the same escape-aware split + greedy
  literal matching.
- CFG-04 empty-segment guard now splits escape-aware so escaped keys are
  not misclassified.
- Tests for every repro shape in the family: #84064 provider model keys,
  #80006 Matrix room IDs, #91095 dotted models under custom_providers list
  index (incl. escaped creation-when-absent), #91607 model_overrides via
  atomic_roundtrip_yaml_update, #99124 dotted leaf keys; plus
  backward-compat coverage. Also fixed the carrier's one stale assertion
  (structured-value coercion landed on main after #84152 branched) and
  removed its dead _MCP_SECRETS_CONFIG fixture flagged in review.
- Docs: 'Dots inside key names' section in website/docs/reference/cli-commands.md.

Fixes #84064, fixes #80006, fixes #91095, fixes #91607, fixes #99124
2026-08-31 12:18:47 -07:00
webtecnica 8951947531 fix(cli): support literal dots in config set/unset key paths (#84064) 2026-08-31 12:18:47 -07:00
the3asic 18ac3c4fb6 fix(state): defer corrupt FTS rebuilds past live operations 2026-08-31 12:08:30 -07:00
Teknium b0acc558b1 fix(cli): supervised gateway launches skip the sticky active_profile redirect
Generalize the HERMES_S6_SUPERVISED_CHILD supervisor-marker mechanism so
ANY supervised gateway launch (systemd, launchd, Windows Scheduled Task,
external supervisor) skips the active_profile redirect in
_apply_profile_override(). Previously only the s6 container marker was
honored, so a systemd-launched default-profile gateway with
HERMES_HOME=<root> followed the sticky active_profile file and silently
assumed another profile's identity — logging under that profile's tree
and connecting with its Telegram bot token (double-polling a token owned
by that profile's own live gateway).

- hermes_cli/main.py: honor HERMES_SUPERVISED_CHILD (new generalized
  marker), HERMES_S6_SUPERVISED_CHILD (back-compat), INVOCATION_ID
  (systemd; gateway commands only, since it leaks into every descendant
  of systemd-launched processes), and HERMES_GATEWAY_EXTERNAL_SUPERVISOR.
- hermes_cli/gateway.py: export HERMES_SUPERVISED_CHILD=1 in generated
  systemd units (user + system) and the launchd plist.
- hermes_cli/gateway_windows.py: export it from the Scheduled-Task cmd/vbs
  launchers and the windowless respawn env overlay.
- hermes_cli/service_manager.py: export it alongside the s6 sentinel.
- tests: regression coverage for all markers + non-gateway INVOCATION_ID
  neutrality + generated-unit marker presence.

Fixes #74872
2026-08-31 12:05:07 -07:00
andyst-dev 4d3e1e4d13 fix(desktop): prevent venv scan timeout on busy Windows hosts 2026-08-31 12:00:33 -07:00
HexLab98 85118cc9d4 test(cli): cover mixed-config ImportError recovery hint on chat startup (#96900) 2026-09-01 00:26:55 +05:30
mbac 3e912874bf fix(models): support OpenRouter preset references 2026-08-31 11:19:26 -07:00
nftpoetrist 4d4cffd118 fix(profiles): make_targz writes to a temp file and renames, not the destination directly
tarfile.open(archive_path, "w:gz") truncates the destination the instant
it opens. If tf.add() fails partway (disk full, permission loss,
interruption), whatever was previously at that path is gone — including
an existing profile or board export the caller chose to overwrite. This
is the same failure shape a7e7de6407 just fixed for the desktop gateway
file-save path, one commit earlier in the same window, but it was never
propagated to this shared archive-writing primitive even though board
export gained a new caller into it in that same window.

make_targz now writes into a sibling temp file (mkstemp, same directory
as the destination so the final step is a same-volume rename) and only
replaces the destination via os.replace() after the archive is fully
written and closed, mirroring the mkstemp+os.replace pattern already
used throughout this codebase (agent/secret_sources/_cache.py,
cron/jobs.py, gateway/status.py, etc). The temp file is unlinked on any
failure.
2026-08-31 11:19:05 -07:00
RickyYii 3145986c20 fix(cli): honour model_aliases api_key, stop cross-provider key leak (#83612)
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes #83612.
2026-08-31 10:59:45 -07:00
Teknium 90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
Ayush Nangia c923b53913 fix(update): refuse gateway ancestor tree-kill on Windows 2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
gebilaowang404 cdd063528f fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).

Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden

Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.

Refs #90471, #89614

Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
2026-08-31 10:41:54 -07:00
phi hu 081030a7dd fix(config): warn for empty platform toolsets 2026-08-31 10:08:31 -07:00
phihu ccd32a9f0b fix(config): warn when a platform_toolsets entry is an empty list
validate_platform_toolsets() accumulated a single valid_count across every
platform, so the "zero valid toolsets" safety net was suppressed as soon as any
one platform carried a valid toolset. A platform wiped to [] — the active one,
typically cli — therefore produced no warning at all.

resolve_enabled_toolsets() honours that empty list verbatim ([] is a list, so
the platform-default fallback is skipped), leaving the agent with zero tool
schemas. The model then has nothing to call and emits the tool call as
assistant text with finish_reason=stop: no error, no warning, no log entry.
That is the silent-failure mode this module was written to prevent (#38798).

Note the asymmetry this leaves intact: a malformed *string* value is not a list,
so it falls back to the platform default and fails open (#78103); an empty list
fails closed. The fail-closed resolution is deliberate (the explicit_empty_
selection contract in tools_config.py, and #82010 wants it persistable), so this
only adds the missing warning and does not change resolution semantics.

Fixes #89050

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:08:31 -07:00