Commit Graph

3613 Commits

Author SHA1 Message Date
Ben Barclay 180291162f feat(telemetry): opt-in shared-metrics exporter (#95278)
feat(telemetry): opt-in shared-metrics exporter
2026-09-02 08:35:36 +10:00
Jeffrey Quesnelle c56f8cdd48 Merge pull request #100667 from NousResearch/feat/local-models-squash
feat: local models — managed llama.cpp runtime with one-click desktop  setup
2026-09-01 17:53:28 -04:00
Pedro Fontana b3576a29c3 Merge pull request #97354 from NousResearch/fix/nous-org-model-policy
fix(nous): honour the org model policy in the model pickers
2026-09-01 18:20:18 -03:00
kshitijk4poor d4611ac837 chore: review follow-ups for salvaged #99954
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
2026-09-02 01:52:11 +05:30
Sahilvishnaliya d0f0afb009 fix(update): post-update state.db guard covers every profile, not just the root
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).

Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.

Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes #97994.
2026-09-02 01:52:11 +05:30
kshitijk4poor 6b46725a17 fix(update): surface leftover update autostashes older than 7 days (#63717)
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
2026-09-02 01:50:21 +05:30
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Mariano Nicolini f48e61bb99 fix(nous): show the policy notice only when the filter narrowed the list 2026-09-01 16:37:36 -03:00
Mariano Nicolini 3c4e84c166 fix(models): peek past expired and superseded pricing entries 2026-09-01 16:35:47 -03:00
Mariano Nicolini d7520b2822 fix(aux): seed the shared Nous catalog entry with the pickers' arguments 2026-09-01 16:34:48 -03:00
Teknium 33797073bb fix: harden startup route salvage — aggregator-slug guard, alias credential ownership, oneshot dedup
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
  routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
  URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
  provider label never carries the vendor token to the alias host (#28660);
  route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
  already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
2026-09-01 12:07:35 -07:00
joaomarcos 4c870951e2 fix(cli): resolve startup model routes before provider defaults
Resolve configured aliases and provider/model inputs before HermesCLI attaches the configured default provider. Keep aggregator namespaces intact, cover oneshot startup, and document the supported CLI forms.
2026-09-01 12:07:35 -07:00
liuhao1024 f82d2f1301 fix(models): scope prefix routing to user-configured providers only 2026-09-01 12:07:35 -07:00
liuhao1024 4033f3fc5f fix(models): honor vendor/model prefix and dict model.aliases in provider detection (#87189) 2026-09-01 12:07:35 -07:00
David Metcalfe e4c35e397c test: pin /api/model/options to _config_profile_scope for selected profiles
Regression for #58576: _profile_scope holds _SKILLS_PROFILE_LOCK across
the payload build, which can block up to 15s on a models.dev cache miss
and starve concurrent /api/config on the same lock. The test records
which scope the handler enters for a selected profile and asserts only
the config-only (contextvar) scope is used.
2026-09-01 12:07:00 -07:00
Teknium ab9866bc64 fix(gateway): survive Windows Job-Object teardown across gateway restarts (#48820)
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.

Three surgical changes:

1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
   The inlined watcher now routes the respawned gateway's stray
   stdout/stderr to the same sidecar log gateway_windows._spawn_detached
   uses (DEVNULL only as fallback), so a gateway killed moments after
   respawn leaves a trace. Direct implementation of the 4th repro's
   hardening suggestion (1).

2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
   canonical _spawn_detached, so the respawned gateway's exit-diag /
   lifecycle records show whether it escaped the parent Job Object — a
   job-teardown kill is no longer indistinguishable from any other silent
   death.

3. Post-update resume verifies liveness before vouching
   (hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
   runs the same provisional-hit + 2s-confirmation liveness poll every
   other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
   with all_profiles= for the fleet) before printing ✓, writes the #91675
   start attestation for the verified PIDs, and fails the resume with a
   "restart could not be verified" warning + recovery hint when no stable
   gateway appears. Suggestion (2) of the 4th repro; closes the last
   silent-success hole in the family (#84185 fixed the cold-start leg,
   #91675 the direct-start leg; this is the relaunch leg).

Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.

Fixes the Bug-1 relaunch-trust leg of #48820.
2026-09-01 11:43:36 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium 28834a2098 test: raise tight wall-clock bounds that flaked on loaded CI runners
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).

Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
2026-09-01 10:52:42 -07:00
kshitijk4poor b81383ec21 fix(auth): compare pool-identity callers against all candidate keys
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:

- _prune_replaced_custom_model_config_credentials skipped only the
  preferred key, so a keyed provider's own legacy-named pool
  (custom:b.ai) was false-pruned of its current model_config credential
  when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
  key, so a legacy-named pool stopped being seeded from model.api_key.

Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.

Follow-up to #100413.
2026-09-01 22:42:27 +05:30
xxxigm 43470980bf test(auth): cover keyed providers.<key> credential-pool lookup
New-style providers store keys under the durable config slug, but
runtime still looks up custom:<display-name> and sends a placeholder.
2026-09-01 22:42:27 +05:30
Teknium 6ddafd34f0 test(proxy): cover multi-line SSE data joins, truthy lastOne, EOF-without-blank-line dispatch 2026-09-01 10:12:21 -07:00
rainbowgits ce7f805869 fix(proxy): append SSE [DONE] when Nous streams omit the sentinel
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:12:21 -07:00
Teknium e9541d213f fix(gateway): never print ✓ for a Windows gateway that dies after the liveness poll
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.

Two layers:

1. _wait_for_gateway_ready now treats the first hit as provisional: the
   gateway must stay visible through a 2s confirmation window
   (_confirm_gateway_stable) before it is reported ready; a death during
   confirmation resumes polling until the deadline. Failure output is an
   honest ✗ with the Job Object explanation and the schtasks /Run
   recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
   state/gateway.start-attestation.json with the vouched-for PIDs. The
   next `gateway start`/`gateway status` invocation checks it — if the
   attested PIDs are gone with no clean-exit record in the lifecycle
   ledger, the CLI reports (once) that the previous ✓ was false and
   prints the schtasks recovery hint. `gateway stop` and a clean
   lifecycle-ledger exit clear the marker silently.

Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.

Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.

Fixes #91675
2026-09-01 10:06:06 -07:00
Teknium 045865377c fix(update): restore user model settings config.yaml rewrites drop during update
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.

Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.

6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.

Fixes #64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
2026-09-01 09:54:22 -07:00
Konstantin Khlopkov aea2daee29 fix(backup): import into the active HERMES_HOME instead of the root home
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.

Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.

Fixes #99839
2026-09-01 09:28:46 -07:00
Teknium ada3c28d45 fix(restore): fail closed on in-process holders before unlinking state.db sidecars
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).

Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).

Part of the #90837 sidecar-unlink audit (wave 6).
2026-09-01 09:27:47 -07:00
Teknium 9f377264aa refactor(update): consolidate the gateway drain triage into one shared helper
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.

No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
2026-09-01 08:34:51 -07:00
salch-cred 49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Teknium 71c4bcf7af fix(cron): surface missed-fire catch-up lateness in hermes cron list / status
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).

- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
  scheduled_at, dispatched_at, lateness_seconds, and kind
  (on_time / late / catch_up, classified against the ticker tolerance and
  the schedule's catch-up grace window). Manual triggers and one-shots are
  not stamped (no scheduled instant to be late against / retired beyond
  grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
  show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
  catch-up, in both the built-in ticker and external provider paths.

CLI surface only — no new tools, no policy engine.

Addresses the visibility half of #99879.
2026-09-01 08:31:37 -07:00
Justin Wilson fb18fedf29 fix(updater): rebuild desktop on Windows hand-off repair path
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.

Fixes #97343
2026-09-01 08:27:58 -07:00
JoaoMarcos44 86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
Teknium 35f4ababe1 test: the managed-dashboard restart now continues the serve scan
test_user_scope_restart_never_falls_back_to_system_or_sudo asserted the
short-circuit (#92145 barrier 5) that this change deliberately removes.
Its real invariant — user scope never falls back to system scope or sudo —
is kept; the scan-continues side is now asserted instead of forbidden.
2026-09-01 07:00:54 -07:00
Teknium 5a677479e9 test: pin the no-authority contract with the serve probe disabled
The PR's manual-only-fleet test assumed no recovery child is spawned, but on
any Linux host with systemctl the serve-unit authority alone now spawns the
child (test_serve_only_fleet_still_spawns_the_recovery_child pins that side).
Disable the probe explicitly so the test states which contract it pins —
this was the red 'Python tests / Run tests' leg on the original PR head.
2026-09-01 07:00:54 -07:00
joaomarcos 27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
joaomarcos 2d70d67b15 fix(update): stop the managed dashboard restart from skipping serve backends
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.

The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
2026-09-01 07:00:54 -07:00
joaomarcos f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00
Teknium be597fc730 fix: extend corrupt-config fail-closed guard to gateway, serve, and cron surfaces
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
  construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
  on every surface
2026-09-01 07:00:22 -07:00
embwl0x 6f85df97fd fix(cli): keep quiet prompts interactive 2026-09-01 07:00:22 -07:00
embwl0x 55e7ecd260 test(cli): cover config guard edge cases 2026-09-01 07:00:22 -07:00
embwl0x c335dc734a fix(cli): reject corrupt config in noninteractive runs 2026-09-01 07:00:22 -07:00
lesseradmin 779482598f fix(desktop): pass --disable-setuid-sandbox on the userns launch path
When chrome-sandbox is present but not root-owned 4755, Chromium can still
abort via setuid_sandbox_host even though the namespace sandbox works.
After the userns probe skips sudo, append --disable-setuid-sandbox so
.desktop/no-TTY launches keep the namespace sandbox without a privilege
prompt. Does not add --no-sandbox.

Fixes #51327
2026-09-01 02:32:39 -07:00
4dlt 3a7f2234a6 fix(cli): use Chromium's namespace sandbox when userns is available on Linux
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).

On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.

Fixes #88032
Fixes #51327
Fixes #58593

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 02:32:39 -07:00
Teknium a1c25d393a feat(desktop): built-in optional-skills catalog in Capabilities → Skills with one-click install
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.

- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
  scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
  install-state spinners off $hubActions, and an OfficialSkillDetail pane
  (hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
2026-09-01 01:39:59 -07:00
liuhao1024 46ac31e84c fix(desktop): sanitized deferral evidence for ledgered manual serve blockers
The Desktop venv-blocker scan (since #99724) defers ledger-verified
serve/dashboard holders to the CLI updater's stop+relaunch rungs, but the
scan output only carried an opaque deferred_backends count — nothing
explained WHICH holders the deferral consumed or why they vanished from
processes.

Add sanitized decision evidence (#98350): deferred_backend_evidence lists
structured ledger identity only (pid, purpose, recorded port) — never the
command line, which can carry tokens or private endpoints. Adds a desktop
parser contract fixture proving the consumer tolerates the diagnostics
while keeping blocked/processes authoritative.

Salvaged from PR #98350; the exemption half of that PR was independently
consolidated on main via #99724 (_is_updater_owned_backend).
2026-09-01 01:30:09 -07:00
kshitijk4poor 801466765b fix: harden temp fallback against predictable-path attacks; neutral error text; docs
Final-diff review findings (/simplify-code on the full 3-commit stack):

- Temp-dir fallback uses a per-uid name (hermes-profile-exports-<uid>) and
  get_profile_export_path refuses a pre-existing symlink or a directory
  owned by another user — a fixed /tmp/hermes-profile-exports is a
  predictable shared path a local attacker could pre-create to receive the
  secret-bearing archive. Regression test mutation-checked.
- Fail-closed message reworded interface-neutrally (the web API surfaces it
  as HTTP 400 detail where '-o' alone made no sense).
- Docs now cover the temp fallback and the fail-closed refusal.
2026-09-01 01:00:23 -07:00
kshitijk4poor 525dd12da7 fix: fail closed when no safe export destination exists; polish salvage edges
- _profile_export_directory(): when the managed store, the home-sibling
  store, AND the temp dir all resolve inside Git checkouts, raise a clear
  ValueError instead of warning and proceeding — a stderr warning would not
  stop a scripted export from staging a secret-bearing archive in a source
  tree, which is the exact #92457 incident class. All three callers already
  surface ValueError cleanly (CLI/TUI print Error: + exit, API returns 400).
- .dockerignore: drop the /default.tar.gz line made redundant by the global
  *.tar.gz pattern this PR adds.
- hermes profile export -o help text: stop advertising the old
  <name>.tar.gz cwd default.
- Tests: cwd-in-unrelated-checkout topology (the second production shape
  from the blocking review) and the fail-closed path. Mutation-checked:
  both fail on the pre-fix helper.
2026-09-01 01:00:23 -07:00
kshitijk4poor 26fb8f60e6 fix: anchor checkout detection to the export path, not cwd
Follow-up to the salvaged #92689:

- _profile_export_directory() now proves safety on the export dir's OWN
  ancestry (_inside_git_checkout) instead of walking Path.cwd(). The old
  heuristic missed the checkout whenever HERMES_HOME sat inside one but
  the process ran from elsewhere (cron, service manager) — the export
  landed back inside the source tree, the exact incident class.
- When every candidate is inside a checkout, warn instead of silently
  violating the invariant.
- CLI/TUI export callers: move get_profile_export_path() inside the try
  and catch OSError too — a bad profile name or read-only home printed a
  raw traceback instead of the clean error main previously gave.
- Tests: bind module objects at call time (importlib) so sibling reload
  pollution in the tests/hermes_cli sweep can't divorce monkeypatches
  from the code under test; add regression tests for the cwd-independent
  topology and the clean-error path.
- Docs: mention the ~/.hermes-profile-exports fallback store.
2026-09-01 01:00:23 -07:00
joaomarcos c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
Ben Barclay 56916841b5 refactor(dashboard-auth): replace PKCE cookie payload with base64url(JSON) codec (#99210)
The PKCE cookie's payload has now needed three serialization fixes at
the same spot: the original flat 'k=v;k=v' string tripped http.cookies'
\073 quoted form (dropped whole by strict cookie parsers like Go's
net/http — #83832 field case), and #99176 URL-encoded the whole flat
payload to stay inside the RFC 6265 cookie-octet set. The stacked
layers (single-encoded next=, ';' joins, whole-payload encoding, legacy
discriminator) were the recurring defect source.

Kill the bug class instead of patching it again: the payload is a dict
end-to-end and goes on the wire as base64url(JSON) — the urlsafe
alphabet is a strict subset of cookie-octets, and JSON framing means no
segment value can ever collide with a delimiter. parse_pkce_payload
keeps a three-rung compatibility ladder (base64url(JSON) -> oldest flat
form split-as-is -> #99176 unquote-then-split) for in-flight cookies
during a rolling upgrade (10-minute TTL); a new cookie hitting an old
server fails the OAuth state check and the user just retries.

The 'next' segment is stored as its plain validated path — no extra
encoding layer, so the post-login redirect Location is byte-for-byte
the original target.

Refs #99176, #84065.
2026-09-01 11:47:38 +10:00