Commit Graph

7225 Commits

Author SHA1 Message Date
Teknium 1a500a43a7 fix(gateway): keep launchd_stop's bootout quiet too; prove the fd-2 invariant with a fake launchctl
Widens #106272 to the one remaining sibling: `launchd_stop()` boots out with check=True and
already handles exit 3/113/125 (job unloaded) and 5/125 (domain unmanageable) by falling through
to the PID kill, yet inherited stderr — so `hermes gateway stop` against an unloaded job printed
"Boot-out failed: 3: No such process" next to "✓ Service stopped". Same `_CAPTURE_TEXT` kwargs as
the sibling calls; an unexpected exit still raises with `e.stderr` populated.

Replaces the contributor's two kwarg-assertion tests (`capture_output is True` on a mocked
`subprocess.run`) with two invariant tests that run a real fake `launchctl` on PATH and read fd 2
through `capfd`: restart-on-unloaded prints only its ↻/✓ lines and drives
kickstart→bootout→bootstrap→kickstart; stop-on-unloaded is silent, while a real bootout failure
(exit 1) still raises with the captured stderr. Both red on origin/main, green here.
2026-09-09 10:42:32 -07:00
Davy ebcc87ad4f fix(gateway): silence expected launchctl bootout/kickstart noise on macOS
Best-effort bootout calls (unloaded-job recovery, stale-EIO retry,
plist refresh, uninstall) and the handled kickstart -k in
launchd_restart inherited the terminal's stderr, so an expected
unloaded job printed raw launchctl errors around the CLI's own lines:

  Could not find service "ai.hermes.gateway" in domain for user gui: 501
  ↻ launchd job was unloaded; reloading
  Boot-out failed: 3: No such process

Capture them with _CAPTURE_TEXT instead. The kickstart error stays
available as e.stderr for the update_cmd failure diagnostic, and the
post-bootstrap kickstart intentionally stays loud (its failure feeds
the domain-unsupported fallback). Same precedent as the reload
helper, which already runs bootout with 2>/dev/null.

[salvage: picked hermes_cli/gateway.py only; the two capture_output kwarg-assertion tests are
replaced by fd-level invariant tests with a fake launchctl in the follow-up commit]
2026-09-09 10:42:32 -07:00
teknium1 8d93081971 fix(desktop): stop flagging local/LAN auxiliary pins as stale
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.

- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
  the verdict of the one canonical classifier
  (`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
  private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
  `staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
  appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
  IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
  a global-scope address (`2607:f8b0::1`) is not local while `::1`,
  ULA and link-local still are via the `ipaddress` scope checks.

Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.

Refs #106228

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 10:33:00 -07:00
teknium1 113304199e refactor(desktop): gate the WSLg D3D12 selection inside the helper and test the real launch env
Move the WSL / /dev/dxg / d3d12_dri.so probes into _prefer_wsl_d3d12 with
the probed paths as module constants, so the launcher call site is a single
line and a test can lay out a fake WSLg host without touching real
/dev or /usr/lib. The two tests now run the real _desktop_launch_env end to
end (selected under WSL+dxg+driver; untouched with an explicit Mesa
override, off WSL, without /dev/dxg, or without the driver file) instead of
unit-testing the helper with a precomputed boolean.

Docs: one paragraph in the Desktop guide on the automatic selection and
the env vars that keep an explicit choice authoritative.

Follow-up to Xipong's fix for #106117 (salvaged from #106118).
2026-09-09 10:28:52 -07:00
Xipong 51fa04d49d fix(desktop): select installed WSL D3D12 driver before Electron exec 2026-09-09 10:28:52 -07:00
teknium1 a785db3672 fix(cli): inline the $HERMES_HOME/npmrc lookup, trim tests to two, document it
Fold the helper into _npm_lifecycle_env itself: the whole fix is one
is_file() check plus a setdefault, so a separate function, the
try/except around get_hermes_home() (it never raises) and the
dict-returning indirection were shape-gate violations. Build the path
with os.fspath so it is correct on Windows too.

Tests: keep the two invariants (file present -> NPM_CONFIG_USERCONFIG
points at it; explicit process/caller value wins and a missing file
sets nothing), fold the other two into the negative test.

Docs: one paragraph in the desktop troubleshooting page next to
ELECTRON_MIRROR describing $HERMES_HOME/npmrc.

Refs #106373
2026-09-09 10:26:06 -07:00
Konstantin Khlopkov ce4a33a9f7 fix(cli): keep $HERMES_HOME/npmrc npm config across updates (#106373)
(cherry picked from commit 339368f2215ffcfd04acf295c81806bf1e850770)
2026-09-09 10:26:06 -07:00
teknium1 d3dcc064df fix(cli): detect ssl.SSLError by type in the Codex login hint; trim tests; add the openssl.cnf snippet to docs
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
  __cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
  SSLEOFError whose text httpx did not repeat still gets the hint. The
  hint now names the TLS 1.2 diagnostic and links the providers docs
  note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
  the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
  classic-groups snippet (EN + existing zh-Hans copy).

Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
2026-09-09 10:14:58 -07:00
liuhao1024 d8eb177c93 fix(cli): keep SSL detail and add middlebox hint on Codex device-login transport errors
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.

- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
  (device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
  workaround hint; KeyboardInterrupt handling is unchanged

(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
2026-09-09 10:14:58 -07:00
Teknium e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
teknium1 02005cfe20 fix(kanban): promote refuses undone parents instead of a false --force success
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.

The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.

- drop `--force` from `promote` (parser, CLI handler, `promote_task`
  kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
  the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
  fake `ready`; the flag no longer parses

Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
2026-09-09 09:21:29 -07:00
Totoro-qaq 5d6d5fb223 fix(cli): refresh TERMINAL_CWD when --in re-homes the session
`--in DIR` only chdir'd. Every cwd consumer (resolve_agent_cwd -> Codex
app-server thread cwd, the terminal tool, context-file discovery) prefers
TERMINAL_CWD over the process cwd, so a value inherited from a parent
Hermes surface, the shell or .env survived the chdir and the session kept
running in the old directory. The local backend was rescued by cli.py's
force-export at import time; docker/ssh backends and the TUI launch path,
which never imports cli.py, were not.

Refresh TERMINAL_CWD to the --in target when it is already set. An unset
variable stays unset so the backends keep deriving from the new process
cwd and no host path is pre-seeded into ssh/container backends.

Fixes #106220
2026-09-09 09:20:12 -07:00
teknium1 060cd7f9bb fix(state): trim the futile-holder FTS diagnostic to shape and fix the remedy text
Slim redo of the mechanism from #106410 on top of its pick (no wrappers, no
persisted "kind" enum, no process-local flag that dies with the process):

- Futility = the SAME holder PID set has blocked >= _FTS_HOLDER_FUTILE_ATTEMPTS
  (10) deferrals over >= _FTS_HOLDER_FUTILE_SECONDS (30 min); tracked as
  holders_since/holders_attempts in the persisted fts_rebuild_deferral record
  and reset whenever the holder set changes. The 3-deferral/60 s escalate
  window is the orphan-reap gate and stays as is.
- ONE escalated ERROR line names each holder pid + cmdline and the remedy that
  can actually be followed from inside a gateway session: stop ONLY the other
  holder; this process's own retry admits the rebuild within 60 s. The old
  "with the gateway stopped" advice was unrunnable from a gateway-hosted
  session (the gateway is the session) and is gone from both log and doctor.
- hermes doctor renders the futile record distinctly.
- retry_deferred_fts_recovery: a capped backoff earned by holder set X no
  longer applies once the live holder set differs from X, so stopping the
  other service is followed by a retry on the next tick, not up to an hour
  later (the issue's 16-min wait).
- Tests trimmed from 5 to 2 invariants (futile line + doctor entry after N
  same-holder deferrals; backoff reset when the holder set changes); the
  contributor's control tests for changing PIDs / orphan reap are covered by
  the existing test_repeated_deferrals_reap_inactive_orphan_then_rebuild.

The "canonical writes and LIKE search remain available" WARNING is kept
because it is true on origin/main: a stale open drops every FTS trigger, so
the messages INSERT succeeds (probed live with a real state.db + a second
process holding it). Writes fail only when a peer re-publishes triggers over
the corrupt index — a separate class, not this diagnostic.

Refs #106393
2026-09-09 09:19:57 -07:00
KoNit-K e7ca9b47ac fix(state): diagnose futile FTS deferral from a permanent holder
A supervised peer never satisfies the orphan reap, so stale-FTS repair retried forever with a misleading "canonical writes remain available" warning.

(cherry picked from commit e57f3a975d311aa44da1e92c5e727eba7c8cff70)
2026-09-09 09:19:57 -07:00
Teknium 19f2f19987 fix(gateway): refuse to uninstall a systemd unit pinned to another HERMES_HOME
Defence at the exact boundary the incident crossed: systemd_uninstall() and
uninstall._remove_systemd_gateway() unlinked whatever get_systemd_unit_path()
returned. Before stop/disable/unlink, read the unit's own
Environment="HERMES_HOME=..." line (the parser status/refresh already use)
and, when it names a different home than this process, warn with both paths
and leave the unit alone. A unit without the line (hand-written) is still
removed as before.
2026-09-09 09:19:36 -07:00
Teknium 1fc033a8ef fix(gateway): s6 slot and multiplexer guards key on the profile id, not the host service suffix
With the previous commit a Docker/custom root (HERMES_HOME=/opt/data) gets a
hashed host-service suffix. Three callers used `_profile_suffix() or
"default"` as the PROFILE id, which is a different question: the s6
supervisor's slot for the root home is `gateway-default` regardless of where
the root lives, and the multiplexer's "am I a named profile" probe must not
treat a hash as a profile name. Route them through hermes_constants.
profile_name_for_home() (root -> "default", <root>/profiles/<name> -> name)
with the service suffix as the fallback for unknown layouts.
2026-09-09 09:19:36 -07:00
Teknium 4746e34448 fix(gateway): foreign HERMES_HOME no longer resolves to the default hermes-gateway unit
_profile_suffix() compared HERMES_HOME against get_default_hermes_root(),
which treats ANY home outside ~/.hermes (Docker /opt/data, a mktemp dir) as
"the root itself". Every such home therefore collapsed to the bare
`hermes-gateway` service name and the default profile's unit path
(~/.config/systemd/user/hermes-gateway.service); the documented
"else a short hash of the path" branch was unreachable.

A parity harness run with HERMES_HOME=$(mktemp -d) called
uninstall_gateway_service(), resolved to the production unit, ran
`systemctl --user stop/disable`, unlinked it and daemon-reloaded. With the
unit gone Restart= could not revive it: all cron jobs and every messaging
platform were down for 6.5 days.

Compare against the platform-native default home (~/.hermes) for the bare
name; keep the profile name for <root>/profiles/<name>; everything else
(temp dirs, Docker /opt/data) gets its sha256[:8] suffix as the docstring
always promised. The Docker image supervises with s6 (`gateway-<profile>`
slots), not systemd/launchd, so the bare host-service name was never load-
bearing there.
2026-09-09 09:19:36 -07:00
kshitijk4poor ee7fc3bc06 refactor(auth): one stat per store file in the heal fingerprint
_size duplicated _mtime_ns with a different attribute and the fingerprint
stat'ed each of the four files twice, so mtime and size could come from two
different versions of the file. _stat_sig returns both from one stat. The
persisted compare goes through a JSON round-trip so nested tuples match the
lists they read back as.
2026-09-09 21:17:35 +05:30
John Paul Soliva 57e03e2be6 perf(auth): persist the forked-OAuth clean mark so a fresh process skips the heal's locks
`_heal_forked_single_use_oauth_grants()` runs on every `load_pool()`, and its
clean mark lived only in memory. Every fresh `hermes` invocation and every new
worker therefore re-took the heal's two nested EXCLUSIVE auth-store locks just
to rediscover a store it had already cleared — 4 acquisitions per process on a
two-provider profile, every one of them finding nothing to consolidate. Behind
a sibling process holding those locks that costs a full
`AUTH_LOCK_TIMEOUT_SECONDS` per provider before the process can do anything at
all: measured 30.1s for two providers.

Persist the mark next to the store it describes (`<profile>/cache/
oauth_heal_clean.json`, 0600, paths and stat data only — no credential
material) and consult it BEFORE taking any lock.

Outliving the process means the mark needs a stronger key than the in-memory
one did:

- The ROOT store joins the fingerprint. This heal consolidates root → profile,
  so root acquiring a counterpart turns a row the heal deliberately KEPT into a
  fork it must strip. The in-memory mark could ignore root because it died with
  the process; a persisted mark would keep skipping a heal that has become
  necessary.
- File sizes join it too, so a metadata-preserving rewrite (`rsync -t`,
  `tar -p`, a restore) cannot leave a stale mark looking current indefinitely
  rather than for one process.

Measured on an isolated HERMES_HOME with two OAuth providers:

    lock acquisitions per fresh process    4     -> 0
    contended load_pool() x2               30.1s -> 0.00s
    mark-file reads per 100 load_pool()    -     -> 2 (one per provider)

The mark stays a cache: absent, unreadable, corrupt or wrong-shaped content all
mean "unknown" and fall through to the locked heal, and a failed write only
means the next process re-runs it — the behaviour before this cache existed.
2026-09-09 21:17:35 +05:30
kshitijk4poor e2bd400233 fix(models): only cache unreachability, not HTTP errors; key the entry via base_url_origin
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.

_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
2026-09-09 21:16:48 +05:30
kshitijk4poor 4e2a871adb fix(models): clear probe negative-cache entry on a successful probe 2026-09-09 21:16:48 +05:30
finn763 48b8528e7c fix(desktop): stop UI freeze on unreachable provider Closes #81123 2026-09-09 21:16:48 +05:30
kshitijk4poor 6acd90d11a refactor(dashboard): drop the redundant hash() pre-check in coalesced_read
pending.get(key) already raises TypeError for an unhashable key.
2026-09-09 21:16:40 +05:30
Gianpietro Dal Zio 653418d842 fix(dashboard): coalesce expensive reads before worker admission 2026-09-09 21:16:40 +05:30
kshitijk4poor 1eaeb73839 fix(models): OpenRouter disk snapshot follows the configured catalog TTL, not the default constant
refresh_interval_seconds() honours model_catalog.ttl_minutes / legacy ttl_hours;
reading DEFAULT_TTL_MINUTES would let the snapshot and the manifest it is
filtered from go stale on different clocks for anyone who changed the TTL.
2026-09-09 21:16:28 +05:30
kshitijk4poor 8b66bc88df refactor(models): OpenRouter disk cache reuses catalog TTL and atomic JSON cache helpers 2026-09-09 21:16:28 +05:30
Hyperion 4c003069f3 fix(models): persist curated OpenRouter catalog to disk so picker opens fast
Re-derived from PR #96099 (f127ec4e) on current main; the :nitro/:floor validate hunk is omitted because main already handles routing suffixes in hermes_cli/models_validate.py.
2026-09-09 21:16:28 +05:30
finn763 a74e76632c perf(picker): read Nous pricing cache-only when building the picker row
_nous_picker_model_ids only uses the ids the Portal unions append — both
unions discard the pricing map (`model_ids, _ = union_with_portal_*`) — yet
it called get_pricing_for_provider("nous") without cached_only, so a cold
pricing cache paid a full /v1/models round-trip (network timeout on a slow
Portal) on the picker-open path for nothing. Pass cached_only=True; the
background pricing prewarm (#101685) fills the same cache for later opens.

Re-derived from #102099 by @finn763: the original patched
hermes_cli/model_switch.py, which 3b1ecfc0a1 decomposed; the live call site is
hermes_cli/model_switch_providers.py.

Based on #102099 by @finn763.
2026-09-09 21:12:22 +05:30
moken627-hub 35af06c009 fix(desktop): coalesce projects/tree sidebar scans and memoize raw config parses
The all-profiles sidebar polls GET /api/profiles/projects/tree, the one
heavy sidebar endpoint that was not wrapped in @_sidebar_singleflight_cache.
Every poll fanned out list_profiles() + _build_project_tree() over every
profile (51 on this box, ~17k SKILL.md files walked), and each profile
resolution re-parsed its config.yaml because read_user_config_raw() ran
uncached. On a 2-vCore VPS running 'hermes serve' for the Desktop remote
backend this pinned both cores (py-spy: 110-116% sustained, HostHighCPU).

- Wrap get_profiles_projects_tree in the existing single-flight cache,
  matching get_profiles_sessions_sidebar (5s TTL, errors[] not cached).
- Memoize read_user_config_raw() on a (st_dev, st_ino, st_size,
  st_mtime_ns) fingerprint under _CONFIG_LOCK, same strategy as
  read_raw_config(). Hits return a deepcopy so write-back round-trips keep
  their fresh-dict semantics; parse errors are never cached; the docstring
  no longer claims 'no caching'.
- Tests: memo semantics (deepcopy isolation, inode-replace reparse, error
  non-caching, parse-count), a wrapper-presence pin for both heavy sidebar
  endpoints, and a cold-cache autouse fixture in the scope tests (the 5s
  TTL otherwise leaks one test's payload into the next).

Measured on the affected box: list_profiles 1.9s -> 0.16s warm, serve CPU
112% -> 22-34%, host CPU 50-70% -> ~26%.
2026-09-09 21:11:56 +05:30
kshitijk4poor ead7e91dab refactor(recovery): stream the salvaged population; table the shape rules
- Pass 1 no longer materialises every classified record (full
  `messages.content` included) until pass 2; `LayoutEvidence` keeps only the
  capped per-position value sets (+ sessions rows for the one cross-column
  invariant) and pass 2 re-streams the lost_and_found tables. A 276 MB
  corrupted store no longer has to fit in memory.
- `_sentinel_holds` / `_text_shape_holds` if-ladders become rule tables.
- Tests trimmed to the three that bind behaviour (upgraded store maps by
  name; verifier refuses when rows matched no layout; replayed history ends
  at the current schema — the drift guard). No behaviour change; reverting
  inference to "no layout" still fails the name-mapping test.
2026-09-09 18:28:57 +05:30
kshitijk4poor 17c43aba06 fix(recovery): infer the salvaged store's physical layout from its schema history
A store's physical column order depends on which schema it was created at and
which ALTER TABLE ADD COLUMNs it lived through; one hardcoded "upgraded"
layout cannot cover them. `session_schema_history` records the declared
schema of the salvaged tables over time; `reachable_physical_layouts`
replays it to enumerate every physical order a store can have. The mapper
infers the layout once per kind from the whole recovered population (one
store wrote all of them), maps cells by name, and counts records whose width
matched no layout; the recovery verifier refuses to report such a salvage as
healthy.

Rebased onto the simplified session modules (c88d60551e, b9b4600cb2,
7e5a1a11d9, 1915a0d27e); no behaviour change from the pre-rebase branch.
2026-09-09 18:28:57 +05:30
Justin Wilson 7704712168 fix(sessions): map lost_and_found cells by physical column names
Upgraded state.db files gain columns via ALTER TABLE ADD COLUMN, so
physical order diverges from SCHEMA_SQL. Prefix-mapping onto the fresh
template put started_at at 0 and shifted titles/models. Insert by name
using known physical layouts; keep the plausibility gate.

Fixes #101409
2026-09-09 18:28:57 +05:30
Teknium 41b4555ed9 feat(plugins): catalog is the sole discovery system — re-port onto main's layout
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
  .hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
  update, and the dashboard/TUI payload builders. plugins_cmd.py only
  gains the hooks (cmd_install catalog branch, cmd_update / dashboard
  update re-pin, dashboard_install_plugin catalog_name + kill list,
  dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
  test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
  published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
  fallback) instead of the unauthenticated GitHub contents API (60 req/h,
  1 request per entry); in-tree and live removals are unioned so a stale
  cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
  limits); _plugin_runtime_status shared from web_server_dashboard.py;
  hub rows carry removed_reason. TUI plugins.manage gains catalog_name
  install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
  (real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
  first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
  plugin-catalog/** so entry merges republish it.
2026-09-09 04:38:01 -07:00
Teknium 4e312cf22d feat(plugins): requires_hermes manifest gate on main's manifest/loader siblings
Re-port of the PR's version gate onto the decomposed layout: the field and
parser live in plugins_manifest.py (with running_hermes_version /
version_satisfies helpers), the load-time skip in plugins_loader.py
before any import. Unsatisfied plugins record an error and never run
register(); one invariant test proves both halves.
2026-09-09 04:38:01 -07:00
Teknium 7e4d02fef5 fix(cron): unpinned jobs run on their creation-snapshot model instead of failing closed
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).

The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.

Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".

Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).

The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
2026-09-09 04:32:13 -07:00
Teknium 48465c3933 fix(profiles): --clone-all no longer copies cron jobs into the new profile
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.

Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.

--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
2026-09-09 04:27:24 -07:00
Teknium d47adec28f Merge origin/main into feat/plugin-catalog
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
2026-09-09 04:15:27 -07:00
buihongduc132 c1ff9390f6 fix(backup): prune old hermes-backup-*.zip after each run, keep last 3
run_backup() previously wrote "hermes-backup-<timestamp>.zip" on every
invocation without deleting old ones. Hourly callers accumulated 157 zips
(14 GiB). Add _prune_run_backups() to keep the newest N (default 3,
configurable via backup.run_backup_keep or --keep CLI flag).
2026-09-09 03:33:14 -07:00
Teknium bf28c2fe1f fix(desktop): local endpoint probes ignore HTTP(S)_PROXY; non-2xx names the status (#63472)
httpx honours the env/system proxy (on Windows, the registry ProxyServer
even with no *_PROXY vars) but never the bypass list, so a system proxy
(Clash, corporate) answered 127.0.0.1 probes from both Desktop validators
with its own error page. That parsed as models=[] and the GUI said
"advertised no models at /v1/models" for a llama.cpp server the CLI
(urllib, honours <local>) saw fine.

Local endpoints (loopback, LAN, Tailscale via is_local_endpoint) now
probe with trust_env=False; public endpoints keep honouring env proxies.
A reachable endpoint answering non-2xx with no model list reports
"<url> answered HTTP <status>." instead of an empty catalog, so the
onboarding card stops telling the user to start a model.

Reimplemented on the decomposed router (the original patched
web_server.py before the split). Diagnosis and fix direction by
Solitud1nem in #63656; Windows registry-proxy confirmation by
Ulysses-Gaia on #63472.

Live repro (real loopback server, HTTP_PROXY=http://127.0.0.1:9):
  before  ok=False reachable=False 'Could not reach .../v1/models'
  after   ok=True  models=['Qwen3.6-35B-A3B-Q5_K_M.gguf']

Co-authored-by: Solitud1nem <76743883+Solitud1nem@users.noreply.github.com>
2026-09-09 03:33:06 -07:00
Teknium 734461d213 fix(models): same-URL custom endpoints stop evicting each other's cached catalog; no-probe picker opens revalidate
Two picker-freshness defects in cached_fetch_api_models():

1. The disk cache row was keyed on base_url only, with the credential
   fingerprint stored inside the row. N custom_providers entries sharing
   one proxy URL with different keys (#106184) took turns overwriting the
   single slot; every sibling then failed the fingerprint check, got an
   empty catalog, and disappeared from the Desktop pickers (which hide
   zero-model rows). Key on url#fingerprint so each credential owns a row.

2. cache_only opens (Desktop model.options without refresh) served a
   past-TTL row for up to 7 days without ever revalidating, so a model
   loaded on a non-current local endpoint stayed invisible until the user
   found "Refresh Models". Serve the stale row AND spawn the same
   off-thread SWR refresh the blocking path uses; the caller still never
   waits on the network.

Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
  GUI no-probe open  before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
                     after  {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
2026-09-09 03:33:06 -07:00
briandevans f4c55323fa fix(cli): back up config.yaml before --reset overwrites it
`hermes setup --reset` calls `save_config(copy.deepcopy(DEFAULT_CONFIG))`,
which writes `get_hermes_home()/config.yaml` — the exact file the backup
block a few lines below copies to `config.yaml.bak.<timestamp>`. Because the
copy ran after the reset, the backup captured the defaults that had just been
written, not the user's config. The one invocation where a backup matters most
produced a worthless one, and the original was unrecoverable.

The block's own comment already claimed it runs "before setup modifies it";
on the --reset path that was false. Move it above the --reset branch so it
captures the true pre-setup state on every path.

Also report the backup location on the --reset path. --reset is destructive
and can leave the wizard early (the non-interactive return exits before the
end-of-setup notice), so a user who just lost their config was never told
where the copy is. The end-of-setup notice is unchanged for the normal path
and is suppressed only when it has already been shown, so no run prints it
twice; the shared wording now lives in one helper.

Behaviour otherwise preserved: `copy2` (config.yaml holds secrets, so mode is
preserved), the try/except fallback to `_backup_path = None`, and the existing
notice for the full-setup path.

Follow-ups deliberately out of scope: pruning accumulated `.bak.*` files, and
printing the notice on the other early-return paths (--portal, section runs).

Refs #3522
2026-09-09 03:32:58 -07:00
Teknium bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium 677e8ed8a4 fix(desktop): SSH remote backend stops following the host's sticky active_profile
A Desktop-owned `hermes serve --isolated --ssh-session-token-file ...` child
is spawned with an explicit `--profile <name>` when the connection names a
remote profile, and with no flag for the remote root home. Without the flag,
`_apply_profile_override` read the remote host's sticky `active_profile`
file and re-homed the backend into whatever profile the user last selected
on that machine's CLI. Settings then read one config.yaml while the remote
gateway wrote another, so model picks and toggles "didn't stick".

Treat the SSH token flag as a fixed-identity marker, the same way
supervisor-launched gateway children are (#74872): a Desktop backend's
profile is chosen by the client, never by the host.

Live repro (before/after, temp HERMES_HOME with active_profile=foo):
  serve --isolated --ssh-session-token-file ...   hermes_home=<root>/profiles/foo -> <root>
  same + --profile foo                             hermes_home=<root>/profiles/foo (unchanged)
  serve (no token file, user CLI)                  hermes_home=<root>/profiles/foo (unchanged)
2026-09-09 02:34:59 -07:00
Teknium 13c580422c fix(auth): carry pool-row lineage into the provider-block heal
With account-identity matching gone, the providers.<id> block consolidation
only fired on shared token material. A historical fork (same copied pool-row
id, profile rotated, both pairs diverged) then healed the pool row into root
but left root's providers.openai-codex block on the spent pair; root's next
load_pool() re-seeds its device_code row FROM that block and undid the heal.

_HealPass now records that a profile pool row matched root by copied id or
shared tokens and passes that verdict to _heal_forked_provider_block, which
accepts it as lineage proof. No account-identity guessing is restored; an
independent same-account grant (no id/token match) is still left alone.

Follow-up to simpolism's #106177.
2026-09-09 01:46:03 -07:00
simpolism 73f9de0c3a fix(auth): preserve independent same-account OAuth grants 2026-09-09 01:46:03 -07:00
kshitijk4poor 8077206073 fix(cli): keep the no-stored-model early return ahead of the route read
CI: tests/cli/test_cli_resume_command.py builds bare HermesCLI objects without .model; the
refactor read self.model before the stored-model check the contributor's code made first.
2026-09-09 12:41:07 +05:30
kshitijk4poor 32273b8118 refactor(cli): one stored_session_route for interactive and one-shot resume
_apply_stored_session_runtime was a line-for-line copy of the first half of
_restore_session_model (stored-model guard, session_gateway_runtime, bare-custom heal,
model/provider-changed check). Extract that pure decision into
cli_model_switch_mixin.stored_session_route and have both resume paths call it; the
one-shot keeps only the _ModelChoice mapping and the drop-ambient-key rule.

main.py stops re-normalising `resume` — _resolve_chat_session_args already did.
Tests trimmed from 20 to 13: near-duplicate unit tests of the private helpers go, the
end-to-end _run_agent contracts (stored runtime + reopen; explicit --model wins) and the
empty-session-keeps-id case stay.
2026-09-09 12:41:07 +05:30
liuhao1024 8aa773af89 fix(cli): restore stored session runtime and reopen ended rows on oneshot resume
Review fixes (#105957):

- A resumed one-shot ignored the session's stored model/provider runtime:
  _resolve_model_and_provider()/resolve_runtime_provider() ran before
  _load_resume_target(), which only loaded the session id + transcript, so an
  ambient config (e.g. openrouter/ambient-model) served the resumed transcript
  instead of the stored route (custom:stored/stored-model). The stored runtime
  is now applied before runtime resolution, with the same contract as the
  interactive _restore_session_model(): stored model/provider/base_url/api_mode
  replace the ambient choice unless --model was passed explicitly, and a
  changed provider drops the ambient api_key so resolution re-fetches
  credentials for the restored endpoint.

- Passing the resumed id to AIAgent did not reopen the already-ended session
  row: end_session() only writes rows whose ended_at is null and the
  existing-row upsert never clears the end fields, so the resumed turn was
  recorded under a session that stayed closed and its new lifecycle boundary
  was lost. _load_resume_target() now reopens the row (best effort), same as
  the interactive resume does before continuing.
2026-09-09 12:41:07 +05:30
liuhao1024 5ff6cb0edb fix(cli): keep the resolved session id when a resumed oneshot session is empty
Review finding on #105957: `_load_resume_target` returned None for a
resolved session with no stored messages, so `hermes -z "hello" -c <title>
--create-if-missing` recorded the turn under a freshly minted session id and
the just-created titled session stayed empty. Preserve `resolved` unconditionally — the interactive /resume path keeps the selected id for an
empty session too; only the history replay is empty. Regression tests pin the
durable id for both a plain empty session and an empty compression-chain head.
2026-09-09 12:41:07 +05:30
liuhao1024 86d606ca34 fix(cli): honor --resume in one-shot mode (#105892)
The -z exit path accepted --resume/-c in the parser but never forwarded
args.resume: every resumed one-shot turn silently started a fresh session,
so each wire request carried only [system, current user] and the model
lost all prior context (reported against Ollama/custom OpenAI-compatible
endpoints, but provider-independent).

Normalize session args (latest/title/--continue/--in + cwd restore) via
the chat path's _resolve_chat_session_args before the oneshot exit path
takes over, then load the resumed transcript in _run_agent through the
same contract the interactive CLI uses (compression-chain redirect,
safe-resume guard, session_meta filtering) and continue the existing
session id instead of creating a new one. An explicit --resume of an
unknown session now fails loudly instead of starting fresh.
2026-09-09 12:41:07 +05:30