Commit Graph

3864 Commits

Author SHA1 Message Date
Teknium 1a500a43a7 fix(gateway): keep launchd_stop's bootout quiet too; prove the fd-2 invariant with a fake launchctl
Widens #106272 to the one remaining sibling: `launchd_stop()` boots out with check=True and
already handles exit 3/113/125 (job unloaded) and 5/125 (domain unmanageable) by falling through
to the PID kill, yet inherited stderr — so `hermes gateway stop` against an unloaded job printed
"Boot-out failed: 3: No such process" next to "✓ Service stopped". Same `_CAPTURE_TEXT` kwargs as
the sibling calls; an unexpected exit still raises with `e.stderr` populated.

Replaces the contributor's two kwarg-assertion tests (`capture_output is True` on a mocked
`subprocess.run`) with two invariant tests that run a real fake `launchctl` on PATH and read fd 2
through `capfd`: restart-on-unloaded prints only its ↻/✓ lines and drives
kickstart→bootout→bootstrap→kickstart; stop-on-unloaded is silent, while a real bootout failure
(exit 1) still raises with the captured stderr. Both red on origin/main, green here.
2026-09-09 10:42:32 -07:00
teknium1 8d93081971 fix(desktop): stop flagging local/LAN auxiliary pins as stale
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.

- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
  the verdict of the one canonical classifier
  (`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
  private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
  `staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
  appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
  IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
  a global-scope address (`2607:f8b0::1`) is not local while `::1`,
  ULA and link-local still are via the `ipaddress` scope checks.

Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.

Refs #106228

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 10:33:00 -07:00
teknium1 113304199e refactor(desktop): gate the WSLg D3D12 selection inside the helper and test the real launch env
Move the WSL / /dev/dxg / d3d12_dri.so probes into _prefer_wsl_d3d12 with
the probed paths as module constants, so the launcher call site is a single
line and a test can lay out a fake WSLg host without touching real
/dev or /usr/lib. The two tests now run the real _desktop_launch_env end to
end (selected under WSL+dxg+driver; untouched with an explicit Mesa
override, off WSL, without /dev/dxg, or without the driver file) instead of
unit-testing the helper with a precomputed boolean.

Docs: one paragraph in the Desktop guide on the automatic selection and
the env vars that keep an explicit choice authoritative.

Follow-up to Xipong's fix for #106117 (salvaged from #106118).
2026-09-09 10:28:52 -07:00
Xipong 51fa04d49d fix(desktop): select installed WSL D3D12 driver before Electron exec 2026-09-09 10:28:52 -07:00
teknium1 a785db3672 fix(cli): inline the $HERMES_HOME/npmrc lookup, trim tests to two, document it
Fold the helper into _npm_lifecycle_env itself: the whole fix is one
is_file() check plus a setdefault, so a separate function, the
try/except around get_hermes_home() (it never raises) and the
dict-returning indirection were shape-gate violations. Build the path
with os.fspath so it is correct on Windows too.

Tests: keep the two invariants (file present -> NPM_CONFIG_USERCONFIG
points at it; explicit process/caller value wins and a missing file
sets nothing), fold the other two into the negative test.

Docs: one paragraph in the desktop troubleshooting page next to
ELECTRON_MIRROR describing $HERMES_HOME/npmrc.

Refs #106373
2026-09-09 10:26:06 -07:00
Konstantin Khlopkov ce4a33a9f7 fix(cli): keep $HERMES_HOME/npmrc npm config across updates (#106373)
(cherry picked from commit 339368f2215ffcfd04acf295c81806bf1e850770)
2026-09-09 10:26:06 -07:00
teknium1 d3dcc064df fix(cli): detect ssl.SSLError by type in the Codex login hint; trim tests; add the openssl.cnf snippet to docs
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
  __cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
  SSLEOFError whose text httpx did not repeat still gets the hint. The
  hint now names the TLS 1.2 diagnostic and links the providers docs
  note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
  the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
  classic-groups snippet (EN + existing zh-Hans copy).

Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
2026-09-09 10:14:58 -07:00
liuhao1024 d8eb177c93 fix(cli): keep SSL detail and add middlebox hint on Codex device-login transport errors
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.

- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
  (device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
  workaround hint; KeyboardInterrupt handling is unchanged

(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
2026-09-09 10:14:58 -07:00
Teknium 00bcef9b5f test(kanban): install the reviewer profile the review-surface fixture hands off to
kanban_request_review now rejects reviewers that are not installed profiles
(#106163); the cross-surface lifecycle test used a bare "reviewer" name with no
profile behind it, which is exactly the phantom the guard exists to catch.
2026-09-09 09:45:13 -07:00
Teknium e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
teknium1 02005cfe20 fix(kanban): promote refuses undone parents instead of a false --force success
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.

The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.

- drop `--force` from `promote` (parser, CLI handler, `promote_task`
  kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
  the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
  fake `ready`; the flag no longer parses

Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
2026-09-09 09:21:29 -07:00
Totoro-qaq 5d6d5fb223 fix(cli): refresh TERMINAL_CWD when --in re-homes the session
`--in DIR` only chdir'd. Every cwd consumer (resolve_agent_cwd -> Codex
app-server thread cwd, the terminal tool, context-file discovery) prefers
TERMINAL_CWD over the process cwd, so a value inherited from a parent
Hermes surface, the shell or .env survived the chdir and the session kept
running in the old directory. The local backend was rescued by cli.py's
force-export at import time; docker/ssh backends and the TUI launch path,
which never imports cli.py, were not.

Refresh TERMINAL_CWD to the --in target when it is already set. An unset
variable stays unset so the backends keep deriving from the new process
cwd and no host path is pre-seeded into ssh/container backends.

Fixes #106220
2026-09-09 09:20:12 -07:00
Teknium b8b2278440 test: import pytest in test_stderr_timestamp (marker needs it)
The spawns_gateway_lookalike marker was added to a module that never imported
pytest; collection failed with NameError in CI.
2026-09-09 09:19:54 -07:00
Teknium ca16cafee4 test(guard): block spawning a real gateway runtime from tests
Tests that exercise the dashboard's gateway-restart path can end up
spawning a REAL `python -m hermes_cli.main gateway restart` child when
the spawn seam is not intercepted. `_spawn_hermes_action` launches it
with start_new_session=True, so it outlives the pytest worker; the
child inherits the pytest-tmp HERMES_HOME, which is not a profile and
hashes to no service suffix, so `get_service_name()` resolves the
DEVELOPER's `hermes-gateway` unit, `systemd_restart` restarts the live
gateway, and without systemd the fallback runs `run_gateway()`
in-process forever and squats the webhook port.

Live repro on this machine (origin/main): an unintercepted spawn of
["gateway", "restart"] from a test restarted the production gateway
(MainPID 136820 -> 1689090, NRestarts=1). On 2026-09-03 a sibling
refactor moved `_spawn_hermes_action` from the `hermes_cli.web_server`
facade to `hermes_cli.web_server_gateway` ~10 minutes before the tests
were repointed; runs in that window patched a name production never
read and left 39 orphans alive for six days.

The live-system guard now rejects any subprocess primitive whose
command line the canonical matcher (`gateway.status.
_gateway_command_subcommand`) classifies as `gateway run|start|restart`.
Argv substrings are never consulted, so `gateway status`, `gateway
--help`, `hermes_cli.main serve`, etc. pass through. Three files that
deliberately spawn and reap a stub child with a gateway-shaped argv
(flock holders, sleep sleepers with an argv tail) opt out with the new
`spawns_gateway_lookalike` marker, which lifts only this check and keeps
os.kill guarded. Two canary tests pin the block and the pass-through.
2026-09-09 09:19:54 -07:00
Teknium 19f2f19987 fix(gateway): refuse to uninstall a systemd unit pinned to another HERMES_HOME
Defence at the exact boundary the incident crossed: systemd_uninstall() and
uninstall._remove_systemd_gateway() unlinked whatever get_systemd_unit_path()
returned. Before stop/disable/unlink, read the unit's own
Environment="HERMES_HOME=..." line (the parser status/refresh already use)
and, when it names a different home than this process, warn with both paths
and leave the unit alone. A unit without the line (hand-written) is still
removed as before.
2026-09-09 09:19:36 -07:00
Teknium 4746e34448 fix(gateway): foreign HERMES_HOME no longer resolves to the default hermes-gateway unit
_profile_suffix() compared HERMES_HOME against get_default_hermes_root(),
which treats ANY home outside ~/.hermes (Docker /opt/data, a mktemp dir) as
"the root itself". Every such home therefore collapsed to the bare
`hermes-gateway` service name and the default profile's unit path
(~/.config/systemd/user/hermes-gateway.service); the documented
"else a short hash of the path" branch was unreachable.

A parity harness run with HERMES_HOME=$(mktemp -d) called
uninstall_gateway_service(), resolved to the production unit, ran
`systemctl --user stop/disable`, unlinked it and daemon-reloaded. With the
unit gone Restart= could not revive it: all cron jobs and every messaging
platform were down for 6.5 days.

Compare against the platform-native default home (~/.hermes) for the bare
name; keep the profile name for <root>/profiles/<name>; everything else
(temp dirs, Docker /opt/data) gets its sha256[:8] suffix as the docstring
always promised. The Docker image supervises with s6 (`gateway-<profile>`
slots), not systemd/launchd, so the bare host-service name was never load-
bearing there.
2026-09-09 09:19:36 -07:00
kshitijk4poor e2bd400233 fix(models): only cache unreachability, not HTTP errors; key the entry via base_url_origin
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.

_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
2026-09-09 21:16:48 +05:30
kshitijk4poor 20762c67eb test(models): trim probe negative-cache tests to two invariants 2026-09-09 21:16:48 +05:30
finn763 48b8528e7c fix(desktop): stop UI freeze on unreachable provider Closes #81123 2026-09-09 21:16:48 +05:30
kshitijk4poor a4113eb994 test(dashboard): keep the read-coalescing tests to the two admission invariants
The contributor's 15-test module pinned encoding details (frozen-arg shapes,
postponed-annotation resolution, monkeypatch seams). The two behaviour
contracts that matter survive: a 12-request /api/profiles burst admits ONE
worker and leaves /api/status responsive on a 2-token pool (this file), and
the kanban board read shares one worker per key (test_kanban_read_admission).
Both go red when the coalescing wrapper is removed from the route.
2026-09-09 21:16:40 +05:30
Gianpietro Dal Zio 653418d842 fix(dashboard): coalesce expensive reads before worker admission 2026-09-09 21:16:40 +05:30
kshitijk4poor 441157f184 test(models): two invariants for the OpenRouter curated-catalog disk cache 2026-09-09 21:16:28 +05:30
finn763 a74e76632c perf(picker): read Nous pricing cache-only when building the picker row
_nous_picker_model_ids only uses the ids the Portal unions append — both
unions discard the pricing map (`model_ids, _ = union_with_portal_*`) — yet
it called get_pricing_for_provider("nous") without cached_only, so a cold
pricing cache paid a full /v1/models round-trip (network timeout on a slow
Portal) on the picker-open path for nothing. Pass cached_only=True; the
background pricing prewarm (#101685) fills the same cache for later opens.

Re-derived from #102099 by @finn763: the original patched
hermes_cli/model_switch.py, which 3b1ecfc0a1 decomposed; the live call site is
hermes_cli/model_switch_providers.py.

Based on #102099 by @finn763.
2026-09-09 21:12:22 +05:30
kshitijk4poor b36f654271 test(profiles): pin projects/tree single-flight as a behaviour contract
Replace the attribute-presence check (asserts a cache_clear attribute the decorator does not
expose) with one invariant: eight concurrent projects/tree requests run the profile scan once
and each caller gets its own copy of the payload. The scope-suite fixture now disables the TTL
instead of calling a non-existent clear hook, which silently no-op'd.
2026-09-09 21:11:56 +05:30
moken627-hub 35af06c009 fix(desktop): coalesce projects/tree sidebar scans and memoize raw config parses
The all-profiles sidebar polls GET /api/profiles/projects/tree, the one
heavy sidebar endpoint that was not wrapped in @_sidebar_singleflight_cache.
Every poll fanned out list_profiles() + _build_project_tree() over every
profile (51 on this box, ~17k SKILL.md files walked), and each profile
resolution re-parsed its config.yaml because read_user_config_raw() ran
uncached. On a 2-vCore VPS running 'hermes serve' for the Desktop remote
backend this pinned both cores (py-spy: 110-116% sustained, HostHighCPU).

- Wrap get_profiles_projects_tree in the existing single-flight cache,
  matching get_profiles_sessions_sidebar (5s TTL, errors[] not cached).
- Memoize read_user_config_raw() on a (st_dev, st_ino, st_size,
  st_mtime_ns) fingerprint under _CONFIG_LOCK, same strategy as
  read_raw_config(). Hits return a deepcopy so write-back round-trips keep
  their fresh-dict semantics; parse errors are never cached; the docstring
  no longer claims 'no caching'.
- Tests: memo semantics (deepcopy isolation, inode-replace reparse, error
  non-caching, parse-count), a wrapper-presence pin for both heavy sidebar
  endpoints, and a cold-cache autouse fixture in the scope tests (the 5s
  TTL otherwise leaks one test's payload into the next).

Measured on the affected box: list_profiles 1.9s -> 0.16s warm, serve CPU
112% -> 22-34%, host CPU 50-70% -> ~26%.
2026-09-09 21:11:56 +05:30
kshitijk4poor ead7e91dab refactor(recovery): stream the salvaged population; table the shape rules
- Pass 1 no longer materialises every classified record (full
  `messages.content` included) until pass 2; `LayoutEvidence` keeps only the
  capped per-position value sets (+ sessions rows for the one cross-column
  invariant) and pass 2 re-streams the lost_and_found tables. A 276 MB
  corrupted store no longer has to fit in memory.
- `_sentinel_holds` / `_text_shape_holds` if-ladders become rule tables.
- Tests trimmed to the three that bind behaviour (upgraded store maps by
  name; verifier refuses when rows matched no layout; replayed history ends
  at the current schema — the drift guard). No behaviour change; reverting
  inference to "no layout" still fails the name-mapping test.
2026-09-09 18:28:57 +05:30
kshitijk4poor 17c43aba06 fix(recovery): infer the salvaged store's physical layout from its schema history
A store's physical column order depends on which schema it was created at and
which ALTER TABLE ADD COLUMNs it lived through; one hardcoded "upgraded"
layout cannot cover them. `session_schema_history` records the declared
schema of the salvaged tables over time; `reachable_physical_layouts`
replays it to enumerate every physical order a store can have. The mapper
infers the layout once per kind from the whole recovered population (one
store wrote all of them), maps cells by name, and counts records whose width
matched no layout; the recovery verifier refuses to report such a salvage as
healthy.

Rebased onto the simplified session modules (c88d60551e, b9b4600cb2,
7e5a1a11d9, 1915a0d27e); no behaviour change from the pre-rebase branch.
2026-09-09 18:28:57 +05:30
Justin Wilson 7704712168 fix(sessions): map lost_and_found cells by physical column names
Upgraded state.db files gain columns via ALTER TABLE ADD COLUMN, so
physical order diverges from SCHEMA_SQL. Prefix-mapping onto the fresh
template put started_at at 0 and shifted titles/models. Insert by name
using known physical layouts; keep the plausibility gate.

Fixes #101409
2026-09-09 18:28:57 +05:30
Teknium f4af457679 test(plugins): trim catalog tests to invariants against the new shapes
538+570+296+216 lines of change-detectors → 4 files of contract tests:
seed catalog valid, bad entries skipped, kill list name-or-repo, live
fallback+union; real-git pinned install + sidecar + re-pin; kill list
blocks CLI/dashboard/TUI with only the CLI bypass; dashboard merge via
sidecar; probe get_config default; extractor live document.
2026-09-09 04:38:01 -07:00
Teknium 4e312cf22d feat(plugins): requires_hermes manifest gate on main's manifest/loader siblings
Re-port of the PR's version gate onto the decomposed layout: the field and
parser live in plugins_manifest.py (with running_hermes_version /
version_satisfies helpers), the load-time skip in plugins_loader.py
before any import. Unsatisfied plugins record an error and never run
register(); one invariant test proves both halves.
2026-09-09 04:38:01 -07:00
Teknium bcff0a920e test(cron): replace fail-closed drift tests with snapshot-as-pin invariants
The old tests encoded the skip (agent never constructed, [drift_skip] text,
alert-once bit). New invariants, proven red on origin/main: an unpinned job
with model_snapshot=A / provider_snapshot=P runs with AIAgent(model=A) and
resolve_runtime_provider(requested=P) after the global default moved to B/Q;
an explicit job pin still beats the snapshot; cron.model / cron.model_provider
still beat the snapshot; a legacy job without a snapshot follows the global
default. test_cron_drift_alert_once.py is deleted with the bit it tested; the
config-notice and impact-summary tests drop guard_enabled / model_drift_guard;
the desktop toast test asserts the informational copy.
2026-09-09 04:32:13 -07:00
Teknium 48465c3933 fix(profiles): --clone-all no longer copies cron jobs into the new profile
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.

Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.

--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
2026-09-09 04:27:24 -07:00
Teknium d47adec28f Merge origin/main into feat/plugin-catalog
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
2026-09-09 04:15:27 -07:00
buihongduc132 c1ff9390f6 fix(backup): prune old hermes-backup-*.zip after each run, keep last 3
run_backup() previously wrote "hermes-backup-<timestamp>.zip" on every
invocation without deleting old ones. Hourly callers accumulated 157 zips
(14 GiB). Add _prune_run_backups() to keep the newest N (default 3,
configurable via backup.run_backup_keep or --keep CLI flag).
2026-09-09 03:33:14 -07:00
Teknium 734461d213 fix(models): same-URL custom endpoints stop evicting each other's cached catalog; no-probe picker opens revalidate
Two picker-freshness defects in cached_fetch_api_models():

1. The disk cache row was keyed on base_url only, with the credential
   fingerprint stored inside the row. N custom_providers entries sharing
   one proxy URL with different keys (#106184) took turns overwriting the
   single slot; every sibling then failed the fingerprint check, got an
   empty catalog, and disappeared from the Desktop pickers (which hide
   zero-model rows). Key on url#fingerprint so each credential owns a row.

2. cache_only opens (Desktop model.options without refresh) served a
   past-TTL row for up to 7 days without ever revalidating, so a model
   loaded on a non-current local endpoint stayed invisible until the user
   found "Refresh Models". Serve the stale row AND spawn the same
   off-thread SWR refresh the blocking path uses; the caller still never
   waits on the network.

Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
  GUI no-probe open  before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
                     after  {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
2026-09-09 03:33:06 -07:00
briandevans f4c55323fa fix(cli): back up config.yaml before --reset overwrites it
`hermes setup --reset` calls `save_config(copy.deepcopy(DEFAULT_CONFIG))`,
which writes `get_hermes_home()/config.yaml` — the exact file the backup
block a few lines below copies to `config.yaml.bak.<timestamp>`. Because the
copy ran after the reset, the backup captured the defaults that had just been
written, not the user's config. The one invocation where a backup matters most
produced a worthless one, and the original was unrecoverable.

The block's own comment already claimed it runs "before setup modifies it";
on the --reset path that was false. Move it above the --reset branch so it
captures the true pre-setup state on every path.

Also report the backup location on the --reset path. --reset is destructive
and can leave the wizard early (the non-interactive return exits before the
end-of-setup notice), so a user who just lost their config was never told
where the copy is. The end-of-setup notice is unchanged for the normal path
and is suppressed only when it has already been shown, so no run prints it
twice; the shared wording now lives in one helper.

Behaviour otherwise preserved: `copy2` (config.yaml holds secrets, so mode is
preserved), the try/except fallback to `_backup_path = None`, and the existing
notice for the full-setup path.

Follow-ups deliberately out of scope: pruning accumulated `.bak.*` files, and
printing the notice on the other early-return paths (--portal, section runs).

Refs #3522
2026-09-09 03:32:58 -07:00
Teknium bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium 677e8ed8a4 fix(desktop): SSH remote backend stops following the host's sticky active_profile
A Desktop-owned `hermes serve --isolated --ssh-session-token-file ...` child
is spawned with an explicit `--profile <name>` when the connection names a
remote profile, and with no flag for the remote root home. Without the flag,
`_apply_profile_override` read the remote host's sticky `active_profile`
file and re-homed the backend into whatever profile the user last selected
on that machine's CLI. Settings then read one config.yaml while the remote
gateway wrote another, so model picks and toggles "didn't stick".

Treat the SSH token flag as a fixed-identity marker, the same way
supervisor-launched gateway children are (#74872): a Desktop backend's
profile is chosen by the client, never by the host.

Live repro (before/after, temp HERMES_HOME with active_profile=foo):
  serve --isolated --ssh-session-token-file ...   hermes_home=<root>/profiles/foo -> <root>
  same + --profile foo                             hermes_home=<root>/profiles/foo (unchanged)
  serve (no token file, user CLI)                  hermes_home=<root>/profiles/foo (unchanged)
2026-09-09 02:34:59 -07:00
kshitijk4poor 32273b8118 refactor(cli): one stored_session_route for interactive and one-shot resume
_apply_stored_session_runtime was a line-for-line copy of the first half of
_restore_session_model (stored-model guard, session_gateway_runtime, bare-custom heal,
model/provider-changed check). Extract that pure decision into
cli_model_switch_mixin.stored_session_route and have both resume paths call it; the
one-shot keeps only the _ModelChoice mapping and the drop-ambient-key rule.

main.py stops re-normalising `resume` — _resolve_chat_session_args already did.
Tests trimmed from 20 to 13: near-duplicate unit tests of the private helpers go, the
end-to-end _run_agent contracts (stored runtime + reopen; explicit --model wins) and the
empty-session-keeps-id case stay.
2026-09-09 12:41:07 +05:30
liuhao1024 8aa773af89 fix(cli): restore stored session runtime and reopen ended rows on oneshot resume
Review fixes (#105957):

- A resumed one-shot ignored the session's stored model/provider runtime:
  _resolve_model_and_provider()/resolve_runtime_provider() ran before
  _load_resume_target(), which only loaded the session id + transcript, so an
  ambient config (e.g. openrouter/ambient-model) served the resumed transcript
  instead of the stored route (custom:stored/stored-model). The stored runtime
  is now applied before runtime resolution, with the same contract as the
  interactive _restore_session_model(): stored model/provider/base_url/api_mode
  replace the ambient choice unless --model was passed explicitly, and a
  changed provider drops the ambient api_key so resolution re-fetches
  credentials for the restored endpoint.

- Passing the resumed id to AIAgent did not reopen the already-ended session
  row: end_session() only writes rows whose ended_at is null and the
  existing-row upsert never clears the end fields, so the resumed turn was
  recorded under a session that stayed closed and its new lifecycle boundary
  was lost. _load_resume_target() now reopens the row (best effort), same as
  the interactive resume does before continuing.
2026-09-09 12:41:07 +05:30
liuhao1024 5ff6cb0edb fix(cli): keep the resolved session id when a resumed oneshot session is empty
Review finding on #105957: `_load_resume_target` returned None for a
resolved session with no stored messages, so `hermes -z "hello" -c <title>
--create-if-missing` recorded the turn under a freshly minted session id and
the just-created titled session stayed empty. Preserve `resolved` unconditionally — the interactive /resume path keeps the selected id for an
empty session too; only the history replay is empty. Regression tests pin the
durable id for both a plain empty session and an empty compression-chain head.
2026-09-09 12:41:07 +05:30
liuhao1024 86d606ca34 fix(cli): honor --resume in one-shot mode (#105892)
The -z exit path accepted --resume/-c in the parser but never forwarded
args.resume: every resumed one-shot turn silently started a fresh session,
so each wire request carried only [system, current user] and the model
lost all prior context (reported against Ollama/custom OpenAI-compatible
endpoints, but provider-independent).

Normalize session args (latest/title/--continue/--in + cwd restore) via
the chat path's _resolve_chat_session_args before the oneshot exit path
takes over, then load the resumed transcript in _run_agent through the
same contract the interactive CLI uses (compression-chain redirect,
safe-resume guard, session_meta filtering) and continue the existing
session id instead of creating a new one. An explicit --resume of an
unknown session now fails loudly instead of starting fresh.
2026-09-09 12:41:07 +05:30
kshitijk4poor a73b750391 fix(cron): the dashboard "Trigger" run-now no longer stamps the next occurrence either
Second entry of the same bug class: POST /api/cron/jobs/{id}/trigger →
_fire_cron_job_for_profile → CronScheduler.fire_due → claim_fire built its claim
without `manual`, so an off-tick run from the web UI stamped the future slot exactly
like the tools path #105704 fixes. fire_due/claim_fire gain `manual` (forwarded only
when set, mirroring `force`, so third-party providers keep working) and the dashboard
trigger passes it when the provider's signature accepts it. Webhook and misfire
catch-up fires run the slot that is due and keep the stamp.

Also drops the base-green tick-stamp test (the same contract is pinned by
tests/cron/test_scheduled_occurrence.py) and documents `manual` vs `force`.
2026-09-09 12:17:13 +05:30
Ben Barclay 5a1246f830 fix(observability): attribute ACP and batch execution surfaces
Fleet telemetry showed "unknown" as the single largest execution_surface
bucket. Two construction paths were mis-attributed, both silently:

1. ACP editor sessions (VS Code / Zed / JetBrains) declare platform="acp",
   but "acp" was absent from EXECUTION_SURFACES, so the contract's
   closed-schema fallback folded every editor session into "other" --
   the bucket meant for genuinely unclassifiable traffic.

2. batch_runner built agents from _AGENT_PASSTHROUGH, which omitted
   "platform" entirely, so every batch task run reported "unknown"
   despite "batch" already being a first-class surface.

Neither is a reporting bug in the exporter: both are declaration gaps at
the construction site. "unknown" must mean "this run genuinely could not
be attributed", not "a construction site forgot to say who it was".

Changes:
- add "acp" to EXECUTION_SURFACES and map it to the "interactive"
  entrypoint alongside cli/desktop/tui
- add "acp" to the v2 wire schema enum (kept in sync by an existing test)
- pass platform through batch_runner: added to _AGENT_PASSTHROUGH, set
  self.platform = "batch" on the runner, and defaulted at the worker call
  site so callers that build a config without it stay attributable

Wire compatibility: the ingest service validates the envelope only and
stores metric bodies verbatim, so packages carrying the new value are
accepted by the already-deployed server. No coordinated deploy needed.

Tests: 12 new behavioural tests. Verified red before the fix (4 failed),
green after. Three fix-mutants confirmed killed:
  M1 revert acp from EXECUTION_SURFACES  -> 3 failed
  M2 revert acp entrypoint mapping only  -> 1 failed
  M3 revert batch passthrough            -> 1 failed
No source-text assertions; every test is a contract between the surfaces
the schema accepts and the surface each path declares. A guard test pins
that a genuinely undeclared run still reports "unknown", so attribution
cannot be "fixed" by inventing a default that hides real gaps.
2026-09-09 11:27:36 +10:00
Teknium 9b0d75ce44 fix: gateway ordering test runs without root access 2026-09-08 13:25:54 -07:00
kshitijk4poor 6f11a3296e refactor(gateway): wait on the target user's bus socket, not the generic control-socket predicate
_user_systemd_socket_ready() accepts systemd/private alone, which is enough for
systemctl --user but not for the systemd-run --user that restart-safe workers
need; systemd_user_bus_env() requires the bus socket. Replace the uid threading
through five helpers with one _wait_for_target_user_bus(uid) that polls
/run/user/<uid>/bus, and move the post-enable wait + restart hint out of
_ensure_linger_enabled into _ensure_system_service_linger so the activity probe
runs only when linger was actually just enabled. Kanban applies the bus env
unconditionally like the cron sibling. Refs #104893.
2026-09-08 22:35:39 +05:30
kshitijk4poor 4c4845f0af fix(gateway): order the system unit after the target user's manager
run_gateway() adopts the user bus once at boot; the generated system unit had
no ordering against user@<uid>.service, so after a reboot the two race and
the adoption can miss until the next gateway restart. Emit After=/Wants=
user@<uid>.service for the unit's User= (uid now returned by
_system_service_identity, which already resolved the account). Existing
system units are flagged outdated once and refreshed on the next
install/restart. Refs #104893.
2026-09-08 22:35:39 +05:30
joaomarcos 748e68432b fix(gateway): provision linger for system-service users
A system unit's `User=` has no login session on a headless host, so
user@<uid>.service never starts and `systemd-run --user --scope` — every
restart-safe cron/Kanban worker — has no bus to reach. `hermes gateway
install --system` runs as root and already knows the target user, so enable
linger for that user on fresh install, on an already-current unit, and on
repair.

- `get_systemd_linger_status()` / `_ensure_linger_enabled()` take the target
  username; root is included (restart-safe workers cross `systemd-run --user`
  regardless of who the gateway runs as).
- After `loginctl enable-linger` succeeds, wait for the TARGET uid's control
  socket (`_wait_for_user_dbus_socket(uid=...)`) — logind starts the user
  manager asynchronously and `--start-now` boots the gateway immediately; the
  caller's own env (root's) says nothing about it and is never adopted.
- Messages and the manual-remediation hint are scope-aware (`sudo systemctl
  restart`, not `systemctl --user`); when the repaired service is already
  active, say that a restart is required — `systemctl start` on an active
  unit is a no-op and the running gateway keeps its bus-less environment.

Refs #104893.
2026-09-08 22:35:39 +05:30
ethernet c8aa5608c2 Merge pull request #101420 from ethernet8023/ethie/desktop-update-tests
test(install): cross-OS install/update E2E matrix (windows, macos, linux)
2026-09-08 10:30:44 -04:00
Teknium 4cd4f395ea test: isolate missing first-party import guard fixture 2026-09-08 03:06:30 -07:00