Widens #106272 to the one remaining sibling: `launchd_stop()` boots out with check=True and
already handles exit 3/113/125 (job unloaded) and 5/125 (domain unmanageable) by falling through
to the PID kill, yet inherited stderr — so `hermes gateway stop` against an unloaded job printed
"Boot-out failed: 3: No such process" next to "✓ Service stopped". Same `_CAPTURE_TEXT` kwargs as
the sibling calls; an unexpected exit still raises with `e.stderr` populated.
Replaces the contributor's two kwarg-assertion tests (`capture_output is True` on a mocked
`subprocess.run`) with two invariant tests that run a real fake `launchctl` on PATH and read fd 2
through `capfd`: restart-on-unloaded prints only its ↻/✓ lines and drives
kickstart→bootout→bootstrap→kickstart; stop-on-unloaded is silent, while a real bootout failure
(exit 1) still raises with the captured stderr. Both red on origin/main, green here.
Best-effort bootout calls (unloaded-job recovery, stale-EIO retry,
plist refresh, uninstall) and the handled kickstart -k in
launchd_restart inherited the terminal's stderr, so an expected
unloaded job printed raw launchctl errors around the CLI's own lines:
Could not find service "ai.hermes.gateway" in domain for user gui: 501
↻ launchd job was unloaded; reloading
Boot-out failed: 3: No such process
Capture them with _CAPTURE_TEXT instead. The kickstart error stays
available as e.stderr for the update_cmd failure diagnostic, and the
post-bootstrap kickstart intentionally stays loud (its failure feeds
the domain-unsupported fallback). Same precedent as the reload
helper, which already runs bootout with 2>/dev/null.
[salvage: picked hermes_cli/gateway.py only; the two capture_output kwarg-assertion tests are
replaced by fd-level invariant tests with a fake launchctl in the follow-up commit]
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.
- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
the verdict of the one canonical classifier
(`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
`staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
a global-scope address (`2607:f8b0::1`) is not local while `::1`,
ULA and link-local still are via the `ipaddress` scope checks.
Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.
Refs #106228
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Move the WSL / /dev/dxg / d3d12_dri.so probes into _prefer_wsl_d3d12 with
the probed paths as module constants, so the launcher call site is a single
line and a test can lay out a fake WSLg host without touching real
/dev or /usr/lib. The two tests now run the real _desktop_launch_env end to
end (selected under WSL+dxg+driver; untouched with an explicit Mesa
override, off WSL, without /dev/dxg, or without the driver file) instead of
unit-testing the helper with a precomputed boolean.
Docs: one paragraph in the Desktop guide on the automatic selection and
the env vars that keep an explicit choice authoritative.
Follow-up to Xipong's fix for #106117 (salvaged from #106118).
Fold the helper into _npm_lifecycle_env itself: the whole fix is one
is_file() check plus a setdefault, so a separate function, the
try/except around get_hermes_home() (it never raises) and the
dict-returning indirection were shape-gate violations. Build the path
with os.fspath so it is correct on Windows too.
Tests: keep the two invariants (file present -> NPM_CONFIG_USERCONFIG
points at it; explicit process/caller value wins and a missing file
sets nothing), fold the other two into the negative test.
Docs: one paragraph in the desktop troubleshooting page next to
ELECTRON_MIRROR describing $HERMES_HOME/npmrc.
Refs #106373
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
__cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
SSLEOFError whose text httpx did not repeat still gets the hint. The
hint now names the TLS 1.2 diagnostic and links the providers docs
note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
classic-groups snippet (EN + existing zh-Hans copy).
Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.
- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
(device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
workaround hint; KeyboardInterrupt handling is unchanged
(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.
The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.
- drop `--force` from `promote` (parser, CLI handler, `promote_task`
kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
fake `ready`; the flag no longer parses
Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
`--in DIR` only chdir'd. Every cwd consumer (resolve_agent_cwd -> Codex
app-server thread cwd, the terminal tool, context-file discovery) prefers
TERMINAL_CWD over the process cwd, so a value inherited from a parent
Hermes surface, the shell or .env survived the chdir and the session kept
running in the old directory. The local backend was rescued by cli.py's
force-export at import time; docker/ssh backends and the TUI launch path,
which never imports cli.py, were not.
Refresh TERMINAL_CWD to the --in target when it is already set. An unset
variable stays unset so the backends keep deriving from the new process
cwd and no host path is pre-seeded into ssh/container backends.
Fixes#106220
Slim redo of the mechanism from #106410 on top of its pick (no wrappers, no
persisted "kind" enum, no process-local flag that dies with the process):
- Futility = the SAME holder PID set has blocked >= _FTS_HOLDER_FUTILE_ATTEMPTS
(10) deferrals over >= _FTS_HOLDER_FUTILE_SECONDS (30 min); tracked as
holders_since/holders_attempts in the persisted fts_rebuild_deferral record
and reset whenever the holder set changes. The 3-deferral/60 s escalate
window is the orphan-reap gate and stays as is.
- ONE escalated ERROR line names each holder pid + cmdline and the remedy that
can actually be followed from inside a gateway session: stop ONLY the other
holder; this process's own retry admits the rebuild within 60 s. The old
"with the gateway stopped" advice was unrunnable from a gateway-hosted
session (the gateway is the session) and is gone from both log and doctor.
- hermes doctor renders the futile record distinctly.
- retry_deferred_fts_recovery: a capped backoff earned by holder set X no
longer applies once the live holder set differs from X, so stopping the
other service is followed by a retry on the next tick, not up to an hour
later (the issue's 16-min wait).
- Tests trimmed from 5 to 2 invariants (futile line + doctor entry after N
same-holder deferrals; backoff reset when the holder set changes); the
contributor's control tests for changing PIDs / orphan reap are covered by
the existing test_repeated_deferrals_reap_inactive_orphan_then_rebuild.
The "canonical writes and LIKE search remain available" WARNING is kept
because it is true on origin/main: a stale open drops every FTS trigger, so
the messages INSERT succeeds (probed live with a real state.db + a second
process holding it). Writes fail only when a peer re-publishes triggers over
the corrupt index — a separate class, not this diagnostic.
Refs #106393
A supervised peer never satisfies the orphan reap, so stale-FTS repair retried forever with a misleading "canonical writes remain available" warning.
(cherry picked from commit e57f3a975d311aa44da1e92c5e727eba7c8cff70)
Defence at the exact boundary the incident crossed: systemd_uninstall() and
uninstall._remove_systemd_gateway() unlinked whatever get_systemd_unit_path()
returned. Before stop/disable/unlink, read the unit's own
Environment="HERMES_HOME=..." line (the parser status/refresh already use)
and, when it names a different home than this process, warn with both paths
and leave the unit alone. A unit without the line (hand-written) is still
removed as before.
With the previous commit a Docker/custom root (HERMES_HOME=/opt/data) gets a
hashed host-service suffix. Three callers used `_profile_suffix() or
"default"` as the PROFILE id, which is a different question: the s6
supervisor's slot for the root home is `gateway-default` regardless of where
the root lives, and the multiplexer's "am I a named profile" probe must not
treat a hash as a profile name. Route them through hermes_constants.
profile_name_for_home() (root -> "default", <root>/profiles/<name> -> name)
with the service suffix as the fallback for unknown layouts.
_profile_suffix() compared HERMES_HOME against get_default_hermes_root(),
which treats ANY home outside ~/.hermes (Docker /opt/data, a mktemp dir) as
"the root itself". Every such home therefore collapsed to the bare
`hermes-gateway` service name and the default profile's unit path
(~/.config/systemd/user/hermes-gateway.service); the documented
"else a short hash of the path" branch was unreachable.
A parity harness run with HERMES_HOME=$(mktemp -d) called
uninstall_gateway_service(), resolved to the production unit, ran
`systemctl --user stop/disable`, unlinked it and daemon-reloaded. With the
unit gone Restart= could not revive it: all cron jobs and every messaging
platform were down for 6.5 days.
Compare against the platform-native default home (~/.hermes) for the bare
name; keep the profile name for <root>/profiles/<name>; everything else
(temp dirs, Docker /opt/data) gets its sha256[:8] suffix as the docstring
always promised. The Docker image supervises with s6 (`gateway-<profile>`
slots), not systemd/launchd, so the bare host-service name was never load-
bearing there.
_size duplicated _mtime_ns with a different attribute and the fingerprint
stat'ed each of the four files twice, so mtime and size could come from two
different versions of the file. _stat_sig returns both from one stat. The
persisted compare goes through a JSON round-trip so nested tuples match the
lists they read back as.
`_heal_forked_single_use_oauth_grants()` runs on every `load_pool()`, and its
clean mark lived only in memory. Every fresh `hermes` invocation and every new
worker therefore re-took the heal's two nested EXCLUSIVE auth-store locks just
to rediscover a store it had already cleared — 4 acquisitions per process on a
two-provider profile, every one of them finding nothing to consolidate. Behind
a sibling process holding those locks that costs a full
`AUTH_LOCK_TIMEOUT_SECONDS` per provider before the process can do anything at
all: measured 30.1s for two providers.
Persist the mark next to the store it describes (`<profile>/cache/
oauth_heal_clean.json`, 0600, paths and stat data only — no credential
material) and consult it BEFORE taking any lock.
Outliving the process means the mark needs a stronger key than the in-memory
one did:
- The ROOT store joins the fingerprint. This heal consolidates root → profile,
so root acquiring a counterpart turns a row the heal deliberately KEPT into a
fork it must strip. The in-memory mark could ignore root because it died with
the process; a persisted mark would keep skipping a heal that has become
necessary.
- File sizes join it too, so a metadata-preserving rewrite (`rsync -t`,
`tar -p`, a restore) cannot leave a stale mark looking current indefinitely
rather than for one process.
Measured on an isolated HERMES_HOME with two OAuth providers:
lock acquisitions per fresh process 4 -> 0
contended load_pool() x2 30.1s -> 0.00s
mark-file reads per 100 load_pool() - -> 2 (one per provider)
The mark stays a cache: absent, unreadable, corrupt or wrong-shaped content all
mean "unknown" and fall through to the locked heal, and a failed write only
means the next process re-runs it — the behaviour before this cache existed.
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.
_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
refresh_interval_seconds() honours model_catalog.ttl_minutes / legacy ttl_hours;
reading DEFAULT_TTL_MINUTES would let the snapshot and the manifest it is
filtered from go stale on different clocks for anyone who changed the TTL.
Re-derived from PR #96099 (f127ec4e) on current main; the :nitro/:floor validate hunk is omitted because main already handles routing suffixes in hermes_cli/models_validate.py.
_nous_picker_model_ids only uses the ids the Portal unions append — both
unions discard the pricing map (`model_ids, _ = union_with_portal_*`) — yet
it called get_pricing_for_provider("nous") without cached_only, so a cold
pricing cache paid a full /v1/models round-trip (network timeout on a slow
Portal) on the picker-open path for nothing. Pass cached_only=True; the
background pricing prewarm (#101685) fills the same cache for later opens.
Re-derived from #102099 by @finn763: the original patched
hermes_cli/model_switch.py, which 3b1ecfc0a1 decomposed; the live call site is
hermes_cli/model_switch_providers.py.
Based on #102099 by @finn763.
The all-profiles sidebar polls GET /api/profiles/projects/tree, the one
heavy sidebar endpoint that was not wrapped in @_sidebar_singleflight_cache.
Every poll fanned out list_profiles() + _build_project_tree() over every
profile (51 on this box, ~17k SKILL.md files walked), and each profile
resolution re-parsed its config.yaml because read_user_config_raw() ran
uncached. On a 2-vCore VPS running 'hermes serve' for the Desktop remote
backend this pinned both cores (py-spy: 110-116% sustained, HostHighCPU).
- Wrap get_profiles_projects_tree in the existing single-flight cache,
matching get_profiles_sessions_sidebar (5s TTL, errors[] not cached).
- Memoize read_user_config_raw() on a (st_dev, st_ino, st_size,
st_mtime_ns) fingerprint under _CONFIG_LOCK, same strategy as
read_raw_config(). Hits return a deepcopy so write-back round-trips keep
their fresh-dict semantics; parse errors are never cached; the docstring
no longer claims 'no caching'.
- Tests: memo semantics (deepcopy isolation, inode-replace reparse, error
non-caching, parse-count), a wrapper-presence pin for both heavy sidebar
endpoints, and a cold-cache autouse fixture in the scope tests (the 5s
TTL otherwise leaks one test's payload into the next).
Measured on the affected box: list_profiles 1.9s -> 0.16s warm, serve CPU
112% -> 22-34%, host CPU 50-70% -> ~26%.
- Pass 1 no longer materialises every classified record (full
`messages.content` included) until pass 2; `LayoutEvidence` keeps only the
capped per-position value sets (+ sessions rows for the one cross-column
invariant) and pass 2 re-streams the lost_and_found tables. A 276 MB
corrupted store no longer has to fit in memory.
- `_sentinel_holds` / `_text_shape_holds` if-ladders become rule tables.
- Tests trimmed to the three that bind behaviour (upgraded store maps by
name; verifier refuses when rows matched no layout; replayed history ends
at the current schema — the drift guard). No behaviour change; reverting
inference to "no layout" still fails the name-mapping test.
A store's physical column order depends on which schema it was created at and
which ALTER TABLE ADD COLUMNs it lived through; one hardcoded "upgraded"
layout cannot cover them. `session_schema_history` records the declared
schema of the salvaged tables over time; `reachable_physical_layouts`
replays it to enumerate every physical order a store can have. The mapper
infers the layout once per kind from the whole recovered population (one
store wrote all of them), maps cells by name, and counts records whose width
matched no layout; the recovery verifier refuses to report such a salvage as
healthy.
Rebased onto the simplified session modules (c88d60551e, b9b4600cb2,
7e5a1a11d9, 1915a0d27e); no behaviour change from the pre-rebase branch.
Upgraded state.db files gain columns via ALTER TABLE ADD COLUMN, so
physical order diverges from SCHEMA_SQL. Prefix-mapping onto the fresh
template put started_at at 0 and shifted titles/models. Insert by name
using known physical layouts; keep the plausibility gate.
Fixes#101409
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
.hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
update, and the dashboard/TUI payload builders. plugins_cmd.py only
gains the hooks (cmd_install catalog branch, cmd_update / dashboard
update re-pin, dashboard_install_plugin catalog_name + kill list,
dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
fallback) instead of the unauthenticated GitHub contents API (60 req/h,
1 request per entry); in-tree and live removals are unioned so a stale
cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
limits); _plugin_runtime_status shared from web_server_dashboard.py;
hub rows carry removed_reason. TUI plugins.manage gains catalog_name
install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
(real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
plugin-catalog/** so entry merges republish it.
Re-port of the PR's version gate onto the decomposed layout: the field and
parser live in plugins_manifest.py (with running_hermes_version /
version_satisfies helpers), the load-time skip in plugins_loader.py
before any import. Unsatisfied plugins record an error and never run
register(); one invariant test proves both halves.
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).
The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.
Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".
Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).
The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.
Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.
--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
run_backup() previously wrote "hermes-backup-<timestamp>.zip" on every
invocation without deleting old ones. Hourly callers accumulated 157 zips
(14 GiB). Add _prune_run_backups() to keep the newest N (default 3,
configurable via backup.run_backup_keep or --keep CLI flag).
httpx honours the env/system proxy (on Windows, the registry ProxyServer
even with no *_PROXY vars) but never the bypass list, so a system proxy
(Clash, corporate) answered 127.0.0.1 probes from both Desktop validators
with its own error page. That parsed as models=[] and the GUI said
"advertised no models at /v1/models" for a llama.cpp server the CLI
(urllib, honours <local>) saw fine.
Local endpoints (loopback, LAN, Tailscale via is_local_endpoint) now
probe with trust_env=False; public endpoints keep honouring env proxies.
A reachable endpoint answering non-2xx with no model list reports
"<url> answered HTTP <status>." instead of an empty catalog, so the
onboarding card stops telling the user to start a model.
Reimplemented on the decomposed router (the original patched
web_server.py before the split). Diagnosis and fix direction by
Solitud1nem in #63656; Windows registry-proxy confirmation by
Ulysses-Gaia on #63472.
Live repro (real loopback server, HTTP_PROXY=http://127.0.0.1:9):
before ok=False reachable=False 'Could not reach .../v1/models'
after ok=True models=['Qwen3.6-35B-A3B-Q5_K_M.gguf']
Co-authored-by: Solitud1nem <76743883+Solitud1nem@users.noreply.github.com>
Two picker-freshness defects in cached_fetch_api_models():
1. The disk cache row was keyed on base_url only, with the credential
fingerprint stored inside the row. N custom_providers entries sharing
one proxy URL with different keys (#106184) took turns overwriting the
single slot; every sibling then failed the fingerprint check, got an
empty catalog, and disappeared from the Desktop pickers (which hide
zero-model rows). Key on url#fingerprint so each credential owns a row.
2. cache_only opens (Desktop model.options without refresh) served a
past-TTL row for up to 7 days without ever revalidating, so a model
loaded on a non-current local endpoint stayed invisible until the user
found "Refresh Models". Serve the stale row AND spawn the same
off-thread SWR refresh the blocking path uses; the caller still never
waits on the network.
Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
GUI no-probe open before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
after {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
`hermes setup --reset` calls `save_config(copy.deepcopy(DEFAULT_CONFIG))`,
which writes `get_hermes_home()/config.yaml` — the exact file the backup
block a few lines below copies to `config.yaml.bak.<timestamp>`. Because the
copy ran after the reset, the backup captured the defaults that had just been
written, not the user's config. The one invocation where a backup matters most
produced a worthless one, and the original was unrecoverable.
The block's own comment already claimed it runs "before setup modifies it";
on the --reset path that was false. Move it above the --reset branch so it
captures the true pre-setup state on every path.
Also report the backup location on the --reset path. --reset is destructive
and can leave the wizard early (the non-interactive return exits before the
end-of-setup notice), so a user who just lost their config was never told
where the copy is. The end-of-setup notice is unchanged for the normal path
and is suppressed only when it has already been shown, so no run prints it
twice; the shared wording now lives in one helper.
Behaviour otherwise preserved: `copy2` (config.yaml holds secrets, so mode is
preserved), the try/except fallback to `_backup_path = None`, and the existing
notice for the full-setup path.
Follow-ups deliberately out of scope: pruning accumulated `.bak.*` files, and
printing the notice on the other early-return paths (--portal, section runs).
Refs #3522
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.
hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.
Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
A Desktop-owned `hermes serve --isolated --ssh-session-token-file ...` child
is spawned with an explicit `--profile <name>` when the connection names a
remote profile, and with no flag for the remote root home. Without the flag,
`_apply_profile_override` read the remote host's sticky `active_profile`
file and re-homed the backend into whatever profile the user last selected
on that machine's CLI. Settings then read one config.yaml while the remote
gateway wrote another, so model picks and toggles "didn't stick".
Treat the SSH token flag as a fixed-identity marker, the same way
supervisor-launched gateway children are (#74872): a Desktop backend's
profile is chosen by the client, never by the host.
Live repro (before/after, temp HERMES_HOME with active_profile=foo):
serve --isolated --ssh-session-token-file ... hermes_home=<root>/profiles/foo -> <root>
same + --profile foo hermes_home=<root>/profiles/foo (unchanged)
serve (no token file, user CLI) hermes_home=<root>/profiles/foo (unchanged)
With account-identity matching gone, the providers.<id> block consolidation
only fired on shared token material. A historical fork (same copied pool-row
id, profile rotated, both pairs diverged) then healed the pool row into root
but left root's providers.openai-codex block on the spent pair; root's next
load_pool() re-seeds its device_code row FROM that block and undid the heal.
_HealPass now records that a profile pool row matched root by copied id or
shared tokens and passes that verdict to _heal_forked_provider_block, which
accepts it as lineage proof. No account-identity guessing is restored; an
independent same-account grant (no id/token match) is still left alone.
Follow-up to simpolism's #106177.
CI: tests/cli/test_cli_resume_command.py builds bare HermesCLI objects without .model; the
refactor read self.model before the stored-model check the contributor's code made first.
_apply_stored_session_runtime was a line-for-line copy of the first half of
_restore_session_model (stored-model guard, session_gateway_runtime, bare-custom heal,
model/provider-changed check). Extract that pure decision into
cli_model_switch_mixin.stored_session_route and have both resume paths call it; the
one-shot keeps only the _ModelChoice mapping and the drop-ambient-key rule.
main.py stops re-normalising `resume` — _resolve_chat_session_args already did.
Tests trimmed from 20 to 13: near-duplicate unit tests of the private helpers go, the
end-to-end _run_agent contracts (stored runtime + reopen; explicit --model wins) and the
empty-session-keeps-id case stay.
Review fixes (#105957):
- A resumed one-shot ignored the session's stored model/provider runtime:
_resolve_model_and_provider()/resolve_runtime_provider() ran before
_load_resume_target(), which only loaded the session id + transcript, so an
ambient config (e.g. openrouter/ambient-model) served the resumed transcript
instead of the stored route (custom:stored/stored-model). The stored runtime
is now applied before runtime resolution, with the same contract as the
interactive _restore_session_model(): stored model/provider/base_url/api_mode
replace the ambient choice unless --model was passed explicitly, and a
changed provider drops the ambient api_key so resolution re-fetches
credentials for the restored endpoint.
- Passing the resumed id to AIAgent did not reopen the already-ended session
row: end_session() only writes rows whose ended_at is null and the
existing-row upsert never clears the end fields, so the resumed turn was
recorded under a session that stayed closed and its new lifecycle boundary
was lost. _load_resume_target() now reopens the row (best effort), same as
the interactive resume does before continuing.
Review finding on #105957: `_load_resume_target` returned None for a
resolved session with no stored messages, so `hermes -z "hello" -c <title>
--create-if-missing` recorded the turn under a freshly minted session id and
the just-created titled session stayed empty. Preserve `resolved` unconditionally — the interactive /resume path keeps the selected id for an
empty session too; only the history replay is empty. Regression tests pin the
durable id for both a plain empty session and an empty compression-chain head.
The -z exit path accepted --resume/-c in the parser but never forwarded
args.resume: every resumed one-shot turn silently started a fresh session,
so each wire request carried only [system, current user] and the model
lost all prior context (reported against Ollama/custom OpenAI-compatible
endpoints, but provider-independent).
Normalize session args (latest/title/--continue/--in + cwd restore) via
the chat path's _resolve_chat_session_args before the oneshot exit path
takes over, then load the resumed transcript in _run_agent through the
same contract the interactive CLI uses (compression-chain redirect,
safe-resume guard, session_meta filtering) and continue the existing
session id instead of creating a new one. An explicit --resume of an
unknown session now fails loudly instead of starting fresh.