Widens #106272 to the one remaining sibling: `launchd_stop()` boots out with check=True and
already handles exit 3/113/125 (job unloaded) and 5/125 (domain unmanageable) by falling through
to the PID kill, yet inherited stderr — so `hermes gateway stop` against an unloaded job printed
"Boot-out failed: 3: No such process" next to "✓ Service stopped". Same `_CAPTURE_TEXT` kwargs as
the sibling calls; an unexpected exit still raises with `e.stderr` populated.
Replaces the contributor's two kwarg-assertion tests (`capture_output is True` on a mocked
`subprocess.run`) with two invariant tests that run a real fake `launchctl` on PATH and read fd 2
through `capfd`: restart-on-unloaded prints only its ↻/✓ lines and drives
kickstart→bootout→bootstrap→kickstart; stop-on-unloaded is silent, while a real bootout failure
(exit 1) still raises with the captured stderr. Both red on origin/main, green here.
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.
- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
the verdict of the one canonical classifier
(`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
`staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
a global-scope address (`2607:f8b0::1`) is not local while `::1`,
ULA and link-local still are via the `ipaddress` scope checks.
Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.
Refs #106228
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Move the WSL / /dev/dxg / d3d12_dri.so probes into _prefer_wsl_d3d12 with
the probed paths as module constants, so the launcher call site is a single
line and a test can lay out a fake WSLg host without touching real
/dev or /usr/lib. The two tests now run the real _desktop_launch_env end to
end (selected under WSL+dxg+driver; untouched with an explicit Mesa
override, off WSL, without /dev/dxg, or without the driver file) instead of
unit-testing the helper with a precomputed boolean.
Docs: one paragraph in the Desktop guide on the automatic selection and
the env vars that keep an explicit choice authoritative.
Follow-up to Xipong's fix for #106117 (salvaged from #106118).
Fold the helper into _npm_lifecycle_env itself: the whole fix is one
is_file() check plus a setdefault, so a separate function, the
try/except around get_hermes_home() (it never raises) and the
dict-returning indirection were shape-gate violations. Build the path
with os.fspath so it is correct on Windows too.
Tests: keep the two invariants (file present -> NPM_CONFIG_USERCONFIG
points at it; explicit process/caller value wins and a missing file
sets nothing), fold the other two into the negative test.
Docs: one paragraph in the desktop troubleshooting page next to
ELECTRON_MIRROR describing $HERMES_HOME/npmrc.
Refs #106373
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
__cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
SSLEOFError whose text httpx did not repeat still gets the hint. The
hint now names the TLS 1.2 diagnostic and links the providers docs
note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
classic-groups snippet (EN + existing zh-Hans copy).
Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.
- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
(device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
workaround hint; KeyboardInterrupt handling is unchanged
(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
kanban_request_review now rejects reviewers that are not installed profiles
(#106163); the cross-surface lifecycle test used a bare "reviewer" name with no
profile behind it, which is exactly the phantom the guard exists to catch.
`hermes kanban promote --force <id>` printed `Promoted <id> -> ready` and
then the very next claim (a human `claim`, or the dispatcher tick seconds
later) demoted the task back to `todo` with `claim_rejected
{parents_not_done}` and returned None (#106195). The non-force refusal
even pointed operators at `--force` as the escape hatch.
The claim gate is deliberate: `claim_task` is the single enforcement point
("never ready -> running with an undone parent, whichever writer set
'ready'", cda20eec0c), and `complete_task`/`request_review` re-check the
same predicate, so a child let through by a forced claim could still never
finish. A promotion override therefore has no honest outcome; the
dependency edge is the real knob.
- drop `--force` from `promote` (parser, CLI handler, `promote_task`
kwarg, the `forced` event field nothing read)
- the refusal message now states why the gate cannot be bypassed and names
the working remedies: complete the parents or `hermes kanban unlink`
- two invariant tests: refusal on an undone parent leaves `todo` with no
fake `ready`; the flag no longer parses
Salvage direction from #75354 by @vyacheslavk (diagnosis of the promote ->
claim gap); the consume-at-claim authorization there is not taken because
the same parent gate also blocks completion of the forced child.
`--in DIR` only chdir'd. Every cwd consumer (resolve_agent_cwd -> Codex
app-server thread cwd, the terminal tool, context-file discovery) prefers
TERMINAL_CWD over the process cwd, so a value inherited from a parent
Hermes surface, the shell or .env survived the chdir and the session kept
running in the old directory. The local backend was rescued by cli.py's
force-export at import time; docker/ssh backends and the TUI launch path,
which never imports cli.py, were not.
Refresh TERMINAL_CWD to the --in target when it is already set. An unset
variable stays unset so the backends keep deriving from the new process
cwd and no host path is pre-seeded into ssh/container backends.
Fixes#106220
Tests that exercise the dashboard's gateway-restart path can end up
spawning a REAL `python -m hermes_cli.main gateway restart` child when
the spawn seam is not intercepted. `_spawn_hermes_action` launches it
with start_new_session=True, so it outlives the pytest worker; the
child inherits the pytest-tmp HERMES_HOME, which is not a profile and
hashes to no service suffix, so `get_service_name()` resolves the
DEVELOPER's `hermes-gateway` unit, `systemd_restart` restarts the live
gateway, and without systemd the fallback runs `run_gateway()`
in-process forever and squats the webhook port.
Live repro on this machine (origin/main): an unintercepted spawn of
["gateway", "restart"] from a test restarted the production gateway
(MainPID 136820 -> 1689090, NRestarts=1). On 2026-09-03 a sibling
refactor moved `_spawn_hermes_action` from the `hermes_cli.web_server`
facade to `hermes_cli.web_server_gateway` ~10 minutes before the tests
were repointed; runs in that window patched a name production never
read and left 39 orphans alive for six days.
The live-system guard now rejects any subprocess primitive whose
command line the canonical matcher (`gateway.status.
_gateway_command_subcommand`) classifies as `gateway run|start|restart`.
Argv substrings are never consulted, so `gateway status`, `gateway
--help`, `hermes_cli.main serve`, etc. pass through. Three files that
deliberately spawn and reap a stub child with a gateway-shaped argv
(flock holders, sleep sleepers with an argv tail) opt out with the new
`spawns_gateway_lookalike` marker, which lifts only this check and keeps
os.kill guarded. Two canary tests pin the block and the pass-through.
Defence at the exact boundary the incident crossed: systemd_uninstall() and
uninstall._remove_systemd_gateway() unlinked whatever get_systemd_unit_path()
returned. Before stop/disable/unlink, read the unit's own
Environment="HERMES_HOME=..." line (the parser status/refresh already use)
and, when it names a different home than this process, warn with both paths
and leave the unit alone. A unit without the line (hand-written) is still
removed as before.
_profile_suffix() compared HERMES_HOME against get_default_hermes_root(),
which treats ANY home outside ~/.hermes (Docker /opt/data, a mktemp dir) as
"the root itself". Every such home therefore collapsed to the bare
`hermes-gateway` service name and the default profile's unit path
(~/.config/systemd/user/hermes-gateway.service); the documented
"else a short hash of the path" branch was unreachable.
A parity harness run with HERMES_HOME=$(mktemp -d) called
uninstall_gateway_service(), resolved to the production unit, ran
`systemctl --user stop/disable`, unlinked it and daemon-reloaded. With the
unit gone Restart= could not revive it: all cron jobs and every messaging
platform were down for 6.5 days.
Compare against the platform-native default home (~/.hermes) for the bare
name; keep the profile name for <root>/profiles/<name>; everything else
(temp dirs, Docker /opt/data) gets its sha256[:8] suffix as the docstring
always promised. The Docker image supervises with s6 (`gateway-<profile>`
slots), not systemd/launchd, so the bare host-service name was never load-
bearing there.
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.
_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
The contributor's 15-test module pinned encoding details (frozen-arg shapes,
postponed-annotation resolution, monkeypatch seams). The two behaviour
contracts that matter survive: a 12-request /api/profiles burst admits ONE
worker and leaves /api/status responsive on a 2-token pool (this file), and
the kanban board read shares one worker per key (test_kanban_read_admission).
Both go red when the coalescing wrapper is removed from the route.
_nous_picker_model_ids only uses the ids the Portal unions append — both
unions discard the pricing map (`model_ids, _ = union_with_portal_*`) — yet
it called get_pricing_for_provider("nous") without cached_only, so a cold
pricing cache paid a full /v1/models round-trip (network timeout on a slow
Portal) on the picker-open path for nothing. Pass cached_only=True; the
background pricing prewarm (#101685) fills the same cache for later opens.
Re-derived from #102099 by @finn763: the original patched
hermes_cli/model_switch.py, which 3b1ecfc0a1 decomposed; the live call site is
hermes_cli/model_switch_providers.py.
Based on #102099 by @finn763.
Replace the attribute-presence check (asserts a cache_clear attribute the decorator does not
expose) with one invariant: eight concurrent projects/tree requests run the profile scan once
and each caller gets its own copy of the payload. The scope-suite fixture now disables the TTL
instead of calling a non-existent clear hook, which silently no-op'd.
The all-profiles sidebar polls GET /api/profiles/projects/tree, the one
heavy sidebar endpoint that was not wrapped in @_sidebar_singleflight_cache.
Every poll fanned out list_profiles() + _build_project_tree() over every
profile (51 on this box, ~17k SKILL.md files walked), and each profile
resolution re-parsed its config.yaml because read_user_config_raw() ran
uncached. On a 2-vCore VPS running 'hermes serve' for the Desktop remote
backend this pinned both cores (py-spy: 110-116% sustained, HostHighCPU).
- Wrap get_profiles_projects_tree in the existing single-flight cache,
matching get_profiles_sessions_sidebar (5s TTL, errors[] not cached).
- Memoize read_user_config_raw() on a (st_dev, st_ino, st_size,
st_mtime_ns) fingerprint under _CONFIG_LOCK, same strategy as
read_raw_config(). Hits return a deepcopy so write-back round-trips keep
their fresh-dict semantics; parse errors are never cached; the docstring
no longer claims 'no caching'.
- Tests: memo semantics (deepcopy isolation, inode-replace reparse, error
non-caching, parse-count), a wrapper-presence pin for both heavy sidebar
endpoints, and a cold-cache autouse fixture in the scope tests (the 5s
TTL otherwise leaks one test's payload into the next).
Measured on the affected box: list_profiles 1.9s -> 0.16s warm, serve CPU
112% -> 22-34%, host CPU 50-70% -> ~26%.
- Pass 1 no longer materialises every classified record (full
`messages.content` included) until pass 2; `LayoutEvidence` keeps only the
capped per-position value sets (+ sessions rows for the one cross-column
invariant) and pass 2 re-streams the lost_and_found tables. A 276 MB
corrupted store no longer has to fit in memory.
- `_sentinel_holds` / `_text_shape_holds` if-ladders become rule tables.
- Tests trimmed to the three that bind behaviour (upgraded store maps by
name; verifier refuses when rows matched no layout; replayed history ends
at the current schema — the drift guard). No behaviour change; reverting
inference to "no layout" still fails the name-mapping test.
A store's physical column order depends on which schema it was created at and
which ALTER TABLE ADD COLUMNs it lived through; one hardcoded "upgraded"
layout cannot cover them. `session_schema_history` records the declared
schema of the salvaged tables over time; `reachable_physical_layouts`
replays it to enumerate every physical order a store can have. The mapper
infers the layout once per kind from the whole recovered population (one
store wrote all of them), maps cells by name, and counts records whose width
matched no layout; the recovery verifier refuses to report such a salvage as
healthy.
Rebased onto the simplified session modules (c88d60551e, b9b4600cb2,
7e5a1a11d9, 1915a0d27e); no behaviour change from the pre-rebase branch.
Upgraded state.db files gain columns via ALTER TABLE ADD COLUMN, so
physical order diverges from SCHEMA_SQL. Prefix-mapping onto the fresh
template put started_at at 0 and shifted titles/models. Insert by name
using known physical layouts; keep the plausibility gate.
Fixes#101409
538+570+296+216 lines of change-detectors → 4 files of contract tests:
seed catalog valid, bad entries skipped, kill list name-or-repo, live
fallback+union; real-git pinned install + sidecar + re-pin; kill list
blocks CLI/dashboard/TUI with only the CLI bypass; dashboard merge via
sidecar; probe get_config default; extractor live document.
Re-port of the PR's version gate onto the decomposed layout: the field and
parser live in plugins_manifest.py (with running_hermes_version /
version_satisfies helpers), the load-time skip in plugins_loader.py
before any import. Unsatisfied plugins record an error and never run
register(); one invariant test proves both halves.
The old tests encoded the skip (agent never constructed, [drift_skip] text,
alert-once bit). New invariants, proven red on origin/main: an unpinned job
with model_snapshot=A / provider_snapshot=P runs with AIAgent(model=A) and
resolve_runtime_provider(requested=P) after the global default moved to B/Q;
an explicit job pin still beats the snapshot; cron.model / cron.model_provider
still beat the snapshot; a legacy job without a snapshot follows the global
default. test_cron_drift_alert_once.py is deleted with the bit it tested; the
config-notice and impact-summary tests drop guard_enabled / model_drift_guard;
the desktop toast test asserts the informational copy.
Cron jobs are scheduled work bound to the source profile and its origin
channel. A clone that inherited cron/jobs.json fired every job twice: two
gateways with identical job ids running the same weekly jobs in parallel
(double spend, duplicate deliveries) until one gateway died.
Root cause: `cron` was not in _CLONE_ALL_HISTORY_EXCLUDE_ROOT, so the
copytree in _clone_all_into carried jobs.json along. Add it to the
per-profile history exclude set (applies to any source, CLI, dashboard
and TUI/desktop RPC all funnel through create_profile), recreate the
_PROFILE_DIRS skeleton after the copy so the clone still has an empty
cron/ (and sessions/), and say so in the CLI summary line and docs.
--clone (config-only) never copied cron; export/backup keep cron as
before (an archive is a portable snapshot, not a second live profile).
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
run_backup() previously wrote "hermes-backup-<timestamp>.zip" on every
invocation without deleting old ones. Hourly callers accumulated 157 zips
(14 GiB). Add _prune_run_backups() to keep the newest N (default 3,
configurable via backup.run_backup_keep or --keep CLI flag).
Two picker-freshness defects in cached_fetch_api_models():
1. The disk cache row was keyed on base_url only, with the credential
fingerprint stored inside the row. N custom_providers entries sharing
one proxy URL with different keys (#106184) took turns overwriting the
single slot; every sibling then failed the fingerprint check, got an
empty catalog, and disappeared from the Desktop pickers (which hide
zero-model rows). Key on url#fingerprint so each credential owns a row.
2. cache_only opens (Desktop model.options without refresh) served a
past-TTL row for up to 7 days without ever revalidating, so a model
loaded on a non-current local endpoint stayed invisible until the user
found "Refresh Models". Serve the stale row AND spawn the same
off-thread SWR refresh the blocking path uses; the caller still never
waits on the network.
Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
GUI no-probe open before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
after {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
`hermes setup --reset` calls `save_config(copy.deepcopy(DEFAULT_CONFIG))`,
which writes `get_hermes_home()/config.yaml` — the exact file the backup
block a few lines below copies to `config.yaml.bak.<timestamp>`. Because the
copy ran after the reset, the backup captured the defaults that had just been
written, not the user's config. The one invocation where a backup matters most
produced a worthless one, and the original was unrecoverable.
The block's own comment already claimed it runs "before setup modifies it";
on the --reset path that was false. Move it above the --reset branch so it
captures the true pre-setup state on every path.
Also report the backup location on the --reset path. --reset is destructive
and can leave the wizard early (the non-interactive return exits before the
end-of-setup notice), so a user who just lost their config was never told
where the copy is. The end-of-setup notice is unchanged for the normal path
and is suppressed only when it has already been shown, so no run prints it
twice; the shared wording now lives in one helper.
Behaviour otherwise preserved: `copy2` (config.yaml holds secrets, so mode is
preserved), the try/except fallback to `_backup_path = None`, and the existing
notice for the full-setup path.
Follow-ups deliberately out of scope: pruning accumulated `.bak.*` files, and
printing the notice on the other early-return paths (--portal, section runs).
Refs #3522
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.
hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.
Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
A Desktop-owned `hermes serve --isolated --ssh-session-token-file ...` child
is spawned with an explicit `--profile <name>` when the connection names a
remote profile, and with no flag for the remote root home. Without the flag,
`_apply_profile_override` read the remote host's sticky `active_profile`
file and re-homed the backend into whatever profile the user last selected
on that machine's CLI. Settings then read one config.yaml while the remote
gateway wrote another, so model picks and toggles "didn't stick".
Treat the SSH token flag as a fixed-identity marker, the same way
supervisor-launched gateway children are (#74872): a Desktop backend's
profile is chosen by the client, never by the host.
Live repro (before/after, temp HERMES_HOME with active_profile=foo):
serve --isolated --ssh-session-token-file ... hermes_home=<root>/profiles/foo -> <root>
same + --profile foo hermes_home=<root>/profiles/foo (unchanged)
serve (no token file, user CLI) hermes_home=<root>/profiles/foo (unchanged)
_apply_stored_session_runtime was a line-for-line copy of the first half of
_restore_session_model (stored-model guard, session_gateway_runtime, bare-custom heal,
model/provider-changed check). Extract that pure decision into
cli_model_switch_mixin.stored_session_route and have both resume paths call it; the
one-shot keeps only the _ModelChoice mapping and the drop-ambient-key rule.
main.py stops re-normalising `resume` — _resolve_chat_session_args already did.
Tests trimmed from 20 to 13: near-duplicate unit tests of the private helpers go, the
end-to-end _run_agent contracts (stored runtime + reopen; explicit --model wins) and the
empty-session-keeps-id case stay.
Review fixes (#105957):
- A resumed one-shot ignored the session's stored model/provider runtime:
_resolve_model_and_provider()/resolve_runtime_provider() ran before
_load_resume_target(), which only loaded the session id + transcript, so an
ambient config (e.g. openrouter/ambient-model) served the resumed transcript
instead of the stored route (custom:stored/stored-model). The stored runtime
is now applied before runtime resolution, with the same contract as the
interactive _restore_session_model(): stored model/provider/base_url/api_mode
replace the ambient choice unless --model was passed explicitly, and a
changed provider drops the ambient api_key so resolution re-fetches
credentials for the restored endpoint.
- Passing the resumed id to AIAgent did not reopen the already-ended session
row: end_session() only writes rows whose ended_at is null and the
existing-row upsert never clears the end fields, so the resumed turn was
recorded under a session that stayed closed and its new lifecycle boundary
was lost. _load_resume_target() now reopens the row (best effort), same as
the interactive resume does before continuing.
Review finding on #105957: `_load_resume_target` returned None for a
resolved session with no stored messages, so `hermes -z "hello" -c <title>
--create-if-missing` recorded the turn under a freshly minted session id and
the just-created titled session stayed empty. Preserve `resolved` unconditionally — the interactive /resume path keeps the selected id for an
empty session too; only the history replay is empty. Regression tests pin the
durable id for both a plain empty session and an empty compression-chain head.
The -z exit path accepted --resume/-c in the parser but never forwarded
args.resume: every resumed one-shot turn silently started a fresh session,
so each wire request carried only [system, current user] and the model
lost all prior context (reported against Ollama/custom OpenAI-compatible
endpoints, but provider-independent).
Normalize session args (latest/title/--continue/--in + cwd restore) via
the chat path's _resolve_chat_session_args before the oneshot exit path
takes over, then load the resumed transcript in _run_agent through the
same contract the interactive CLI uses (compression-chain redirect,
safe-resume guard, session_meta filtering) and continue the existing
session id instead of creating a new one. An explicit --resume of an
unknown session now fails loudly instead of starting fresh.
Second entry of the same bug class: POST /api/cron/jobs/{id}/trigger →
_fire_cron_job_for_profile → CronScheduler.fire_due → claim_fire built its claim
without `manual`, so an off-tick run from the web UI stamped the future slot exactly
like the tools path #105704 fixes. fire_due/claim_fire gain `manual` (forwarded only
when set, mirroring `force`, so third-party providers keep working) and the dashboard
trigger passes it when the provider's signature accepts it. Webhook and misfire
catch-up fires run the slot that is due and keep the stamp.
Also drops the base-green tick-stamp test (the same contract is pinned by
tests/cron/test_scheduled_occurrence.py) and documents `manual` vs `force`.
Fleet telemetry showed "unknown" as the single largest execution_surface
bucket. Two construction paths were mis-attributed, both silently:
1. ACP editor sessions (VS Code / Zed / JetBrains) declare platform="acp",
but "acp" was absent from EXECUTION_SURFACES, so the contract's
closed-schema fallback folded every editor session into "other" --
the bucket meant for genuinely unclassifiable traffic.
2. batch_runner built agents from _AGENT_PASSTHROUGH, which omitted
"platform" entirely, so every batch task run reported "unknown"
despite "batch" already being a first-class surface.
Neither is a reporting bug in the exporter: both are declaration gaps at
the construction site. "unknown" must mean "this run genuinely could not
be attributed", not "a construction site forgot to say who it was".
Changes:
- add "acp" to EXECUTION_SURFACES and map it to the "interactive"
entrypoint alongside cli/desktop/tui
- add "acp" to the v2 wire schema enum (kept in sync by an existing test)
- pass platform through batch_runner: added to _AGENT_PASSTHROUGH, set
self.platform = "batch" on the runner, and defaulted at the worker call
site so callers that build a config without it stay attributable
Wire compatibility: the ingest service validates the envelope only and
stores metric bodies verbatim, so packages carrying the new value are
accepted by the already-deployed server. No coordinated deploy needed.
Tests: 12 new behavioural tests. Verified red before the fix (4 failed),
green after. Three fix-mutants confirmed killed:
M1 revert acp from EXECUTION_SURFACES -> 3 failed
M2 revert acp entrypoint mapping only -> 1 failed
M3 revert batch passthrough -> 1 failed
No source-text assertions; every test is a contract between the surfaces
the schema accepts and the surface each path declares. A guard test pins
that a genuinely undeclared run still reports "unknown", so attribution
cannot be "fixed" by inventing a default that hides real gaps.
_user_systemd_socket_ready() accepts systemd/private alone, which is enough for
systemctl --user but not for the systemd-run --user that restart-safe workers
need; systemd_user_bus_env() requires the bus socket. Replace the uid threading
through five helpers with one _wait_for_target_user_bus(uid) that polls
/run/user/<uid>/bus, and move the post-enable wait + restart hint out of
_ensure_linger_enabled into _ensure_system_service_linger so the activity probe
runs only when linger was actually just enabled. Kanban applies the bus env
unconditionally like the cron sibling. Refs #104893.
run_gateway() adopts the user bus once at boot; the generated system unit had
no ordering against user@<uid>.service, so after a reboot the two race and
the adoption can miss until the next gateway restart. Emit After=/Wants=
user@<uid>.service for the unit's User= (uid now returned by
_system_service_identity, which already resolved the account). Existing
system units are flagged outdated once and refreshed on the next
install/restart. Refs #104893.
A system unit's `User=` has no login session on a headless host, so
user@<uid>.service never starts and `systemd-run --user --scope` — every
restart-safe cron/Kanban worker — has no bus to reach. `hermes gateway
install --system` runs as root and already knows the target user, so enable
linger for that user on fresh install, on an already-current unit, and on
repair.
- `get_systemd_linger_status()` / `_ensure_linger_enabled()` take the target
username; root is included (restart-safe workers cross `systemd-run --user`
regardless of who the gateway runs as).
- After `loginctl enable-linger` succeeds, wait for the TARGET uid's control
socket (`_wait_for_user_dbus_socket(uid=...)`) — logind starts the user
manager asynchronously and `--start-now` boots the gateway immediately; the
caller's own env (root's) says nothing about it and is never adopted.
- Messages and the manual-remediation hint are scope-aware (`sudo systemctl
restart`, not `systemctl --user`); when the repaired service is already
active, say that a restart is required — `systemctl start` on an active
unit is a no-op and the running gateway keeps its bus-less environment.
Refs #104893.