Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.
Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
role (unreadable/non-object/id-less payloads still reject rather than
block the queue) and now records the raw install_id in
sent_install_id; _body rewrites the payload's install_id from that
frozen column, keeping byte-identical resends anchored to one
recorded value.
Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.
Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.
Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.
258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
Seventh review found the claim-token fix incomplete, and its
reproduction is exact: the pre-POST check was READ-ONLY. A claimant
whose lease expired while suspended still passes it when it wakes
BEFORE anyone reclaims - its token is still in the row - and then a
second process legitimately reclaims while the first one's POST is in
flight. Both send. Reproduced at 60addb16e2: posts ['B', 'A'], both
reporting 'sent'. This is the check-to-POST expiry race, not the
documented mid-POST residual: A's lease was already dead before its
authority check passed.
The check is now an atomic RENEWAL (single CAS UPDATE): it requires the
token to match, the row to be pending, AND the current lease to be
unexpired, and only then extends next_attempt_at a fresh lease into the
future. rowcount == 1 is the only grant. A claimant that wakes past its
own lease fails the unexpired condition and yields even though its
token was never replaced - expiry alone means another process may
claim at any moment, so waking stale is disqualifying regardless of
whether anyone has taken the row yet. The renewed lease (300s) covers
the POST (30s timeout) with margin, and renewal runs before every
retry, not just the first attempt.
Regressions: the reviewer's exact ordering (expired wake before any
reclaim -> zero POSTs, row stays claimable), plus a healthy-claimant
renewal test. Mutation-checked: dropping the lease-unexpired condition
or the token condition each fails the suite.
The at-least-once scope note on _send_one stands: a suspension landing
mid-POST remains client-unfixable; the fixable window is now closed on
both sides (before the check, and between check and POST).
277 tests pass; ruff + footguns clean; staging E2E 202.
Responds to the independent PR review (andrexibiza). Both P1s were
checked against current HEAD rather than taken on authority - the
review was written against 613849c190, before the interval-model
consent replacement landed.
P1-1 (same-UTC-day revoke/re-enable releases refused data): already
fixed by the interval model. The reviewer's exact reproduction - opt in
06:00, revoke 12:00, package collected 18:00, re-enable 20:00 same day
- was re-run at HEAD: the off-window package stays local, and a full-day
aggregate straddling the revocation boundary also stays local (period
containment, timestamp precision). The consent-windows harness already
pins both. The reviewer's related ask that consent-ledger persistence
failures fail closed also holds structurally now: reconciliation derives
state rather than recording transitions, so a lost write means a shorter
confirmed horizon - less is released, never more.
P1-2 (lease has no owner) was VALID at head. Reproduced exactly as
described: A claims, is suspended past the 300s lease, B reclaims and
POSTs, A resumes and POSTs again - and the ingest key is minute-
prefixed, so the duplicate lands as a DISTINCT stored object, making
this worse than a benign idempotent overwrite.
Fix: every claim now mints a claim_token (additive nullable column,
schema version unchanged). Ownership is revalidated immediately before
every external POST, and every settlement, rejection, and backoff write
is compare-and-set on (package_id, claim_token, pending). A lapsed
claimant that resumes yields without transmitting, and its stale
backoff cannot move next_attempt_at under the live claim's lease.
Two deterministic regressions ship with it: expiry -> reclaim -> resume
(the reviewer's schedule), and the subtler stale-backoff-clobber case.
Honest scope, documented on _send_one: delivery remains at-least-once.
The token closes the claim->POST gap; a suspension landing mid-POST
(bytes already on the wire) is not client-revocable. The residual
duplicate is byte-identical content; collapsing it fully needs
package_id-keyed dedupe at the ingest service.
275 tests pass; ruff + footguns clean; staging E2E 202.
Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.
The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.
Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
call. Honest heartbeats never bind it; a machine off for months
catches up in a few hook fires (fail-closed latency only); one insane
sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
raw stamp lets an honest clock at revoke time pull a poisoned horizon
back to the true revoke moment. A rolled-back clock at close time
only closes earlier - fail-closed.
Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
(the harness re-implemented the insert; deleting the production line
survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
skip was dead code - the store constructor creates the directory
before the exists() check ran. The probe now checks the default path
without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
upgrade (fail-closed; deliberate).
New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.
273 tests pass; ruff and windows-footguns clean; staging E2E 202.
Structural fix after five review rounds put four blockers in the same
subsystem. The root cause was representational: consent history is a
sequence of on/off intervals, but it was stored as ONE moving day-stamp
plus a revoked flag. Every fix had to mutate that scalar at exactly the
right moment from exactly the right place, and each round the mutation
was missing from some reachable path (write-once stamp in R3; recorded
inside a loop that never runs when sending is off in R4; dead code
whenever collection was off in R5).
Consent is now recorded as explicit intervals (send_consent_windows) and
eligibility is a pure derivation: a package is sent only when its whole
period falls inside a recorded window. One writer -
reconcile_send_consent - derives window state from an observation of
(config, now). It is idempotent and order-independent, so the wizard,
the relay, and the mid-pass check all call the same function and cannot
disagree; there are no edges to detect and no ordering between writers
to get wrong. The relay reconciles once per process BEFORE the
collection gate, which fixes round-5 D1 (enabled:false made the only
idle-path observer unreachable). The claim reads the table and never
writes it, removing the read-path mutation (D2's rewrite vector).
Timestamp discipline, each rule load-bearing and mutation-tested:
- 'obs' high-water mark: monotonic, advanced only by observations;
confirms an open window forward (last_confirmed_at).
- 'data' high-water mark: advanced only by stored package period_end;
clamps window OPENS so a rolled-back clock cannot slide a window
under refused packages already on disk (round-5 D2).
- A close stamps last_confirmed_at, never "now": consent is asserted
only for observed time, so a hand-edited config with no process
running for 90 days fails closed (round-5 D1 strongest form).
- The gate requires period containment, not period_start >=, so an
intra-day revoke/re-enable holds back the day package (round-5 D3).
- Unlike the day-stamp, a revoke/re-enable cycle no longer destroys the
undelivered backlog from the earlier consented window (round-5 D4).
The redesign was validated BEFORE implementation against all 13
reproduced defect scenarios on a real store; the first two drafts each
failed scenarios in that harness (v1 leaked the unobserved-gap case by
closing at "now"; v2 leaked refused windows by letting data stamps
confirm consent). The harness ships as
tests/hermes_cli/test_shared_metrics_consent_windows.py.
Deleted: OPT_IN_PERIOD_KEY, SEND_REVOKED_KEY, LAST_SEEN_SEND_KEY,
opt_in_period(), record_revoked(), the relay edge detector body, and the
setup wizard's key bookkeeping (~170 lines of transition machinery).
Schema: two additive tables, version deliberately unchanged; verified
against a copy of the real production DB (13 rows intact, reopen no-op).
Also kills round-5's M8 survivor: the seen-exclusion mutation now fails
the suite. New mutation sweep: 8/8 killed, including one vacuous test of
my own this round (obs-mark monotonicity was covered only by
coincidence of the data mark; now pinned directly).
Documented cost: a fresh package waits at most one process start after
its period completes before release (fail-closed direction).
270 tests pass; ruff and windows-footguns clean. Staging E2E re-run
through the interval gate: both packages 202.
Fourth independent review. Two more consent leaks, both reproduced through
the real relay entry point before and after the fix. Both are failures of
my own round-3 fix, which recorded revocation in the wrong place.
BLOCKER 1 - revoking while idle recorded nothing. _record_revocation lived
inside send_pending's loop, but _send_exported_packages returns early when
send is false, before a sender is ever constructed. The dominant case is a
user turning sending off while no pass is running, so the loop that was
meant to observe the revocation could never run. Reproduced: 6 periods
collected during a refused window were transmitted on re-enable.
The window now closes on the observed config EDGE, before the early return.
Last-seen send state is persisted because each hook fires in a fresh
process, so a true->false transition is only visible by comparison. The
rising edge also opens the window explicitly: the sender only runs when
there is something to send, so a user who opts in and out before any
package exists would otherwise have no window for record_revoked to close.
BLOCKER 2 - turning COLLECTION off never recorded revocation. The
not-enabled branch in setup.py force-set send=false and returned without
calling _record_send_consent_change, so `hermes tools` -> disable shared
metrics silently dropped consent while leaving the window open. Same
retroactive release on re-enable. Both consent surfaces now record, and
setup keeps the relay's edge detector in step.
Also, from the same review's mutation sweep:
- the scheme check is now pinned as an allowlist. Replacing the http test
with `if True` survived the entire suite, because every non-http case
targeted a REMOTE host where the loopback branch rejects anyway. Only a
non-http scheme on loopback distinguishes the two. Shipped behaviour was
already correct; nothing guarded it.
- A.3 no longer claims rotation bounds long-term linkability outright.
Measured against 11 real packages: resource is a stable low-entropy
tuple and periods are contiguous across a rotation, so for a RARE
configuration those can bridge windows. The honest claim is that
rotation raises the cost, not that it makes correlation impossible.
Two mutants are documented as unkillable rather than papered over with
tests that only appear to cover them: the _defer clamp is unreachable from
any current caller, and widening the falling-edge check to an
unconditional else is behaviourally equivalent because record_revoked is
idempotent and no-ops without an open window.
An earlier version of the anti-spurious-revocation test could not fail
either - it used a never-consented store, where record_revoked no-ops
regardless. Rewritten to opt in, revoke, re-enable, and then assert that a
steady enabled state does not re-close the reopened window.
259 tests pass. Staging E2E re-run: both packages 202.
Third independent review. Both blockers reproduced against a real store
before and after the fix.
BLOCKER 1 — head-of-line starvation. The claim query is LIMIT 1, and a
package already handled this pass was rejected AFTER the fetch, so
_claim_next returned None and send_pending read that as 'queue empty'.
Any row that sorts first and becomes eligible again mid-pass therefore
terminated the pass. This is reachable normally: a 429 with a short
Retry-After, or a pass outliving the 15-minute failure backoff (a legal
pass runs ~1900s). Measured: 10 of 19 healthy packages silently dropped.
The seen-set is now excluded IN SQL, so None genuinely means no eligible work.
Same scenario now delivers 19 of 19.
BLOCKER 2 — revoking consent leaked once it was re-granted. opt_in_period
was write-once, so packages collected while the user had send: false
still had period_start >= the ORIGINAL opt-in day; re-enabling released
the whole refused window. Reproduced: 5 packages from a 5-day opted-out
window transmitted on re-enable. Turning sending off now closes the
consent window, and the next enabled pass opens a new one from that day.
Recorded both in the setup wizard and in the sender itself, because
config.yaml can be hand-edited where the wizard never sees it.
Also: a send_attempts ceiling (a poisoned head row burned ~160 requests
over 30 days, unbounded), _defer clamps to >= 1s so it cannot write a
past deadline, and the dead skipped_not_due field is removed.
Test-quality fixes, since vacuous tests have been the recurring problem:
- the lease test asserted only 'in the future', passing for a 1s lease;
it now requires the lease to outlast one package's worst legal case
- test_shutdown_joins_the_send_thread grepped getsource for a method
name — a change-detector AGENTS.md rejects — and is now behavioural
- gzip determinism was unguarded: both retries in one pass compress in
the same second, so removing mtime=0 was caught by nothing. Now
compares output across a real second boundary.
All five new regressions are mutation-verified: reintroducing each bug
fails its test. The first attempt-ceiling test SURVIVED its mutation
(the seeded row was excluded by another predicate) and was rewritten to
drive the real loop.
251 tests pass. Staging E2E re-run: both packages 202.
Second independent review found the lease fix incomplete. Reproduced
each finding before fixing.
BLOCKER — the batch lease expired mid-pass. _claim took up to 20 rows
under ONE shared lease, but a single package can legally consume ~96s
(three 30s timeouts plus 1s+5s backoff), so a full batch runs ~1900s
against a 180s lease. Later rows' leases expired while this pass still
held them, and another process re-sent them. Reproduced: 192s elapsed,
pkg-2 POSTed twice.
Packages are now claimed ONE AT A TIME, immediately before being sent,
so a lease only has to cover the package actually in flight. Verified:
same scenario now sends each package exactly once.
HIGH — revoking consent did not stop a running pass. The runtime read
send consent once before starting the thread, so a pass could keep
transmitting for minutes after a user set send: false, contradicting
the documented promise that it 'stops transmission immediately'.
Consent is now re-read before every package and fails CLOSED if it
cannot be established.
MEDIUM — all non-429 4xx were treated as permanent, discarding data.
403 is the ingest service's own origin guard: a Transform Rule or edge
misconfiguration would have permanently dropped every package sent
during the incident. Only 400 (malformed envelope) and 413 (over the
1 MiB cap) are terminal now; everything else retries.
MEDIUM — valid JSON that is not an object blocked the whole queue.
json.loads('["a"]') succeeds, then .get() raised AttributeError inside
the claim transaction, rolling it back and starving every healthy
package behind it. Payload shape and install_id are now validated, and
an unusable row is rejected individually.
LOW — the clock-rollback comment and test name claimed the opposite of
the code. The behaviour is right (a future issued_at means the recorded
age is untrustworthy, so reissue); the wording is now honest about it.
LOW — removed the stale HERMES_TELEMETRY_ENDPOINT reference left in
config_defaults after the override was deleted.
247 tests pass (was 234). Staging E2E re-run: both packages 202.
Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.
The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.
The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.
Also from review:
- shutdown() never joined the send thread; the join was only wired into
deactivate(). A short-lived CLI therefore killed an in-flight send at
exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
secrets, and a behavioural override here was a consent hazard: an
inherited variable could silently redirect telemetry a user agreed to
send to Nous. The staging E2E writes the endpoint into its throwaway
profile instead, which also exercises the real config path.
- Added the shared-metrics toggle that AGENTS.md requires
as the third opt-in surface, delegating to the setup prompt so the
consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
so a wrong path or oversized body retried every 15 minutes for 30
days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
send pass, which silently dropped the opt-in day whenever the next
export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
package differ on the wire, so the 'byte-identical retry' E2E was
comparing parsed bodies and could not have caught it. It now compares
raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
said no remote-delivery path exists.
233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
Step 6 of the shared-metrics exporter, plus a loopback E2E.
_export now triggers an opt-in send pass on a daemon thread. The hook
runs on finish_task — the user's interactive path — so a 30s network
timeout there would be felt directly; the thread keeps that latency off
the caller. A test asserts _export returns in under a second while a
send is deliberately blocked.
At most one pass is in flight per process: a queued second pass would
add nothing, because the next hook fire picks up whatever is still
pending. Shutdown joins the thread for at most two seconds, then lets
it go — the packages remain in SQLite and go out on the next run, so
blocking a user's exit on a slow network is the wrong trade.
Sending is resolved per pass from the profile's own config, so turning
it off takes effect at the next hook fire without a restart.
E2E (tests/hermes_cli/test_shared_metrics_sender_e2e.py): the real
sender against a real HTTPServer on loopback — actual urllib, gzip,
headers and sockets rather than an injected fake. Covers delivery and
sent-state, 400/429/5xx handling, a retry sending byte-identical
bytes, gzip shrinking a realistic 120-metric package and the server
parsing it back, install_id never crossing the wire, the outbox file
staying untouched, and a dead server deferring without raising.
Wiring tests: 12, all negative-space properties — no send without
opt-in, no blocking, no pile-up, no crash propagation.
Steps 4, 5 and 7 of the shared-metrics exporter: the send logic, the
consent gate, and backoff plus multi-process claiming. These arrive
together because the sender is not correct without all three.
Contract handling: 202 marks sent; 400 is permanent and never retried;
429 honours Retry-After (clamped to a day so a bogus value cannot park a
package); 5xx, timeouts and transport errors retry three times in-process
with 1s/5s/25s full-jitter backoff, then defer to a later pass.
Consent is gated on the package's PERIOD, not its creation time. A period
is split across packages created on different days, so a created-at gate
would send a period's tail while dropping its head and silently
undercount the opt-in day — data that looks complete and is wrong. The
opt-in day is recorded once and never moves, so toggling sending off and
on does not re-open the pre-consent backlog.
Rows are claimed in a write transaction, which is what stops two Hermes
processes sharing one database from sending the same package twice.
next_attempt_at persists backoff across restarts, so a hard-down service
is not retried on every task completion.
The body is recomputed from payload_json rather than stored a second
time: json.dumps is deterministic here (verified against the real outbox
— 11 of 11 files reproduce byte-for-byte), and the only mutable input,
the derived identity, is frozen on the row at first attempt. That keeps
retries byte-identical across a salt rotation for ~36 bytes instead of a
duplicate ~11 KB payload.
The outbox directory is never written to or deleted from. A 202 updates
SQLite only, because those files are the user's 30-day local history and
retention already owns their lifecycle.
Tests: 33. Two of them caught real defects in this commit — an
unreadable row aborted the claim transaction and blocked every package
behind it, and the compression assertions were passing through an
injected fake that bypassed the code under test.
Step 3 of the shared-metrics exporter.
The shared-metrics doc commits that a remote exporter 'must not reuse
the persistent local identifier by default'. install_id is therefore
never transmitted: each package carries
HMAC-SHA256(local-only rotation salt, install_id) instead.
Within a 30-day rotation window the value is stable, so distinct
installs remain countable — the first question the data has to answer.
Across windows it changes, bounding long-term linkability. The
derivation is one-way, so the service cannot recover install_id.
The salt lives in telemetry_state next to install_id, so removing the
shared-metrics directory resets both together and the documented reset
behaviour keeps working with no second cleanup path.
Rotation is deliberately not a bare 'age > interval' check: a clock
that jumps backwards must not read as an expired salt, and an
unparseable issued-at reissues instead of raising.
substitute_install_id replaces exactly one field and copies rather than
mutating, so payload schema evolution stays a sender-side concern.
Tests: 19, including that install_id never survives substitution, that
no other field changes, and — the property that keeps retries
contract-compliant — that a package rebuilt from a FROZEN derived id is
byte-stable across a salt rotation while a fresh derivation is not.
Step 1+2 of the shared-metrics exporter.
Config: telemetry.shared_metrics.send (default false) and .endpoint
(default production), resolved by a new shared_metrics_send_config
module. Precedence is HERMES_TELEMETRY_ENDPOINT > config > default; the
env var exists so the live staging E2E never has to mutate a user's
config. send requires enabled and never implies it — that combination
is a misconfiguration the user believes is working, so it logs an ERROR
once per process rather than silently doing nothing. Plaintext
endpoints are refused unless the host is loopback, so a typo cannot
send telemetry in clear text.
Per AGENTS.md, outbound telemetry needs a user-facing opt-in, so
setup_telemetry now prompts for sending as a second, separate question
and force-disables send when collection is turned off.
Storage: six additive nullable columns on package_outbox for send
bookkeeping. The store schema version deliberately does NOT move —
_ensure_schema_in_transaction raises on any version it does not
recognise and has no forward-compatibility branch, so bumping it would
hard-fail an older Hermes, a second profile on an older build, or a
rollback, against the same file. Old readers select named columns and
never SELECT *, so the additions are invisible to them.
Also corrects the two places that promised telemetry is never uploaded
(config_defaults comment and cli-config.yaml.example); leaving them
would make them false privacy statements once sending ships.
Tests: 26 covering config precedence, the enabled/send relationship,
transport safety, fresh-database creation, upgrade from a pre-send
database (rows preserved, version pinned, idempotent), and that the
shipped export query still runs. Mutation-checked: bumping the schema
version fails 5 of them.
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue #9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.
This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:
- the options
- the renderers for config.yaml, .env and the documents
- the activation body
- the command lines of the processes
nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.
`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.
The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.
The change also makes four corrections that apply to both modules:
- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
had only the gateway. But Hermes Desktop and the web dashboard connect
to a different process, so six of the configurations in public repos
add a second unit by hand. serve and dashboard are one entry point with
one flag of difference, and you can run only one of them. Thus the
option is an enum. The NixOS module asserts against container mode with
a backend, and does not make a unit that cannot start.
- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
installs into the working directory, which is correct for AGENTS.md but
wrong for SOUL.md and memories/. Hermes reads those files from
HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
made a workspace file that Hermes never loaded as the identity. The
documentation said this in prose, but two directory diagrams showed the
opposite. This change corrects both. A key in either option can now
contain subdirectories.
- `documents` needs an explicit `workingDirectory`. The default of that
option is bad on both modules. It is the home directory of the user on
Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
workspace files without a directory therefore gets a place that the
user did not select. The place is also different on each module. The
modules now refuse that combination.
The test is on the priority of the option and not on its value. An
option that nothing sets keeps the priority of its own default, and
each definition from a user is stronger. Thus a directory with the same
text as the default still counts as a selection, and so does a
mkDefault. A comparison of values detects neither case.
- Each activation writes .env again from a base in the Nix store, and
does not add to the file that exists. Thus a second activation cannot
put the same secret in the file two times, and a removed
environmentFile goes away. environmentFiles keeps the type `listOf
str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
the Nix store, which all users can read.
- HERMES_MANAGED and the .managed marker now hold the name of the system
that manages the install. Thus a refusal says "managed by home-manager"
and not "managed by NixOS", and `hermes update` gives the Nix guidance
for both shapes. The CLI does not print a rebuild command for each
system. It names the owner, and the user knows their own tool. A bare
`true` and an empty marker still mean NixOS, so this does not change an
existing install.
Verification. Six new checks, all built:
nixos-module evaluates the module with evalModules and the
NixOS module list. It asserts both units, one
HERMES_HOME, and that the module refuses
container mode with a backend.
home-manager-module evaluates the module with the
homeManagerConfiguration function of
home-manager. The process assertions run against
systemd units on Linux and launchd agents on
Darwin.
module-option-parity asserts that each shared option is on both
modules, and that the two exclusion lists name
only options that exist.
env-file-assembly runs the real .env script and checks the
contents, the mode, that a second run gives the
same bytes, and that a removed file goes away.
workspace-files-need-a-directory
checks that the module refuses `documents`
without a directory, and accepts a directory
that has the same text as the default.
service-argv runs each command line that the modules build
through the real parser of the CLI, with one
sentinel flag added, and requires that argparse
refuses only the sentinel.
`nix flake check` passes, with 21 checks in total.
The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.
Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:
- a lost --no-open
- a backend that runs the gateway
- an overwritten config.yaml
- documents in the wrong directory
- a different HERMES_HOME on the two processes
- a lost HERMES_HOME export
- a missing backend unit
- a removed assertion
- an .env file that grows at each activation
- an install that reports NixOS
- an empty .managed marker
- an option on the NixOS module only
- a stale entry in an exclusion list
- a renamed subcommand
- an unknown flag
- the workspace-files assertion always passes
- the assertion compares values instead of priorities
- an off-by-one that lets an untouched default through
- the assertion also fires for hermesHomeFiles
- a mkDefault no longer counts as a selection
- the Home Manager module stops wiring the assertion
- the NixOS module stops wiring the assertion
The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.
These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:
- test_git_probe_tree_kill.py (2 tests)
- test_update_import_guard.py (1 test)
- test_telegram_media_read_timeout.py (2 tests)
- test_teams.py (a collection error)
Closes#9056
# Conflicts:
# hermes_cli/main.py
# hermes_cli/update_cmd.py
# hermes_cli/web_server.py
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
Four hot-path consumers paid a full config deepcopy per read:
- telemetry gate relay_shared_metrics.enabled() — runs 2-3x per agent
turn (2x per API call from lifecycle hooks + 1x per tool call) and
called read_raw_config(), which deepcopies the whole raw config every
call. New read_raw_config_readonly() serves the cached dict directly:
248 us -> 4.6 us per call (54x) on Teknium's real 77-key config.
- interruptible_streaming_api_call local-endpoint stale-timeout branch
called load_config() once per API call for every local-model user.
- gateway get_inbound_media_max_bytes() + _get_ephemeral_system_ttl_default()
called load_config() on per-message paths. All three switched to
load_config_readonly() (345 us -> 12 us; PR #28866 lineage).
Together these account for ~90% of the ~1,900 deepcopy primitives per
turn measured in the 26-call stubbed-LLM profile.
read_raw_config_readonly() keeps the (mtime_ns, size) freshness key so
config edits are picked up next call, and preserves the identity
invariant (cache-miss returns the same object later hits serve) —
regression-tested with 'is', per the PR #28866 identity-bug lesson.
The mutable read_raw_config() is unchanged for save-path callers.
581 targeted tests green (config, relay metrics x2, ephemeral reply,
platform base, new readonly suite).