Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.
The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.
Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
call. Honest heartbeats never bind it; a machine off for months
catches up in a few hook fires (fail-closed latency only); one insane
sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
raw stamp lets an honest clock at revoke time pull a poisoned horizon
back to the true revoke moment. A rolled-back clock at close time
only closes earlier - fail-closed.
Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
(the harness re-implemented the insert; deleting the production line
survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
skip was dead code - the store constructor creates the directory
before the exists() check ran. The probe now checks the default path
without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
upgrade (fail-closed; deliberate).
New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.
273 tests pass; ruff and windows-footguns clean; staging E2E 202.
Structural fix after five review rounds put four blockers in the same
subsystem. The root cause was representational: consent history is a
sequence of on/off intervals, but it was stored as ONE moving day-stamp
plus a revoked flag. Every fix had to mutate that scalar at exactly the
right moment from exactly the right place, and each round the mutation
was missing from some reachable path (write-once stamp in R3; recorded
inside a loop that never runs when sending is off in R4; dead code
whenever collection was off in R5).
Consent is now recorded as explicit intervals (send_consent_windows) and
eligibility is a pure derivation: a package is sent only when its whole
period falls inside a recorded window. One writer -
reconcile_send_consent - derives window state from an observation of
(config, now). It is idempotent and order-independent, so the wizard,
the relay, and the mid-pass check all call the same function and cannot
disagree; there are no edges to detect and no ordering between writers
to get wrong. The relay reconciles once per process BEFORE the
collection gate, which fixes round-5 D1 (enabled:false made the only
idle-path observer unreachable). The claim reads the table and never
writes it, removing the read-path mutation (D2's rewrite vector).
Timestamp discipline, each rule load-bearing and mutation-tested:
- 'obs' high-water mark: monotonic, advanced only by observations;
confirms an open window forward (last_confirmed_at).
- 'data' high-water mark: advanced only by stored package period_end;
clamps window OPENS so a rolled-back clock cannot slide a window
under refused packages already on disk (round-5 D2).
- A close stamps last_confirmed_at, never "now": consent is asserted
only for observed time, so a hand-edited config with no process
running for 90 days fails closed (round-5 D1 strongest form).
- The gate requires period containment, not period_start >=, so an
intra-day revoke/re-enable holds back the day package (round-5 D3).
- Unlike the day-stamp, a revoke/re-enable cycle no longer destroys the
undelivered backlog from the earlier consented window (round-5 D4).
The redesign was validated BEFORE implementation against all 13
reproduced defect scenarios on a real store; the first two drafts each
failed scenarios in that harness (v1 leaked the unobserved-gap case by
closing at "now"; v2 leaked refused windows by letting data stamps
confirm consent). The harness ships as
tests/hermes_cli/test_shared_metrics_consent_windows.py.
Deleted: OPT_IN_PERIOD_KEY, SEND_REVOKED_KEY, LAST_SEEN_SEND_KEY,
opt_in_period(), record_revoked(), the relay edge detector body, and the
setup wizard's key bookkeeping (~170 lines of transition machinery).
Schema: two additive tables, version deliberately unchanged; verified
against a copy of the real production DB (13 rows intact, reopen no-op).
Also kills round-5's M8 survivor: the seen-exclusion mutation now fails
the suite. New mutation sweep: 8/8 killed, including one vacuous test of
my own this round (obs-mark monotonicity was covered only by
coincidence of the data mark; now pinned directly).
Documented cost: a fresh package waits at most one process start after
its period completes before release (fail-closed direction).
270 tests pass; ruff and windows-footguns clean. Staging E2E re-run
through the interval gate: both packages 202.
Fourth independent review. Two more consent leaks, both reproduced through
the real relay entry point before and after the fix. Both are failures of
my own round-3 fix, which recorded revocation in the wrong place.
BLOCKER 1 - revoking while idle recorded nothing. _record_revocation lived
inside send_pending's loop, but _send_exported_packages returns early when
send is false, before a sender is ever constructed. The dominant case is a
user turning sending off while no pass is running, so the loop that was
meant to observe the revocation could never run. Reproduced: 6 periods
collected during a refused window were transmitted on re-enable.
The window now closes on the observed config EDGE, before the early return.
Last-seen send state is persisted because each hook fires in a fresh
process, so a true->false transition is only visible by comparison. The
rising edge also opens the window explicitly: the sender only runs when
there is something to send, so a user who opts in and out before any
package exists would otherwise have no window for record_revoked to close.
BLOCKER 2 - turning COLLECTION off never recorded revocation. The
not-enabled branch in setup.py force-set send=false and returned without
calling _record_send_consent_change, so `hermes tools` -> disable shared
metrics silently dropped consent while leaving the window open. Same
retroactive release on re-enable. Both consent surfaces now record, and
setup keeps the relay's edge detector in step.
Also, from the same review's mutation sweep:
- the scheme check is now pinned as an allowlist. Replacing the http test
with `if True` survived the entire suite, because every non-http case
targeted a REMOTE host where the loopback branch rejects anyway. Only a
non-http scheme on loopback distinguishes the two. Shipped behaviour was
already correct; nothing guarded it.
- A.3 no longer claims rotation bounds long-term linkability outright.
Measured against 11 real packages: resource is a stable low-entropy
tuple and periods are contiguous across a rotation, so for a RARE
configuration those can bridge windows. The honest claim is that
rotation raises the cost, not that it makes correlation impossible.
Two mutants are documented as unkillable rather than papered over with
tests that only appear to cover them: the _defer clamp is unreachable from
any current caller, and widening the falling-edge check to an
unconditional else is behaviourally equivalent because record_revoked is
idempotent and no-ops without an open window.
An earlier version of the anti-spurious-revocation test could not fail
either - it used a never-consented store, where record_revoked no-ops
regardless. Rewritten to opt in, revoke, re-enable, and then assert that a
steady enabled state does not re-close the reopened window.
259 tests pass. Staging E2E re-run: both packages 202.
Second independent review found the lease fix incomplete. Reproduced
each finding before fixing.
BLOCKER — the batch lease expired mid-pass. _claim took up to 20 rows
under ONE shared lease, but a single package can legally consume ~96s
(three 30s timeouts plus 1s+5s backoff), so a full batch runs ~1900s
against a 180s lease. Later rows' leases expired while this pass still
held them, and another process re-sent them. Reproduced: 192s elapsed,
pkg-2 POSTed twice.
Packages are now claimed ONE AT A TIME, immediately before being sent,
so a lease only has to cover the package actually in flight. Verified:
same scenario now sends each package exactly once.
HIGH — revoking consent did not stop a running pass. The runtime read
send consent once before starting the thread, so a pass could keep
transmitting for minutes after a user set send: false, contradicting
the documented promise that it 'stops transmission immediately'.
Consent is now re-read before every package and fails CLOSED if it
cannot be established.
MEDIUM — all non-429 4xx were treated as permanent, discarding data.
403 is the ingest service's own origin guard: a Transform Rule or edge
misconfiguration would have permanently dropped every package sent
during the incident. Only 400 (malformed envelope) and 413 (over the
1 MiB cap) are terminal now; everything else retries.
MEDIUM — valid JSON that is not an object blocked the whole queue.
json.loads('["a"]') succeeds, then .get() raised AttributeError inside
the claim transaction, rolling it back and starving every healthy
package behind it. Payload shape and install_id are now validated, and
an unusable row is rejected individually.
LOW — the clock-rollback comment and test name claimed the opposite of
the code. The behaviour is right (a future issued_at means the recorded
age is untrustworthy, so reissue); the wording is now honest about it.
LOW — removed the stale HERMES_TELEMETRY_ENDPOINT reference left in
config_defaults after the override was deleted.
247 tests pass (was 234). Staging E2E re-run: both packages 202.
Independent review found the claim mechanism did not work. Reproduced
against the real store: two senders POSTed the same package.
The claim wrote next_attempt_at = now, but selection requires
next_attempt_at <= now, so a concurrent pass matched the same row
immediately. It now writes a LEASE INTO THE FUTURE
(_CLAIM_LEASE_SECONDS), which is what actually excludes another pass,
and expires by itself if a process dies mid-send. _mark is additionally
guarded on send_state so a straggler whose lease lapsed cannot
overwrite a completed send back to pending.
The old concurrency test could not fail: it raised AssertionError from
inside a transport, and _send_one catches every exception as a
retryable transport error. It now records what the second pass saw.
Also from review:
- shutdown() never joined the send thread; the join was only wired into
deactivate(). A short-lived CLI therefore killed an in-flight send at
exit, on the only cadence this feature has.
- Removed HERMES_TELEMETRY_ENDPOINT. AGENTS.md reserves HERMES_* for
secrets, and a behavioural override here was a consent hazard: an
inherited variable could silently redirect telemetry a user agreed to
send to Nous. The staging E2E writes the endpoint into its throwaway
profile instead, which also exercises the real config path.
- Added the shared-metrics toggle that AGENTS.md requires
as the third opt-in surface, delegating to the setup prompt so the
consent rules stay in one place.
- Non-429 4xx (401/403/404/413/422) are now permanent. Only 400 was,
so a wrong path or oversized body retried every 15 minutes for 30
days until retention pruned it.
- The opt-in day is stamped when the user consents, not on the first
send pass, which silently dropped the opt-in day whenever the next
export crossed midnight UTC.
- gzip now uses mtime=0. The embedded timestamp made two sends of one
package differ on the wire, so the 'byte-identical retry' E2E was
comparing parsed bodies and could not have caught it. It now compares
raw request bytes.
- Reconciled the three stale claims in relay-shared-metrics.md that
said no remote-delivery path exists.
233 tests pass (was 213). Staging E2E re-run through the config path:
both packages 202, and the service logged both objects written to S3.
Step 6 of the shared-metrics exporter, plus a loopback E2E.
_export now triggers an opt-in send pass on a daemon thread. The hook
runs on finish_task — the user's interactive path — so a 30s network
timeout there would be felt directly; the thread keeps that latency off
the caller. A test asserts _export returns in under a second while a
send is deliberately blocked.
At most one pass is in flight per process: a queued second pass would
add nothing, because the next hook fire picks up whatever is still
pending. Shutdown joins the thread for at most two seconds, then lets
it go — the packages remain in SQLite and go out on the next run, so
blocking a user's exit on a slow network is the wrong trade.
Sending is resolved per pass from the profile's own config, so turning
it off takes effect at the next hook fire without a restart.
E2E (tests/hermes_cli/test_shared_metrics_sender_e2e.py): the real
sender against a real HTTPServer on loopback — actual urllib, gzip,
headers and sockets rather than an injected fake. Covers delivery and
sent-state, 400/429/5xx handling, a retry sending byte-identical
bytes, gzip shrinking a realistic 120-metric package and the server
parsing it back, install_id never crossing the wire, the outbox file
staying untouched, and a dead server deferring without raising.
Wiring tests: 12, all negative-space properties — no send without
opt-in, no blocking, no pile-up, no crash propagation.
Older nemo-relay bindings reject metadata= on scope.pop, which aborted
turn finalization and left scopes open. Filter kwargs to what the live
binding accepts so close paths can complete.
Four hot-path consumers paid a full config deepcopy per read:
- telemetry gate relay_shared_metrics.enabled() — runs 2-3x per agent
turn (2x per API call from lifecycle hooks + 1x per tool call) and
called read_raw_config(), which deepcopies the whole raw config every
call. New read_raw_config_readonly() serves the cached dict directly:
248 us -> 4.6 us per call (54x) on Teknium's real 77-key config.
- interruptible_streaming_api_call local-endpoint stale-timeout branch
called load_config() once per API call for every local-model user.
- gateway get_inbound_media_max_bytes() + _get_ephemeral_system_ttl_default()
called load_config() on per-message paths. All three switched to
load_config_readonly() (345 us -> 12 us; PR #28866 lineage).
Together these account for ~90% of the ~1,900 deepcopy primitives per
turn measured in the 26-call stubbed-LLM profile.
read_raw_config_readonly() keeps the (mtime_ns, size) freshness key so
config edits are picked up next call, and preserves the identity
invariant (cache-miss returns the same object later hits serve) —
regression-tested with 'is', per the PR #28866 identity-bug lesson.
The mutable read_raw_config() is unchanged for save-path callers.
581 targeted tests green (config, relay metrics x2, ephemeral reply,
platform base, new readonly suite).