The PR shipped nineteen tests across six files, most of them variations of
one boundary. Keep the three that pin distinct behaviour:
- a routed desktop-ticker fire runs under multiplex semantics for exactly
its scope: a scope miss returns None instead of the launch credential and
the parent os.environ is byte-identical afterwards;
- a routed no_agent child never sees a launch-only name, whether the launch
.env defined it or a launch external source supplied it (applied or lost
to a pre-existing process value), while its own values come through;
- administrator-managed keys keep policy precedence over the routed
profile's own value.
Everything else was either a positive control of the same seam, a
set-membership check on a module-level constant, or a re-statement through
a different entry point.
Review findings on f5f88d5058. Three are defects the previous round introduced.
Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.
Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.
Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.
Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.
Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.
(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
Review findings on d8c467f223, each reproduced through its production path.
Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.
Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.
Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).
Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.
(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
`/*.db-wal` / `/*.db-shm` only match at the checkout root, so on a flat
install `git stash push --include-untracked` still swept
`cron/executions.db-wal` / `-shm` while the scheduler held the WAL-mode
database open (cron.executions._connect opens it via open_db in WAL mode).
The base file stayed put but its WAL vanished under a live writer, so the
next `_connect()` failed with `disk I/O error` (review finding on #111175).
Use `/cron/executions.db*` like the gateway recovery db rule already does,
covering -wal/-shm/-journal and retired-WAL dirs. Every other non-root db
rule in the block already uses the glob. The regression tuple gains the two
sidecars, and the stash test gets the reviewer's repro against the real
`_stash_local_changes_if_needed`: an open WAL connection with one committed
row must still be readable from a fresh connection afterwards.
The flat-install block listed state.db and kanban.db sidecars one by one,
so any other root-level SQLite store (response_store.db, a future ledger)
and its -wal/-shm/-journal sidecars would still be swept by the updater's
`git stash push --include-untracked` and unlinked under the running
gateway (#110648). Replace the per-file lines with root-anchored globs
(`/*.db`, `/*.db-wal`, ...). No tracked root-level *.db exists, and
`git ls-files -ci --exclude-standard` is unchanged before/after, so the
globs newly ignore nothing that is committed.
Also add the rest of the flat-install runtime roots the previous fold
missed: the credential siblings from
gateway/platforms/base.py::_ROOT_CREDENTIAL_PATHS (.anthropic_oauth.json,
google_token.json, google_oauth_pending.json, auth/,
webhook_subscriptions.json), the active pairing location platforms/
(gateway/pairing.py), kanban/, gateway_state.json, processes.json,
cron.pid, the channel directory/alias and feishu pairing stores,
pending_messages/, checkpoints/, plugin-data/, hooks/, and the Discord
message-recovery db under gateway/ (gateway/ itself is tracked, so only
that file pattern is ignored). `/.credentials/` had no producer -- the
real dir is `credentials/` (web_routers/files.py, _ROOT_CREDENTIAL_PATHS)
-- so it is replaced. `/state-snapshots/` is dropped: the existing
unanchored `*-snapshots/` rule already matches it.
The test tuple now carries one representative per ignored class and its
comment no longer claims _ROOT_CREDENTIAL_PATHS enumerates the sidecar
set (that is `_sqlite_files`).
The same `git stash push --include-untracked` sweep that took state.db
on a flat install (#110648) also takes every other untracked file at the
$HERMES_HOME root: config.yaml, auth.json/auth.lock, memories/,
profiles/, .credentials/, mcp-tokens/ and pairing/. Losing those on a
declined or failed restore strands the user's credentials and profile
config just as badly as losing the session store.
Extend the root-anchored block with those paths (none are tracked or
already ignored on main) and append them to the test's
FLAT_INSTALL_RUNTIME_STATE list so the existing stash invariant covers
them without a new test.
The per-profile job store lives at HERMES_HOME/cron/jobs.json, so a
flat install keeps it beside executions.db inside the checkout-root
stash domain. Without an ignore rule the untracked autostash of
hermes update sweeps it away with the rest of the runtime state.
Absorb the path (noted in #110670) and its regression assertion into
the runtime-state carrier.
(cherry picked from commit 66282e6dc3ac5f276cdee5b92a22856f146c9a45)
On a flat install (checkout root == $HERMES_HOME) the untracked autostash of
`hermes update` sweeps the live state.db/-wal, snapshots, cron ledger and
lock/pid files into the stash and unlinks them under the running gateway; the
restart recreates an empty store at the same path and the declined restore
leaves the profile with no transcripts (#110648).
Root-anchor the runtime state set in .gitignore, mirroring the
.hermes-bootstrap-complete (#38529) and /.install_method (#66189) precedent
and the $HERMES_HOME-root enumeration in the platforms base module, so the
stash step is never entered for runtime state alone. Regression test runs the
exact stash command against a real repo carrying the tracked .gitignore.
(cherry picked from commit 6c74c24e6131af189801e54f1e11b260a529b74a)
The timeout branch of _run_hook_callback_bounded unconditionally added
gate_key to _hook_abandoned. A worker that finishes between done.wait()
returning False and the caller taking the lock has already popped its
token via _release_token, so nothing would ever clear that entry: the
callback stayed blocked for every later call id until reload with no
thread behind it. Guard the insert on the worker still being registered.
The new test makes the race deterministic by swapping the module's
threading.Event for one whose wait() lets the worker finish and then
reports a timeout, and asserts a fresh call id still runs.
Also pass tool_call_id inline from terminal_tool_result instead of the
conditional dict plumbing: an empty id is already treated as "no
identity" by _hook_call_identity and unknown fields are withheld from
narrow-signature callbacks (same shape as _fire_approval_hook). Update
the stale "(hook_name, id(cb))" comment above _hook_running_callbacks.
The same-call negative control fired two sequential calls with a 0.1 s
timeout, so the first call timed out and the second was dropped by the
60 s suppression window; the identity gate itself was never exercised.
Mirror the positive test instead: a 5 s timeout, two threads with the
same tool_call_id while the callback is held on an Event.
Add one test for the abandoned-worker gate: a hung callback followed by
a call with a fresh tool_call_id, with suppression zeroed, must start
exactly one worker. Reverting the gate change makes it fail.
Concurrent invocations of the same tool in one session collapsed into a single
busy key (hook_name, id(cb)): the second invocation was reported as 'still
running' and dropped. For pre_tool_call a drop is a fail-closed block, so the
gate silenced itself on an ordinary, healthy callback.
Measured on a busy profile: 3574 skip lines and 0 timeout lines in one hour —
every skip was the 'while still running' branch, i.e. pure key collision, not
slowness.
The gate now keys on the call identity that is already in the payload
(tool_call_id, else turn_id, else none — the last case behaves exactly as
before). Suppression stays keyed coarsely on (hook_name, id(cb)): a hung
callback is a fact about the callback, so its back-off must not be diluted
per call.
Refs #98382. Independent of #107894 (that one releases the slot on timeout;
this one stops healthy concurrency from colliding).
(cherry picked from commit 53b3dacd008418fcdf5fa6dfcadde575a35a776e)
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.
/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
`hermes gateway migrate --multiplex` ran its one fallible step LAST (install +
start the default gateway) with nothing around it. On a fleet whose secondary
ran a system unit as root (#110850) that step raised, leaving the flag on, the
secondary's unit removed and no gateway anywhere, and the re-run hit the
"already multiplexing (flag on)" short-circuit over an empty fleet.
- apply_migration(): the default bring-up runs inside a rollback. On failure the
manifest written before the first destructive step restores the flag and
reinstalls every recorded per-profile gateway with its recorded User=.
- MigrationPlan.interrupted: flag on + manifest present + no live default
gateway is a half-applied migration, not "already multiplexed"; the re-run
resumes from the manifest (target manager and User= read from it, since the
units themselves are gone) instead of refusing. Flag off + leftover manifest
refuses to overwrite it and points at --standalone.
- ProfileGateway.services records EVERY installed unit (user and system) and the
manifest carries them; apply stops/uninstalls all of them and rollback
reinstalls all of them, so a second owner is never left live beside the
multiplexer. The unattended hook treats a two-unit profile as an ambiguous
topology and refuses (review finding on #110205).
- gateway_identity(): an unresolvable User= on a system unit stays None instead
of borrowing the profile directory's owner; the unattended hook treats the
unknown principal as a boundary (review finding on #110205).
- auto_migration_opted_out(): reads the effective config (load_config_readonly
under the default home), so a managed `false` wins over a user `true` and a
YAML string "false" is an opt-out, not a truthy value (review finding on
#110205).
Builds on KoNit-K's #110854 (run_as_user threaded through install, preserved
from the removed system unit).
Pin that hermes config set no longer treats gateway.auto_migrate as a known key after DEFAULT_CONFIG uses auto_multiplex_migration.
Co-authored-by: Cursor <cursoragent@cursor.com>
Rename DEFAULT_CONFIG to the documented key so the reader, docs, and defaults agree. Do not reintroduce auto_migrate.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema
One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.
MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.
Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.
`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.
Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.
`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.
The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.
* wip(desktop): connection.request store, resume restore, card routing for MCP targets
Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.
* fix(config): hermes update turns on the connections toolset for saved toolset lists
`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.
Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.
`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.
* refactor: anti-slop pass on the desktop slice; shorten added comments
Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.
* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed
The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.
session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.
lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.
vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.
* chore: drop __pycache__ files swept in by an over-broad git add
* fix(desktop): correlate the connection.request row with the model's tool call by reason
The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.
* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds
* fix(connections): settle reason derives from target state, never from the renderer
A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.
* fix(desktop): a pending connection card re-arms on resume and activate
The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.
* style: literal wording in added comments, docstrings and docs
* fix: shared gateway-event contract and config-schema category for the connection events
connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.
* style: import order (perfectionist) in the desktop and shared files this PR touches
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
Review findings on #110914 (@ehz0ah): the psutil leg compared watched
abspath against the kernel-resolved path psutil reports, so a symlinked
HERMES_HOME on macOS returned no holders and let the fallback probe mint
replacement sidecars under a live writer. Both sides now realpath.
doctor's `file:{path}?mode=ro` truncated at '?'/'#' in a home name;
build the URI with as_uri().
`hermes doctor --fix`'s WAL checkpoint and `repair_state_db_schema`'s
preflight documented themselves as fail-OPEN: `live_writer_holds_db` only
refused on unknown/deleted/uninspectable holders and then trusted a
`BEGIN IMMEDIATE` probe, which is blind to a `journal_mode=DELETE` reader
(SHARED only) and cannot run on a malformed file — exactly the states repair
and checkpoint get invoked in. A repair in a second process then REINDEXed /
VACUUMed a file the gateway still held (#103339 item 2).
- `hermes_state_holders.live_writer_holds_db`: any foreign holder of the DB or
a sidecar is a live holder; the probe is only an additional positive signal.
- doctor `--fix`: the checkpoint runs on `_exclusive_repair_db_guard`'s
connection instead of a bare writable `sqlite3.connect`, so an opener
arriving after the scan is refused, not joined; `_session_count` is a
`mode=ro` reader.
- Normal SessionDB writers are untouched: gateway + dashboard in two processes
both keep writing (a process-wide flock on the write path — PR #109270's
shape — would break that).
Tests: the two-process repair race test releases the test process's own
header-probe fd (it is a genuine holder now); the mid-repair writer fixture
opens its connection after staging starts (a pre-existing holder is refused up
front, which is the point).
Refs #103339#100896
The budget test's frozen clock bound nothing (a fresh per-round deadline
passed it too): it now drives two rounds with an advancing clock and
asserts round 2 waits only the remaining budget (600 then 300, not 600
and 600). New: the timeout-with-drained-texts stop (one wait, texts still
run), and the finalize skip (_wait_for_oneshot_background_completions
must not re-wait once _quiet_notify_linger_done is set).
drain_notifications pops every owned event off the shared queue; the
salvaged loop kept only type=completion texts, so an owned
async_delegation result was consumed and silently dropped — neither
injected as a follow-up turn nor requeued for another consumer.
Every drained event type renders formatted text, so the loop now injects
all of them. Regression: test_quiet_notify_loop_injects_owned_async_delegation_events.
The salvaged loop called wait_for_pending_completions(None) with a fresh
default 600s deadline on every round, and _finalize_single_query then
re-waited the full timeout on the same stuck notify_on_complete child:
a hung child blocked a quiet one-shot 2x-9x longer than before.
One deadline now covers the whole run — the loop passes the remaining
budget each round, stops after draining once a wait times out (a timed-out
process never fires this run), and the finalize pass skips its re-wait
when the loop already consumed the budget.
Regression: test_quiet_notify_loop_shares_one_linger_budget (one wait
call, full budget, on a stuck child).
Review follow-ups on the identity binding:
- The sentinel's `start_time` is `time.time()` at `record_startup`, seconds
after the process was born once imports finish, so comparing it with
psutil's create_time within 2 s would have read every real gateway as
undecidable and silently stopped the #109538 cold-start. `record_startup`
now stamps `create_time` (psutil birth via the existing
`process_identity._process_create_time`), `mark_exited` carries it, and
the attestation compares birth to birth. A sentinel from a gateway older
than the stamp falls back to the PID-only rule.
- A resume token written by pre-generation code and resumed by this code
probes the marker again instead of skipping the spawn.
- Horizon allows a 60 s backwards clock step; the unused `now` parameter is
gone; the create-time tolerance is a named constant; the read-then-unlink
in `_consume_start_attestation` is documented as best-effort.
`_attested_pid_exited_cleanly` matched the lifecycle sentinel by numeric PID
only, so a stale marker for PID 111 flipped from "clean exit" to "crash" once
an unrelated PID 222 lifecycle overwrote the sentinel, and a reused PID's clean
exit could vouch for a different life (#110020 review, gateway_windows.py:937).
`_write_start_attestation` now records `create_times: {pid: create_time}` via
the existing `process_identity._process_create_time`; `mark_exited` carries the
running sentinel's `start_time` onto the exited sentinel; the attested probes
fail closed for a bound PID whenever the sentinel cannot be shown to describe
that incarnation (other PID, start time off by > 2s, or no start time) —
"unknown" never reads as "dead". A missing sentinel still reads as dead, and
markers without `create_times` keep the PID-only rule.
Tests: two attestation tests (stale marker vs. moved-on sentinel → no
authority; own incarnation keeps authority / clean exit / legacy marker) and a
ledger test for the carried `start_time`. Mutation: with HEAD's prod files the
no-authority test and the ledger test fail.
`attested_death_generation` matched the clean-exit sentinel by PID only and never looked at
the marker's `ts`, so a historical marker (e.g. one left behind before a Desktop-owned era)
could later override Desktop ownership and authorize a duplicate gateway (#76129).
The read-only probe now treats a marker whose `ts` is missing, unparsable, in the future or
older than `START_ATTESTATION_MAX_AGE_S` (24h — the marker only bridges the seconds between a
✓ and the next CLI invocation) as no authority: `None`, fail closed, marker left unconsumed.
Partial follow-up to #110020 review thread (d); binding to per-PID process create_time is left
for a follow-up (needs `mark_exited` to carry `start_time` into the exited sentinel).
The resume token only recorded `cold_start_if_installed: bool` and execution re-read the
mutable one-shot start-attestation marker to decide whether a Desktop-owned install still owed
a cold-start. A concurrent `hermes gateway status`/`start` (`check_start_attestation`) consumes
that marker between plan and execution, so the spawn was skipped and the token cleared with no
gateway running.
The marker now carries a `generation` nonce. The plan records the generation whose unclean
death authorized the cold-start on the token (`attested_generation`); execution authorizes the
spawn from the token, still re-checks live gateway PIDs, and consumes the marker only while it
is still that generation — a newer marker written by a concurrent start keeps its own report.
`attested_gateway_died` becomes `attested_death_generation` (`None` = undecidable, fail closed).
Follow-up to #110020 review thread (a).
_cold_start_windows_gateway_after_update cleared the dead attestation as soon
as _spawn_detached() returned a PID, before _wait_for_gateway_ready() proved
the gateway survived. When readiness failed, the RuntimeError registered the
retry, but the retry then saw Desktop lifecycle ownership with no marker and
returned success without spawning anything — the silent outage of #109538
came back through the retry path (#110020 review). The marker is now
consumed after readiness is confirmed, so a failed spawn leaves the retry
its recovery obligation.
_attested_pids_from kept any `isinstance(p, int)` item, so `[true]`, `[0]`
and `[-1]` each survived as a "PID" and made attested_gateway_died([]) read
True — a malformed marker could override Desktop gateway-lifecycle ownership
and spawn a duplicate gateway (#110020 review). PIDs are now `type(p) is int
and p > 0`, and one malformed item fails the whole list closed: the writer
never emits such values, so a partially-bad list is not trustworthy either.
Review follow-ups on the refresh transaction:
- `_provider_state_transaction` takes a `timeout_seconds` applied to BOTH the
active and the root lock. The refresh passes max(default, refresh timeout
+ 5 s); before, only the profile lock used that budget and root's lock kept
the 15 s default, so the waiting profile raised TimeoutError instead of
adopting whenever the peer's POST ran long. The regression test now holds
the endpoint past the lock floor and fails without the passthrough.
- Peer adoption requires a stored access token as well as a rotated refresh
token; an incomplete stored pair falls through to the refresh.
- `_save_codex_tokens` keeps its body: the per-path lock is reentrant, so the
refresh calls it inside the open transaction (as the CLI-recovery path
already did) instead of a split-out helper.
Follow-up to #110024 (ehz0ah's review thread). `_refresh_codex_auth_tokens` POSTed the
single-use refresh token to OpenAI first and only then entered
`_provider_state_transaction("openai-codex")` for the write-back, so root's lock covered the
save alone. `resolve_codex_runtime_credentials` holds only the caller's own profile lock, so
two profiles borrowing the same ROOT grant could both submit `old-rt`; last root save won and
OpenAI answered `refresh_token_reused` / revoked the family — the failure #87503 exists to
prevent.
The transaction now spans re-read -> endpoint refresh -> write-back:
- enter `_provider_state_transaction` first; the yielded state is root's, re-read under root's
lock. If its refresh token already differs from the one we were about to submit, a peer
rotated it: adopt the stored pair and return without touching the endpoint.
- otherwise POST and write back through `_store_codex_tokens_in`, the body of
`_save_codex_tokens` split out so it can run inside an already-open transaction.
`_save_codex_tokens` keeps its signature for the login/import/CLI-recovery callers.
Holding the advisory flock across the network call is safe here and already the established
shape: `resolve_codex_runtime_credentials` holds the active-store lock across the same POST,
and every waiter's timeout is `max(AUTH_LOCK_TIMEOUT_SECONDS, refresh_timeout + 5)`, i.e. it
outlives one full endpoint timeout. `_load_auth_store` readers never take the lock, so
readers are not blocked; `_file_lock` is reentrant per thread per path, so the nested
transaction inside the caller's lock and the CLI-recovery save inside the transaction both
re-enter cleanly. A release-POST-retake variant would reopen the window it is meant to close.
Test: two refreshers with the same stale pre-read pair against a rotate-once endpoint that
rejects any replay — the endpoint sees `old-rt` exactly once, both callers end with the
rotated pair, root holds it, the profile store stays unshadowed. Red on origin/main
(`refresh_token_reused` surfaces for the second caller).
Following the grant's source on every save made a fresh device-code
login (or `hermes auth import`) under a profile that had been borrowing
root's Codex grant overwrite root's account instead of creating the
profile's own. Redirecting a save into another file is the exception, so
it is opt-in: the refresh path passes write_through=True; login, import
and recovery keep saving locally. The two save branches collapse into one
(store, path, set_active) triple.
Test: root discovery on Windows comes from LOCALAPPDATA — set it so the
fixture's root is the resolved root on every host.
The picked tests monkeypatched _auth_file_path/_global_auth_file_path
directly and leaned on a HOME override to dodge the pytest seat belt.
Isolate the way the rest of tests/hermes_cli does instead: Path.home ->
tmp_path and HERMES_HOME -> <root>/profiles/<name>, so the fixture drives
the same get_default_hermes_root() resolution production uses. Drop the
classic-mode test (no new behaviour: source == active store is the
pre-existing save path). Two invariants remain: root-borrowed refresh
lands in root (singleton + pool) with no profile shadow; profile-owned
grant stays local with root untouched.
Codex refresh tokens are single-use with rotation-family reuse
detection. _save_codex_tokens resolved the state via the profile's
root fallback but always persisted into the ACTIVE (profile) store, so
a profile-scoped refresh left the global store holding the consumed
refresh token — the next process to read it replayed it and OpenAI
revoked the whole rotation family, forcing a manual device-code
re-auth (#87503; observed four times on one multi-profile deployment).
Mirror the xAI source-aware save (#43589/#74339): resolve the state
with _load_provider_state_with_source; when the grant came from the
global root, write the rotated chain back to root only — singleton AND
credential_pool entries, under the root store's own lock, without
creating a shadowing profile key. Best-effort, with the same pytest
seat belt as the xAI path.
Fixes#87503
The launcher guard rejected every python shebang, so a pip/uv console
script pinned to the running venv (#!<venv>/bin/python) was also
discarded in favour of `python -m hermes_cli.main`. Reuse
linux_desktop_entry._shebang_escapes_running_env, which already knows
that `env` shebangs escape and a shebang inside the running
interpreter's directory does not; only the escaping launcher loses the
venv. Also drops the second shebang classifier the fix had introduced.
attested_gateway_died() re-ran find_gateway_pids() (current profile only)
although both callers had just proven the process table empty with
all_profiles=True, and it re-implemented check_start_attestation's
liveness rule. Callers now pass the liveness they hold (current_pids=[])
and both probes share _attested_dead(), so the consuming and read-only
twins cannot drift.
Four new tests overlapped: plan-time and spawn-time attested-death overrides
both exercised the same predicate via monkeypatched lambdas. Collapse to:
- one end-to-end test using a real attestation marker in a tmp home: dead
attested gateway keeps the plan under Desktop ownership, survives the
spawn-time re-check, and the marker is consumed by the spawn;
- one probe test: no marker / null or non-list pids / non-dict / non-JSON all
read False (fail closed), alive and clean-exit read False, read-only when
it does read True.
Existing #76129 tests keep their attested_gateway_died=False pins unchanged.
A Desktop self-update hand-off exits the app before the updater runs and can
kill the messaging gateway in those same seconds (#109538), so the updater's
discovery finds no live PID while the one-shot start attestation still
vouches for the dead one. Both Desktop-ownership checks then read "nothing
running" as "nothing to restore" and the bot stayed down until a manual
start.
Consult the attestation non-destructively before Desktop-owned lifecycle
suppresses a cold-start: a vouched-for PID gone without a clean ledger exit
keeps the plan and is restored; no attested death preserves the #76129 skip
unchanged.
Two remaining halves of #89184 (Desktop Settings saves rewriting unrelated
config):
- The `fallback_providers` structured editor normalized every entry down to
`{provider, model}`, so any edit (remove a row, pick a model) re-emitted a
hand-written local-gateway chain without its `base_url` / `api_key` /
`key_env` / `api_mode` — the next autosave persisted bare pairs and the
fallbacks silently routed to the public provider. Entries now carry every
key through; the editor only owns the two selects.
- `PUT /api/model/moa` did `cfg = load_config(); cfg["moa"].update(...);
save_config(cfg)`: the whole default-expanded snapshot went back to disk,
so a Desktop MoA autosave re-persisted every other section too (the
2026-09-10 repro: `fallback_providers: []` written alongside the MoA block
the user had just edited). It now saves `{"moa": ...}` with
`merge_existing=True`, the same section-scoped write every other sparse
writer uses since #110535. Hand-edited moa keys (#58819) still survive.
The `model.default not persisted / base_url cleared` symptom from the 0.20.4
report no longer reproduces on main through the real REST path (Config page
diffs against a baseline since 5361867c6d32; `_denormalize_config_from_web`
keeps the on-disk `model:` block).
SessionDB(read_only=True) cannot create a missing store, so a fresh install
running hermes insights / /insights errored instead of reporting no data
(reported by @ehz0ah on #110718; guard shape from @kshitijk4poor's #110026).
Co-authored-by: kshitijk4poor <kshitijk4poor@users.noreply.github.com>