- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
.hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
update, and the dashboard/TUI payload builders. plugins_cmd.py only
gains the hooks (cmd_install catalog branch, cmd_update / dashboard
update re-pin, dashboard_install_plugin catalog_name + kill list,
dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
fallback) instead of the unauthenticated GitHub contents API (60 req/h,
1 request per entry); in-tree and live removals are unioned so a stale
cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
limits); _plugin_runtime_status shared from web_server_dashboard.py;
hub rows carry removed_reason. TUI plugins.manage gains catalog_name
install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
(real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
plugin-catalog/** so entry merges republish it.
Re-port of the PR's version gate onto the decomposed layout: the field and
parser live in plugins_manifest.py (with running_hermes_version /
version_satisfies helpers), the load-time skip in plugins_loader.py
before any import. Unsatisfied plugins record an error and never run
register(); one invariant test proves both halves.
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
run_backup() previously wrote "hermes-backup-<timestamp>.zip" on every
invocation without deleting old ones. Hourly callers accumulated 157 zips
(14 GiB). Add _prune_run_backups() to keep the newest N (default 3,
configurable via backup.run_backup_keep or --keep CLI flag).
httpx honours the env/system proxy (on Windows, the registry ProxyServer
even with no *_PROXY vars) but never the bypass list, so a system proxy
(Clash, corporate) answered 127.0.0.1 probes from both Desktop validators
with its own error page. That parsed as models=[] and the GUI said
"advertised no models at /v1/models" for a llama.cpp server the CLI
(urllib, honours <local>) saw fine.
Local endpoints (loopback, LAN, Tailscale via is_local_endpoint) now
probe with trust_env=False; public endpoints keep honouring env proxies.
A reachable endpoint answering non-2xx with no model list reports
"<url> answered HTTP <status>." instead of an empty catalog, so the
onboarding card stops telling the user to start a model.
Reimplemented on the decomposed router (the original patched
web_server.py before the split). Diagnosis and fix direction by
Solitud1nem in #63656; Windows registry-proxy confirmation by
Ulysses-Gaia on #63472.
Live repro (real loopback server, HTTP_PROXY=http://127.0.0.1:9):
before ok=False reachable=False 'Could not reach .../v1/models'
after ok=True models=['Qwen3.6-35B-A3B-Q5_K_M.gguf']
Co-authored-by: Solitud1nem <76743883+Solitud1nem@users.noreply.github.com>
Two picker-freshness defects in cached_fetch_api_models():
1. The disk cache row was keyed on base_url only, with the credential
fingerprint stored inside the row. N custom_providers entries sharing
one proxy URL with different keys (#106184) took turns overwriting the
single slot; every sibling then failed the fingerprint check, got an
empty catalog, and disappeared from the Desktop pickers (which hide
zero-model rows). Key on url#fingerprint so each credential owns a row.
2. cache_only opens (Desktop model.options without refresh) served a
past-TTL row for up to 7 days without ever revalidating, so a model
loaded on a non-current local endpoint stayed invisible until the user
found "Refresh Models". Serve the stale row AND spawn the same
off-thread SWR refresh the blocking path uses; the caller still never
waits on the network.
Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
GUI no-probe open before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
after {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
`hermes setup --reset` calls `save_config(copy.deepcopy(DEFAULT_CONFIG))`,
which writes `get_hermes_home()/config.yaml` — the exact file the backup
block a few lines below copies to `config.yaml.bak.<timestamp>`. Because the
copy ran after the reset, the backup captured the defaults that had just been
written, not the user's config. The one invocation where a backup matters most
produced a worthless one, and the original was unrecoverable.
The block's own comment already claimed it runs "before setup modifies it";
on the --reset path that was false. Move it above the --reset branch so it
captures the true pre-setup state on every path.
Also report the backup location on the --reset path. --reset is destructive
and can leave the wizard early (the non-interactive return exits before the
end-of-setup notice), so a user who just lost their config was never told
where the copy is. The end-of-setup notice is unchanged for the normal path
and is suppressed only when it has already been shown, so no run prints it
twice; the shared wording now lives in one helper.
Behaviour otherwise preserved: `copy2` (config.yaml holds secrets, so mode is
preserved), the try/except fallback to `_backup_path = None`, and the existing
notice for the full-setup path.
Follow-ups deliberately out of scope: pruning accumulated `.bak.*` files, and
printing the notice on the other early-return paths (--portal, section runs).
Refs #3522
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.
hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.
Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
A Desktop-owned `hermes serve --isolated --ssh-session-token-file ...` child
is spawned with an explicit `--profile <name>` when the connection names a
remote profile, and with no flag for the remote root home. Without the flag,
`_apply_profile_override` read the remote host's sticky `active_profile`
file and re-homed the backend into whatever profile the user last selected
on that machine's CLI. Settings then read one config.yaml while the remote
gateway wrote another, so model picks and toggles "didn't stick".
Treat the SSH token flag as a fixed-identity marker, the same way
supervisor-launched gateway children are (#74872): a Desktop backend's
profile is chosen by the client, never by the host.
Live repro (before/after, temp HERMES_HOME with active_profile=foo):
serve --isolated --ssh-session-token-file ... hermes_home=<root>/profiles/foo -> <root>
same + --profile foo hermes_home=<root>/profiles/foo (unchanged)
serve (no token file, user CLI) hermes_home=<root>/profiles/foo (unchanged)
With account-identity matching gone, the providers.<id> block consolidation
only fired on shared token material. A historical fork (same copied pool-row
id, profile rotated, both pairs diverged) then healed the pool row into root
but left root's providers.openai-codex block on the spent pair; root's next
load_pool() re-seeds its device_code row FROM that block and undid the heal.
_HealPass now records that a profile pool row matched root by copied id or
shared tokens and passes that verdict to _heal_forked_provider_block, which
accepts it as lineage proof. No account-identity guessing is restored; an
independent same-account grant (no id/token match) is still left alone.
Follow-up to simpolism's #106177.
CI: tests/cli/test_cli_resume_command.py builds bare HermesCLI objects without .model; the
refactor read self.model before the stored-model check the contributor's code made first.
_apply_stored_session_runtime was a line-for-line copy of the first half of
_restore_session_model (stored-model guard, session_gateway_runtime, bare-custom heal,
model/provider-changed check). Extract that pure decision into
cli_model_switch_mixin.stored_session_route and have both resume paths call it; the
one-shot keeps only the _ModelChoice mapping and the drop-ambient-key rule.
main.py stops re-normalising `resume` — _resolve_chat_session_args already did.
Tests trimmed from 20 to 13: near-duplicate unit tests of the private helpers go, the
end-to-end _run_agent contracts (stored runtime + reopen; explicit --model wins) and the
empty-session-keeps-id case stay.
Review fixes (#105957):
- A resumed one-shot ignored the session's stored model/provider runtime:
_resolve_model_and_provider()/resolve_runtime_provider() ran before
_load_resume_target(), which only loaded the session id + transcript, so an
ambient config (e.g. openrouter/ambient-model) served the resumed transcript
instead of the stored route (custom:stored/stored-model). The stored runtime
is now applied before runtime resolution, with the same contract as the
interactive _restore_session_model(): stored model/provider/base_url/api_mode
replace the ambient choice unless --model was passed explicitly, and a
changed provider drops the ambient api_key so resolution re-fetches
credentials for the restored endpoint.
- Passing the resumed id to AIAgent did not reopen the already-ended session
row: end_session() only writes rows whose ended_at is null and the
existing-row upsert never clears the end fields, so the resumed turn was
recorded under a session that stayed closed and its new lifecycle boundary
was lost. _load_resume_target() now reopens the row (best effort), same as
the interactive resume does before continuing.
Review finding on #105957: `_load_resume_target` returned None for a
resolved session with no stored messages, so `hermes -z "hello" -c <title>
--create-if-missing` recorded the turn under a freshly minted session id and
the just-created titled session stayed empty. Preserve `resolved` unconditionally — the interactive /resume path keeps the selected id for an
empty session too; only the history replay is empty. Regression tests pin the
durable id for both a plain empty session and an empty compression-chain head.
The -z exit path accepted --resume/-c in the parser but never forwarded
args.resume: every resumed one-shot turn silently started a fresh session,
so each wire request carried only [system, current user] and the model
lost all prior context (reported against Ollama/custom OpenAI-compatible
endpoints, but provider-independent).
Normalize session args (latest/title/--continue/--in + cwd restore) via
the chat path's _resolve_chat_session_args before the oneshot exit path
takes over, then load the resumed transcript in _run_agent through the
same contract the interactive CLI uses (compression-chain redirect,
safe-resume guard, session_meta filtering) and continue the existing
session id instead of creating a new one. An explicit --resume of an
unknown session now fails loudly instead of starting fresh.
Second entry of the same bug class: POST /api/cron/jobs/{id}/trigger →
_fire_cron_job_for_profile → CronScheduler.fire_due → claim_fire built its claim
without `manual`, so an off-tick run from the web UI stamped the future slot exactly
like the tools path #105704 fixes. fire_due/claim_fire gain `manual` (forwarded only
when set, mirroring `force`, so third-party providers keep working) and the dashboard
trigger passes it when the provider's signature accepts it. Webhook and misfire
catch-up fires run the slot that is due and keep the stamp.
Also drops the base-green tick-stamp test (the same contract is pinned by
tests/cron/test_scheduled_occurrence.py) and documents `manual` vs `force`.
Fleet telemetry showed "unknown" as the single largest execution_surface
bucket. Two construction paths were mis-attributed, both silently:
1. ACP editor sessions (VS Code / Zed / JetBrains) declare platform="acp",
but "acp" was absent from EXECUTION_SURFACES, so the contract's
closed-schema fallback folded every editor session into "other" --
the bucket meant for genuinely unclassifiable traffic.
2. batch_runner built agents from _AGENT_PASSTHROUGH, which omitted
"platform" entirely, so every batch task run reported "unknown"
despite "batch" already being a first-class surface.
Neither is a reporting bug in the exporter: both are declaration gaps at
the construction site. "unknown" must mean "this run genuinely could not
be attributed", not "a construction site forgot to say who it was".
Changes:
- add "acp" to EXECUTION_SURFACES and map it to the "interactive"
entrypoint alongside cli/desktop/tui
- add "acp" to the v2 wire schema enum (kept in sync by an existing test)
- pass platform through batch_runner: added to _AGENT_PASSTHROUGH, set
self.platform = "batch" on the runner, and defaulted at the worker call
site so callers that build a config without it stay attributable
Wire compatibility: the ingest service validates the envelope only and
stores metric bodies verbatim, so packages carrying the new value are
accepted by the already-deployed server. No coordinated deploy needed.
Tests: 12 new behavioural tests. Verified red before the fix (4 failed),
green after. Three fix-mutants confirmed killed:
M1 revert acp from EXECUTION_SURFACES -> 3 failed
M2 revert acp entrypoint mapping only -> 1 failed
M3 revert batch passthrough -> 1 failed
No source-text assertions; every test is a contract between the surfaces
the schema accepts and the surface each path declares. A guard test pins
that a genuinely undeclared run still reports "unknown", so attribution
cannot be "fixed" by inventing a default that hides real gaps.
_user_systemd_socket_ready() accepts systemd/private alone, which is enough for
systemctl --user but not for the systemd-run --user that restart-safe workers
need; systemd_user_bus_env() requires the bus socket. Replace the uid threading
through five helpers with one _wait_for_target_user_bus(uid) that polls
/run/user/<uid>/bus, and move the post-enable wait + restart hint out of
_ensure_linger_enabled into _ensure_system_service_linger so the activity probe
runs only when linger was actually just enabled. Kanban applies the bus env
unconditionally like the cron sibling. Refs #104893.
run_gateway() adopts the user bus once at boot; the generated system unit had
no ordering against user@<uid>.service, so after a reboot the two race and
the adoption can miss until the next gateway restart. Emit After=/Wants=
user@<uid>.service for the unit's User= (uid now returned by
_system_service_identity, which already resolved the account). Existing
system units are flagged outdated once and refreshed on the next
install/restart. Refs #104893.
A system-level gateway unit has no ordering against user@<uid>.service and
linger may be enabled after boot, so the bus can appear after the one-shot
adoption in run_gateway() ran. Derive XDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESS
fresh for the availability probe and every scoped spawn (cron worker, Kanban
worker, PTY/pipe terminal spawns, scope cleanup) so the 60s failure TTL can
actually recover. Refs #104893.
A system unit's `User=` has no login session on a headless host, so
user@<uid>.service never starts and `systemd-run --user --scope` — every
restart-safe cron/Kanban worker — has no bus to reach. `hermes gateway
install --system` runs as root and already knows the target user, so enable
linger for that user on fresh install, on an already-current unit, and on
repair.
- `get_systemd_linger_status()` / `_ensure_linger_enabled()` take the target
username; root is included (restart-safe workers cross `systemd-run --user`
regardless of who the gateway runs as).
- After `loginctl enable-linger` succeeds, wait for the TARGET uid's control
socket (`_wait_for_user_dbus_socket(uid=...)`) — logind starts the user
manager asynchronously and `--start-now` boots the gateway immediately; the
caller's own env (root's) says nothing about it and is never adopted.
- Messages and the manual-remediation hint are scope-aware (`sudo systemctl
restart`, not `systemctl --user`); when the repaired service is already
active, say that a restart is required — `systemctl start` on an active
unit is a no-op and the running gateway keeps its bus-less environment.
Refs #104893.
The docs batch (#105782-#105787, tracking #105788) fixed the website docs but
flagged two in-code strays it could not touch:
- hermes_cli/tips.py:121 still said delegate_task "spawns up to 3 concurrent
sub-agents"; the default has been 10 since the v33 config migration
(config_defaults.py:1256, config_migrations.py:613).
- hermes_cli/config_defaults.py:1778 said "remote backends run per-call",
contradicted by tools/code_kernel_remote.py (session kernels on remote
backends, failing open to per-call only when the backend cannot spawn one)
since #96991.
Comment-only changes: tips.py list entry updated, config_defaults.py comment
rewritten to match the actual kernel lifecycle. No behaviour change.
Validated: test_tips.py 6 passed; ruff, check-windows-footguns, and
check_compat_pointers all clean. Refs #105788 (crumb noted in the issue body).
Preserve Ben Barclay's diagnosis and replace the staging-file approach with
SQLite-coordinated writes and a five-second backup callback deadline. A
main-file replacement can replay an abandoned destination WAL; immutable
source reads can miss committed source WAL. Neither raw copy nor replacement
is safe when the destination is locked.
Refuse unavailable auth databases without raw-copy fallback, retaining the
existing close-browser-and-retry flow. Keep two invariant tests for lock
refusal/recovery and source-versus-destination WAL contents. Convert existing
text masquerading as database fixtures into real SQLite fixtures.
Related: #105754
Related: #96659
A `browser_exec` call could park a thread in `sqlite3_sleep` permanently
while mirroring Chrome's auth DBs, holding the agent's turn open. The turn
never reaches its `finally`, so no `session.info running=false` settle is
emitted and the Desktop composer latches busy — every later message queues
and never sends. Captured live: one thread stuck 24+ minutes across two
dumps, turn accepted at 15:03 with no `tui turn finished` 46 minutes later.
Root cause is the DESTINATION, not the source. `Connection.backup()` retries
a busy destination internally and ignores the connection's busy timeout, so
`sqlite3.connect(dst, timeout=5)` cannot bound it. A destination left locked
by an earlier hung mirror therefore blocks the next mirror forever — and
because the tool-level 420s timeout abandons the thread without interrupting
a C-level lock wait, the lock is never released and every subsequent launch
re-hangs the same way. Self-perpetuating.
Two changes:
- Back up into a fresh `<dst>.new` and `os.replace()` it into place. No other
process can hold a file we just created, so there is nothing to contend on,
and the swap stays atomic. Measured against a live Chrome with a
deliberately locked destination: 0.0006s vs an indefinite hang.
- Drop the `mode=ro` (no `immutable=1`) source fallback. 8e746668ba added
`immutable=1` to fix exactly this hang but left `mode=ro` as a fallback,
keeping the unbounded path one exception away; sqlite's busy timeout does
not cover lock negotiation, so nothing bounds it. `immutable=1` is also the
semantically correct mode — a committed snapshot of a file another process
owns. The bounded plain-copy fallback is unchanged.
Tests: three regressions, all mutation-checked (fail on base, pass here).
The locked-destination test runs the copy on a worker with a join deadline so
the unfixed behaviour fails fast instead of hanging the suite. 199 passing
across the browser real-profile and CLI suites.
Fence owner discovery by profile, lease and loopback endpoint, and hand the authenticated URL to the existing Ink transport without acquiring a competing lease. Distinguish lease age from turn activity in unsupported-owner recovery.
Client slice only: requires the integration runtime to advertise shared_runtime_url and provide the session-attach handshake. Classic CLI attachment remains an integration gap.
`key_cmd` (#86891) authenticates a provider with a SHORT-LIVED bearer minted
by a command — SSO/OIDC brokers, cloud IAM, internal auth proxies. The
request path has honoured it since it landed, but the picker resolved probe
credentials from `api_key`/`key_env` ONLY, so a key_cmd provider probed
`/v1/models` with an EMPTY key.
Against an authenticated endpoint the probe 401s, discovery returns nothing,
and the provider falls back to its single configured default model. The
picker shows ONE model, indistinguishable from an endpoint that genuinely
serves one — while inference keeps working, because that path mints
correctly. Reproduced against a LiteLLM gateway behind Entra OIDC: 0 models
discovered with an empty key, 26 with the minted token.
Both picker probe sites already funnel through `_entry_credentials()`, so
the fix lands in one place: it now reports a `cmd:<key_cmd>` identity, and
each site falls back to `resolve_probe_token()` after api_key/key_env. An
explicit static key still wins, so existing configs are unaffected.
The identity is keyed on the COMMAND, never the minted token: the token
rotates on every refresh, so keying on its value would change the group
fingerprint constantly and force a re-probe on every open. Two entries on
one URL with different helpers still get distinct rows.
`resolve_probe_token()` lives in agent.command_token_source, which already
owns key_cmd minting, and shares the CommandTokenSource cache with the
request path — a cache read, not a fresh sign-in. Fail-closed: a helper
needing an interactive sign-in degrades to today's empty-key behaviour
rather than taking down every other provider's row.
`_model_flow_named_custom` (the `hermes model` setup flow) is the sibling
path — it builds its own `Authorization: Bearer` from the same incomplete
resolution — and is fixed the same way, with one ordering constraint: the
value persisted to config.yaml is computed BEFORE the mint, so a short-lived
bearer can never be written back to shadow the key_cmd meant to re-mint it.
Tests drive the real code paths and assert on the credential each probe
receives rather than on function source, so a semantics-preserving refactor
does not fail them. Verified they fail with the fix reverted.
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.
Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
Carry the owning task's notification subscriptions independently of dependency
edges, within the creation transaction. Prefer its durable session over worker
and request-local sessions while preserving explicit overrides. Cover worker
CLI create and built-in decomposition, and retain conversation route anchors.
Auto-subscribe no longer upgrades an inherited passive subscription.
Slim adaptation of Christopher-Schulze's session-precedence fix in #85687,
expanded to durable subscription provenance and sibling creation paths.
Related: #85575, #85687
Validation: strict RED/GREEN (7 failing cases before; 7 passing after), then
58 Kanban test files: 383 passed, 2 skipped. Real dispatcher-spawn subprocess
probe covers direct, linked, unlinked, explicit-session, worker CLI, built-in
children and a plain CLI negative control, with recording transport only.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
A system-level unit (/etc/systemd/system, User=<someone>) is exec'd with neither
XDG_RUNTIME_DIR nor DBUS_SESSION_BUS_ADDRESS, and a process environment is fixed
at exec time. 'systemd-run --user --scope' therefore fails for the whole lifetime
of that gateway even after the user manager is up and /run/user/<uid>/bus is
reachable. That is the seam every restart-safe worker crosses
(restart_safe_gateway_child_argv), and it fails closed by design — so on headless
systemd installs every agent-driven cron job and every Kanban dispatch died at
launch, ~26ms in, with nothing but 'error' on the job row.
_ensure_user_systemd_env() already derives both values from our own uid and adopts
them only when the runtime dir is really ours and the socket really exists; it was
just wired exclusively to the systemctl management paths, never to the gateway's
own boot. Call it from run_gateway() — the single in-process boot every entry point
goes through — so the adoption precedes every worker-environment snapshot (cron
builds its env after the scope check, Kanban before it, so fixing this at the
dispatch seam would only fix one of them).
The fail-closed posture is unchanged: with no user manager at all the probe still
reports unavailable and dispatch still refuses. That refusal now names the remedy
in the message the operator actually reads (it is stored as the cron execution's
error), instead of only the symptom.
Fixes#104893
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.
Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
``gated_cache`` bypassed both the fresh-hit and the stale-while-revalidate branches of
``cached_provider_model_ids`` whenever the disk entry contained Astra. For an entitled
account that is every entry: live discovery rewrites the entry with Astra → next call is gated
again → a blocking /models round-trip on every picker open, for exactly the providers the user
is most likely on. The parallel prefetch's staleness check didn't know about the gate either, so
it skipped the slug and the blocking call landed in the serial picker loop the prefetch exists to
avoid.
Gate on provenance, not contents: only live discovery ever writes Astra into a same-credential
entry (the static/offline paths filter it out), so a fresh entry IS the entitlement record. The
one filter that matters stays — a stale entry served because the refresh failed drops Astra, and
the entry itself is left intact so the next successful fetch restores it.
The regression test now pins both halves: fresh entry served with Astra and zero live calls;
failed refresh past the fresh window serves the entry minus Astra.
The slug pair {"gpt-6-astra", "gpt-6-astra-900k"} was spelled out five times
(reasoning_effort, transports/codex, models.py x2, codex_models) with five hand-rolled
``.strip().lower().rsplit("/", 1)[-1]`` normalisations. ``is_astra_model`` in
agent/reasoning_effort.py (the module the other four already import) is now the only home, so a
new Astra alias is one edit.
``_finalize_codex_models(..., allow_astra=)`` was a control-coupling flag set True by the two
live callers and left False by the two static ones. The filter now lives where the untrusted
inputs are — ``_drop_undiscovered_astra`` over config.toml default + models_cache.json in
``get_codex_model_ids`` — and ``_finalize_codex_models`` is back to its one-line original.
DEFAULT_CODEX_MODELS never contains Astra, so the static catalog needs no gate.
``_openai_catalog`` appends whatever Astra ids live discovery returned instead of a name-keyed
``if "gpt-6-astra" in live_lower`` branch with a literal fallback that could never fire.