Commit Graph

133 Commits

Author SHA1 Message Date
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium c5179a1d55 refactor(agent/credential_pool): inline single-use refresh dispatch; rebind entry before POST as base did 2026-09-02 19:11:28 -07:00
Teknium 522e114109 refactor(agent/credential_pool): split refresh/seed god methods, unify sync + status helpers (-1107 LOC) 2026-09-02 18:57:14 -07:00
Teknium 3312947e14 fix(auth): concurrent Nous 401 recovery adopts a peer's refresh instead of re-rotating the shared grant
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.

- resolve_nous_runtime_credentials(stale_access_token=): under the store
  lock, skip the refresh POST when the on-disk token differs from the one
  that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
  pass the failed bearer through, and treat a lock TimeoutError as
  'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
  refresh/0 unrecovered.
2026-09-02 09:33:05 -07:00
Teknium 37f3ba110a fix(auth): auto-heal single-use OAuth grants already forked across profiles (#100339)
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.

`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.

Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.

Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
2026-09-02 00:58:29 -07:00
Teknium 3038493ee6 fix(auth): never fork single-use OAuth grants across profiles (#100339)
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:

1. `hermes profile create --clone-all` and the dashboard/TUI
   `mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
   verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
   drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
   `providers.<id>` device-code blocks, and the PKCE singleton file; API
   keys are still copied. The clone reads the root grant through the
   existing credential-pool root fallback.

2. A named profile with no local rows BORROWS the root grant via
   `read_credential_pool()`'s fallback, but every persist
   (`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
   rows into the profile's own auth.json — materializing a fork on the first
   rotation. `persist_pool_entries()` now routes borrowed single-use rows
   back to the root store (update-only, under the root lock; never falls
   back to a local copy). A borrowed `hermes_pkce` rotation commits its
   singleton to the root `.anthropic_oauth.json`, the borrower never prunes
   root-seeded rows it cannot see the backing file for, and
   `hermes -p <profile> auth add` persists only the profile's own rows.

Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.

Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).

Closes #100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 00:58:29 -07:00
kshitijk4poor b81383ec21 fix(auth): compare pool-identity callers against all candidate keys
Two callers of get_custom_provider_pool_key compared against its single
preferred key and broke when the pool held the other identity:

- _prune_replaced_custom_model_config_credentials skipped only the
  preferred key, so a keyed provider's own legacy-named pool
  (custom:b.ai) was false-pruned of its current model_config credential
  when the preferred key resolved to the bare slug (b-ai).
- _seed_custom_pool seeded only when the pool key equaled the preferred
  key, so a legacy-named pool stopped being seeded from model.api_key.

Both now compare against the full custom_provider_pool_key_candidates
set. Also drops a redundant get_custom_provider_pool_key call from
_try_resolve_from_custom_pool (it returned candidates[0], doubling the
config traversal) and updates the two test files that monkeypatched the
removed module attribute.

Follow-up to #100413.
2026-09-01 22:42:27 +05:30
xxxigm 0bee5ff408 fix(auth): look up keyed custom providers by durable pool slug
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
2026-09-01 22:42:27 +05:30
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00
Teknium b4403a942a fix(auth): carry the spent-rotation verdict across processes via a durable sidecar registry
The consumed-but-uncommitted rotation verdict was process-local
(_SPENT_ROTATION_FINGERPRINTS), while the credential it protects is
explicitly cross-process: ~/.claude/.credentials.json is shared by every
Hermes profile and process. A fresh interpreter could lease the stale
access token or re-POST the already-spent single-use refresh token and
burn the credential family into invalid_grant.

- Persist non-secret one-way fingerprints to a sidecar registry next to
  the shared singleton source (claude_code / hermes_pkce), written under
  the same path-keyed cross-process lock that serializes refreshes.
- Consult the sidecar in the pool resolver, the pool refresh path, and
  the direct claude_code resolver/refresh before leasing or POSTing.
- Two-process regression: A rotates and loses the commit; B (fresh
  interpreter, empty local registry) must neither lease the stale pair
  nor POST the spent refresh token. Plus a no-verdict control.

Closes the remaining P1 from the exact-head review of f228439b on
PR #87891.
2026-08-29 18:34:35 -07:00
joaomarcos b7a9db8b9a fix(auth): keep the borrowed claude_code row out of token authority and carry the spent-rotation verdict through resolution
Two runtime blockers from the exact-head review of c057ef5.

1. A sanitized `claude_code` pool row was treated as token authority.

`claude_code` is a borrowed source: it is absent from the owned-source
allowlist, so `sanitize_borrowed_credential_payload` strips `access_token`
and `refresh_token` before the row reaches `auth.json`. `load_pool()`
re-hydrates the live pair from the singleton on every load, which is what
makes `~/.claude/.credentials.json` — not the pool store — authoritative
for this source.

`_sync_anthropic_entry_from_pool_store()` re-read that persisted row during
refresh. Being token-less, it "differed" from the live entry, so it was
adopted as a rotation performed by another process: `_refresh_entry()`
replaced a usable credential with an empty one and returned it before
`_claude_code_credentials_lock()` and the authoritative re-read were ever
entered. The empty OAuth entry then stayed selectable, because the
empty-runtime-key guard in `_available_entries()` covered API-key rows only.

Repairs: the pool-store sync refuses borrowed sources outright (plus a
defensive refusal of any token-less row, for future sources that sanitize on
write); the `claude_code` branch of `_refresh_entry()` now runs before the
generic adopt-and-return shortcut, so the path-keyed lock and the
authoritative re-read are always entered before deciding to POST or adopt;
and an OAuth entry with no access token is never leased.

2. A failed commit still fell through to the same spent credential.

`_refresh_oauth_token()` correctly returns None when the refresh POST
rotated the single-use token but the replacement could not be committed.
That verdict did not survive the caller: `resolve_anthropic_token()`
continued to `_resolve_anthropic_pool_token()`, which enumerates read-only
(`clear_expired=False, refresh=False`) over a pool that `load_pool()` had
just re-seeded from the unchanged singleton — so the pair whose refresh half
was already spent came back as a healthy token, and
`_refresh_provider_credentials("anthropic")` reported success and evicted
its cached clients.

Repair: every commit-failure path records the consumed pre-rotation pair as
non-reversible fingerprints (bounded, process-local), and both the Claude
Code file resolver and the pool resolver refuse a credential whose
fingerprint is on that list. `_refresh_provider_credentials("anthropic")`
consequently returns False when the spent family is the only credential,
while genuinely independent pool credentials stay eligible.

Coverage: `test_anthropic_borrowed_row_authority.py` starts from
`load_pool()` reading an actually persisted, actually sanitized row, forces
a refresh, and asserts the full pair survives with exactly one POST and one
commit, that the shared-file lock is entered, and that no empty OAuth entry
can be leased. `test_anthropic_spent_rotation_verdict.py` takes the full
resolver path: successful POST plus failed commit must make
`resolve_anthropic_token()` return None, make
`_refresh_provider_credentials("anthropic")` return False, and keep the
spent fingerprint out of every lease — with a control proving a successful
commit quarantines nothing and an independent credential still resolving.
Five of the seven new borrowed-row tests fail on the previous head, and the
three resolution tests fail with the verdict disabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
2026-08-29 18:34:35 -07:00
joaomarcos 7cbffdd125 refactor(anthropic): split the adapter godfile into four modules
`agent/anthropic_adapter.py` was 3,423 lines and this PR adds another auth
boundary to it. Split along the seams that were already there, so the
credential surface this PR changes has a single owner instead of being
interleaved with request building:

- `agent/anthropic_endpoints.py` (258) — base-URL/endpoint-family predicates.
  Pure functions over a URL string, which is what lets both of the modules
  below depend on it without a cycle.
- `agent/anthropic_message_convert.py` (1,225) — OpenAI-style to Anthropic
  Messages payload conversion: model ids, tool schemas, content/thinking
  blocks, tool_use pairing, cache_control, screenshot eviction, blank-block
  scrubbing.
- `agent/anthropic_credentials.py` (910) — credential sources, the OAuth
  flows, and the refresh commit (`CredentialPersistError` and both singleton
  writers).
- `agent/anthropic_adapter.py` (1,215) — client construction and the Messages
  API call, re-exporting every name from the three modules above so existing
  `from agent.anthropic_adapter import ...` imports keep resolving. The
  re-export surface was diffed against the pre-split module: nothing dropped.

Call sites that read a moved name through the adapter's namespace at runtime
(`credential_pool._refresh_entry_impl`, `auxiliary_client`) now import it from
the defining module, so there is one patchable seam rather than two bindings
that can disagree. The tests that monkeypatched those seams were retargeted to
match; no assertion was changed.

No behavior change.
2026-08-29 18:34:35 -07:00
joaomarcos 07faed33bb fix(auth): make the Anthropic refresh commit part of the transaction
Anthropic OAuth refresh tokens are single-use: the POST that returns a new
pair invalidates the one that was sent. The replacement therefore only
becomes real once it reaches its authoritative store -
~/.claude/.credentials.json for claude_code entries,
~/.hermes/.anthropic_oauth.json for hermes_pkce ones. Both writers caught
OSError/IOError, logged at debug level and returned nothing, so no caller
could tell a durable commit from a failed one.

That let a refresh spend the only refresh token, report success, and leave
the consumed pre-rotation pair on disk. _seed_from_singletons() re-reads
those files on every load_pool(), so the next process seeded the spent pair
back over the fresh pool row and the following refresh replayed a consumed
token (invalid_grant / refresh_token_reused) - exactly the failure this PR
set out to remove.

- _write_claude_code_credentials() and _write_hermes_oauth_credentials()
  now raise CredentialPersistError instead of swallowing the write error.
- _refresh_oauth_token() treats a failed commit as a failed refresh and
  returns None rather than handing back an access token whose refresh half
  was lost.
- _refresh_entry_impl() fails closed on both the primary and the recovery
  path: the rotated pair is never marked, persisted or returned, and the
  entry is quarantined DEAD with a credential_persist_failed reason so it
  leaves rotation and surfaces as an explicit re-auth instead of a silent
  fallback to another provider. The retry path now commits to the singleton
  before persisting the pool row.
- _upsert_entry() no longer treats re-seeding a borrowed source as a
  rotation. Borrowed rows (claude_code, env-backed) are written to auth.json
  without their secret, so comparing the re-seeded token against the empty
  stored value reported a rotation on every load and cleared the DEAD state
  the previous process had just written - resurrecting the quarantined,
  already-consumed credential on restart. It now compares the incoming
  token against the row's secret_fingerprint.

Adds failure-injection coverage for both writers, the direct resolver, the
claude_code and hermes_pkce pool paths and the retry path, each asserting
that a reload cannot bring the pre-refresh pair back as a usable credential.
2026-08-29 18:34:35 -07:00
joaomarcos 0099f250c2 fix(auth): close Anthropic OAuth review gaps 2026-08-29 18:34:35 -07:00
joaomarcos e1a210652a fix(auth): harden claude_code refresh lock and remove dashboard Anthropic OAuth
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).

Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
2026-08-29 18:34:35 -07:00
joaomarcos 739dc6d198 fix(auth): close Anthropic OAuth CSRF gap, cross-process refresh race, and API-key shadowing
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.

A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.

Fixes #87887, #87888, #87889.
2026-08-29 18:34:35 -07:00
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
Jack Lau f57209bc9f fix(agent): carry the ambiguity of Anthropic's 'out of extra usage' 400 through classification, cooldown, and terminal surfaces
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:

- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
  status-less paths now attach error_context {billing_unverified,
  possible_content_filter}. Reason stays FailoverReason.billing (rotation +
  fallback remain the right recovery either way); ClassifiedError grows a
  billing_unverified property.

- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
  unverified billing exhaustion gets the short transient cooldown instead of
  the one-hour bench, regardless of pool size: a content-filter rejection
  leaves the credential healthy and fails identically on every key, and the
  hour-long sole-credential latch is what replayed the stored error and made
  real fixes look ineffective. A true 402 keeps the full bench. The marker
  persists with the entry so a restart cannot upgrade it back to a bench.

- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
  threads billing_unverified and hands the pool 'billing_unverified' as the
  persisted failure_reason.

- agent/conversation_loop.py: the fallback-switch status, max-retries status,
  terminal label, and both structured terminal results hedge when the verdict
  is unverified. New _billing_terminal_label + _billing_failure_result build
  the returned terminal response in one place; the result dict now carries
  billing_unverified and the billing_block gains 'unverified': true. The
  confirmed-billing path (a real 402 or an API-key credit depletion) keeps
  the original assertive wording, so the caveat no longer dilutes it.

Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.

Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
2026-08-14 21:54:56 -07:00
686f6c61 bf7c716648 fix(agent): rebind pool entry id after env credential refresh
Per-turn .env adoption could rewrite agent.api_key while leaving
_credential_pool_entry_id on a previously rotated fallback. The next 429
then marked the healthy fallback exhausted via credential_id precedence
(#79156).

- Sync pool entry id after a successful env credential refresh
- First look does not stomp a pool-rotated key with the env primary
- mark_exhausted_and_rotate prefers api_key_hint when it disagrees with
  credential_id

Fixes #79156
2026-08-08 19:17:02 -07:00
Brooklyn Nicholson 9cd0338688 fix(credential-pool): bench a billing 403 fully, even as the sole key
The sole-credential cooldown sized the bench from the raw HTTP status, but
403 is overloaded: error_classifier maps OpenRouter's "key limit exceeded"
and xAI's spending-limit block to FailoverReason.billing, while an edge
throttle with the same status is transient. Only 402 was excluded from the
short cooldown, so a spent account on a single key retried every 60 seconds
and re-failed forever.

Thread the classified reason from recover_with_credential_pool through
mark_exhausted_and_rotate to _exhausted_ttl. Billing keeps the full bench
regardless of status; everything else transient still recovers in 60s. The
verdict is stored on the entry (_EXTRA_KEYS, so it persists to auth.json) —
without that a restart would re-read a bare 403 and downgrade the bench.

Tests: sole billing-403 stays benched, survives reload, unclassified 403
still recovers; call-site coverage that the reason actually reaches the pool.
Three existing kwargs assertions updated for the new argument.
2026-08-04 23:33:39 +05:30
kshitij d1eb08fcf3 fix: thread sole_credential into next_available_at sibling site
next_available_at() was computing the full 1-hour TTL for a sole
credential on a 429, contradicting the 60s cooldown in _available_entries.
The fallback restore gate (agent_runtime_helpers) uses next_available_at
to decide when to switch back from fallback to primary — so the agent
stayed on fallback for an hour instead of ~60s.

Add sole_credential computation in next_available_at mirroring
_available_entries, and a test verifying the short cooldown propagates.
2026-08-04 23:33:39 +05:30
A.Alzaro dcd7504349 fix(credential-pool): short cooldown for sole credential on transient throttle
A pool with only one usable (non-DEAD) credential has nothing to rotate to.
On a transient throttle (429 rate-limit, 403 edge-throttle, 5xx) the offending
key was benched for a full hour (EXHAUSTED_TTL_429/DEFAULT), so single-key /
no-fallback setups got an hour of hard failures for a throttle that resets in
seconds. The pool already special-cases 401 to recover quickly for single-key
setups; extend that to transient throttles when the credential is the sole
non-DEAD entry. 402 (billing/quota) keeps the full bench — a quick retry can't
help. Provider-supplied reset_at still overrides.

Adds tests covering sole 429/403 recovery, 402 full-bench, and multi-key
(no early recovery).
2026-08-04 23:33:39 +05:30
Jeongseok Kang 2d70f56327 fix(agent): adopt .env credential/base-url edits at the turn boundary (#67843)
* fix(agent): adopt .env credential/base-url edits at the turn boundary

A Settings save (desktop PUT /api/env, hermes setup) updates .env and
the saving process's os.environ, but a live session worker keeps the
base_url/api_key captured at agent init until restart — an open chat
silently kept calling the old endpoint (e.g. a local-server key sent to
api.openai.com, failing with an opaque 401).

Add AIAgent._try_refresh_env_client_credentials(), called at the start
of each conversation turn: re-resolve the provider's env-sourced
credentials (load_env() is mtime-memoized, so an unchanged file costs
one stat()) and rebuild the client via the existing
_replace_primary_openai_client machinery when the user edited them.

The refresh reacts only to env edits — resolved values changed since
the last look — never to mere divergence from the agent's current
values: credential-pool rotation and failover legitimately move the
session off the env credential, and stomping those back would flap.
Config model.base_url / pool custom endpoints keep precedence: edits
are only adopted while the session still runs on the registry default
or the previously-seen env value.

Lift _get_env_prefer_dotenv out of _seed_from_env to module level
(get_env_prefer_dotenv) so both the pool seeder and the per-turn
refresh share the same .env-over-os.environ resolution, including the
op:// indirection handling.

Fixes #67821

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(agent): address sweeper review on env credential refresh

- Cover named custom providers (#67935): provider="custom" has no
  PROVIDER_REGISTRY entry, so resolve the config block's key_env through
  the same lookup the runtime resolver uses.
- Make the edit baseline transactional: a failed client rebuild rolls the
  agent back and leaves _env_creds_seen un-advanced so the unchanged edit
  is retried next turn.
- Recompute route-derived TLS material and default headers on a base-url
  change, via a _reapply_route_client_config helper shared with
  credential-pool rotation so the two paths cannot drift.
- Rebase onto main: get_env_prefer_dotenv keeps the scoped _get_secret
  semantics from the profile-isolation fix (no raw os.environ reads).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: map jskang@lablup.com to rapsealk

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
2026-08-04 17:53:17 +00:00
Aleksandr Pasevin fe859a1f55 fix(credential-pool): clear exhaustion state on key rotation (#22622)
* fix(credential-pool): clear exhaustion state on key rotation

When a user rotates an API key (e.g. via `hermes setup` after hitting a
rate limit), _upsert_entry updates the access_token on the existing pool
entry but preserves the stale last_status=exhausted from the old key.
On the next session the pool finds the entry, sees it exhausted, and
returns no usable credentials — even though the new key is valid.

Fix: when access_token changes on an existing entry, reset last_status,
last_error_code, last_error_reason, last_error_message, and
last_error_reset_at. The exhaustion state belongs to the old key, not
the new one.

* chore: add pasevin@gmail.com to AUTHOR_MAP

* fix: clear last_status_at on key rotation, remove unused pytest import

Address review feedback from teknium1 on PR #22622:
- Add last_status_at=None to the reset block (matches all other
  token-sync reset paths in credential_pool.py)
- Assert last_status_at is None in the regression test
- Remove unused pytest import flagged by ruff + ty
2026-08-04 11:35:06 -06:00
kshitijk4poor 4075c8fd5a fix(credential-pool): lock the quarantine read-modify-write of _entries
#71775 moved deferred single-use-token refreshes outside the pool lock
(correct — they hold a cross-process flock plus network I/O). But
_refresh_entry_impl's three terminal-auth-failure quarantine paths do a
bare read-modify-write of self._entries. Those used to run with the
caller holding self._lock; on the deferred path they run unlocked, so a
concurrent mutation between the read and the write is silently lost.

Wrap all three in 'with self._lock' (an RLock, so locked callers
re-enter safely) and correct the _refresh_pending_entries docstring,
which claimed the mutations were already self-locking.

Post-merge gate-sweep finding on the #71775 salvage (#77714).
Sibling to the acquire_lease re-select fix.
2026-08-04 13:20:15 +05:30
kshitijk4poor db0bd42119 fix(credential-pool): re-select in acquire_lease after a deferred refresh
select() re-selects once deferred single-use-token refreshes complete;
acquire_lease() performed the refresh but returned its pre-refresh
answer. Since _acquire_lease_under_lock returns early exactly when a
refresh is pending (if not available: return None, pending_refresh),
a pool whose entries all needed a refresh always returned None — the
caller failed an answerable request right after the refresh succeeded.

Retry once, only when the first pass was empty and a refresh ran.

Post-merge gate-sweep finding on the #71775 salvage (#77714).
2026-08-04 13:12:24 +05:30
kshitijk4poor 82019e7c1b fix(credential_pool): unpack the tuple in next_available_at's gate
Cross-PR interaction fix: #77714 (salvage of #71775) changed
_available_entries to return (available, pending_refresh) while #77631
(salvage of #67642) added next_available_at() which still truthiness-
tests the bare return. A non-empty tuple is always truthy — even
([], []) — so the reset-aware gate silently returned None ('no wait
info') for every exhausted pool, disabling the feature #77631 shipped.
Unpack the tuple and test the available list.

Also adapts the lock-probe test for the RLock introduced by #77714
(same-thread non-blocking acquire always succeeds on an RLock; probe
from a helper thread instead).
2026-08-03 19:32:50 +05:30
kshitijk4poor 687dd632a1 fix(credential_pool): serialize deferred-refresh pool mutations
Review folds on the #71775 salvage (dossier findings 1+2):

- self._lock becomes an RLock and the mutation primitives
  (_replace_entry, _persist) are now self-locking, so the deferred
  single-use-token refresh path — which deliberately runs its
  cross-process flock + OAuth network I/O OUTSIDE the pool lock —
  still serializes its pool mutations against concurrent
  select()/rotation. In-lock callers re-acquire reentrantly.
- Dropped _refresh_pending_entries' redundant second _replace_entry:
  _refresh_entry already merges the refreshed entry internally.

Adds tests/agent/test_credential_pool_deferred_refresh.py pinning both
invariants: select() must NOT hold the lock during the refresh window
(the PR's whole point), and the post-refresh mutations MUST contend on
the lock (blocking-thread probe).
2026-08-03 19:08:09 +05:30
dsad 9cd605f303 fix(credential_pool): defer single-use-token refresh outside threading lock
select() and acquire_lease() held self._lock during the entire
_available_entries() loop, which for openai-codex and xai-oauth providers
includes a cross-process file lock (_auth_store_lock) plus OAuth token
refresh HTTP POST.  The lock timeout can exceed 20 seconds, blocking all
credential pool consumers across every gateway thread and subagent.

Collect single-use-token refresh entries under the lock, then execute the
refreshes outside it.  On success the refreshed entry is merged back into
the pool and re-selected.  Non-single-use providers (anthropic, nous)
continue refreshing inside the lock since their refresh is a simple HTTP
POST with no cross-process coordination.
2026-08-03 19:08:09 +05:30
kshitijk4poor 4c2d473a80 fix(credential_pool): run next_available_at under the pool lock
Review fold on the #67642 salvage: next_available_at() called
_available_entries() — which prunes DEAD entries, syncs tokens, and
persists — and iterated self._entries with no lock, racing concurrent
select()/rotation exactly as has_available()'s comment warns. Wrap the
method body in self._lock and pin it with a non-blocking-acquire probe
test.
2026-08-03 19:02:12 +05:30
WojtekMR3 6611d87003 feat: reset-aware primary restore — stay on fallback until the rate-limit window resets
restore_primary_runtime retries the primary every turn once the 60s
transient cooldown clears. For subscription-window limits (Claude
Pro/Max 5h windows, Codex weekly caps) the reset is hours or days away,
so every retry is a guaranteed failure costing two provider switches
and two prompt-cache invalidations per turn.

Add CredentialPool.next_available_at() (earliest reset across exhausted
entries; None when available now or no reset info) and gate the restore
on it: skip while the primary's pool says nobody can serve, restore on
the first turn after the reset elapses. Fail-open: any gate error or
missing reset info falls through to the existing per-turn retry, so
recovery can never be later than today. Cross-provider fallbacks
consult the PRIMARY's pool (not the attached fallback pool), reusing
the loaded pool for the existing rebind to keep auth reads at one per
restore.
2026-08-03 19:02:12 +05:30
kshitijk4poor 536ed6a33e fix(credential_pool): classify copilot sources by exact match
Review fold on the #76341 salvage: the substring test ('gh' in
source.lower()) classified GH_TOKEN and GITHUB_TOKEN as gh_cli, so a
user's env-var-specific suppression was silently bypassed (and
suppressing gh_cli silently dropped env tokens). Pre-existing bug on
main, but the PR's early gate makes the classification decide whether
the exchange runs at all. Match resolve_copilot_token's exact
'gh auth token' sentinel instead.

Adds 3 regression tests: env-var suppression gates the exchange,
gh_cli suppression doesn't swallow env tokens, all-sources suppression
skips the resolve subprocess entirely. Also corrects the ~13s comment
(actual worst case ~35s: 3x10s timeouts + 4.5s backoff).
2026-08-03 15:56:39 +05:30
wangyunyou 77404ce086 perf(credential_pool): skip gh subprocess when all copilot sources suppressed
The all-sources suppression gate now runs before resolve_copilot_token(),
which shells out to `gh auth token` (~30ms) on every pool load. A user
who suppressed every copilot source (hermes auth remove copilot gh_cli
suppresses gh_cli + all env variants) still paid the subprocess spawn on
every load — model picker open, /model, agent startup.

Enumerate the same source space credential_sources._remove_copilot_gh
suppresses and bail before any work when all are suppressed. Measured:
model.options payload build drops from ~0.46s to ~0.26s cold for an
all-suppressed user; resolve_copilot_token() is no longer called at all.
2026-08-03 15:56:39 +05:30
wangyunyou 0a2a69d80b fix(credential_pool): check copilot suppression before token exchange
The copilot branch of _seed_from_singletons ran the suppression gate
_after get_copilot_api_token(), which retries the network exchange 3x
with backoff (~13s worst case). A source the user already suppressed
(hermes auth remove copilot gh_cli) still burned the full exchange dead
time on every pool load — model picker open, /model, agent startup —
only to have the entry discarded afterwards.

Move the _is_suppressed() gate ahead of the network call, matching the
early-gate pattern every other singleton branch uses. Suppressed copilot
sources now skip the exchange entirely. Measured: model.options payload
build drops from ~13s to ~0.2-0.4s for a user with copilot suppressed.

Add regression test test_load_pool_skips_exchange_for_suppressed_copilot
asserting the exchange is never invoked for a suppressed source.
2026-08-03 15:56:39 +05:30
kshitij 7380b48589 fix: widen source guard to include manual:device_code entries
The _sync_codex_entry_from_auth_store source guard returned early for
source='manual:device_code', which is the recommended quarantine-safe
configuration (hermes auth add openai-codex produces SOURCE_MANUAL_DEVICE_CODE).
The PR's fix for refresh_token adoption was unreachable for these entries.

Widen the guard to accept both 'device_code' and 'manual:device_code'.

Follow-up to #70111. Issue reporter (imgyf) confirmed this caused a
12-of-16 fleet outage on Aug 1.

Co-authored-by: imgyf <imgyf@users.noreply.github.com>
2026-08-03 00:28:48 +05:30
JonthanaHanh 0ab4cdc27d fix(codex): adopt refresh_token from auth.json even without access_token (#70097)
Two defects in the openai-codex credential pool recovery path:

Defect 1 — adoption path silently no-ops when store_access is empty

_sync_codex_entry_from_auth_store() skipped adoption when the auth
store had no access_token (only last_refresh).  When another process
rotated the token pair, the stale profile's entry kept the consumed
refresh_token and replayed it, getting refresh_token_reused and going
terminally DEAD.

Fix: also adopt when store_refresh differs from entry_refresh, even
when store_access is empty.  Keep the entry's existing access_token
in that case (store_access or entry.access_token).

Defect 2 — false 'auth refreshed' success log

_try_refresh_codex_client_credentials() returned True whenever
resolve_codex_runtime_credentials() returned any non-empty credentials,
including the same stale token when the underlying refresh failed.
The conversation loop then logged 'auth refreshed after 401' right
before the retry failed with the identical token_expired.

Fix: compare the access token before/after the refresh.  If unchanged,
return False so the 401-retry path logs the truth.

Fixes #70097
2026-08-03 00:28:48 +05:30
praneshnikhar 4e6299af48 fix(credential_pool): use source-path-based write-through to root (#74339)
_sync_device_code_entry_to_auth_store used key-presence on the profile
store to decide whether to write-through rotated tokens to the global
root.  _store_provider_state unconditionally creates that key, so every
refresh after the first self-disabled the write-through — root kept a
revoked refresh token and every other profile died with
refresh_token_reused / invalid_grant.

Fix: use _load_provider_state_with_source to learn where the grant was
resolved from.  When the source is the global root, write back only to
root and skip _store_provider_state so the profile never accrues a
shadowing providers.<id> key that blocks future root fallback.

Add regression test verifying write-through fires on refresh 2+, not
just the first call.
2026-07-31 22:32:48 -07:00
dstkwll 7779409a76 fix(copilot): recover from stale/degraded token 400 AND expired IDE-token 401
Copilot degrades in two related ways that both abort a turn as non-retryable
and only clear on a gateway restart (a cold process re-runs the token exchange):

1. HTTP 400 model_not_available_for_integrator / model_not_supported — a
   raw/degraded token routes to the restricted copilot-language-server
   integrator whose allowlist omits enterprise-only models (e.g.
   claude-opus-4.8). Because it is a 400 (not 401), the existing 401 refresh
   path never fired. Prevented (retry-with-backoff exchange + on-disk JWT
   persistence + header guard at the client chokepoint) and self-healed at
   runtime (single-shot forced re-exchange + client rebuild + retry before
   fallback).

2. HTTP 401 'IDE token expired: unauthorized: token expired' — the short-TTL
   *exchanged* IDE token expires mid-turn. The clean-401 path DID fire and call
   _try_refresh_copilot_client_credentials(), but that method only re-resolved
   the stable raw ghu_ token and rebuilt the client — it never evicted the
   cached exchanged JWT or forced a fresh exchange, so the retry put the SAME
   expired token back on the wire, 401'd again, and the single-shot guard
   aborted the turn. Fix: force a fresh IDE-token exchange (evict cached JWT via
   evict_cached_exchanged_token + re-mint via get_copilot_api_token) before the
   client rebuild, mirroring the merged auxiliary-path recovery (#59837) and the
   400 recovery in this same PR. Graceful fallback to the resolved token if the
   exchange endpoint is unreachable; picks up the enterprise base_url on
   re-exchange.

Brings main-loop clean-401 recovery to parity with the merged auxiliary path
(#59837), using the newer on-disk-aware evict helper. Companion context: #58743
(this PR, expanded), #51313, #63204 (which assumed the 401 path already
recovered — it reached the method but the method was too weak).

Tests: exchange retry/persist round-trip, restart-blip disk reuse, stale-cred
400 classifier, 400 recovery, and 3 new 401 cases (fresh exchanged token on the
wire; network-blip fallback to resolved token). 58 copilot tests green on
current main.
2026-07-31 22:31:09 -07:00
JonthanaHanh 431d2a628c fix: break unbounded 401 retry loop in credential pool OAuth path
When api_key_hint from a 401 response doesn't match any pool entry
(common with OAuth tokens where runtime_api_key rotates), the pool
rotated without marking anything exhausted and handed back a fresh
selection. Because nothing was ever marked, the pool could never reach
the "no available entries" state — the caller retried the same dead
token forever (~6 attempts/sec), starving the event loop so /stop was
never processed; only killing the gateway ended it.

Rebased onto the identity-tracking rework that landed on main
(73c4b5a045): the single-entry escape from that commit already stops
the most common OAuth case, so this fix bounds the REMAINING gap —
multi-entry pools ping-ponging A->B->A with an unmatched hint. Cap
consecutive no-mark rotations at one full lap of the available
entries, then return None so the error surfaces / fallback activates.

Deliberately does NOT mark innocent entries exhausted (the original
PR's approach): that would quarantine a healthy key for the full
cooldown TTL on a hint that provably matches nothing. No cooldown is
written by the escape, so healthy keys stay available next turn --
bounded without hammering.

The streak resets when a rotation identifies a real entry and on any
successful normal select(), so only genuinely consecutive unmatched
rotations trip the bound.

Fixes #70401
2026-07-24 15:50:36 -07:00
Gille 73c4b5a045 fix(auth): stop stale-key credential recovery loops
Track the selected credential by stable pool entry ID so token refreshes and shared cursor movement cannot detach failures from the entry that issued them. Stop unmatched single-entry pools from reporting a no-op rotation as successful recovery.

Co-authored-by: Maxim Esipov <maksesipov@gmail.com>
2026-07-24 22:58:47 +05:30
Teknium d9165d7a67 fix: resolve current entry unlocked in try_refresh_matching no-hint branch
Follow-up to the #62614 salvage: try_refresh_matching (added by the
#69843 salvage after this PR's base) calls self.current() while already
holding the now-locking non-reentrant pool lock — a guaranteed deadlock
that git merges silently (no textual conflict). Use _current_unlocked()
and cover the method in the no-deadlock test.
2026-07-23 09:31:58 -07:00
solyanviktor-star 769381fb3e fix(credential_pool): complete the locking boundary across the public pool surface
Follow-up to review feedback:

- Acquire self._lock in the remaining public pool-state methods:
  has_credentials, reset_statuses, remove_index, resolve_target, and
  add_entry. All of them read or rebind self._entries (and the mutating
  ones persist auth.json), so they now hold the same lock as select()
  and the query methods. None are called from within the lock, so no
  unlocked helpers are needed.
- Make the blocking test deterministic: an instrumented lock records the
  acquire attempt, and the test first waits for the worker to actually
  reach self._lock before asserting it blocks. Previously an unlocked
  method could pass if the worker thread was scheduled late.
- Extend the lock test matrix to all nine public methods; the five newly
  locked ones fail the test without this fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 09:31:58 -07:00
solyanviktor-star 5b794c984e fix(credential_pool): acquire the pool lock in has_available/peek/current/entries
`has_available()`, `peek()`, `current()` and `entries()` read (and, via
`_available_entries()`, mutate and persist) `self._entries` without holding
`self._lock`, while every other entry point — `select()`,
`mark_exhausted_and_rotate()`, `acquire_lease()`, `try_refresh_current()` —
guards the exact same access with the lock.

`_available_entries()` is not read-only: it prunes aged-out DEAD manual
entries (rebinding `self._entries` at the prune step) and calls `_persist()`
(writes auth.json). The gateway runs platform adapters in threads and cron
runs jobs in a ThreadPoolExecutor, so a status probe via `has_available()`
or `peek()` can race a concurrent `select()`/rotation: torn iteration of
`self._entries`, interleaved auth.json writes, or a lost token rotation.

Fix: take `self._lock` in all four query methods. Because the lock is
non-reentrant and `peek()` composes `current()` + `_available_entries()`,
add a lock-free `_current_unlocked()` helper and route the already-locked
internal callers (`_select_unlocked`, `mark_exhausted_and_rotate`,
`_try_refresh_current_unlocked`) through it to avoid self-deadlock.

Added regression tests: a no-deadlock check (peek re-entrancy) and a
lock-held-blocks-the-call check for each of the four methods.
2026-07-23 09:31:58 -07:00
Blade 3d67f00fe1 fix(credential-pool): stop lost-update cooldown erasure and wrong-key quarantine
Two related races in credential-pool cooldown state:

1. Lost update across processes: write_credential_pool merged only
   entries missing from the caller's snapshot; for entries present on
   both sides the caller's in-memory copy won wholesale. A process
   holding a snapshot taken before another process marked a key
   exhausted would, on its next persist (e.g. a round-robin rotation),
   write the key back as healthy — erasing the cooldown so every
   process resumes hammering a rate-limited key. Merge status fields by
   last_status_at recency: adopt the on-disk status only when it is
   strictly newer AND still binding (DEAD, or EXHAUSTED with an
   unexpired cooldown), and never onto re-authed (token-changed)
   entries, so legitimate expiry-clears and fresh logins are preserved.

2. Wrong-key quarantine: when mark_exhausted_and_rotate received an
   api_key_hint that matched no entry, it fell through to
   current()/_select_unlocked() — on a freshly loaded pool that selects
   the NEXT healthy key and benches it for the full cooldown TTL,
   punishing an innocent credential. When a hint is provided but
   unmatched, rotate without marking anything instead of guessing.

Includes regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 09:22:02 -07:00
Teknium 1f07fae6df perf(credential-pool): persist same-key sibling exhaustion once
Follow-up to the #68565 salvage: batch the sibling _mark_exhausted calls
behind a single _persist() instead of one auth.json write per sibling.
2026-07-23 09:08:28 -07:00
李航 0e15805e25 fix(credential-pool): exhaust all entries sharing a failed API key on 402
A 402/429/401 is an API-key–level failure (account out of balance,
rate-limited, or key rejected), but the same key can back more than one
pool entry — e.g. an explicit pool entry plus a `model_config` entry
auto-seeded from `model.api_key`, both carrying the identical
`runtime_api_key`.

`mark_exhausted_and_rotate(api_key_hint=...)` only marked the *first*
matching entry, leaving the sibling OK. `_select_unlocked()` then kept
handing back the same depleted key, so the billing-recovery `continue`
loop in the conversation retry path never converged: the request hung
until the client disconnected (~2.5min observed against DeepSeek),
emitting only `response.created` with no 402 ever surfaced to the user.

Mark every entry sharing the failed key so the pool can reach the
"no available entries" state and let the error propagate immediately.

Adds a regression test covering two entries backed by the same key.
2026-07-23 09:08:28 -07:00
Teknium fea838c9f2 fix(auth): detect upstream Codex quota resets and lift stale pool cooldowns (#69494)
When Codex returns 429 usage_limit_reached, Hermes persists the provider's
reset_at on the pool entry and freezes the credential until it elapses --
which can be days out for weekly windows. But the upstream window can
reopen EARLY: the user redeems a banked rate-limit reset (Codex CLI /
ChatGPT UI), upgrades their plan, or OpenAI resets the window. Hermes
never re-checked, so it kept erroring with 'Codex provider quota
exhausted (429); retry after Ns' until a manual re-auth rewrote the
tokens (issue #43747, externally-reset variant).

- hermes_cli/auth.py: add _probe_codex_quota_restored() -- a throttled
  (5 min/token) GET of the Codex /usage endpoint; quota counts as
  restored when every reported window is <100% used. Add
  clear_codex_pool_quota_cooldowns() to lift 429/quota-shaped cooldowns
  from persisted pool entries (DEAD and auth-shaped entries untouched).
- resolve_codex_runtime_credentials(): before surfacing a pool-only
  cooldown as 'quota exhausted', probe upstream; on a positive probe
  clear the cooldown and return the pool credential.
- agent/credential_pool.py: _available_entries() probes frozen
  openai-codex entries (clear_expired path only) and unfreezes them when
  upstream confirms the reset.
- agent/account_usage.py: a successful /usage reset redemption now
  clears persisted pool cooldowns immediately.

Negative paths preserved: probe 429/exhausted/indeterminate keeps the
cooldown; read-only enumeration never probes; non-JWT tokens never
probe (no network in hermetic tests).
2026-07-22 10:58:22 -07:00
izumi0uu 0583692c2d fix(secrets): scope BWS-injected provider keys
Snapshot values applied by external secret sources per resolved HERMES_HOME so a later profile cannot replace an earlier profile scope through shared os.environ.

Keep provider and credential-pool fallback reads on the active secret scope, and fail closed on unscoped multiplex reads.

Tests: scripts/run_tests.sh tests/test_env_loader_secret_sources.py tests/test_env_loader_op_bootstrap.py tests/agent/test_secret_scope.py tests/agent/test_credential_pool.py tests/tools/test_credential_pool_env_fallback.py tests/hermes_cli/test_xiaomi_provider.py tests/cron/test_run_one_job.py tests/hermes_cli/test_api_key_providers.py tests/gateway/test_multiplex_credential_isolation.py -q (395 passed)
2026-07-22 04:39:17 -07:00
Brooklyn Nicholson 9b8b054c2d perf: fast model picker + dialogs — config-load hot path, model.options off the reader thread, off-screen turns skip rendering
Third profiling round (after #66033 / #66347), targeting the composer
model picker and dialog opens (worktree dialog etc.), measured over CDP
on real 1000+-message sessions.

Backend — model.options took 4.8s cold / 1.8s warm per call, and the
desktop model pill/picker blocks on it every open:

- agent/credential_pool: _load_config_safe uses load_config_readonly().
  Every consumer only reads, and the per-call deepcopy was the dominant
  cost — list_authenticated_providers calls load_pool() per provider
  row, and each load_pool loaded (and deep-copied) the full config
  again via get_pool_strategy.
- hermes_cli/config: memoize ensure_hermes_home() per home path. It
  runs inside the config lock on EVERY load_config(), paying ~14
  mkdir/chmod syscalls per call. The fast path still re-checks that the
  home dir exists, so a deleted home is recreated as before; profile
  switches hit the new path and re-run. Tests cover both.
- tui_gateway/server: add model.options to _LONG_HANDLERS. It measured
  seconds inline on the WS reader thread — while it ran, prompt.submit
  and session.interrupt sat unread (same class as #21123).

Together: model.options RPC 4825/1842ms → 426/230ms (measured on the
live desktop backend); build_models_payload in isolation 6.2s → 0.97s
cold, 0.27s warm.

Desktop — every Radix dialog/popover open forced a whole-document style
recalc (Presence reads getComputedStyle on mount), which on a
1300-message transcript cost ~650-730ms per open (CPU profile:
getAnimationName 483ms self). The worktree dialog (⌘⇧B) paid it on
every single open:

- thread/list: content-visibility:auto + contain-intrinsic-size on the
  per-turn group wrappers. Off-screen turns now skip style recalc,
  layout, and paint entirely; never-rendered turns hold a placeholder
  height (auto: remembered real size once rendered) so scrollbar and
  anchoring stay stable. Verified over CDP: worktree dialog open 656-
  730ms → ~200ms on the same session; stick-to-bottom pin, scroll-to-
  top rendering, and sticky human bubbles all intact.

Also: profile-session-switch harness accepts CDP_HTTP (Chrome tends to
squat on 9222).

Verification:
- scripts/run_tests.sh: config, credential-pool, inventory,
  model-switch routing, tui_gateway protocol, profiles suites green
  (test_profiles has one pre-existing failure on main, unrelated);
  new tests for the ensure_hermes_home memo.
- apps/desktop: tsc clean, eslint/prettier clean, thread + session
  suites green (326 tests).
- E2E over CDP on the live app: numbers above, plus scroll/pin sanity.
2026-07-17 15:01:54 -04:00
HexLab 594308d4bb fix(credential-pool): throttle "no available entries" log to stop Windows log-lock storm (contributes to #62698) (#66338)
* fix(credential-pool): throttle "no available entries" log to stop Windows log-lock storm

Credential selection runs on a hot path (every model call plus auxiliary
tasks), so an empty/exhausted pool logged "no available entries" at INFO on
*every* selection. On Windows, where multiple Hermes processes share one
rotating log guarded by concurrent-log-handler's cross-process lock, that
per-selection volume storms the lock (RuntimeError: Cannot acquire lock after
20 attempts), pegs a core, and stalls the asyncio event loop long enough that
the Desktop backend readiness probe times out ("Timed out connecting to Hermes
backend after 15000ms") even though the backend already announced
HERMES_BACKEND_READY.

Log the condition at most once per 60s window, re-arming on a successful
selection so recovery->re-exhaustion still surfaces promptly. Same fix class as
the warn-once dedup in #58265.

* test(credential-pool): cover no-available-entries log throttle

Assert the empty-pool INFO line logs at most once per throttle window, logs
again after the window elapses, and re-arms on a successful selection so a
recover->re-exhaust transition surfaces promptly. Uses a deterministic fake
monotonic clock (no sleeps, no network).
2026-07-17 13:08:46 -04:00