fix: sibling Nous 401 recovery adopts a peer's refresh instead of rotating again

The stampede fix added a `stale_access_token` hint to
resolve_nous_runtime_credentials() so a process whose bearer just 401'd
adopts a token a sibling already rotated instead of re-POSTing the shared
grant — but only the credential-pool caller passed it. The main agent's
401 path (run_agent._try_refresh_nous_client_credentials), the auxiliary
client rebuild, and the proxy adapter all called force_refresh=True with
no hint, so `_already_rotated_by_peer` could never fire: N subagents
hitting hourly expiry still issued N serialized refreshes, each one
invalidating the token a sibling had just adopted.

Live 12-process A/B against a fake Portal: 12 refresh POSTs / 9 distinct
final tokens before, 1 POST / 1 token after.
This commit is contained in:
Teknium
2026-09-02 21:01:11 -07:00
parent 97f3229dfd
commit 116ca1db1e
5 changed files with 46 additions and 5 deletions
+3
View File
@@ -6423,9 +6423,12 @@ class AIAgent:
try:
from hermes_cli.auth import resolve_nous_runtime_credentials
# Pass the bearer that just 401'd so a refresh already done by a
# sibling process is adopted instead of rotating the grant again.
creds = resolve_nous_runtime_credentials(
timeout_seconds=env_float("HERMES_NOUS_TIMEOUT_SECONDS", 15),
force_refresh=force,
stale_access_token=self.api_key or None,
)
except Exception as exc:
logger.debug("Nous credential refresh failed: %s", exc)