fix(auth): make hermes auth reset actually clear a binding cooldown

Port onto the split layout: persist_pool_entries now forwards
status_cleared_ids to BOTH writers — write_credential_pool (hermes_cli/
auth.py) and the root-store row merge for single-use-refresh providers
(_update_root_pool_rows), which the original fix predates. Both skip the
disk recency merge for deliberately-cleared entries; the kwarg is only
sent when non-empty so upstream fakes with the old signature keep
working. reset_statuses clears failure_reason (lives in extra, replace()
cannot reach it) alongside the status fields.

tests/agent/test_credential_pool.py: 62 passed; refresh-race suite: 3
passed (mutation: the reset/binding tests fail with the fix stashed).
The dashboard-auth-gate / user-providers failures are pre-existing on
pristine upstream/main (reproduced with changes stashed).
This commit is contained in:
rodrigogs
2026-09-04 23:13:16 -03:00
committed by kshitij
parent 2e24e06e55
commit 8378551311
3 changed files with 234 additions and 17 deletions
+20 -4
View File
@@ -924,13 +924,23 @@ def _entry_ids(entries: Iterable[Any]) -> Dict[str, Dict[str, Any]]:
def write_credential_pool(
provider_id: str, entries: List[Dict[str, Any]], *, removed_ids: Optional[Iterable[str]] = None,
provider_id: str, entries: List[Dict[str, Any]], *,
removed_ids: Optional[Iterable[str]] = None,
status_cleared_ids: Optional[Iterable[str]] = None,
) -> Path:
"""Persist one provider's credential pool under auth.json.
Final disk-boundary sanitizer for borrowed credentials (callers may pass raw dicts). Entries on
disk but missing from *entries* (added concurrently) are merged back unless in *removed_ids*,
so a rotation/exhaustion rewrite never drops a concurrent credential."""
so a rotation/exhaustion rewrite never drops a concurrent credential.
Pass ``status_cleared_ids`` for entries whose status the caller intentionally
cleared — the same problem one step further in. The recency merge cannot tell
an operator's ``hermes auth reset`` from a stale snapshot: a clear sets
``last_status_at`` to None, which compares as epoch 0 and so always loses to
the on-disk timestamp, and the cooldown was copied straight back. Declaring
the intent is what separates "I have not seen the newer status" from "I have
seen it and I am dropping it"."""
removed = {rid for rid in (removed_ids or ()) if rid}
with _auth_store_lock():
auth_store = _load_auth_store()
@@ -942,9 +952,15 @@ def write_credential_pool(
existing_list = existing_list if isinstance(existing_list, list) else []
existing_by_id = _entry_ids(existing_list)
new_ids = set(_entry_ids(sanitized))
status_cleared = {cid for cid in (status_cleared_ids or ()) if cid}
merged: List[Dict[str, Any]] = [
_merge_disk_cooldown_state(e, existing_by_id.get(e.get("id")), provider_id)
if isinstance(e, dict) else e
e
if isinstance(e, dict) and e.get("id") in status_cleared
else (
_merge_disk_cooldown_state(e, existing_by_id.get(e.get("id")), provider_id)
if isinstance(e, dict)
else e
)
for e in sanitized]
for disk_entry in existing_list:
disk_id = disk_entry.get("id") if isinstance(disk_entry, dict) else None