Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Independent review of the first fix found a credential-identity takeover:
_adopt_nous_key_before_expiry() called the singleton resolver
unconditionally, so an agent running on an explicitly supplied (or
pool-selected) account-A key near expiry was moved onto the logged-in
account-B key from auth.json before the A key had even failed, and the next
real SDK request went out as B. That silently changes who is billed.
The adoption now reads the `sub` claim of the key in hand and passes it as
require_account; _try_refresh_nous_client_credentials refuses any
replacement whose `sub` differs (logged at INFO, current key kept). A key
with no `sub` is not adopted proactively at all. The reactive 401 path is
unchanged, and same-account adoption (the keepalive's or a peer's fresh key)
still works.
Verified with the reviewer's own live probe (real AIAgent -> real
prepare_iteration -> real auth.json transaction -> real OpenAI SDK ->
loopback capture): explicit-account case account-A -> account-A,
adopted_fresh=False (was A -> B); same-account still adopts; 12 concurrent
near-expiry agents still 0 x 401.
Tests (2 new): a fresh key for a different account is never adopted; a key
without an account claim is left alone without touching the store.
The Nous agent key lives 3,599 s. In the 1,393-agent refactor run every
in-process agent learned about the hourly expiry from its own 401: 620
authentication_error 401s in the logged window (177 in one hour), each a
failed attempt the model never saw, and the credential pool benched the
sole credential for all of them at once. At 08:30 the storm took the
parent process down.
Two gaps. The proactive refresher (hermes_cli/nous_auth_keepalive.py) is
started only by the gateway and the web server; the CLI process, and every
subagent built inside it, never started it. And even with a fresh key in
the store, nothing adopted it before a request: a request went out with
whatever key the agent was constructed with until it 401'd.
Now _finalize_routing starts the keepalive (idempotent, process-wide,
daemon) whenever an agent resolves onto provider "nous", and
prepare_iteration calls _adopt_nous_key_before_expiry(): the agent key is
a JWT, its exp is read locally, and inside a 180 s skew the store is
re-read under the auth-store lock with force_refresh=False, so the
keepalive's (or a peer's) fresh key is adopted without a POST; when none
exists, ONE refresh runs there instead of N reactive ones after N 401s.
_try_refresh_nous_client_credentials no longer rebuilds the client when
the store returns the key already in hand.
Live A/B (local server: 401s any bearer but FRESH; store patched to hold
FRESH; 40 agents holding a JWT that expires in 30 s fire concurrently):
main 40 x 401 then recover, branch 0 x 401.
Tests (4): far from expiry the store is not touched; inside the skew the
store's fresh key is adopted with force_refresh=False; the same key back
from the store is not re-adopted; a real AIAgent routed to nous starts
the keepalive and one routed to openrouter does not.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
Adolanium review §3 (Medium/Low). BASE shape at every steer/redirect/persist/agent-cache
lock site was: lock = getattr(obj, "_x_lock", None); if lock is not None: with lock: <direct
attribute read>; else: <getattr fallback for object.__new__ test stubs>. The refactor rewrote
several of these as 'with getattr(...) or nullcontext(): getattr(slot, None)', which (a) turned a
missing slot under the lock from a loud AttributeError into silent None and (b) in the gateway
peek helper read the cache without any lock when the lock attribute was absent.
Restored BASE semantics at:
- agent/agent_runtime_helpers.py::_requeue_pending_steer
- agent/interrupt_control.py: steer, redirect, clear_interrupt, _has_pending_redirect,
_drain_pending_redirect, _drain_pending_steer (new _ic_slot helper: direct read under lock)
- agent/session_persistence.py::_persist_lock (explicit None check + BASE rationale)
- gateway/slash_commands.py::_cached_agent_for (BASE callers read ONLY under the lock; no lock -> None)
- gateway/run_agent_cache.py::_evict_cached_agent (BASE: self._agent_cache direct under lock)
- agent/client_lifecycle.py::_is_openai_client_closed: BASE body + docstring verbatim (outer
is_closed first; inner _client.is_closed only when _client exists; else False)
- agent/stream_delivery.py::_ensure_stream_writer_state: restore BASE rationale that the lock is
created unconditionally in agent_init (_STREAM_STATE) and the lazy path is stub-only
A/B (/tmp/rf/rev/ab_lock_fallbacks.py, ab_client_closed.py) is byte-identical BASE vs HEAD on
the reviewer's inputs + edge cases. tests/agent/test_lock_fallback_base_semantics.py pins it
(7 of 25 cases fail on the pre-fix tree).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.