1608a46884
Widen the #70773 fix beyond the three in-request cleanup sites removed in the cherry-picked commit: every remaining path that swaps out the shared OpenAI client could still hard-close its pool from a thread that doesn't own the in-flight sockets (credential rotation/refresh on the turn thread, dead-connection cleanup, gateway cache eviction, transport recovery) — the same FD-recycle corruption vector, just rarer. Add AIAgent._retire_shared_openai_client(): shutdown(SHUT_RDWR) all pooled sockets (FD-safe from any thread, unblocks in-flight readers) but never call client.close() — FD release is deferred to GC, which cannot run until every borrowing thread has unwound its SSL BIO. Refcounting is the ownership handshake; with no borrowers the FDs are released immediately. Wired into: - _replace_primary_openai_client (rotation/refresh/dead-conn cleanup) - try_recover_primary_transport (primary_recovery) - release_clients (gateway cache_evict) agent.close() keeps the hard close: full teardown is a real session boundary where no request may be in flight. Tests: new tests/run_agent/test_70773_shared_client_fd_corruption.py covers the three watchdog/retry sites plus retire semantics; existing close-assertions updated to pin retire-not-close.