Commit Graph

1 Commits

Author SHA1 Message Date
joaomarcos e8c41568ca fix(state): bound state.db read connections per FILE, and stop opening two gateway handles
Issue #98573 reports a long-lived gateway holding ~20 `state.db` descriptors
that never shrink, walking into the 256 soft RLIMIT_NOFILE a launchd/systemd
service manager hands the process. The cause named there — a per-thread
`threading.local()` read connection — is already gone (87aedbe7b6 pooled the
read connections, 0472c31aa1 added the peak permit). Measured on main: one
SessionDB with 40 concurrent reader threads peaks at 9 live connections, not 40.

The symptom survives one layer up. `_READ_POOL_MAX` was enforced by a
BoundedSemaphore owned by each SessionDB, which bounds the wrong noun: the
descriptors are spent on a FILE, so every additional handle on one state.db got
its own allowance and peak scaled as `instances x (1 + _READ_POOL_MAX)`.

Two changes, both needed:

* The permits move to a per-path `_PathReadBudget`, shared by every SessionDB
  in the process that points at that file. A permit miss first reclaims an IDLE
  pooled connection from a peer handle before degrading to the writer lock —
  without that, whichever handle warmed up first would pin the whole budget and
  permanently demote every later one (a cron job's transient handle, a second
  profile's store) to the locked writer connection.

* `GatewayRunner` borrows `SessionStore`'s handle instead of opening its own.
  Both caches resolve the same `_default_db_path()`, so the process was holding
  two writer connections and two read pools against one file for no reason, and
  doubling again per profile on a multiplexed gateway. The store owns the
  connection and sweeps it at shutdown; the runner's cache now holds only the
  async wrapper and its sweep skips borrowed handles.

Measured peak live connections against one file, 40 reader threads, by handle
count 1/2/4/8:

  before: 9 / 18 / 36 / 51   (51 not 72 only because the sample window ended
                              before every pool filled)
  after:  9 / 10 / 12 / 16   (read connections capped at 8 in total; the
                              remainder is one writer per handle, and the
                              gateway's per-profile pair is now one)

Fixes #98573

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 11:23:48 -07:00