hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
Follow-up to the per-file budget in the previous commit, which closed the
scaling axis it measured and left three others open.
* A per-file ceiling still lets the cost grow with the PROFILE count: a
multiplexed gateway serves N profiles from one process and each has its own
state.db, so `_READ_POOL_MAX` bounded each file while the process total went
unbounded — the per-instance bug one level out. `_READ_POOL_PROCESS_MAX`
(three files' worth) now bounds the process, and a miss reclaims an idle
connection from ANY path before degrading: a profile quiet for an hour must
not hold descriptors the profile being served right now needs.
* Hermes's SQLite descriptors are only ever a share of the fd table. In #98573
the ~20 state.db handles were not the whole 256 — they were the share that
pushed httpx sockets and terminal subprocess pipes over, and the EMFILE
surfaced in tools/terminal_tool.py rather than here. New read connections are
now refused when the process is within `_FD_HEADROOM_RESERVE` of its soft
RLIMIT_NOFILE, measured from /proc/self/fd or /dev/fd and cached briefly. The
guard fails OPEN where it cannot measure (Windows has neither the fd
directory nor RLIMIT_NOFILE, and a CRT limit in the thousands) and CLOSED on
evidence — including a probe that could not get a descriptor of its own.
`_read_open_denied_fd_headroom` makes it diagnosable from a running process.
* Writer connections cannot be rationed the way read connections can: a
SessionDB without one cannot write. Their only real bound is not opening
redundant handles, so a process that accumulates more than
`_HANDLES_PER_PATH_WARN` handles on one file now says so once, and the next
duplicate is visible before it is an incident instead of inferred from an
lsof after one.
`_READ_POOL_MAX` itself is deliberately unchanged at 8. Retuning that constant
is #98585's subject; with a process ceiling above it and the headroom guard
in front of it, the value is no longer the binding constraint.
Refs #98573
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Issue #98573 reports a long-lived gateway holding ~20 `state.db` descriptors
that never shrink, walking into the 256 soft RLIMIT_NOFILE a launchd/systemd
service manager hands the process. The cause named there — a per-thread
`threading.local()` read connection — is already gone (87aedbe7b6 pooled the
read connections, 0472c31aa1 added the peak permit). Measured on main: one
SessionDB with 40 concurrent reader threads peaks at 9 live connections, not 40.
The symptom survives one layer up. `_READ_POOL_MAX` was enforced by a
BoundedSemaphore owned by each SessionDB, which bounds the wrong noun: the
descriptors are spent on a FILE, so every additional handle on one state.db got
its own allowance and peak scaled as `instances x (1 + _READ_POOL_MAX)`.
Two changes, both needed:
* The permits move to a per-path `_PathReadBudget`, shared by every SessionDB
in the process that points at that file. A permit miss first reclaims an IDLE
pooled connection from a peer handle before degrading to the writer lock —
without that, whichever handle warmed up first would pin the whole budget and
permanently demote every later one (a cron job's transient handle, a second
profile's store) to the locked writer connection.
* `GatewayRunner` borrows `SessionStore`'s handle instead of opening its own.
Both caches resolve the same `_default_db_path()`, so the process was holding
two writer connections and two read pools against one file for no reason, and
doubling again per profile on a multiplexed gateway. The store owns the
connection and sweeps it at shutdown; the runner's cache now holds only the
async wrapper and its sweep skips borrowed handles.
Measured peak live connections against one file, 40 reader threads, by handle
count 1/2/4/8:
before: 9 / 18 / 36 / 51 (51 not 72 only because the sample window ended
before every pool filled)
after: 9 / 10 / 12 / 16 (read connections capped at 8 in total; the
remainder is one writer per handle, and the
gateway's per-profile pair is now one)
Fixes#98573
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review of this PR was right that maxsize=8 bounds the wrong thing. The
LifoQueue caps how many connections are RETURNED; _checkout_read_conn opened
unconditionally on a miss, so N readers arriving on a cold pool all missed, all
opened, and peaked at N. The surplus was closed on release, so nothing
accumulated forever -- but EMFILE is a peak-instant condition and the burst
that empties the pool is exactly the burst that exhausts the fd table, so the
original wedge was still reachable. Measured on the previous commit: 64
concurrent readers held 64 live connections at once.
A connection now holds a permit for its whole lifetime -- acquired in
_get_read_conn() before the open, released in _close_read_conn() after the
close -- so open+checked-out is bounded together. A pool hit costs no permit
because the connection it hands back already holds one, which leaves
_get_read_conn() as the only place that can open. The acquire is non-blocking:
past the ceiling readers fall back to the locked writer connection rather than
queueing, since blocking would convert descriptor exhaustion into a stall,
which is the same outage with a different stack trace. Same burst now peaks at
8. BoundedSemaphore rather than Semaphore so an unpaired release raises instead
of silently widening the ceiling.
Two latent leaks in the same function, found while doing this:
- a CJK extension load that failed after a successful open returned None
without closing the connection, leaking a descriptor the tracking registry
still counted -- the same leak shape one level down;
- any non-sqlite3.Error between open and return stranded a permit
permanently, which would ratchet the ceiling down to zero and silently
demote every later read to the writer lock.
On the test: the existing one joins every worker before counting, so it
measures the pool at rest and structurally cannot observe peak -- which is why
this got through. The new one uses a barrier so all 64 workers hold their
connections until every worker has checked out, making the count taken at that
moment the actual simultaneous peak. Verified it fails against the previous
commit (64 checked out, 65 live) and passes at 8/9. Also covers the
writer-connection fallback, permit recovery after a failed open, and that
close() releases exactly the permits it drained.