fix(gateway): keep the drain-marker read off the event loop

_drain_control_watcher polls drain_requested() once a second, and that
does a synchronous Path.read_text() on the marker file. At a 1s cadence
that is ~86k blocking disk reads a day sitting on the event loop.

It stays invisible until the host is under I/O pressure, at which point
a single read can stall for 30s+ and take every platform heartbeat down
with it. Observed as repeated "discord.gateway: Shard ID None heartbeat
blocked for more than 30 seconds" followed by the Discord websocket
dropping, with the gateway still reporting active.

Move the call to asyncio.to_thread so the read blocks a worker thread
instead of the loop. Semantics are unchanged: drain_control never
raises, the surrounding try/except still covers it, and the sequential
await means at most one probe is in flight at a time.
This commit is contained in:
RFingAdam
2026-08-31 20:43:47 -04:00
committed by kshitij
parent 04224b2f82
commit b6f42c667a
+6 -1
View File
@@ -9876,7 +9876,12 @@ class GatewayRunner(GatewayAuthorizationMixin, GatewayKanbanWatchersMixin, Gatew
while self._running:
try:
if drain_requested():
# drain_requested() does a synchronous read_text() on the
# marker file. At this 1s cadence that puts a blocking disk
# read on the event loop ~86k times a day; when the host is
# under I/O pressure a single read can stall for 30s+ and
# take every platform heartbeat down with it. Off-thread it.
if await asyncio.to_thread(drain_requested):
self._enter_external_drain()
# API and cron work live outside messaging's
# _running_agents map. Refresh the aggregate while an