fix(gateway): keep the drain-marker read off the event loop
_drain_control_watcher polls drain_requested() once a second, and that does a synchronous Path.read_text() on the marker file. At a 1s cadence that is ~86k blocking disk reads a day sitting on the event loop. It stays invisible until the host is under I/O pressure, at which point a single read can stall for 30s+ and take every platform heartbeat down with it. Observed as repeated "discord.gateway: Shard ID None heartbeat blocked for more than 30 seconds" followed by the Discord websocket dropping, with the gateway still reporting active. Move the call to asyncio.to_thread so the read blocks a worker thread instead of the loop. Semantics are unchanged: drain_control never raises, the surrounding try/except still covers it, and the sequential await means at most one probe is in flight at a time.
This commit is contained in:
+6
-1
@@ -9876,7 +9876,12 @@ class GatewayRunner(GatewayAuthorizationMixin, GatewayKanbanWatchersMixin, Gatew
|
||||
|
||||
while self._running:
|
||||
try:
|
||||
if drain_requested():
|
||||
# drain_requested() does a synchronous read_text() on the
|
||||
# marker file. At this 1s cadence that puts a blocking disk
|
||||
# read on the event loop ~86k times a day; when the host is
|
||||
# under I/O pressure a single read can stall for 30s+ and
|
||||
# take every platform heartbeat down with it. Off-thread it.
|
||||
if await asyncio.to_thread(drain_requested):
|
||||
self._enter_external_drain()
|
||||
# API and cron work live outside messaging's
|
||||
# _running_agents map. Refresh the aggregate while an
|
||||
|
||||
Reference in New Issue
Block a user