fix(cron): a raising profile gate ticks nothing; housekeeping respawns a dead ticker (#111010)
Follow-up on the salvaged #111034: - `_start_multiplex` published the enumerated home list before the gate had filtered it, so a raising `profile_gate` (the Desktop stand-down probe from #100489) kept the thread alive but ticked every profile UNGATED — racing the gateway that owns them for the same cron store. The list is now assigned only after gating; a gate failure yields zero ticks for that cycle. - `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor" chore that respawns a ticker that ended without a stop request and logs the outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing outside it could notice a thread that had already ended. - Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt `executions.db` (no patched recover) no longer kills the ticker; a raising gate keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and leaves a stopped one alone. Root cause: the unguarded pre-loop `recover_interrupted_executions()` + `record_ticker_heartbeat()` were added byd9dd05b69d(#61791, "truthful execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried #107485's `completed_occurrence()` in the due scan, which opens the same ledger on every tick — the first traceback in their errors.log is that in-loop hit (caught); the restart then hit the SAME corrupt ledger from the pre-loop recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
This commit is contained in:
@@ -543,16 +543,15 @@ class InProcessCronScheduler(CronScheduler):
|
||||
# See #87644.
|
||||
_cycle_exc: BaseException | None = None
|
||||
# Enumeration and gating run on the ticker thread; a raising gate callable must
|
||||
# fail THIS cycle (logged, no heartbeats), not end the thread (#111010).
|
||||
# fail THIS cycle (logged, no heartbeats, NO ticks), not end the thread (#111010).
|
||||
# Publish the list only once the gate has filtered it: a partial assignment would
|
||||
# tick the ungated set — the exact stand-down the Desktop gate exists for (#100489).
|
||||
cycle_homes: list = []
|
||||
try:
|
||||
cycle_homes = [
|
||||
_profile_entry(e) for e in _existing_profile_homes(profile_homes)
|
||||
]
|
||||
enumerated = [_profile_entry(e) for e in _existing_profile_homes(profile_homes)]
|
||||
if profile_gate is not None:
|
||||
cycle_homes = [
|
||||
(name, home) for name, home in cycle_homes if profile_gate(name, home)
|
||||
]
|
||||
enumerated = [(name, home) for name, home in enumerated if profile_gate(name, home)]
|
||||
cycle_homes = enumerated
|
||||
except BaseException as e:
|
||||
logger.error("Cron profile enumeration error: %s", e, exc_info=True)
|
||||
_tick_error = f"{type(e).__name__}: {e}"
|
||||
|
||||
@@ -0,0 +1,53 @@
|
||||
"""Supervised daemon thread for the in-process cron ticker.
|
||||
|
||||
The gateway (and the Desktop backend) run ``InProcessCronScheduler.start`` on a bare daemon thread.
|
||||
Every guard inside ``start`` keeps the loop alive on a per-tick failure, but nothing outside it could
|
||||
notice a thread that had already ended — the gateway kept serving while ``ticker_heartbeat`` froze and
|
||||
no job fired again until a restart (#111010). The supervisor is the missing outer layer: the
|
||||
housekeeping loop asks it once a cycle to respawn a ticker that died while shutdown was not requested.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import threading
|
||||
from typing import Any, Callable, Mapping
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class SupervisedTickerThread:
|
||||
"""``threading.Thread``-shaped handle whose ``restart_if_dead`` respawns a dead ticker."""
|
||||
|
||||
def __init__(self, target: Callable[..., Any], *, args: tuple = (), kwargs: Mapping[str, Any] | None = None,
|
||||
stop_event: threading.Event, name: str = "cron-scheduler") -> None:
|
||||
self._target, self._args, self._kwargs = target, args, dict(kwargs or {})
|
||||
self._stop_event, self._name = stop_event, name
|
||||
self._thread = self._spawn()
|
||||
self.restarts = 0
|
||||
|
||||
def _spawn(self) -> threading.Thread:
|
||||
return threading.Thread(target=self._target, args=self._args, kwargs=self._kwargs, daemon=True, name=self._name)
|
||||
|
||||
def start(self) -> None:
|
||||
self._thread.start()
|
||||
|
||||
def is_alive(self) -> bool:
|
||||
return self._thread.is_alive()
|
||||
|
||||
def join(self, timeout: float | None = None) -> None:
|
||||
self._thread.join(timeout)
|
||||
|
||||
def restart_if_dead(self) -> bool:
|
||||
"""Respawn the ticker when it ended without ``stop_event``; True when a restart happened."""
|
||||
if self._stop_event.is_set() or self._thread.is_alive():
|
||||
return False
|
||||
self.restarts += 1
|
||||
logger.error(
|
||||
"Cron ticker thread %r died without a stop request; restarting (restart #%d). Scheduled jobs "
|
||||
"did not fire while it was down — check errors.log for the escaping exception.",
|
||||
self._name, self.restarts,
|
||||
)
|
||||
self._thread = self._spawn()
|
||||
self._thread.start()
|
||||
return True
|
||||
Reference in New Issue
Block a user