bb597e1c02
* fix(gateway): pass live adapters to cron fire webhook's fire_due
The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).
Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.
Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).
* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)
The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.
Restore the invariant that the GATEWAY owns cron execution:
- Dashboard route: after verifying the NAS JWT and resolving the job's
profile, FORWARD the fire to the gateway api_server's own
/api/cron/fire on loopback, NAS bearer preserved (the gateway
re-verifies the JWT — defense in depth, no new trust link), and pass
the gateway's response through. Gateway unreachable → 503 so NAS
retries per the Chronos contract (non-2xx = retryable; the store CAS
de-dupes the eventual double fire). Deliberately NO local-execution
fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
per target profile (config.yaml extra.port → API_SERVER_PORT from
process env or the profile's .env → 8642), with /p/<profile>/ prefix
routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
first boot when absent (never overwrites an operator value), so the
loopback api_server passes its startup guard on hosted images. The
fire route itself is NAS-JWT-authed; the key gates the rest of the
api_server surface. The listener binds 127.0.0.1 by default and the
Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
topology and the 503-retry semantics.
Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.
* fix(cron): read the profile api_server port via the canonical config loader
CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).
Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.
* fix(gateway): only messaging platforms count for the scale-to-zero arm gate
The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).
The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
373 lines
14 KiB
Python
373 lines
14 KiB
Python
"""Watcher-level tests for scale-to-zero: the idle watcher's dormant sequence and
|
|
the arm-gate wiring, exercised against the real GatewayRunner methods bound onto
|
|
a lightweight stand-in (booting a full gateway is unnecessary for this logic and
|
|
would be slow/flaky).
|
|
|
|
These cover the parts gateway/test_scale_to_zero.py (pure helpers) can't: that
|
|
the watcher calls the relay adapter's go_dormant() exactly when idle+armed,
|
|
respects the cooldown, and skips when busy — the F7/D3 + D12 behaviour.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import time
|
|
|
|
import pytest
|
|
|
|
from gateway.run import GatewayRunner
|
|
|
|
|
|
class _FakeRelayAdapter:
|
|
def __init__(self):
|
|
self.go_dormant_calls = 0
|
|
|
|
async def go_dormant(self):
|
|
self.go_dormant_calls += 1
|
|
return True
|
|
|
|
|
|
def _runner_with(monkeypatch, *, idle, armed_adapter=True):
|
|
"""Build a GatewayRunner without booting it, stubbing just what the watcher
|
|
touches. Real methods (_scale_to_zero_is_idle composition, the watcher body)
|
|
run; only their dependencies are stubbed."""
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
r._running = True
|
|
r._scale_to_zero_cooldown_until = 0.0
|
|
r._last_inbound_at = time.time()
|
|
r._running_agents = {}
|
|
r._background_tasks = set()
|
|
adapter = _FakeRelayAdapter() if armed_adapter else None
|
|
|
|
monkeypatch.setattr(r, "_scale_to_zero_is_idle", lambda: idle, raising=False)
|
|
monkeypatch.setattr(r, "_relay_adapter_for_dormancy", lambda: adapter, raising=False)
|
|
monkeypatch.setattr(r, "_scale_to_zero_idle_timeout_seconds", lambda: 300.0, raising=False)
|
|
monkeypatch.setattr(r, "_update_runtime_status", lambda *a, **k: None, raising=False)
|
|
return r, adapter
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_watcher_goes_dormant_when_idle(monkeypatch):
|
|
r, adapter = _runner_with(monkeypatch, idle=True)
|
|
# Run one iteration: stop after the first sleep so the loop exits cleanly.
|
|
task = asyncio.create_task(r._scale_to_zero_watcher(interval=0.01))
|
|
await asyncio.sleep(0.1)
|
|
r._running = False
|
|
await asyncio.wait_for(task, timeout=2)
|
|
assert adapter.go_dormant_calls >= 1
|
|
# After driving dormant, a re-arm cooldown is set (0.F).
|
|
assert r._scale_to_zero_cooldown_until > time.time()
|
|
|
|
|
|
# No exception, loop exits cleanly — nothing to assert beyond survival.
|
|
|
|
|
|
def test_bg_work_blocks_idle_via_background_tasks(monkeypatch):
|
|
"""_scale_to_zero_has_live_background_work() reports True when a tracked
|
|
background task is still live (D3/F7) — the guard that keeps a gateway with
|
|
an in-flight backgrounded subagent/terminal awake."""
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
|
|
async def _never():
|
|
await asyncio.sleep(0.2)
|
|
|
|
loop = asyncio.new_event_loop()
|
|
try:
|
|
t = loop.create_task(_never())
|
|
r._background_tasks = {t}
|
|
# process_registry has nothing active in this fresh process.
|
|
assert r._scale_to_zero_has_live_background_work() is True
|
|
t.cancel()
|
|
finally:
|
|
loop.run_until_complete(asyncio.gather(t, return_exceptions=True))
|
|
loop.close()
|
|
|
|
|
|
def test_real_inbound_after_dormancy_restores_running_status(monkeypatch):
|
|
"""Once a dormant gateway receives real inbound after wake, the runtime
|
|
lifecycle must not remain stuck in the watcher-written `draining` state."""
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
r._last_inbound_at = 0.0
|
|
r._scale_to_zero_cooldown_until = time.time() + 60.0
|
|
status_updates = []
|
|
monkeypatch.setattr(
|
|
r,
|
|
"_update_runtime_status",
|
|
lambda state=None, *a, **k: status_updates.append(state),
|
|
raising=False,
|
|
)
|
|
|
|
r._scale_to_zero_note_real_inbound()
|
|
|
|
assert r._last_inbound_at > 0.0
|
|
assert status_updates == ["running"]
|
|
|
|
|
|
# ── _scale_to_zero_should_arm: the CALL SITE feeds config.platforms (the F25 bug) ──
|
|
#
|
|
# config.platforms is pre-seeded with a DISABLED placeholder PlatformConfig for every
|
|
# known platform, so list(config.platforms.keys()) is always the full ~20-entry catalog
|
|
# regardless of what the instance runs. The arm check must filter to ENABLED platforms
|
|
# (mirroring the connect loop) before asking messaging_is_relay_only_or_absent — passing
|
|
# the bare placeholder keys made it see disabled `discord`/`telegram`/… as live direct
|
|
# platforms and refuse to arm on a real relay-only instance. The pure-helper tests in
|
|
# test_scale_to_zero.py pass bare names so they never exercised this call site.
|
|
|
|
|
|
def _arm_runner(monkeypatch, platform_states, *, enabled=True, wake_url="https://wake.example"):
|
|
"""Build a GatewayRunner stand-in whose config.platforms mirrors a real load:
|
|
`platform_states` is {Platform: enabled_bool}; everything runs the REAL
|
|
_scale_to_zero_should_arm. Only the env flag + wake_url resolution are stubbed."""
|
|
from types import SimpleNamespace
|
|
|
|
from gateway.config import PlatformConfig
|
|
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
platforms = {p: PlatformConfig(enabled=en) for p, en in platform_states.items()}
|
|
r.config = SimpleNamespace(platforms=platforms)
|
|
|
|
monkeypatch.setattr("gateway.scale_to_zero.scale_to_zero_enabled", lambda *a, **k: enabled)
|
|
monkeypatch.setattr("gateway.relay.relay_wake_url", lambda: wake_url)
|
|
return r
|
|
|
|
|
|
def test_arm_true_for_relay_only_with_disabled_placeholders(monkeypatch):
|
|
"""The F25 regression test: relay ENABLED, every other platform present but
|
|
DISABLED (the real load_gateway_config() shape). Must arm — the disabled
|
|
placeholders must NOT count as live direct-socket platforms."""
|
|
from gateway.platforms.base import Platform
|
|
|
|
r = _arm_runner(
|
|
monkeypatch,
|
|
{
|
|
Platform.TELEGRAM: False,
|
|
Platform.DISCORD: False,
|
|
Platform.SLACK: False,
|
|
Platform.MATRIX: False,
|
|
Platform.RELAY: True,
|
|
},
|
|
)
|
|
assert r._scale_to_zero_should_arm() is True
|
|
|
|
|
|
def test_no_arm_when_a_direct_platform_is_actually_enabled(monkeypatch):
|
|
"""A genuinely-enabled direct-socket platform (real Discord token) DOES disarm —
|
|
the filter must not over-broaden to 'ignore everything but relay'."""
|
|
from gateway.platforms.base import Platform
|
|
|
|
r = _arm_runner(
|
|
monkeypatch,
|
|
{Platform.DISCORD: True, Platform.RELAY: True},
|
|
)
|
|
assert r._scale_to_zero_should_arm() is False
|
|
|
|
|
|
|
|
# ── the self-suspend step: fires only after a clean quiesce, in order ─────────
|
|
#
|
|
# The gateway owns the suspend (Fly Proxy autostop is inbound-only/job-blind and
|
|
# no longer held open by outbound sockets), so the watcher must (a) suspend only
|
|
# AFTER go_dormant succeeded — the relay flip precedes the freeze, closing the
|
|
# buffered-event black hole — and (b) never suspend when the quiesce failed or
|
|
# inbound landed mid-quiesce.
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_watcher_self_suspends_after_dormant(monkeypatch):
|
|
r, adapter = _runner_with(monkeypatch, idle=True)
|
|
calls = []
|
|
|
|
async def fake_suspend():
|
|
calls.append(("suspend", adapter.go_dormant_calls))
|
|
r._running = False # stop the loop after the first full sequence
|
|
|
|
monkeypatch.setattr(r, "_scale_to_zero_self_suspend", fake_suspend, raising=False)
|
|
task = asyncio.create_task(r._scale_to_zero_watcher(interval=0.01))
|
|
await asyncio.wait_for(task, timeout=2)
|
|
# Suspend fired exactly once, and only AFTER go_dormant ran (flip-before-freeze).
|
|
assert calls == [("suspend", 1)]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_watcher_skips_suspend_when_dormant_fails(monkeypatch):
|
|
r, adapter = _runner_with(monkeypatch, idle=True)
|
|
|
|
async def broken_dormant():
|
|
raise RuntimeError("quiesce failed")
|
|
|
|
adapter.go_dormant = broken_dormant
|
|
suspend_calls = []
|
|
|
|
async def fake_suspend():
|
|
suspend_calls.append(1)
|
|
|
|
monkeypatch.setattr(r, "_scale_to_zero_self_suspend", fake_suspend, raising=False)
|
|
task = asyncio.create_task(r._scale_to_zero_watcher(interval=0.01))
|
|
await asyncio.sleep(0.1)
|
|
r._running = False
|
|
await asyncio.wait_for(task, timeout=2)
|
|
# A failed quiesce means an UNFLIPPED relay — suspending would black-hole
|
|
# inbound events. Must stay awake.
|
|
assert suspend_calls == []
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_watcher_skips_suspend_when_inbound_lands_mid_quiesce(monkeypatch):
|
|
r, adapter = _runner_with(monkeypatch, idle=True)
|
|
# First idle check (loop gate) True, second (post-quiesce re-check) False.
|
|
reads = iter([True, False, False, False, False, False])
|
|
monkeypatch.setattr(
|
|
r, "_scale_to_zero_is_idle", lambda: next(reads, False), raising=False
|
|
)
|
|
suspend_calls = []
|
|
|
|
async def fake_suspend():
|
|
suspend_calls.append(1)
|
|
|
|
monkeypatch.setattr(r, "_scale_to_zero_self_suspend", fake_suspend, raising=False)
|
|
task = asyncio.create_task(r._scale_to_zero_watcher(interval=0.01))
|
|
await asyncio.sleep(0.15)
|
|
r._running = False
|
|
await asyncio.wait_for(task, timeout=2)
|
|
assert adapter.go_dormant_calls == 1
|
|
assert suspend_calls == []
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_self_suspend_noop_off_fly(monkeypatch):
|
|
"""Off-Fly (no flaps socket/identity) the helper is a silent no-op —
|
|
dormancy without platform suspend, never an error."""
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
monkeypatch.setattr(
|
|
"gateway.scale_to_zero.self_suspend_available", lambda *a, **k: False
|
|
)
|
|
called = []
|
|
monkeypatch.setattr(
|
|
"gateway.scale_to_zero.suspend_self",
|
|
lambda *a, **k: called.append(1) or True,
|
|
)
|
|
await r._scale_to_zero_self_suspend()
|
|
assert called == []
|
|
|
|
|
|
# ── non-messaging platforms must not disarm (the api_server-key regression) ──
|
|
#
|
|
# The Docker stage2 hook now generates API_SERVER_KEY for every container, and
|
|
# key presence force-enables the api_server platform (gateway/config.py). The
|
|
# arm gate counted every enabled platform, so `api_server` (a loopback
|
|
# listener, not a messaging socket) made messaging_is_relay_only_or_absent
|
|
# False on EVERY hosted instance — silently disarming scale-to-zero. The gate
|
|
# must only count messaging platforms (excluding LOCAL/API_SERVER/WEBHOOK,
|
|
# mirroring _connect_platforms' messaging_platforms exclusion set).
|
|
|
|
|
|
def test_arm_true_with_api_server_enabled(monkeypatch):
|
|
from gateway.platforms.base import Platform
|
|
|
|
r = _arm_runner(
|
|
monkeypatch,
|
|
{
|
|
Platform.RELAY: True,
|
|
Platform.API_SERVER: True,
|
|
Platform.TELEGRAM: False,
|
|
},
|
|
)
|
|
assert r._scale_to_zero_should_arm() is True
|
|
|
|
|
|
def test_arm_true_with_all_non_messaging_surfaces_enabled(monkeypatch):
|
|
from gateway.platforms.base import Platform
|
|
|
|
r = _arm_runner(
|
|
monkeypatch,
|
|
{
|
|
Platform.RELAY: True,
|
|
Platform.API_SERVER: True,
|
|
Platform.WEBHOOK: True,
|
|
Platform.LOCAL: True,
|
|
},
|
|
)
|
|
assert r._scale_to_zero_should_arm() is True
|
|
|
|
|
|
def test_direct_platform_still_disarms_alongside_api_server(monkeypatch):
|
|
"""The messaging-only filter must not over-broaden: a genuinely enabled
|
|
direct-socket platform still disarms even with api_server also enabled."""
|
|
from gateway.platforms.base import Platform
|
|
|
|
r = _arm_runner(
|
|
monkeypatch,
|
|
{
|
|
Platform.RELAY: True,
|
|
Platform.API_SERVER: True,
|
|
Platform.DISCORD: True,
|
|
},
|
|
)
|
|
assert r._scale_to_zero_should_arm() is False
|
|
|
|
|
|
# ── supervised watchers must NOT count as live background work (staging bug) ──
|
|
#
|
|
# _spawn_supervised parks every permanent watcher task (session-expiry, kanban,
|
|
# reconnect, the scale-to-zero watcher ITSELF, ...) in _background_tasks. The
|
|
# bg-work check counted them, so an armed gateway considered itself busy
|
|
# forever and never went dormant — verified live on staging 2026-08-12 (armed
|
|
# at 05:25, fully idle 25+ min, zero "going dormant" lines). Fly's coarse
|
|
# autostop masked this until the gateway took ownership of the suspend.
|
|
# These tests exercise the REAL _spawn_supervised path — the earlier tests
|
|
# stubbed _background_tasks and missed the call site (same trap as F25).
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_supervised_watchers_do_not_block_idle():
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
r._running = True
|
|
r._background_tasks = set()
|
|
|
|
async def _forever():
|
|
await asyncio.sleep(3600)
|
|
|
|
# Spawn like production does — through _spawn_supervised.
|
|
for name in ("session_expiry", "kanban", "scale_to_zero_watcher"):
|
|
r._spawn_supervised(lambda: _forever(), name)
|
|
await asyncio.sleep(0) # let tasks start
|
|
try:
|
|
assert r._scale_to_zero_has_live_background_work() is False
|
|
finally:
|
|
for t in r._background_tasks:
|
|
t.cancel()
|
|
await asyncio.gather(*r._background_tasks, return_exceptions=True)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_transient_background_task_still_blocks_idle():
|
|
"""A plain (untagged) task in _background_tasks — startup-resume events,
|
|
ad-hoc work — must still count as live background work."""
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
r._running = True
|
|
|
|
async def _work():
|
|
await asyncio.sleep(3600)
|
|
|
|
t = asyncio.create_task(_work())
|
|
r._background_tasks = {t}
|
|
try:
|
|
assert r._scale_to_zero_has_live_background_work() is True
|
|
finally:
|
|
t.cancel()
|
|
await asyncio.gather(t, return_exceptions=True)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_done_supervised_watcher_is_ignored_either_way():
|
|
r = GatewayRunner.__new__(GatewayRunner)
|
|
r._running = True
|
|
|
|
async def _quick():
|
|
return None
|
|
|
|
t = asyncio.create_task(_quick())
|
|
await t
|
|
r._background_tasks = {t}
|
|
assert r._scale_to_zero_has_live_background_work() is False
|