Commit Graph

14 Commits

Author SHA1 Message Date
pefontana b2c493c173 Assert the skip branch ran in the no-quiesce watcher test
The three existing assertions are absence checks, so the test also
passed when the loop got no iteration inside the sleep window (with
interval=5.0 it passes without the gate ever executing). Checking
_scale_to_zero_no_suspend_logged proves the branch was taken.
2026-08-26 13:02:04 -03:00
pefontana 85816595d2 Merge remote-tracking branch 'origin/main' into fix/scale-to-zero-no-pointless-quiesce 2026-08-26 13:02:03 -03:00
IAvecilla 14833bcc56 Trim comments 2026-08-21 18:10:33 -03:00
IAvecilla 9d5745d0a0 Fix tests 2026-08-21 15:23:21 -03:00
Ben Barclay 0215930526 review: fail-awake work accounting + rename is_idle param to active_work_count
Address sol-reviewer findings:
- The shared shutdown-drain counters swallow exceptions to 0 — fine for
  a drain, unsafe for a suspend predicate (a transient read failure made
  live work look idle, reopening the mid-job freeze). The suspend path
  now reads both sources itself and treats an unreadable source as work
  (sentinel 1, fail-awake) with a debug log. A MISSING api_server
  adapter remains a normal not-work state.
- is_idle()'s parameter renamed running_agent_count -> active_work_count:
  it receives the broad aggregate, and the old name invited future
  callers to pass only agents again.
- New failure-path tests: unreadable cron source and unreadable API
  source each hold the machine awake (both fail against the fail-open
  shape); missing adapter stays idle-capable.
2026-08-21 07:37:11 +10:00
Ben Barclay 743dc935f5 fix(gateway): count cron and API-server work in the scale-to-zero idle predicate
_scale_to_zero_is_idle() consumed _running_agent_count(), but cron jobs
run through a standalone AIAgent on the scheduler's own thread pool and
API-server runs live on the adapter — both outside _running_agents (the
same blind spot the #60432 shutdown-drain fix addressed with
_active_work_count()). The idle predicate therefore read True DURING a
running cron job; a suspend at that moment freezes the job mid-flight.
Observed live on staging 2026-08-20: is_idle held True throughout the
10:45:04-22 cron run — only watcher-tick timing (next tick 9s after
completion) avoided a mid-job freeze.

Use _active_work_count() (agents + cron + API runs). New tests cover a
running cron job and an active API run each blocking idle, plus the
all-quiet True case; both blocking tests fail without the fix.
2026-08-20 20:53:16 +10:00
pierrenode a1ddb54840 fix(gateway): tag the loop-liveness and heartbeat-poll tasks as permanent supervised watchers (#84558)
#84327 excluded _spawn_supervised's permanent watchers (session-expiry,
kanban, reconnect, the scale-to-zero watcher itself, ...) from
_scale_to_zero_has_live_background_work() via a _hermes_supervised_watcher
tag, because counting them made an armed gateway consider itself busy
forever and never go dormant.

Two more permanent, infinite-loop tasks are added to _background_tasks
OUTSIDE _spawn_supervised and were untagged:

- _loop_heartbeat_task (loop_heartbeat_forever, #66892): a `while True`
  loop started unconditionally in start() on every gateway boot. Extracted
  the inline spawn block into _start_loop_heartbeat_task() so it's
  independently testable, matching the existing _start_heartbeat_poller()
  pattern.
- _heartbeat_poll_task (_poll_loop in _start_heartbeat_poller): also a
  `while True` loop, started the first time a session registers a
  heartbeat watch, and then permanent for the rest of the process.

Because _loop_heartbeat_task starts on every boot, it alone made
_scale_to_zero_has_live_background_work() return True forever on every
armed instance, regardless of the #84327 fix -- confirmed empirically
against the real method with the exact untagged-task shape this task has.

Two new regression tests spawn each task through its real production
entry point and assert the busy check returns False; both fail against
the unfixed code (missing method / real assertion failure).

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
Co-authored-by: Ben Barclay <ben@nousresearch.com>
2026-08-20 20:24:40 +10:00
Ben Barclay bb597e1c02 fix(cron): managed-cron fires execute in the gateway process (live adapters + dashboard forwarder) (#84339)
* fix(gateway): pass live adapters to cron fire webhook's fire_due

The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).

Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.

Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).

* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)

The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.

Restore the invariant that the GATEWAY owns cron execution:

- Dashboard route: after verifying the NAS JWT and resolving the job's
  profile, FORWARD the fire to the gateway api_server's own
  /api/cron/fire on loopback, NAS bearer preserved (the gateway
  re-verifies the JWT — defense in depth, no new trust link), and pass
  the gateway's response through. Gateway unreachable → 503 so NAS
  retries per the Chronos contract (non-2xx = retryable; the store CAS
  de-dupes the eventual double fire). Deliberately NO local-execution
  fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
  per target profile (config.yaml extra.port → API_SERVER_PORT from
  process env or the profile's .env → 8642), with /p/<profile>/ prefix
  routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
  first boot when absent (never overwrites an operator value), so the
  loopback api_server passes its startup guard on hosted images. The
  fire route itself is NAS-JWT-authed; the key gates the rest of the
  api_server surface. The listener binds 127.0.0.1 by default and the
  Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
  compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
  topology and the 503-retry semantics.

Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.

* fix(cron): read the profile api_server port via the canonical config loader

CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).

Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.

* fix(gateway): only messaging platforms count for the scale-to-zero arm gate

The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).

The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
2026-08-12 17:04:44 +10:00
Ben Barclay 5fffe56066 fix(gateway): exclude permanent supervised watchers from the scale-to-zero busy check (#84327)
_scale_to_zero_has_live_background_work() counted every task in
_background_tasks — but _spawn_supervised parks all permanent watchers
there (session-expiry, kanban, reconnect, the scale-to-zero watcher
itself, ...). An armed gateway therefore considered itself busy forever
and never went dormant or suspended. Verified live on staging
(hermes-agent-stg-test-6698, 2026-08-12): armed at 05:25, fully idle for
25+ minutes, zero 'going dormant' lines. Fly's coarse proxy autostop used
to mask the bug; once the gateway took ownership of the suspend (#84295)
it became load-bearing.

_spawn_supervised now tags its tasks and the busy check skips them.
Transient tasks (startup-resume events, delegation, tracked processes)
still block suspend. New tests exercise the REAL _spawn_supervised path
rather than a stubbed _background_tasks set — the stubbing is exactly why
the earlier tests missed this (same call-site trap as the F25 arm bug);
the key test fails on main and passes with the fix.
2026-08-12 16:39:01 +10:00
Ben Barclay 356c702b55 fix(gateway): scale-to-zero gateway self-suspends via flaps socket instead of relying on Fly autostop (#84295)
Fly Proxy autostop judges idle exclusively on inbound proxied connections.
It cannot see an in-flight agent turn (outbound-only LLM traffic), and since
Fly's mid-2026 proxy change an open outbound socket (the relay WS) no longer
holds a machine awake. With autostop:"suspend", Fly suspended machines while
they were still processing long-running jobs, and could suspend before the
gateway flipped the relay destination (the buffered-event black hole).

The scale-to-zero watcher now owns the suspend: after the idle predicate
holds (no running agents, no live background work, inbound-quiet) and the
go_dormant() quiesce completes (relay drained + flipped), it POSTs
/v1/apps/{app}/machines/{id}/suspend on the local /.fly/api flaps socket.
Suspend is skipped when the quiesce fails or inbound lands mid-quiesce
(flip-before-freeze), and off-Fly the step is a no-op (fail-awake).

Pairs with the NAS change that provisions scale-to-zero machines with
autostop:"off" (gateway-owned suspend); wake is unchanged (Fly-proxied
wakeUrl poke + autostart).
2026-08-12 14:45:48 +10:00
Teknium 39975613b1 test: prune wave 2 + speed fixes — 28,106 → 19,757 test functions, suite wall 315s → 294s
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.

Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
  (DEFAULT_CONFIG smart-approval leaked in) — pinned approval
  mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
  silently probed live provider catalogs (~2s/test) — stubbed
  cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
  shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
  MagicMock agents — interrupt.side_effect now clears _running_agents
  (22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
  (24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
  patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
  5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
  with rationale comment — the watchdog-dump assertion needs the loop
  blocked well past deadline+grace under parallel load (flaked once in
  the 40-worker verification run at a 0.2s margin).

Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
2026-07-29 13:39:40 -07:00
Ben Barclay dedf5643d8 fix(gateway): scale-to-zero never armed — arm-gate counted disabled placeholder platforms (#52831)
The scale-to-zero idle watcher never started on a correctly-opted-in,
relay-only instance, so the gateway never ran its idle decision, never called
go_dormant(), and never sent going_idle to the connector. Fly's autostop still
suspended the machine on traffic-idle, but the connector never flipped the
instance to buffered-only — so an inbound DM took the live delivery path,
found no live session for the suspended machine, and was dropped fail-closed
with no wake poke. The machine slept and never woke.

Root cause: _scale_to_zero_should_arm() passed list(config.platforms.keys())
to messaging_is_relay_only_or_absent(). config.platforms is pre-seeded with a
DISABLED placeholder PlatformConfig for every known platform (telegram,
discord, slack, matrix, …), so the key set is always the full ~20-entry
catalog regardless of what the instance actually runs. The relay-only check
discarded "relay", saw the disabled placeholders as live direct-socket
platforms, and returned False — so should_arm() was False and the watcher was
never created. Verified live on a staging instance: config.platforms keys =
[telegram, discord, slack, mattermost, matrix, relay] with only relay
enabled=True; should_arm() = False.

Fix: filter config.platforms to ENABLED entries before the relay-only check,
mirroring the adapter-connect loop which already gates on
`if not platform_config.enabled: continue`. This arms off the same notion of
"active platform" the rest of start() already uses — no parallel concept.

Also add a one-line not-armed diagnostic: when an instance IS opted in (the
HERMES_SCALE_TO_ZERO stamp is set) but the watcher still doesn't arm, log why
(relay_only_or_absent, the enabled platforms, wake_url present/missing). A
non-opted instance stays silent. The arm path previously logged only on
success, so a failed arm was invisible.

Tests: the existing pure-helper tests passed bare names so they never
exercised the call site that feeds the placeholder-laden config. Add
behaviour-contract tests against the REAL _scale_to_zero_should_arm with a
realistic config.platforms (relay enabled + others disabled). The F25
regression test (relay-only + disabled placeholders must arm) and the
no-platform case are RED without this fix, GREEN with it; the
genuinely-enabled-direct-platform / not-opted-in / no-wake-url cases stay
correctly non-arming so the filter can't over-broaden.

Wake mechanism itself verified healthy independently (direct wakeUrl GET
resumed a suspended staging instance in 1.15s, clean resume signature).
2026-06-26 14:01:48 +10:00
Ben Barclay d6269da7fd fix(gateway): harden scale-to-zero dormancy guards (#52359)
Block scale-to-zero suspend while background async delegations are active, and restore runtime status to running on real inbound after a dormant wake.\n\nAdd regression coverage for both review findings.
2026-06-25 20:41:03 +10:00
Ben d1cac0e5ef feat(gateway): scale-to-zero idle detection + dormant-quiesce (Phase 0)
The gateway-side BEHAVIOUR layer that consumes the relay scale-to-zero
primitives (gateway-gateway Phase 5): the gateway decides it is idle and
drives the relay transport dormant so the platform (Fly autostop:"suspend")
can suspend the now-traffic-idle machine, which wakes on the connector's
wakeUrl poke (decisions.md Q3=C', D1-D13).

- gateway/scale_to_zero.py: pure helpers — scale_to_zero_enabled (the NAS
  Labs HERMES_SCALE_TO_ZERO stamp, D11/Q8=A), parse_idle_timeout_seconds
  (config.yaml gateway.scale_to_zero.idle_timeout_minutes, D2),
  messaging_is_relay_only_or_absent (F6/D1), should_arm (D1/D11/§3.4(1)),
  is_idle (D2/D3/F7).
- gateway/run.py: _last_inbound_at clock stamped on user inbound in
  _handle_message (F13); the arm-gate + idle predicate + the
  _scale_to_zero_watcher dormant sequence (mark draining -> adapter
  go_dormant() -> cooldown), started only when armed. Deliberately NOT the
  stop path and NOT mark_resume_pending (F12/D13).
- tools/process_registry.py: has_any_active() for the bg-work guard (D3/F7).
- hermes_cli/config.py: gateway.scale_to_zero.idle_timeout_minutes default 5.

Tests: 38 pure-logic + 6 watcher (incl. bg-work regression guard proven RED).
Full relay + scale-to-zero suites: 184 passed. The 20 unrelated failures in
the broader run are PRE-EXISTING on origin/main (custom-provider/tools tests),
confirmed via a pristine baseline worktree.
2026-06-24 18:47:18 -07:00