10 Commits

Author SHA1 Message Date
Teknium 4fde117f4b simplify(compat): tests — repoint 638 web_server.<name> references (321 attr, 217 monkeypatch/patch.object, 60 from-imports, 40 patch() strings) across 65 test files to the owning modules 2026-09-03 14:21:52 -07:00
berg 357062f39e fix(dashboard): multiplex cron fire URL uses the default profile's listener port
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
2026-09-02 06:27:24 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
Shannon Sands 4822156923 fix(cron): stop retry storms when the gateway is deliberately stopped (OOF-266)
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).

Split the unreachable path by durable operator intent:

- desired_state == "stopped" (written only by the s6 lifecycle commands;
  the same intent signal container-boot reconciliation trusts) -> drop
  the fire with 200 + a structured log line, mirroring NAS's own
  instance_stopped drop. Jobs are not lost: the Chronos provider
  reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
  file without desired_state) -> keep the retryable 503, now stamped
  with Retry-After: 60 so a scheduler that honors it spaces retries
  past the wake/restart window instead of exhausting them inside it.
  The gateway's own pass-through 503s (draining) get the same hint.

The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
2026-08-16 19:50:07 -07:00
Evgenii 1a8625abee fix(cron): harden gateway fire admission and provider compatibility
- The gateway api_server fire webhook acknowledges 202 only after a
  durable claim + execution row exist (admission failure stays retryable
  as 503; a live claim answers 200 duplicate), then dispatches the
  claimed snapshot with the live runner adapters (delivery parity with
  the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
  split hooks) keep being driven through their own hook. Capability
  detection now credits claim_fire AND fire_claimed overrides, so
  Chronos is correctly classified split-aware (its re-arm lives in
  fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
  unscoped reconcile would disarm other profiles' armed one-shots in the
  shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
  through every entry point, composing with upstream's manual-run
  heartbeat (#76502) and background dispatch.

Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
2026-08-14 20:46:50 -07:00
Ben Barclay bb597e1c02 fix(cron): managed-cron fires execute in the gateway process (live adapters + dashboard forwarder) (#84339)
* fix(gateway): pass live adapters to cron fire webhook's fire_due

The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).

Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.

Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).

* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)

The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.

Restore the invariant that the GATEWAY owns cron execution:

- Dashboard route: after verifying the NAS JWT and resolving the job's
  profile, FORWARD the fire to the gateway api_server's own
  /api/cron/fire on loopback, NAS bearer preserved (the gateway
  re-verifies the JWT — defense in depth, no new trust link), and pass
  the gateway's response through. Gateway unreachable → 503 so NAS
  retries per the Chronos contract (non-2xx = retryable; the store CAS
  de-dupes the eventual double fire). Deliberately NO local-execution
  fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
  per target profile (config.yaml extra.port → API_SERVER_PORT from
  process env or the profile's .env → 8642), with /p/<profile>/ prefix
  routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
  first boot when absent (never overwrites an operator value), so the
  loopback api_server passes its startup guard on hosted images. The
  fire route itself is NAS-JWT-authed; the key gates the rest of the
  api_server surface. The listener binds 127.0.0.1 by default and the
  Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
  compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
  topology and the 503-retry semantics.

Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.

* fix(cron): read the profile api_server port via the canonical config loader

CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).

Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.

* fix(gateway): only messaging platforms count for the scale-to-zero arm gate

The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).

The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
2026-08-12 17:04:44 +10:00
Teknium 39975613b1 test: prune wave 2 + speed fixes — 28,106 → 19,757 test functions, suite wall 315s → 294s
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.

Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
  (DEFAULT_CONFIG smart-approval leaked in) — pinned approval
  mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
  silently probed live provider catalogs (~2s/test) — stubbed
  cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
  shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
  MagicMock agents — interrupt.side_effect now clears _running_agents
  (22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
  (24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
  patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
  5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
  with rationale comment — the watchdog-dump assertion needs the loop
  blocked well past deadline+grace under parallel load (flaked once in
  the 40-worker verification run at a 0.2s margin).

Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
2026-07-29 13:39:40 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
teknium1 901165b5a4 fix(cron): complete plugins.cron_providers rename in 2 missed test files
uperLu's #50958 renamed plugins/cron → plugins/cron_providers but left
two test files patching the now-gone plugins.cron.chronos.verify path,
which would fail collection. Point them at plugins.cron_providers.*.
Add uperLu to release.py AUTHOR_MAP.
2026-06-23 23:39:22 -07:00
Ben c34840e22e fix(cron): serve /api/cron/fire on the dashboard app (hosted-agent surface)
Live-test finding: the Chronos fire webhook was only on the APIServerAdapter
(aiohttp), but hosted agents expose `hermes dashboard` (the FastAPI web_server
app on :9119) as their public URL — NOT the api_server adapter. So NAS's relay
callback to {callback_url}/api/cron/fire could never reach the verifier on a
hosted agent (the exact target environment). Two layers were wrong:

1. Wrong server: /api/cron/fire didn't exist on the dashboard app. Added
   cron_fire_webhook there, alongside the existing /api/cron/* dashboard routes.
   It resolves the job's profile (_find_cron_job_profile) and runs fire_due via
   the resolved provider under the cron-profile retarget lock
   (_fire_cron_job_for_profile, mirroring _call_cron_for_profile) so the CAS
   claim + run_one_job operate on the right profile's jobs.json. Runs with no
   live adapters (delivery falls back to the per-platform send path, like the
   desktop cron path). 202 + background so a long turn never trips NAS's
   timeout; the store CAS de-dupes a NAS retry. job-not-found -> 200 "gone".

2. Auth gate: the dashboard auth middleware 401s any non-cookie request before
   the handler runs. Added /api/cron/fire to the shared PUBLIC_API_PATHS so the
   NAS bearer-JWT callback reaches the verifier — the JWT (purpose=cron_fire),
   not the cookie, is the real gate. One shared frozenset feeds both the
   loopback and OAuth middlewares, so no drift.

Kept the APIServerAdapter route too (valid self-host api_server surface).
Contract doc updated to name the dashboard app as the hosted-agent callback
surface.

Tests: test_cron_fire_dashboard (6) — route registered on the dashboard app,
in PUBLIC_API_PATHS, 401 on bad token WITH the cookie gate engaged (proves it's
reachable past the gate + JWT is the gate), 400 missing job_id, 200 gone for
unknown job, 202 + fire_due invoked for the resolved profile on a valid token.
Full hermes_cli + cron + chronos + webhook suites green (7637).

Why the original tests missed it: the api_server webhook test built an
APIServerAdapter client directly and never asserted which server the hosted
public URL exposes — green-but-wrong-integration. The new test pins the route
to the dashboard app.
2026-06-19 12:43:30 +10:00