On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
Since the managed-cron redesign (#84339, v2026.8.13) the dashboard fire
webhook forwards fires to the gateway process and returns 503 when it is
unreachable so NAS/QStash retries. Correct for transient windows — but an
operator-STOPPED gateway can never be fixed by retrying: every fire on
every job burns the full scheduler retry budget, NAS converts each 503 to
a retryable 502, and the resulting storms page on-call for a non-incident
(OOF-266 and its five duplicate tickets; +93% relay callback failures as
the fleet adopted v2026.8.13).
Split the unreachable path by durable operator intent:
- desired_state == "stopped" (written only by the s6 lifecycle commands;
the same intent signal container-boot reconciliation trusts) -> drop
the fire with 200 + a structured log line, mirroring NAS's own
instance_stopped drop. Jobs are not lost: the Chronos provider
reconciles and re-arms every job on the next gateway start.
- Anything else (crash loop, scale-to-zero wake, restart, legacy state
file without desired_state) -> keep the retryable 503, now stamped
with Retry-After: 60 so a scheduler that honors it spaces retries
past the wake/restart window instead of exhausting them inside it.
The gateway's own pass-through 503s (draining) get the same hint.
The intent check fails open (any parse/resolution error -> retryable
path) and is only consulted when the gateway is actually unreachable, so
a stale state file can never shadow a live gateway.
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
* fix(gateway): pass live adapters to cron fire webhook's fire_due
The Chronos fire webhook (/api/cron/fire) called
provider.fire_due(job_id, adapters=None, loop=loop), so every
externally-triggered fire delivered through the standalone path even
with a live gateway in-process. E2EE platforms and relay-fronted
logical platforms (whose ONLY send path is the live relay adapter — no
native credential exists on the box) failed every external fire with
"platform 'X' not configured/enabled", while the same job delivered
fine under the built-in ticker (gateway/run.py passes runner.adapters).
Resolve the runner (self.gateway_runner → app['gateway_runner'] →
_gateway_runner_ref(), the same chain the drain check uses) and forward
its adapters. No runner → adapters=None, preserving the historical
standalone path byte-identically.
Note: does not by itself fix Fly-hosted scale-to-zero deployments where
NAS's callback lands on the DASHBOARD process (internal_port 9119) —
_fire_cron_job_for_profile there has no gateway runner in-process. That
topology needs a separate fire handoff (design pending).
* fix(cron): dashboard forwards Chronos fires to the gateway (503 when unreachable)
The dashboard's /api/cron/fire executed cron jobs in the DASHBOARD
process via _fire_cron_job_for_profile with adapters=None. On hosted
deployments (Fly proxy exposes only the dashboard's port) that made
every managed-cron fire deliver through the standalone send path, which
cannot serve relay-fronted logical platforms (their only sender is the
live relay adapter in the gateway process — no native credential exists
on the box) or E2EE rooms. It also ran the whole agent turn inside the
dashboard: wrong process for memory/session ownership and fire-claim
attribution.
Restore the invariant that the GATEWAY owns cron execution:
- Dashboard route: after verifying the NAS JWT and resolving the job's
profile, FORWARD the fire to the gateway api_server's own
/api/cron/fire on loopback, NAS bearer preserved (the gateway
re-verifies the JWT — defense in depth, no new trust link), and pass
the gateway's response through. Gateway unreachable → 503 so NAS
retries per the Chronos contract (non-2xx = retryable; the store CAS
de-dupes the eventual double fire). Deliberately NO local-execution
fallback.
- Endpoint resolution mirrors gateway/config.py's api_server load order
per target profile (config.yaml extra.port → API_SERVER_PORT from
process env or the profile's .env → 8642), with /p/<profile>/ prefix
routing under multiplex.
- docker/stage2-hook.sh: generate a strong API_SERVER_KEY into .env on
first boot when absent (never overwrites an operator value), so the
loopback api_server passes its startup guard on hosted images. The
fire route itself is NAS-JWT-authed; the key gates the rest of the
api_server surface. The listener binds 127.0.0.1 by default and the
Fly service exposes only the dashboard port.
- _fire_cron_job_for_profile kept but deprecated (late-binding seam
compatibility); no route calls it.
- docs/chronos-managed-cron-contract.md: document the two-hop inbound
topology and the 503-retry semantics.
Depends on the previous commit (fire webhook passes live adapters to
fire_due) — together they make NAS→dashboard→gateway fires deliver over
relay end to end.
* fix(cron): read the profile api_server port via the canonical config loader
CI guard test_config_read_guard flagged the new _gateway_fire_endpoint
for a raw yaml.safe_load of the profile's config.yaml — the exact drift
class the guard exists to kill (raw reads miss the managed-scope
overlay, ${ENV_VAR} expansion, and root-model normalization).
Read through load_config() under a HERMES_HOME override scoped to the
target profile instead (the same pattern the deprecated
_fire_cron_job_for_profile uses for its store scope), and pull the port
with cfg_get. Test updated to stub load_config rather than write a raw
config.yaml.
* fix(gateway): only messaging platforms count for the scale-to-zero arm gate
The stage2 hook now generates API_SERVER_KEY for every Docker container,
and key presence force-enables the api_server platform. The scale-to-zero
arm gate counted every enabled platform, so the loopback api_server
listener made messaging_is_relay_only_or_absent False on every hosted
instance — silently disarming the feature (the not-armed log would show
enabled platforms=['relay','api_server']).
The arm gate and the not-armed logger now share one helper that filters
to enabled MESSAGING platforms, excluding LOCAL/API_SERVER/WEBHOOK —
the same non-messaging exclusion set _connect_platforms already uses.
A genuinely enabled direct-socket platform (Discord/Telegram) still
disarms. Two of the three new tests fail without this fix.
Two follow-ups to the off-loop move, from external review (both verified,
the second larger than reported):
- Config read-modify-write handlers moved to worker threads could now
interleave — _CONFIG_LOCK covers each load/save individually, never the
span between them; the event loop used to serialize these accidentally.
New _CONFIG_MUTATION_LOCK (worker-threads only, so it can never block
the loop) held across the whole load→mutate→save span in all seven RMW
handlers. update_config_raw skipped: it's a full-document replace with
no server-side read, so a lock cannot close its client-side window.
- The review flagged two skills routes still taking _SKILLS_PROFILE_LOCK
on the event loop; a systematic audit of hermes_cli/web_routers/ found
24 on-loop routes (skills 5, mcp 9, tools 10, cron 1). All moved to the
same inner-_run + asyncio.to_thread pattern, mutating ones under the
mutation lock, uniform lock order (_SKILLS_PROFILE_LOCK →
_CONFIG_MUTATION_LOCK). Await-safe _config_profile_scope routes, plain
def routes, and already-threaded routes unchanged.
Regression tests: concurrent theme+font updates both survive (fails with
the lock nulled: "theme write lost to a concurrent font write"); event
loop stays responsive while the profile lock is held during GET
/api/skills. 214 tests passing across the touched suites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to the salvaged registration contract:
- share one _raise_if_cron_registration_error() helper for the two
byte-identical dashboard 424 except-blocks (web_server + cron router,
via the existing late() seam)
- add endpoint-level 424 coverage for /api/cron/blueprints/instantiate
(previously only the sync worker was tested)
- give chat/CLI surfaces a human-facing user_message() (job name, no
exception class name) and add a recovery hint (pause/resume or update
re-registers via provider reconcile) to the model/REST message
- consolidate five inline provider test doubles into one ABC-subclassing
make_cron_provider conftest factory; the web_server test double now
subclasses CronScheduler so an ABC rename fails loudly
- narrow the wrapper facade to keyword-only (**kwargs) and route the
tool's partial-failure return through tool_error()