Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.
Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.
Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.
Co-authored-by: holny <holny@foxmail.com>
The salvaged regression from #86582 predates the claim_job_for_fire
owner-fencing that landed with #70638; mock the claim and heartbeat so
the healthy job actually runs through the fenced flow.
- claim_job_for_fire returns the atomically claimed snapshot with a unique
fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
terminal writes by expected_fire_owner so a stale worker cannot record
over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
event (ownership loss OR external cancel) into run_job; the agent path
is interrupted cooperatively and script-based jobs (no_agent + pre-run
scripts) are hard-stopped with a process-tree kill (POSIX killpg
SIGTERM then SIGKILL for surviving group members; Windows
taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
the bare job ID, so a replacement run of the same job never consumes a
stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
completed one-shot retention (#80624), blocked_config preflight
(T1-26), and the advance_next_runs batch on top of current main.
Salvaged from PR #57342 by @liuhao1024 (with the injection-scan half
from PR #57360 by @ghedeselmabot): cronjob(action='run', prompt=...)
silently discarded the prompt argument — per-run context never
reached the spawned cron session.
The prompt is now threaded as extra_prompt through the whole chain
(cronjob run action → _try_dispatch_background_run/_execute_job_now →
_run_claimed_job → run_one_job → run_job → _build_job_prompt) and
appended to the stored prompt under a '## Run Context' header for
that single fire only — never persisted to the job definition. It
passes the same strict _scan_cron_prompt injection scan as stored
prompts before firing, and works identically on the background and
sync fallback paths.
Test fakes across tests/cron/ updated to accept the new kwargs
(sibling-test blast radius from the signature change).
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
The scheduler's pre-dispatch loop called advance_next_run per due job —
one full load_jobs() + one full save_jobs() of the jobs file each — so
N due jobs cost N reads + N writes of the whole file (gateway-restart
catch-up or co-scheduled bursts). advance_next_runs() does one load +
at most one save for the whole due set with identical per-job semantics;
advance_next_run() is now a thin wrapper over it.
Measured (50 due recurring jobs, real jobs file): 107.9 ms -> 2.5 ms
(45x; 50 loads + 50 saves -> 1 + 1).
Tests: batch advances recurring and skips one-shots, single load + save
I/O pin (fails pre-fix — no such function), no save when nothing
advances, and per-job wrapper semantics unchanged. Related: #60946 and
#75833 both restructure this loop's call site for correctness — neither
addresses the I/O cost, and this batch primitive composes with either
dispatch design; happy to rebase onto whichever lands first.
Defense-in-depth alongside the interpreter-shutdown guard: run_job closed
the cron agent's async resources (agent.close + cleanup_stale_async_clients)
in its finally block BEFORE run_one_job called _deliver_result, so a live
delivery could race a torn-down async client. run_job now accepts an optional
defer_agent_teardown holder; when set it hands the live agent back instead of
closing it, and run_one_job tears it down (via the extracted _teardown_cron_agent
helper) only AFTER delivery — in a finally so a failed run never leaks. Default
path (holder=None) is unchanged, so every existing caller keeps inline teardown.
Reorder approach based on #58777 by @LavyaTandel; reworked to keep a single
delivery site in run_one_job and add regression coverage.
Co-authored-by: LavyaTandel <lavya@loom.local>
Fully removes the cron per-job 'profile' arg added in #28124: the
cronjob tool schema field, CLI --profile flags on cron create/edit,
job-record storage/validation, the scheduler's _job_profile_context
wrapper, and the script-runner env override. Sequential-partition
logic reverts to workdir-only.
The context-local HERMES_HOME override in hermes_constants and the
subprocess bridging in tools/environments/local.py are kept — they
now have other consumers (dashboard multi-profile, TUI gateway).
Follow-up on the parallel-dispatch decoupling: the sequential pass for
workdir/profile jobs still ran inline in the ticker thread, so a long
workdir/profile job reintroduced the exact starvation #37312 describes,
just for env-mutating jobs. And the MCP orphan sweep ran immediately
after dispatch in sync=False mode — before jobs finished — defeating its
own 'runs after every job' contract and racing jobs still spawning MCP
children.
- Sequential jobs now queue to a persistent single-thread cron-seq pool
(preserves one-at-a-time ordering across ticks, never blocks the tick).
- Same in-flight dedup guard now covers sequential jobs.
- MCP orphan sweep runs via a done-callback after the LAST dispatched job
completes in async mode; inline after as_completed in sync mode.
Verified E2E: tick(sync=False) returns in ~1ms with a 1.5s sequential job
in flight; sweep fires only after that job ends.
PR #13021 fixed serial starvation by adding ThreadPoolExecutor to tick(),
but kept as_completed(timeout=600) which still blocks the ticker thread
until the slowest job finishes. This causes the same starvation pattern:
when one job runs long (15+ min), other jobs' next_run_at expires past the
grace window and they get perpetually fast-forwarded instead of running.
This PR decouples dispatch from completion:
- Persistent ThreadPoolExecutor (reused across ticks, no auto-join)
- Fire-and-forget dispatch: tick submits and returns immediately
- Running-job guard: prevents re-dispatching active jobs
- sync parameter: defaults to True (backward compatible), callers opt
into sync=False for non-blocking behavior
- atexit shutdown handler for clean pool teardown
- gateway/run.py: production ticker opts into sync=False
Refs #33315 (complementary — that issue's PRs fix grace handling in
jobs.py; this PR prevents the grace from expiring in the first place)