Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.
Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
Persist paused state, timestamp, reason and no first trigger in the original
locked creation write. Forward the same boolean contract across CLI, tool,
gateway API and dashboard API, validating at the store boundary. Preserve
explicit operator force-run behavior and normal enabled creation.
The live CLI probe also caught the command shim dropping failure return codes;
forward them so invalid creation reports exit 1 rather than success.
Credit earlier atomic-creation work in #78935 and #94952 and the focused
implementation in #104578. The broader manifest staging layer is not imported.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Chloé DuPont <321112755+misschloedupont@users.noreply.github.com>
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.
Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Slim redo of #104546 and #104551: scan newest-first, match suppression only before payload separators, and keep error context. Covers wake gates and empty outputs without reading every historical file twice.
Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.
Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.
Refs #103974, #104058, #104904
Canonicalize the accepted raw next_run_at value before fast-forward rather
than its timezone-interpreted datetime. Legacy naive values remain runnable
but cannot establish an exact UTC identity. Preserve repaired aware slots.
Extend the existing identity invariant with the naive-slot control, and use
real ledger creation in the provider ordering test instead of an invented
execution ID that cannot pass the owner-fenced occurrence setter.
Live validation: actual builtin script run was RED (invented UTC identity)
and is now GREEN (NULL identity). Repeated builtin/provider/worker rollback
A/B remains 2 writes on base versus 1 on head, with distinct/manual controls.
Canonical cron regression rerun is queued under the campaign lock.
Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.
Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.
Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.
Co-authored-by: holny <holny@foxmail.com>
The TTL was three copies of the literal 300 (claim_job_for_fire default,
rearm_oneshot, and the new stale-error guard). Hoisted so the three lanes
cannot drift; incident narrative in the guard comment cut to the WHY.
Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs)
re-armed a job every tick while a long run in another process was still
heartbeating its fire_claim, producing claim-fight churn and killing the live
run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
run_job's finally now asks the worker Future whether it is still running before deciding who tears the session down; the heartbeat test's stand-in future lacked done().
The deferral helpers from the previous commit were appended to the cron/scheduler.py facade and
carried ~70 lines of fallbacks (getattr/callable checks, result() waits, a teardown_registered
flag) for hypothetical Future doubles; the only producer is _cron_pool.submit, always a real
concurrent.futures.Future whose done()/add_done_callback() cannot raise. Also drops the
_finalize_cron_session_db passthrough. Behaviour unchanged; the test now asserts the
finalize+teardown contract on the real Future instead of a patched wrapper.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
Tests did monkeypatch.setattr(<facade module>, name) where name is now defined in a
sibling module and the production path reads the sibling's binding. Where production
reads through BOTH bindings the setattr is duplicated onto the defining module (import
added next to the existing alias import); where only the sibling reads it the target is
repointed. Seams whose production readers go through the facade are left alone.
tools/approval.py no longer re-exports sibling names (approval_context/prompt/floors/detection/
human_wait/smart/gateway_wait); it imports only what it uses. Siblings reference sibling-defined
names directly (module-attribute reads on tools.approval_context so patching the defining module
still works); only facade-owned state (_lock, _gateway_queues, _permanent_approved, _denied,
_denial_breaker_addendum, _gateway_notify_cb) is still read back through tools.approval.
approval_detection calls its own _command_detection_variants instead of late-binding through the facade.
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
cron/scheduler.py no longer re-exports the split modules (scheduler_delivery /
_script / _prompt / _preflight); it imports only the 19 names it calls itself
(bottom-of-file, E402 kept for the import cycle). Dropped the shim-only
`import shutil` and the F401 note on windows_hide_flags (still used by
scheduler.py). Split modules now call same-module helpers directly, reach
sibling split modules via late-bound module refs (_delivery/_script/_preflight)
next to _sched, and import windows_hide_flags themselves; origin-resident names
(load_config, Path, _SCRIPT_TIMEOUT, heartbeat_run_claim, ...) still go through
_sched. Callers/tests import + patch the defining module.
tools/terminal_tool.py: drop 30 pure re-export names (lifecycle/config/backends/
sudo/guards/result/interrupt/utils/_DockerEnvironment/is_managed_tool_gateway_ready)
and the noqa-F401 comments on the 25 names the facade itself uses. Sibling modules
(terminal_tool_backends/_result/_sudo/_lifecycle, environments/base, process_registry)
that read removed names through the facade now import from the defining module.
tools/environments/base.py: drop 11 re-exports (base_output/base_session_env/
path_utils) and the BaseEnvironment.stop() compat alias (no in-tree caller; the
lifecycle hasattr(env, 'stop') fallback stays for third-party envs).
tools/environments/docker.py: drop 1 re-export + the re-export comment.
Callers/tests repointed to tools.terminal_tool_{lifecycle,backends,sudo,config,
guards,result}, tools.interrupt, tools.environments.{base_output,base_session_env,
path_utils}.
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).
agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).
toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.
providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.
agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.
model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.
Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
Review fold-in on the salvage of #101877:
- `_deliver_result` routed to the durable queue whenever the worker's
`_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
delivering. A worker whose script dispatches another job in-process
(`hermes cron run <other>`) inherits that env and would have queued the
nested job's message under the outer execution id — `INSERT OR IGNORE`
then drops it silently. Match the marker against the delivering job's own
`execution_id`, as `run_one_job` already does. Regression test added
(mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
reported as a delivery error, so `mark_job_run` recorded
`last_status=delivery_failed` for a message the next gateway's drain goes
on to send, and nothing ever corrects the job record. Log and return
success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
of a second hardcoded terminal set.
Follow-up to the salvaged restart-safe worker (#101877):
- delivery_queue: a row still `pending` at the worker's wait timeout was
marked `failed` and never drained, so any gateway outage longer than the
300s budget (e.g. a restart that runs `hermes update`) silently lost the
delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
every transaction (each `get_status` poll paid for it; terminalizing
paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
(bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
`hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
process with a 1s timeout instead — the worker commits its terminal row
before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
`mark_execution_running -> None`, which now means "ownership lost, return
before run_job" — the test passed without ever reaching the path it
names. Mocking `{}` restores it (mutation-checked).
Byte-identical bodies moved out of hermes_cli/kanban_db.py into four sibling
modules, re-exported from the origin so kanban_db.<name> keeps resolving and
stays the single monkeypatch target; origin-resident helpers are reached via a
late-bound _kb namespace. AST-identity verified for all 329 moved symbols; SQL
statement multiset parity vs base. Three source-inspection tests repointed at
kanban_db_dispatch; the _add_column_if_missing alias test now imports the real
owner (hermes_cli.sqlite_util).
The referenced-script walk in cron/lifecycle_guard.py capped each file
(1 MiB) and the recursion depth (8) but not the walk: a command referencing
hundreds of scripts, or one enormous shlex token, held the GIL for minutes
on every gateway terminal call (#78398).
Add a per-walk _LifecycleScanBudget (bytes, lines, longest line, unique
paths, remote reads) charged BEFORE any text reaches shlex, and cap each
referenced read at the remaining byte budget so an oversized file is never
read whole. Exhaustion fails closed (the existing contract for one oversized
file) and is logged at WARNING so operators can tell it from a genuine
lifecycle block. Limits are sized so real wrapper graphs never hit them:
a 200-script benign graph is allowed and a restart hidden behind it is
still caught.
tools/terminal_tool.py gates its optional launchctl pre-scan (which also
tokenizes) on the same budget; the full guard still runs afterwards.
Redesigned from #83821 by @Riccardo-Vecchi, which introduced the budget
idea but blocked benign wide graphs (64-path cap) and bundled a suffix
classification change that is left out here.
Refs #78398
cron/scheduler.py (8535 -> 6353):
- _deliver_result split into per-target helpers: _resolve_target_transport (live/relay/standalone
transport + enablement), _deliver_via_live_adapter (_live_route_metadata for Telegram DM-topic
vs forum routing, _live_send_text with the cancel()-based timeout disambiguation,
_live_send_media, _seed_live_delivery_sessions), _standalone_send/_deliver_standalone. The
three interpreter-shutdown skip branches and the repeated log+append+continue pattern collapse
into one _note_target_error / one shutdown message; _TargetDelivery carries per-target state.
- Thread/channel session seeding unified into _seed_cron_session (was two near-identical
functions).
- run_job decomposed into _run_no_agent_job, _apply_monitor_gate, _load_cron_job_config
(_CronJobConfig), _resolve_job_runtime, _check_model_drift, _open_cron_session_db,
_run_agent_with_watchdog, _finalize_cron_session, plus one _run_doc_header and one _audit
closure for the success/failure paths.
- _run_one_job_body: ownership-lost bookkeeping, delivery composition and outcome classification
extracted; tick: _acquire_tick_lock/_release_tick_lock, _maybe_reap_dead_owners,
_sweep_stale_inflight_for_tick, _process_due_job, _submit_with_guard.
- _build_job_prompt: context_from injection and skill loading extracted; one
_prepend_context_block for the four fenced-data blocks.
- One _start_heartbeat_thread for the script- and fire-claim heartbeat threads.
- Dropped unreachable return in SharedRouteAdapters.get; boolean-return and nested-if shapes
collapsed.
- Comments/docstrings compacted by hand, rationale kept (fd-leak reason for the late SessionDB
close callback, title-persistence rules, no_agent classification gate, inactivity-vs-provider
timeout ordering, stale-claim force-release, interruption token keying).
Follow-up to the failure_deliver salvage (#100375):
- _preflight_check_delivery also checks the failure lane, so a typo'd
failure_deliver platform blocks at config-validation time instead of
surfacing only when a failure occurs — exactly when the notice must
not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
deliver (text normalization; empty clears the optional override
instead of coalescing), closing the one update path that could write
an unnormalized value into jobs.json.
4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
Review findings (Salt, NS-788):
B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.
S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.
S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.
Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.
- cron/scheduler.py: factor the preflight's primary-config route loader into
`_primary_profile_routes_for_current_home()` (one owner for both halves,
so route semantics cannot drift) and add `SharedRouteAdapters`, a
read-only view over the primary adapter map that resolves an adapter for
a (platform, target) ONLY when an enabled primary route with a
chat_id/thread_id maps that exact target to the current profile —
using the same `ProfileRoute.matches` predicate as inbound routing.
`_deliver_result` resolves the transport per target from it; everything
else (unmatched chat, disabled route, route for another profile, no
primary adapter, guild-only route) is a miss and never uses the primary
bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
fallback: with no matching route the view is falsy and delivers nothing.
Fixes#101113
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:
1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
the caller already has a running loop — the desktop dashboard shape) ran
the standalone sender on a fresh thread with NO profile ContextVars: the
home override and secret scope were gone, so the sender resolved the
process default's home/token (or, fail-closed under multiplex, raised
UnscopedSecretError). Wrap the submit in copy_context().run like the
session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.
2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
whose own gateway (with live adapters) is running; winning the tick-lock
race meant the adapter-less desktop ticker delivered standalone. The
multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
the desktop wires it to `_check_gateway_running(home)` so such profiles
are neither ticked nor heartbeated by the dashboard while their gateway
is alive (re-evaluated every cycle, no restart needed).
Fixes#100489
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:
- startup recovery loop: a per-profile exception (e.g. an unreadable
executions.db raising sqlite3.DatabaseError) was uncaught and killed the
ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
exception escaped to the cycle-wide handler, skipping every remaining
profile that cycle and marking all of them failed.
Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).
Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).
Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.
Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
Upgrading from a UTC-scheduling build to one that honours the profile
timezone (Europe/Brussels) left daily cron jobs sitting in jobs.json with
pre-migration instants — e.g. next_run_at "2026-09-02T04:00:00+00:00" for
expr "0 4 * * *". _ensure_aware normalizes that to 06:00+02, which the
expression excludes, so the stale-expression guard (#93049) read it as a
direct jobs.json edit, logged exactly that, and re-anchored to tomorrow
without firing. The due occurrence disappeared with no error anywhere.
The guard only asked "is the stored instant an occurrence of the current
expr?", never "why not?" — and the two possible answers demand opposite
actions. Add _classify_stale_cron_next_run, which distinguishes them by
whether normalization itself moved the wall clock:
* expr_edit — wall clock unchanged (or the stored wall clock is
not an occurrence either): the instant is genuinely
excluded by the current expression. Re-anchor
without firing, exactly as before.
* timezone_migration — the stored value's own wall clock IS a legal
occurrence and it only left the lattice because
_ensure_aware converted it to a different offset.
Fall through and fire the overdue run once.
Because every value written by this build carries the configured offset, a
real expr edit leaves the wall clock untouched and can never be reclassified
as a migration, so the #93049 protection is intact. At-most-once is
unchanged: the fire flows through the normal due path and the usual
advance_next_run / mark_job_run re-anchor rewrites next_run_at in the
current offset, so the legacy instant is never read again. Future local
wall-clock occurrences are untouched — not-yet-due rows never reach the
guard, and the #28934 offset-repair branch still runs first for a
still-future stored wall clock.
The migration case is classified explicitly rather than retried broadly: it
logs cron.timezone_migration.catch_up with the stored and normalized
instants plus both offsets, and increments a probe-visible counter
(get_timezone_migration_catchup_stats, timezone_migration_catchups.jsonl)
kept separate from catch_up_occurrences so an operator can tell "the upgrade
backlog is draining" from "runs are missing their grace window".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
The #58262 assertion lived in test_scheduler.py against a harness that has
since moved; re-home it in the delivery-confirmation suite alongside the
positive-evidence tests, and widen it to the forum-topic route and the media
route so the marker cannot drift out of any lane.