Commit Graph

666 Commits

Author SHA1 Message Date
Teknium 444fa8166a fix(cron): routed-profile cron delivers through the shared bot for guild-scoped routes and profiles without a platforms block
SharedRouteAdapters.get called ProfileRoute.matches without guild_id, so the
documented Discord route shape (guild_id + chat_id) never authorized a cron
target and the satellite fell to standalone delivery ("DISCORD_BOT_TOKEN is
not set" every fire). A cron target has no inbound guild anchor; the route's
own guild_id is passed so its target-exact discriminators decide.

_resolve_target_transport then vetoed the authorized shared transport on the
SATELLITE's platforms.<p>.enabled (absent block or enabled: false), although
that block describes a connector the satellite never runs. The shared hit now
builds the transport directly (keeping the satellite's non-credential platform
settings), and a live native adapter with no config block is no longer read as
"disabled" (#89302) — same normalization the relay path already had.

Fixes #89302
Co-authored-by: web3blind <264741654+web3blind@users.noreply.github.com>
2026-09-11 15:28:37 -07:00
fangliquanflq c17629a0a2 fix(cron): scope restart-safe worker environment
(cherry picked from commit e57f719f942736104ee7ff79999f41b8f2a23d63)
2026-09-11 15:28:37 -07:00
joaomarcos d294a655f3 fix(cron): serialize lifecycle script reads with sqlite locks 2026-09-11 06:21:48 -07:00
Teknium 08830efd96 fix(secrets): secret-source re-pull no longer latches an empty snapshot or wipes sibling profiles
Symptom (#102041): under a multiplex gateway the default profile's vault/1Password/
Bitwarden/plugin-sourced credentials vanished for the rest of the process after the first
cron fire or the post-discovery plugin refresh; with the key already in the process env
(systemd EnvironmentFile=) the scope was empty from boot. Every get_secret() read then
failed closed ("No usable credentials", every Telegram sender rejected).

Why: _apply_external_secret_sources marked the home applied after any real fetch, but only
snapshotted names in report.provenance — the NEWLY applied ones. On a re-apply the previous
apply's own write-back makes every key `skipped_existing`, so the snapshot latched to {} and
_hydrate_profile_secret_sources returned that empty snapshot forever. Separately,
reset_secret_source_cache() was process-wide, so one profile's cron re-pull dropped every
sibling's hydrated snapshot (1aa62ceb45 isolated the routed reload but not the reset).

Change:
- env_loader: snapshot every name a source supplied (provenance + skipped_existing) from
  the home's effective environment, so a shadowed re-apply keeps the values it had.
- reset_secret_source_cache(hermes_home=None): optional per-home reset; global clear kept
  for tests/config edits.
- cron per-fire re-pull and plugins._refresh_secret_sources_after_discovery reset + reload
  only the home they resolve to.
- Two invariant tests (red on base) in tests/test_env_loader_secret_sources.py; docs note in
  secret-source-plugin.md.

Reported-by: luochen1990
Addresses #102041
2026-09-10 18:13:16 -07:00
kshitijk4poor 4a76e99f87 docs(cron): drop stale mirror-scope wording left by the home lane
- `_target_matches_origin` docstring still claimed fan-out targets are
  "deliberately NOT mirrored" and mirroring is origin-only; that has been
  false since origin_fallback landed and is doubly so with the home lane.
  Point to `_target_mirror_eligible` for the policy instead.
- cron.md DM-only bullet said "origin DM session"; the config comment in
  the same change already says "target DM". Align.
- Remove two `_expand_routing_tokens` unit assertions bolted into the
  eligibility test; the `" ALL , slack "` parametrize covers expansion
  end-to-end through `_resolve_delivery_targets`.
2026-09-10 15:32:47 +05:30
Victor Kyriazakos 22356075f1 fix(cron): make user-written bare-platform home targets continuable
A managed cron with `deliver: "slack"` (no captured origin) delivers to the
Slack home channel, but the brief was never mirrored into that session and
the in_channel seed never fired: bare-platform targets resolved with no
`_resolved_from` provenance, so `_target_mirror_eligible` treated them like
`all` broadcast expansions.

- `_resolve_delivery_targets` derives `from_broadcast` from the raw token
  and passes it to `_resolve_single_delivery_target`; a user-written bare
  platform token is tagged `_resolved_from: "home"`, `all` expansions stay
  untagged (fan-out is never continuable).
- `_target_mirror_eligible`: `home` eligible under the same flags as
  `origin_fallback` (per-job attach_to_session wins, else cron.mirror_delivery).
- `_MIRROR_PROVENANCE_RANK`: `home` == `origin_fallback` so token order
  through dedup cannot strip eligibility.
- Docs, tool schema text, cron/gateway AGENTS.md updated; existing
  exact-dict pins carry the new key.

Squashed from PR #101819 (commits 5fe612c9, 65e36926, 10e0af1d, f2cc4639):
they predate the cron/scheduler.py -> scheduler_delivery.py split and do
not cherry-pick individually onto current main.
2026-09-10 15:32:47 +05:30
Teknium ff76e65e14 fix(cron): a fire_claim whose same-host owner pid has exited is stale immediately
claim_job_for_fire refused any claim younger than FIRE_CLAIM_TTL_SECONDS (300s)
regardless of whether its owner was still alive, so a `hermes cron run` killed
mid-flight (timeout, Ctrl-C, OOM) blocked the next manual run with "already being
fired" for up to five minutes. The executions table already reaps dead owners on
sight; the job-record fire_claim now does the same via gateway.status._pid_exists
when the claim's `by` names a pid on this host. Foreign hosts, explicit
HERMES_MACHINE_ID values, and probe failures keep the TTL (fail safe).
2026-09-09 13:48:27 -07:00
Konstantin Khlopkov 111aea1a01 fix(cron): re-derive repeat defaults when a schedule update flips the kind 2026-09-09 09:20:20 -07:00
Teknium 7e4d02fef5 fix(cron): unpinned jobs run on their creation-snapshot model instead of failing closed
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).

The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.

Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".

Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).

The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
2026-09-09 04:32:13 -07:00
kshitijk4poor a73b750391 fix(cron): the dashboard "Trigger" run-now no longer stamps the next occurrence either
Second entry of the same bug class: POST /api/cron/jobs/{id}/trigger →
_fire_cron_job_for_profile → CronScheduler.fire_due → claim_fire built its claim
without `manual`, so an off-tick run from the web UI stamped the future slot exactly
like the tools path #105704 fixes. fire_due/claim_fire gain `manual` (forwarded only
when set, mirroring `force`, so third-party providers keep working) and the dashboard
trigger passes it when the provider's signature accepts it. Webhook and misfire
catch-up fires run the slot that is due and keep the stamp.

Also drops the base-green tick-stamp test (the same contract is pinned by
tests/cron/test_scheduled_occurrence.py) and documents `manual` vs `force`.
2026-09-09 12:17:13 +05:30
Phil Mossman ac10770894 fix(cron): don't stamp the next occurrence on an off-tick manual run
claim_job_for_fire() derives the occurrence identity from next_run_at
before the same function advances it. On a scheduler tick next_run_at is
the occurrence being run, which is correct; on an off-tick manual run it
is the NEXT occurrence, so the execution is stamped with the identity of
a slot that has not happened yet. _job_is_due() then finds a completed
execution carrying that identity and skips the real slot, returning
before the last_dispatch write — no error, no log line, no dispatch
record.

The manual flag already guards this and both _job_is_due() and
claim_job_for_fire() honour it; the agent-facing run-now path never
declared itself. Add a keyword-only manual= parameter and pass it from
_claim_for_manual_run(). Deliberately not force=True: force also calls
_activate_job_record(), which would resume a paused or disabled job, and
the run-now tool depends on continuing to refuse those.

The local flag is renamed to manual_fire so the new parameter is not
shadowed inside the apply closure, which would raise UnboundLocalError.

Three existing tests in tests/tools/ pinned the old call signature via
assert_called_once_with; they now pin manual=True, so dropping the flag
again fails loudly rather than silently reintroducing the skip.

Restores the intent stated in #104790 — the column records the scheduled
instant an execution was claimed for, and an off-tick manual run was
claimed for none.

Fixes #105690

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 12:17:13 +05:30
ckomma be9c2bd19c fix(cron): derive the user bus env per probe and scoped spawn
A system-level gateway unit has no ordering against user@<uid>.service and
linger may be enabled after boot, so the bus can appear after the one-shot
adoption in run_gateway() ran. Derive XDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESS
fresh for the availability probe and every scoped spawn (cron worker, Kanban
worker, PTY/pipe terminal spawns, scope cleanup) so the 60s failure TTL can
actually recover. Refs #104893.
2026-09-08 22:35:39 +05:30
Teknium 5280fe9987 fix: cron and local DMs reach an open Desktop Bot Chat
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.

Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
2026-09-07 16:48:29 -07:00
Teknium 39ed610f8c feat(cron): create paused jobs without a scheduling race
Persist paused state, timestamp, reason and no first trigger in the original
locked creation write. Forward the same boolean contract across CLI, tool,
gateway API and dashboard API, validating at the store boundary. Preserve
explicit operator force-run behavior and normal enabled creation.

The live CLI probe also caught the command shim dropping failure return codes;
forward them so invalid creation reports exit 1 rather than success.

Credit earlier atomic-creation work in #78935 and #94952 and the focused
implementation in #104578. The broader manifest staging layer is not imported.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Chloé DuPont <321112755+misschloedupont@users.noreply.github.com>
2026-09-07 08:27:50 -07:00
Teknium e1a161538a fix(cron): make failed runs diagnosable without verbose delivery errors
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.

Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:18:56 -07:00
Teknium e3710c1593 fix(cron): preserve continuity across silent audit ticks
Slim redo of #104546 and #104551: scan newest-first, match suppression only before payload separators, and keep error context. Covers wake gates and empty outputs without reading every historical file twice.

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 08:18:20 -07:00
Teknium b578261584 fix: keep Kanban worker scope out of descendant processes
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.

Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.

Refs #103974, #104058, #104904
2026-09-07 07:10:28 -07:00
Teknium 5dfea72f75 refactor: extract Kanban graph persistence into topical sibling 2026-09-07 07:09:59 -07:00
fangliquanflq e24c8499f2 fix(cron): isolate ledger helpers from stale execution modules 2026-09-07 06:01:31 -07:00
liuzikaii 5b2c417db8 fix(cron): serialize delivery deduplication with terminal retention 2026-09-07 05:58:24 -07:00
Teknium 96a903c9d2 fix: keep ambiguous cron occurrence snapshots nullable
Canonicalize the accepted raw next_run_at value before fast-forward rather
than its timezone-interpreted datetime. Legacy naive values remain runnable
but cannot establish an exact UTC identity. Preserve repaired aware slots.

Extend the existing identity invariant with the naive-slot control, and use
real ledger creation in the provider ordering test instead of an invented
execution ID that cannot pass the owner-fenced occurrence setter.

Live validation: actual builtin script run was RED (invented UTC identity)
and is now GREEN (NULL identity). Repeated builtin/provider/worker rollback
A/B remains 2 writes on base versus 1 on head, with distinct/manual controls.
Canonical cron regression rerun is queued under the campaign lock.
2026-09-07 05:57:26 -07:00
Teknium 7d6d0709cd fix(cron): skip completed scheduled occurrences after snapshot rollback
Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.

Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.

Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.

Co-authored-by: holny <holny@foxmail.com>
2026-09-07 05:57:26 -07:00
kshitijk4poor 193f05dec5 refactor(cron): one FIRE_CLAIM_TTL_SECONDS for claiming, one-shot re-arm, and stale-error recovery
The TTL was three copies of the literal 300 (claim_job_for_fire default,
rearm_oneshot, and the new stale-error guard). Hoisted so the three lanes
cannot drift; incident narrative in the guard comment cut to the WHY.
2026-09-07 00:39:44 +05:30
r3x443 eb40bf060f fix(cron): skip stale-error re-arm while another process holds a live fire_claim
Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs)
re-armed a job every tick while a long run in another process was still
heartbeating its fire_claim, producing claim-fight churn and killing the live
run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
2026-09-07 00:39:44 +05:30
kshitijk4poor 40488a4e54 refactor(cron): detached-worker teardown as a sibling module, no guards around a real Future
The deferral helpers from the previous commit were appended to the cron/scheduler.py facade and
carried ~70 lines of fallbacks (getattr/callable checks, result() waits, a teardown_registered
flag) for hypothetical Future doubles; the only producer is _cron_pool.submit, always a real
concurrent.futures.Future whose done()/add_done_callback() cannot raise. Also drops the
_finalize_cron_session_db passthrough. Behaviour unchanged; the test now asserts the
finalize+teardown contract on the real Future instead of a patched wrapper.
2026-09-06 22:56:51 +05:30
joaomarcos aad74f26f9 fix(state): coordinate SessionDB teardown with active writers
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.

Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.

Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.

The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.

Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.

Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.

Fixes #102827

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
2026-09-06 22:56:51 +05:30
kshitijk4poor fa646582ae docs(cron): point the disarm comment at _teardown_cron_agent after the helper extraction 2026-09-06 20:47:44 +05:30
RHODIZ IT 8236f37687 fix(cron): avoid reopening finalized session on teardown 2026-09-06 20:47:44 +05:30
Teknium 4441a2a28d docs(agents): split AGENTS.md into root + per-area files (≤8k each, the subdirectory-hint cap)
Root AGENTS.md 100,797 → 29,295 chars: what applies everywhere (invariants, rubric, footprint ladder, layout + shape rules, commit/PR, testing) plus a routing table. Area rules move to agent/, hermes_cli/, gateway/, tools/, plugins/, tui_gateway/, web/, skills/, cron/, apps/desktop/src/ AGENTS.md (3–9k each; ceiling is now 32k after d61cff60e3, target ~8k). Long-form process-identity and skin key tables go to website/docs/developer-guide/cli-internals.md. Zero rule loss; map in /tmp/rf/agents_md_zero_loss.md. Stale Bot Mode test paths corrected to apps/desktop/src/plugins/hermes-bots/*.test.ts.
2026-09-04 02:12:35 -07:00
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium a9b0dd6742 simplify(compat): cron scheduler — drop 82 re-exports, repoint 8 callers + 33 test files
cron/scheduler.py no longer re-exports the split modules (scheduler_delivery /
_script / _prompt / _preflight); it imports only the 19 names it calls itself
(bottom-of-file, E402 kept for the import cycle). Dropped the shim-only
`import shutil` and the F401 note on windows_hide_flags (still used by
scheduler.py). Split modules now call same-module helpers directly, reach
sibling split modules via late-bound module refs (_delivery/_script/_preflight)
next to _sched, and import windows_hide_flags themselves; origin-resident names
(load_config, Path, _SCRIPT_TIMEOUT, heartbeat_run_claim, ...) still go through
_sched. Callers/tests import + patch the defining module.
2026-09-03 13:38:36 -07:00
Teknium 707161e77b simplify(compat): file_tools/send_message_tool/cronjob_tools/process_registry — drop 34 re-exports, repoint 5 callers + 27 test files 2026-09-03 13:30:27 -07:00
Teknium fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium fb14bc4e11 review-fix(whitespace): strip trailing whitespace and EOF blank lines introduced by this PR
Trailing-whitespace-only edits so 'git diff --check BASE HEAD' is clean
(16 diagnostics across 11 files). No code changes.
2026-09-03 09:31:54 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium 23b9ffc4fa fix(integration): restore subprocess stdin=DEVNULL / utf-8 encoding guards and windows-footgun gates dropped by round-3 compaction
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
2026-09-03 02:46:19 -07:00
Teknium e1487be830 fix(integration): cron preflight reads primary config.yaml via read_user_config_raw()
Compaction put the 'config.yaml' literal within the config-read guard's
proximity window of a raw yaml.safe_load. Use the canonical raw reader with an
explicit path (same semantics: raw primary file, no defaults/overlay).
2026-09-03 02:08:03 -07:00
kshitijk4poor c3e9b28a42 fix(cron): key worker deliveries by the job's own attempt; deferred send is not a failure
Review fold-in on the salvage of #101877:

- `_deliver_result` routed to the durable queue whenever the worker's
  `_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
  delivering. A worker whose script dispatches another job in-process
  (`hermes cron run <other>`) inherits that env and would have queued the
  nested job's message under the outer execution id — `INSERT OR IGNORE`
  then drops it silently. Match the marker against the delivering job's own
  `execution_id`, as `run_one_job` already does. Regression test added
  (mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
  reported as a delivery error, so `mark_job_run` recorded
  `last_status=delivery_failed` for a message the next gateway's drain goes
  on to send, and nothing ever corrects the job record. Log and return
  success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
  of a second hardcoded terminal set.
2026-09-03 12:40:26 +05:30
kshitijk4poor b440a492b3 fix(cron): keep unsent worker deliveries queued, cheapen handoff polling
Follow-up to the salvaged restart-safe worker (#101877):

- delivery_queue: a row still `pending` at the worker's wait timeout was
  marked `failed` and never drained, so any gateway outage longer than the
  300s budget (e.g. a restart that runs `hermes update`) silently lost the
  delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
  queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
  every transaction (each `get_status` poll paid for it; terminalizing
  paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
  (bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
  `hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
  owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
  ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
  process with a 1s timeout instead — the worker commits its terminal row
  before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
  created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
  runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
  `mark_execution_running -> None`, which now means "ownership lost, return
  before run_job" — the test passed without ever reaching the path it
  names. Mocking `{}` restores it (mutation-checked).
2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 83efdf5e5e [verified] fix(cron): close restart handoff races 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 29e5172487 [verified] fix(cron): harden gateway restart handoff 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 3373e97693 [verified] fix(cron): preserve active runs across gateway restart 2026-09-03 12:40:26 +05:30
Teknium caf7fc0686 refactor(cron): compact scheduler.py docstrings by hand (rules/invariants kept) 2026-09-02 20:40:14 -07:00
Teknium ce64929054 refactor(cron): join split string literals in scheduler.py (AST-identical) 2026-09-02 20:31:54 -07:00
Teknium a432b10ad9 Merge branch 'simp/r3-09-w1' into simp/r3-09 2026-09-02 20:28:44 -07:00
Teknium 3dde5e86a7 refactor(cron): jobs.py — collapse repeated record-mutation shapes into dict.update 2026-09-02 20:22:19 -07:00
Teknium 686d8d926e Merge branch 'simp/r3-09-w2' into simp/r3-09 2026-09-02 20:21:17 -07:00