Clicking the Cronjobs tile shifts focus onto the tile itself, momentarily
dropping bot-chat workspace ownership — syncRoutinesPane then unregistered
the pane out from under the user's own click, with no way back. Keep the
tile while Bot Mode is on screen and the tile is the focused surface;
leaving Bot Mode still unregisters as designed. Live-verified.
Hiding the Sessions/Bots strip, or tapping the header of a lone docked
tile (Cronjobs and any plugin pane beside the workspace), left no mouse
path back: the restore menu lived on chrome the gesture just unmounted,
and a row-collapsed rail could size to 0px.
Treat hide-only chrome as stranded so `never` cannot hide those chips.
Stop collapsing on header tap (chevron only). Size a minimized zone to
MINIMIZED_TRACK and keep the horizontal strip when two or more tabs
remain.
job.get("schedule", {}).get("kind") crashes with AttributeError when
schedule is present but explicitly None (disk corruption edge case).
Use (job.get("schedule") or {}).get("kind") instead, which safely
returns False for None. This pattern is already used at other sites
in the file (e.g. cron/scheduler.py line 170).
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:
* state=completed: a one-shot that genuinely has no more occurrences,
ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
fails to produce a next occurrence (e.g. the croniter package is
missing at runtime). _mark_job_run_locked's own comment is explicit:
"Recurring jobs must NEVER be silently disabled" (issue #16265) — the
job is left enabled=True specifically so it keeps being a live,
recoverable job once the underlying issue resolves.
Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:
* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
its own is_terminal_job() check) never runs, because the check itself
skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
and claim steps the scheduler's own dispatch loop calls immediately
after get_due_jobs() for anything that DOES make it into the due list —
both also refuse the job, so even a manually-recovered next_run_at
would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
broken recurring job through the normal path.
The only way out was deleting the job and recreating it.
Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.
Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.
New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.
Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
Adds contributor email→username mappings via the new
contributors/emails/ system (one file per email) for two
compression contributors whose salvage PRs need attribution CI.
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Automatic Desktop ends (ws_orphan_reap, disconnect, idle, LRU, shutdown)
now drop the local runtime but keep the state.db row and durable-key
delegations when another live lease still owns the session.
Co-authored-by: metamindedu <metamind@kakao.com>
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.
Co-authored-by: metamindedu <metamind@kakao.com>
- The stale tick-socket sweep's os.kill(pid, 0) liveness probe sits inside
an explicit os.name == 'posix' gate (AF_UNIX nodes never exist on
Windows) — suppress with the standard inline marker.
- contributors/emails/: rodrigo.smscom@gmail.com -> rodrigogs (author of
the salvaged #92315 commits), unblocking check-attribution.
Rebase onto today's main (#94775 salvage merged): launchd_restart's drain
now goes through _graceful_restart_via_sigusr1 before any exit-wait. The
two composed witness tests feed the REAL launchd_restart os.getpid(), so
the unmocked helper delivered an actual SIGUSR1 to the pytest process
(rc=158, killed at test 18). Mock it (and _wait_for_launchd_service_pid)
in _launchd_harness + the inline harness, and accept either drain-event
shape instead of pinning the pre-#94775 ("drain", 180.0) tuple.
Addresses the review on #92315:
- Windows behavior made explicit: AF_UNIX event-loop support doesn't exist
there, so the witness is permanently absent, the payload records
loop_tick_socket=False, and stale-file probes classify UNKNOWN, never
WEDGED — deliberate fail-safe (graceful drain remains the backstop).
WSL2, the #90502 incident environment, is Linux and arms normally.
- New test asserts the default tick_timeout/tick_strikes/tick_gap_s math
stays inside the documented probe budget so retuning can't silently
blow past the 10s subprocess query tier.
Ported from #95808 (@rtcopenclawgh, closed as duplicate of this PR):
once shutdown starts loading the loop, a heartbeat task that keeps
refreshing state/gateway.heartbeat can make a draining gateway look
healthy to external probes. Cancel it alongside the watchdog and floor
timer.
The other additive piece in #95808 — passing a confirm_s confirmation
window at the update_cmd.py probe call sites — is not needed here: this
PR's probe_gateway_loop_liveness already builds the sustained window in
(tick_strikes=3 consecutive socket misses), so every caller gets it
without a new parameter.
Co-authored-by: rtcopenclawgh <rafael@amora.com.br>
/simplify-code follow-ups on the 90502 salvage:
- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
byte-identical — saw_node was effectively write-only. Collapsed to one
arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
from dead PIDs at arm time (POSIX-only liveness probe; Windows never
creates AF_UNIX nodes) so state/ does not accumulate nodes across
os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
was DISPROVED for this call site — asyncio's create_unix_server
os.remove()s an existing node before binding — but the contract is now
pinned by test_producer_rebinds_over_stale_socket_node (a live
producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
longer leaks a directory per test run.
The witness tests bind real UNIX sockets under HERMES_HOME; pytest's
default tmp_path on macOS exceeds the ~104-byte sockaddr_un limit and
bind() raises 'AF_UNIX path too long' (6 failures locally, invisible on
ubuntu CI). Module-local tmp_path override uses a short mkdtemp.
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.
WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.
New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
the loop is frozen for longer than one tick timeout but shorter than
the wedge window -> probe is ALIVE and launchd_restart drains, never
escalates; loop frozen for longer than the window -> WEDGED.
The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.
The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).
The classifier is now two-witness:
- socket answers -> ALIVE (file age irrelevant: a stalled write
or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
agree the loop stopped scheduling)
- legacy payload (no flag) -> unchanged single-witness contract: the
legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity -> UNKNOWN, never escalate
Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
`loop_heartbeat_forever` wrote the heartbeat inline on the gateway loop. That
write ends in `atomic_json_write` -> `os.fsync`, and on a stalling filesystem the
fsync blocks whichever thread runs it — which was the loop the liveness watchdog
exists to monitor. So the watchdog timed out its probe
(DEFAULT_LOOP_WATCHDOG_TIMEOUT_S = 10s, MAX_STRIKES = 3, a ~90-120s budget) and
took the hard exit, for a loop that was unresponsive because it was blocked
inside the watchdog's own liveness write.
Measured on the install that prompted this: a WSL2 VHDX under io pressure
(/proc/pressure/io full avg300 = 7.60) stalled a trivial stat-and-fsync probe at
p99 31s and max 112s — longer than the entire watchdog budget. Two of that day's
three gateway restarts carry byte-identical stack dumps parked in the heartbeat
writer.
The write now goes to a thread. Awaited, not fire-and-forget, and that distinction
is the whole design: the docstring promises that a frozen loop lets the file age,
because that staleness is how an external supervisor notices. The loop still
initiates and awaits the write, so a wedged loop still stops refreshing the file —
while a blocked fsync no longer stops the loop from answering the probe. One write
in flight at a time, so a 112s stall cannot queue a thread per interval behind it.
Scope: the timeout and strike defaults are untouched. Raising them would only
delay the same kill, and picking a budget above this box's p99 is a deployment
decision, not a fix.
2 tests. The behavioural one patches a slow write and asserts the loop still
completes ~15 ticks while it blocks; it is bounded by a fixed sleep rather than an
Event handshake so a regression fails on the tick count instead of hanging. The
second pins that the write stays awaited, since fire-and-forget would pass the
first test while destroying the staleness signal.
Verified by mutation: reverting only gateway/shutdown_watchdog.py fails both new
tests and leaves the 5 existing ones green. 15 passed across
test_loop_liveness_watchdog / test_shutdown_watchdog /
test_systemd_watchdog_lifecycle / test_watchdog_review_76354.
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.
Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.
Credit: @Finn763
Unquoted 2070 as a providers: key or custom_providers name must list, mark
current, activate, and delete instead of 500/404.
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
PyYAML loads unquoted names like provider: 2070 as int. GET /api/model/options
then called .lower() on the providers dict key and 500ed, and activate/delete
looked up "2070" and missed. Stringify identity fields so Desktop can assign
and remove those endpoints.
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Mistral ignores response_format=text and returns the full transcription
object. Desktop was dumping that JSON into the composer; pull .text out
so dictation shows the spoken words instead.
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.
Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.
Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
_stdio_children_dead() returned True ('all children dead') when a child
process was ALIVE — the liveness predicate was inverted:
for pid in pids:
if not psutil.pid_exists(pid):
continue # dead — skip
return True # BUG: an ALIVE child reported as 'all dead'
Consequence: every tools/call on a stdio MCP server with a healthy
subprocess failed instantly (~0.01-0.3s) with
'MCP stdio subprocess has exited; failing the call fast' (#81995
fast-fail path), while hermes mcp test kept passing (it never reaches
tools/call). Servers appeared dead regardless of restarts.
Fix: return False as soon as one tracked child is alive; True only when
every child has exited:
for pid in pids:
if psutil.pid_exists(pid):
return False # at least one child alive
return True
Also: on a genuinely-dead stdio (session object still present), signal
reconnect instead of a bare fast-fail TimeoutError, so the manager
respawns the subprocess instead of stranding the call slot.
Verified: alive child -> False, all exited -> True, no tracked pids ->
False (unknown, don't fail fast).
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):
- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
ignored mode'); with it, every ZIP update was refused as 'could not
check the working tree'. The mocked tests could not see this — a new
real-git test creates an actual repo + .gitignore and asserts the guard
runs clean, blocks on an ignored user file, and exempts ignored
preserved entries. --ignored=matching also reports an ignored dir as
one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
rename/copy status codes. Porcelain v1 does not quote plain filenames
with spaces, so an ignored file literally named 'venv -> node_modules'
parsed as two preserved tops and slipped past the guard into the
destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
instead of a comment-synced duplicate set (change-detector test added).
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.
Credit: @JoaoMarcos44, whose #87392 included this hardening.
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.
Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.
Original PR #32352 by @hbentel, salvaged onto current main.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).
Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.
Fixes#93820
Review follow-up on the #94033 salvage: guard the equality gate against a
future 'helpful' datetime normalization — any rewrite of next_run_at must
invalidate the marker, and normalizing would weaken that.
Carried from #94034 (closed as duplicate of #94033): explicit regression
that a jobs.json-edited stale next_run_at with NO manual_run_at marker
still re-anchors without firing — the direct #93049 protection case.
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
The generated unit only counted restart_drain_timeout, so a default
cron drain (30s + 10s cleanup) could still be inside budget when
systemd SIGKILLed the cgroup. Size TimeoutStopSec from max(drain,
cron floor + reserve) plus headroom so an in-budget stop is not killed.
cli-config.yaml.example said macOS is "always held at FULL regardless",
which reads as "your setting is ignored on this platform" and would talk
an operator out of choosing EXTRA. _apply_synchronous_pragma only refuses
values BELOW FULL on Darwin; EXTRA is applied normally.
The existing doc guard only asserted the key's presence, so it could not
have caught this. Pin the distinction instead.
Reported by @Enough1122 in review.
`apply_database_pragmas()` reads five sizing pragmas from `database:` and no
durability one, and `_enforce_macos_synchronous_full()` returns early when
`sys.platform != "darwin"`. Between them, nothing in the process ever executes
`PRAGMA synchronous` against state.db on Linux or Windows.
The effective level there is therefore `SQLITE_DEFAULT_WAL_SYNCHRONOUS`, a
compile-time constant of whichever SQLite the interpreter links. Debian and
Ubuntu builds commonly ship it as NORMAL; the bundled build and a plain
source build use FULL. So the durability of state.db is decided by which
python3 the installer found, is invisible from config, and cannot be pinned.
#90837 is three weeks of corruption forensics conducted on Ubuntu under the
stated premise `synchronous=FULL`, with every other cause eliminated live.
That premise is not something the reporter could have verified from config,
because there was no config key to set and no log line to read back.
- `resolve_synchronous_level()` maps the spellings operators actually write
(OFF/NORMAL/FULL/EXTRA, any case, or 0-3) to the PRAGMA integer, and
returns None for anything else. Kept out of the sizing loop on purpose: an
unrecognised `cache_size` harmlessly falls back to a default, an
unrecognised durability level must not.
- `_apply_synchronous_pragma()` applies it, and on Darwin refuses to lower
below FULL. `_enforce_macos_synchronous_full()` runs during
`apply_wal_with_fallback()`, which is earlier than `apply_database_pragmas()`,
so without an explicit floor a configured NORMAL would silently undo #64355
by the accident of running last. Raising to EXTRA on macOS is allowed.
- Unset changes nothing, so no existing install moves.
Tests: 34, covering the parser, application, the unset path, the typo path,
a guardrail that #77630's five keys still apply, and the Darwin floor in both
directions. Removing the wiring fails 6; removing the floor alone fails 1.
Related to #90837
apply_wal_with_fallback treats an on-disk WAL database as authoritative and
says so twice: it never live-downgrades one. The mirror case had no
protection at all. When the on-disk mode is DELETE and the configured mode
is wal, the function flips the database and logs nothing.
journal_mode is a property of the FILE, so that rewrites the header and
persists after the process exits. Setting the mode directly on the file is
something operators do; it was the documented mitigation for the SQLite
3.50.4 WAL-reset bug. A config key that makes the choice durable already
exists (database.journal_mode, #68545), but nothing named it at the moment
the PRAGMA was being undone.
#89293 reports the cost: after upgrading past the vulnerable SQLite,
is_sqlite_wal_reset_vulnerable() stopped short-circuiting into
_apply_delete_for_wal_reset_bug, the flip path went live, and 4 of 5
databases silently returned to WAL with no log line anywhere.
Add a deduped WARNING at both points where the switch succeeds, decided
before the pragma runs since both inputs are only readable while the file is
still in its original state. Log-only: the flip still happens, the return
value is unchanged, and the never-live-downgrade rule is untouched.
WARNING rather than ERROR is deliberate. The reverse direction is ERROR
because dropping to DELETE costs concurrency; this direction is normally the
desirable one (managed_uv treats a database stuck on DELETE as a bug worth
repairing on update). The problem was never the change, it was that the
change was invisible.
The page_count guard is the load-bearing half. A brand-new database also
reports journal_mode=delete and is also about to be switched to WAL, and
every opener applies WAL before creating schema, so without it the warning
would fire on the first run of every install.
Refs #89293
When database.journal_mode=delete is configured but the on-disk DB is already
WAL, apply_wal_with_fallback honors the never-live-downgrade rule and keeps WAL.
That is correct (a live downgrade under open connections causes mixed-mode
corruption), but the operator's configured mode silently has no effect, and on
a WAL-incompatible filesystem (virtiofs/NFS/SMB) the DB then corrupts on the
next crash/sleep exactly what they configured to prevent.
Two code paths return WAL in this situation; both now emit a once-per-process-
per-db_label ERROR telling the operator the config did not apply and they must
convert the DB header offline (stop connections, PRAGMA journal_mode=DELETE):
1. The WAL-reset-vulnerable path (_apply_delete_for_wal_reset_bug): previously
warned only about the vulnerability with an "upgrade SQLite" remedy, which
does not help when the real cause is the filesystem. Emitted after that
warning so the actionable message is last.
2. The read-only probe path (non-vulnerable runtime): previously returned WAL
with no signal at all.
The never-live-downgrade behavior is unchanged (existing test now also asserts
the warning). New tests cover both paths, the per-db_label dedup, and the
require_wal=True edge case.
Real-world impact: a Hermes deployment with state.db on a Podman virtiofs
bind-mount (or any NFS/SMB home) that upgrades across a version where WAL was
the default, then sets journal_mode=delete, sees no corruption protection
until the DB header is converted. This makes the gap visible. See #68545.