b4d7cf735d
A `browser_exec` call could park a thread in `sqlite3_sleep` permanently
while mirroring Chrome's auth DBs, holding the agent's turn open. The turn
never reaches its `finally`, so no `session.info running=false` settle is
emitted and the Desktop composer latches busy — every later message queues
and never sends. Captured live: one thread stuck 24+ minutes across two
dumps, turn accepted at 15:03 with no `tui turn finished` 46 minutes later.
Root cause is the DESTINATION, not the source. `Connection.backup()` retries
a busy destination internally and ignores the connection's busy timeout, so
`sqlite3.connect(dst, timeout=5)` cannot bound it. A destination left locked
by an earlier hung mirror therefore blocks the next mirror forever — and
because the tool-level 420s timeout abandons the thread without interrupting
a C-level lock wait, the lock is never released and every subsequent launch
re-hangs the same way. Self-perpetuating.
Two changes:
- Back up into a fresh `<dst>.new` and `os.replace()` it into place. No other
process can hold a file we just created, so there is nothing to contend on,
and the swap stays atomic. Measured against a live Chrome with a
deliberately locked destination: 0.0006s vs an indefinite hang.
- Drop the `mode=ro` (no `immutable=1`) source fallback. 8e746668ba added
`immutable=1` to fix exactly this hang but left `mode=ro` as a fallback,
keeping the unbounded path one exception away; sqlite's busy timeout does
not cover lock negotiation, so nothing bounds it. `immutable=1` is also the
semantically correct mode — a committed snapshot of a file another process
owns. The bounded plain-copy fallback is unchanged.
Tests: three regressions, all mutation-checked (fail on base, pass here).
The locked-destination test runs the copy on a worker with a join deadline so
the unfixed behaviour fails fast instead of hanging the suite. 199 passing
across the browser real-profile and CLI suites.