3a0e7df799
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR to a reader on a perfectly healthy database: a mode=ro connection cannot perform the WAL recovery the read needs, because recovery writes the -shm index and read-only mode refuses. The window is millisecond-scale. Today that one-shot error escapes the SessionDB read-only constructor, and GET /api/sessions turns it into a 500 the desktop reads as an authoritative empty list. Retry it, bounded, in the constructor so every read-only opener is covered — the sidebar poll, cross-profile aggregation, recall, browse — rather than at one route. A persistent IOERR still exhausts the budget and propagates. Remaining transient failures answer 503, so the client keeps the list it has. On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before the callback runs. That one is safe to retry on the same connection because nothing has been mutated; once the callback starts, settlement is unknown and the error propagates. Never close()+reopen to heal it — close() cancels this process's POSIX advisory locks on the file for every sibling connection, and a list poll's reader must stay disposable so a replaced state.db is observed and the pre-repair forensic backup stays reachable. Fixes #100436 Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com> Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>