10 Commits

Author SHA1 Message Date
Teknium fdd32f81d4 fix(web): report a corrupt state.db as a throttled 503 status instead of a traceback per poll
The dashboard polls /api/analytics/usage and /api/analytics/models every
few seconds. When state.db is malformed the read raised straight through
the handler, so uvicorn logged a full traceback at ERROR on every poll —
one fleet host wrote ~520K identical journal entries in 24 h.

Wrap both analytics handlers in `corrupt_store_as_status`: a corrupt-image
sqlite3.DatabaseError (is_malformed_db_error) becomes a 503 with an
explicit `state_db_corrupt` payload pointing at `hermes doctor`, and the
warning is gated per store path via `{path: monotonic}` (>=300 s), then
debug. Busy/locked and every other error propagate unchanged, and the file
is never renamed or quarantined from the dashboard — repair stays with
`hermes doctor` / `hermes sessions repair`. `_session_db_path_for_profile`
is split out of `_open_session_db_for_profile` so the router can name the
store without opening it.

Refs #96591
Reported-by: #96591
2026-09-11 06:37:27 -07:00
Teknium 44a583fcc8 fix: retain profile idle activity after session removal 2026-09-07 04:56:38 -07:00
Teknium 2f090fbdec fix(serve): run idle skill maintenance on the existing timer
Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
2026-09-07 04:56:38 -07:00
kshitijk4poor b114641c88 refactor(state): fold the speculative-open retry into acquire's loop; keep real inode-swap tests
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
  the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
  type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
  those files carry no windows_only marker so they never run on Windows, and the monkeypatched
  predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
  db.close() already releases a registry-shared handle, so the described NameError leak only
  existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
  teardown; replacement not published before the last close settles) plus the raising-close
  settlement; comments say the WHY once.
2026-09-06 22:56:51 +05:30
joaomarcos aad74f26f9 fix(state): coordinate SessionDB teardown with active writers
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.

Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.

Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.

The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.

Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.

Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.

Fixes #102827

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
2026-09-06 22:56:51 +05:30
Teknium 5f1feb5344 simplify(compat): web_server — drop 221 re-exports (config/status/shutil/run_in_threadpool, lifecycle, 13 web_server_<concern> blocks, 47 route-handler legacy re-exports); web_deps.late()/LateState() take an owning-module arg; concern modules import each other directly (62 lazy sites) 2026-09-03 14:21:32 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 92e52684d1 refactor(web): tools _bad_request helper, skills hub helpers, sessions flag table, docstring compaction
- tools: _bad_request() for 16 status_code=400 raises; payload built in worker;
  drop _terminal_backend_names/_model_catalog_section one-shot helpers
- sessions: _RENAME_FLAG_SETTERS table (4 flags incl. unread), keyset export
  loop tightened, compression_root/lineage_tip compacted
- web_server_chat: single try/except for both PTY bridge imports,
  contextlib.suppress for swallow-all blocks
- docstrings/comments compacted by hand (every WHY kept); AST-neutral packing
2026-09-02 22:21:05 -07:00
Teknium 4d9007cd17 refactor(web): simplify sessions/tools/skills/chat routers and session-db helpers
- sessions: table-driven prune filters, shared scope kwargs, _project_for_display,
  _compact_json, _is_compression_edge, dedup 404 detail, flag-setter table
- web_server_sessions: descendants CTE via db._conn + dict(zip) rows,
  _is_stale_schema_error, compact docs
- tools: _BACKEND_PROBES dispatch table, _env_value/_category_providers helpers
- skills: _hub_sources/_resolve_hub_skill/_hub_lookup/_flag helpers, _API_SOURCE_IDS
- web_server_chat: _ws_client_is_allowed delegates to _ws_client_reason,
  _server_internal_ws_url unifies gateway/sidecar URL builders, _reject/_stamp_identity
- layout compaction (AST-neutral), docstring compaction keeping every WHY
2026-09-02 21:12:22 -07:00
Teknium 21f11b4d29 refactor(web_server): extract session-DB access helpers to web_server_sessions 2026-09-02 15:43:37 -07:00