db339f0051
A gateway process opened state.db from ~12 call sites, each minting its own writer connection, self._lock, close-time WAL checkpoint, and token-writer thread. With N independent writers on one WAL file, one connection's close-time checkpoint could race another's growth — the lost/reordered-page-write signature across 11+ incidents (#90837). Adds hermes_state_registry.py: a process-wide, per-path, refcounted shared registry owning the writer boundary. - acquire(path): same resolved path returns the same instance (one writer connection, one lock, one token-writer thread) for every long-lived in-process caller (gateway runner, SessionStore, per-agent lazy recall, cron per-job, mirror, channel_directory, slash_commands, shutdown_flush, session_search, react_to_message, delegate, mcp_serve, auto_archive, tui_gateway). - close() on a shared instance is a NO-OP — the registry owns the lifecycle, so one caller's close can never tear down a writer other callers still hold. - Generation-aware retirement on inode change: a replaced state.db RETIRES the live generation (never lent again) but keeps it alive for existing holders; release is object-keyed so holders of the old generation drain it independently of the new one. The old generation's own write path still fails with the typed StateDbReplacedError (existing protection, unchanged). - Replacement-open failure leaves NO registry entry for the path — the next acquire retries fresh, never hands out a closed stale object. - All teardown runs OUTSIDE the registry lock: a final release's WAL checkpoint can never stall acquisition for every state.db. - close_shared_session_dbs() at gateway shutdown drains every generation (live + retired) as the final safety net. CLI one-shots, recovery flows, and read-only cross-profile opens keep using SessionDB() directly with their own close() — only long-lived in-process sites route through the registry. References #90837 (root-cause tracker stays open: the #10 EOF signature and the WAL-lifecycle A/B verdict remain under investigation there).
61 lines
2.4 KiB
Python
61 lines
2.4 KiB
Python
"""Standalone regression test for cron runtime request_overrides forwarding.
|
|
|
|
Split out of tests/cron/test_scheduler.py as an independent, upgrade-safe file
|
|
(registered in the local-patch ledger's allowed_untracked list) so the local
|
|
request_overrides forwarding hotfix no longer collides with upstream inserting
|
|
new tests around its anchor in the large test_scheduler.py module.
|
|
|
|
The functional change under test lives in cron/scheduler.py (run_job forwards
|
|
runtime['request_overrides'] into the ephemeral AIAgent); that hunk stays in the
|
|
source patch. Only this test moved here.
|
|
"""
|
|
|
|
from unittest.mock import patch, MagicMock
|
|
|
|
from cron.scheduler import run_job
|
|
|
|
|
|
class TestRunJobRequestOverrides:
|
|
def test_run_job_forwards_runtime_request_overrides_to_agent(self, tmp_path):
|
|
# runtime_provider may resolve provider-specific request_overrides;
|
|
# cron must pass them into the ephemeral AIAgent or scheduled jobs
|
|
# regress to SDK-default request settings.
|
|
job = {
|
|
"id": "request-overrides-job",
|
|
"name": "request-overrides",
|
|
"prompt": "hello",
|
|
}
|
|
fake_db = MagicMock()
|
|
overrides = {
|
|
"extra_headers": {
|
|
"User-Agent": "codex_cli_rs/0.138.0 (Windows 10.0.26100; x86_64)"
|
|
}
|
|
}
|
|
|
|
with patch("cron.scheduler._hermes_home", tmp_path), \
|
|
patch("cron.scheduler._resolve_origin", return_value=None), \
|
|
patch("dotenv.load_dotenv"), \
|
|
patch("hermes_state.get_shared_session_db", return_value=fake_db), \
|
|
patch(
|
|
"hermes_cli.runtime_provider.resolve_runtime_provider",
|
|
return_value={
|
|
"api_key": "test-key",
|
|
"base_url": "https://example.invalid/v1",
|
|
"provider": "custom",
|
|
"api_mode": "codex_responses",
|
|
"request_overrides": overrides,
|
|
},
|
|
), \
|
|
patch("run_agent.AIAgent") as mock_agent_cls:
|
|
mock_agent = MagicMock()
|
|
mock_agent.run_conversation.return_value = {"final_response": "ok"}
|
|
mock_agent_cls.return_value = mock_agent
|
|
|
|
success, _output, final_response, error = run_job(job)
|
|
|
|
assert success is True
|
|
assert error is None
|
|
assert final_response == "ok"
|
|
kwargs = mock_agent_cls.call_args.kwargs
|
|
assert kwargs["request_overrides"] == overrides
|