Files
hermes-agent/tests/cron/test_cron_request_overrides.py
T
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30

61 lines
2.4 KiB
Python

"""Standalone regression test for cron runtime request_overrides forwarding.
Split out of tests/cron/test_scheduler.py as an independent, upgrade-safe file
(registered in the local-patch ledger's allowed_untracked list) so the local
request_overrides forwarding hotfix no longer collides with upstream inserting
new tests around its anchor in the large test_scheduler.py module.
The functional change under test lives in cron/scheduler.py (run_job forwards
runtime['request_overrides'] into the ephemeral AIAgent); that hunk stays in the
source patch. Only this test moved here.
"""
from unittest.mock import patch, MagicMock
from cron.scheduler import run_job
class TestRunJobRequestOverrides:
def test_run_job_forwards_runtime_request_overrides_to_agent(self, tmp_path):
# runtime_provider may resolve provider-specific request_overrides;
# cron must pass them into the ephemeral AIAgent or scheduled jobs
# regress to SDK-default request settings.
job = {
"id": "request-overrides-job",
"name": "request-overrides",
"prompt": "hello",
}
fake_db = MagicMock()
overrides = {
"extra_headers": {
"User-Agent": "codex_cli_rs/0.138.0 (Windows 10.0.26100; x86_64)"
}
}
with patch("cron.scheduler._hermes_home", tmp_path), \
patch("cron.scheduler._resolve_origin", return_value=None), \
patch("dotenv.load_dotenv"), \
patch("hermes_state.get_shared_session_db", return_value=fake_db), \
patch(
"hermes_cli.runtime_provider.resolve_runtime_provider",
return_value={
"api_key": "test-key",
"base_url": "https://example.invalid/v1",
"provider": "custom",
"api_mode": "codex_responses",
"request_overrides": overrides,
},
), \
patch("run_agent.AIAgent") as mock_agent_cls:
mock_agent = MagicMock()
mock_agent.run_conversation.return_value = {"final_response": "ok"}
mock_agent_cls.return_value = mock_agent
success, _output, final_response, error = run_job(job)
assert success is True
assert error is None
assert final_response == "ok"
kwargs = mock_agent_cls.call_args.kwargs
assert kwargs["request_overrides"] == overrides