Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the
hindsight_retain tool schema, threaded into the retain item's timestamp
field. When absent, the item timestamp defaults to the configured event
clock (base from PR #82928 by @ragingbulld, authorship preserved) so the
Hindsight server can resolve relative time phrases; previously no item
timestamp was ever sent and temporal memories landed with null
occurred_start/occurred_end.
Fixes#93568. Salvages #82928.
Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.
The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).
- Add optional tier fields to PricingEntry: tier_threshold_tokens,
input/output/cache_read_cost_per_million_above (None = flat, falls
back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
request once usage.prompt_tokens (input + cache read + cache write)
exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
above 200k).
- Flat entries are untouched: no threshold means no behavior change.
Reported and tier-field shape designed by @tornike14 (#93469).
Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).
- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
that mismatch a declared prefix, logging a WARNING naming the env var and
expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
are never rejected. Valid env keys still win over the pool (precedence
unchanged).
Fixes#93593
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
lifetime budget; the cap trips exactly at the Nth DELIVERED match and
promotes to notify_on_complete with the watch_disabled summary queued
right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
the global breaker drops the final match, so the user always learns why
watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
was already updated by #93532).
Refs #93513
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).
Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.
pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.
Fixes#93518.
Salvage hardening on top of #93601 (with #93597 covering the same core
mechanisms) for #93590:
- _hermes_cli(): after the venv-sibling check (hermes.exe on win32),
try shutil.which('hermes') before the bare-name fallback, so
environments with a PATH but no venv sibling resolve exactly what an
interactive shell would. Platform test switched os.name -> sys.platform
('win32') per repo convention.
- tui_gateway/methods_bot_relay.py deliver: pin encoding='utf-8',
errors='replace' on both subprocess.run sites — without them the
child's UTF-8 output is decoded with the locale codec (cp1252/GBK on
Windows), mangling non-ASCII replies or raising on undecodable bytes.
- Regression tests: shutil.which resolution step, bare-name fallback
with which=None, and encoding-pin assertions in the deliver transport
test.
Refs #93590, #93597, #93601
CI runners have a real hermes sibling next to the venv python, so
local_delivery_command now resolves an absolute path there — the exact
argv filters in the retry-policy fakes and the relay-methods pins must
match by basename instead of the literal "hermes", mirroring the
_delivery_lock matcher.
Two failures on a Windows desktop install relaying to a remote gateway
(#93590):
1. waiter_command embeds the reply path in generated python -c source
with !r. repr escapes each backslash, but the Windows execution layer
folds \\ back to \, so \U in C:\Users\... parses as a unicode escape
and SyntaxErrors the whole waiter script. Raw-string literals keep
the folded single backslash a literal; POSIX paths have no
backslashes so the prefix is a no-op there, and \' inside a raw
literal still cannot terminate the string, keeping the #93091
injection defense intact.
2. local_delivery_command hardcoded "hermes", relying on PATH — absent
in service contexts (systemd units, desktop launchers, non-login SSH
shells), so delivery died with ENOENT. It now resolves the CLI next
to this gateway's own interpreter (venv bin/Scripts sibling,
hermes.exe on Windows) with a bare-name fallback. The #93091
per-profile turn-lock recognition in bot_mode_dm now matches the CLI
element by basename (split on both separators) so resolved absolute
paths still take the lock instead of silently bypassing it.
Fixes#93590
The re-exec'd venv child spawned by
_reexec_dependency_sync_off_windows_shim completes every update step —
the receipt records success / "completed at command boundary" — but then
hangs in interpreter shutdown on a leftover non-daemon thread, freezing
the PowerShell window for minutes after "Update complete!". On the
hand-off path only (HERMES_UPDATE_REEXEC=1), after the receipt is
finalized, the update lock released, and stdio restored, flush and
os._exit(code) instead of unwinding — the same treatment #79040's cron
workaround applies. SystemExit codes (including early refusals)
propagate to the hard exit; real exceptions keep the normal raise path
so tracebacks still print. Non-hand-off invocations are untouched: the
marker env is set solely when the shim spawns the child.
Fixes#93581
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.
Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries
Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.
Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.
Builds on RelaxJonh's #93410. Fixes#93406
Complete the #93527 fix: the tombstone now carries the human prompt
text via _entry_prompt_text (handling flat prompt, ShareGPT, and
chat-style shapes), _scan_completed_prompts_by_content counts
discarded rows as completed instead of only reading ShareGPT
conversations, and the merge step reports excluded tombstones in the
combined-count summary. Adds a dedicated regression suite covering the
tombstone round-trip, the all-discarded-batch resume path, and merge
exclusion.
Salvaged from #93579 (issue reporter's PR), building on #93542.
Fixes#93527
The no-reasoning discard branch in _process_batch_worker continued
before writing any JSONL row, so run(resume=True) — which filters
solely via _scan_completed_prompts_by_content over batch_*.jsonl —
never saw discarded prompts and re-ran them at full cost on every
resume. Write a tombstone row on discard, exclude tombstones from the
trajectories.jsonl merge, and report discarded_no_reasoning in
final statistics.
Salvaged from #93542.
Fixes#93527
Adjust the #93423 max_tokens-only regression test to the merged policy:
max_tokens stays as an explicit LAST-RESORT fallback (some local servers
report nothing else) instead of being dropped entirely, and add coverage
that _reconcile_local_cached_context_length rewrites a cache entry
poisoned by the old probe (393216) upward to the real window (1048576)
once the probe is fixed.
Co-authored-by: pju-hoge <grkt@ppmz.com>
Co-authored-by: re-ITRT <1940428933@qq.com>
The local-endpoint context probe (_query_local_context_length_uncached)
treated max_tokens — an output-completion cap — as a candidate for the
model's context window. For OpenAI-compatible gateways that advertise a
1M context via context_size / max_input_tokens alongside a smaller
max_tokens output cap (e.g. TokenHub serving deepseek-v4-flash:
context_size=1048576, max_input_tokens=1048576, max_tokens=393216),
Hermes mis-detected the window as 393,216 and — because loopback
endpoints are reconciled against a live probe — actively overwrote a
previously-correct 1M cache entry.
- Add context_size and max_input_tokens to both /v1/models probe
candidate lists (single-model detail and list branches).
- Remove max_tokens from the context-length candidates; it remains
handled separately as an output cap (_MAX_COMPLETION_KEYS).
Adds regression tests covering context_size/max_input_tokens priority
over max_tokens and the max_tokens-only (no real context key) case.
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.
Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:
max_input_tokens = context window (e.g. 1M for claude-fable-5)
max_tokens = max output (e.g. 128k)
That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).
Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
PairingStore(profile=None) resolved its storage directory from the
module-level PAIRING_DIR constant, which was computed exactly once, at
module import time. A long-lived process (the gateway, started once at
container/process boot) can import this module before HERMES_HOME or a
profile's context is fully established, freezing PAIRING_DIR to a wrong
value for the rest of that process's lifetime -- even though a freshly
started, short-lived process (e.g. the `hermes pairing` CLI) re-imports
the module later with the environment already correct.
That asymmetry is exactly what made pending pairing codes issued by the
gateway process unrecoverable (the pending-code write landed under the
stale, wrong directory) while CLI-invoked writes to the same nominal
directory kept working -- see #93449 for the full writeup and a live
reproduction. tests/hermes_cli/test_dashboard_admin_endpoints.py already
carried a comment acknowledging this exact staleness in passing ("the
module-level PAIRING_DIR is bound at import"), and
TestProfileScopedStorage::test_default_store_uses_global_dir's own
comment describes working around it rather than it being intentional
behavior -- this fixes the underlying cause both were compensating for.
The profile-scoped branch already resolved its directory lazily inside
__init__ (matching this docstring's claim that resolution is lazy); this
brings the non-profile branch in line with it.
Fix keeps PAIRING_DIR as the same test seam already used throughout the
test suite (`patch("gateway.pairing.PAIRING_DIR", tmp_path)`, ~30 call
sites) unchanged: it's now a None sentinel instead of an eagerly computed
path, and a new _default_pairing_dir() helper resolves it fresh on every
call, honoring a patched (non-None) value when one is set. No existing
test needed to change.
Added a regression test that does not patch PAIRING_DIR directly and
instead exercises the real lazy-resolution path across two different
HERMES_HOME values in the same process -- confirmed it fails on the
pre-fix code (gets stuck with whatever the first PairingStore() call in
the test session happened to see) and passes with the fix.
Verified: tests/gateway/test_pairing.py (39, incl. the new one),
tests/hermes_cli/test_pairing.py, and tests/tools/test_pr_6656_regressions.py
all pass unmodified.
--reasoning takes a value (metavar=LEVEL in _parser.py) but was absent
from _TOP_LEVEL_VALUE_FLAGS (used by _first_positional_argv) and from
_apply_profile_override's value_flags set. Every invocation like
"hermes --reasoning high chat ..." therefore misclassified "high" as the
first positional, and _plugin_cli_discovery_needed() forced full eager
plugin CLI discovery at argparse-setup time - the documented startup
cost paid on every use of the reasoning override.
Add --reasoning to both sets, and add a parser-derived parity regression
test so future drift between the hand-maintained sets and
build_top_level_parser() fails CI instead of silently degrading startup
(the exact drift class AGENTS.md bans).
Fixes#93530
(cherry picked from commit 9280617ab8b308f9a8cf947f1617276a6bc8eb4f)
#93392 was not just one pattern: every hardline rule with a bare \b anchor
fired on its token anywhere in the command line, including inside quoted
prose handed to echo, git commit -m, or gh --body. Anchor the
command-name-token rules and quote-mask the positionless ones:
- dd-to-block-device and kill -1 get the same _CMDPOS anchor as the
format/rm/shutdown families, keeping their argument tails.
- redirect-to-block-device and the fork bomb have no command-name token to
anchor (`>` appears mid-command; the bomb is a function definition), so
they now match a quote-masked variant (_mask_quoted_prose) where quoted
string content is blanked. $() and backtick spans inside double quotes
stay raw (the shell executes them), and any command whose command-position
words include a shell carrier (sh/bash/zsh/ksh/dash -c, eval, source, .)
is scanned unmasked -- quoting is not a bypass. bash/sh -c payloads also
still surface as raw detection variants via _execution_flag_findings.
Regression tests cover both directions for every touched pattern: quoted
prose passes, and every true-positive shape (bare, ; && | separators,
sudo/env prefix, $(), backticks, sh -c/bash -c/eval payloads) stays on the
unconditional floor.
mkfs was the only HARDLINE_PATTERNS entry without a _CMDPOS anchor, so
the unconditional floor blocked any command that merely mentioned the
token inside quoted prose — `echo "does this workflow use mkfs
anywhere?"` was refused outright (#93392) instead of running the echo.
Anchor mkfs to command position like every sibling entry (rm root-
delete, shutdown family, dd): it matches at the start of a command,
after separators, or behind sudo/env/exec/nohup/setsid wrappers, and
no longer fires on argument-position mentions. The quote-aware
_mark_command_starts pass already keeps separators inside quoted
strings from looking like command starts, and \b still protects
mkfs_helper-style names.
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):
- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
_sig_secret helper; empty scoped group list previously meant "drop all
groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
though its sibling dm/group reads already used the scoped _getenv
All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.
Fixes#93522
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.
Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.
Fixes#93522.
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
Decode Ghostty/Kitty enhanced selection and cancellation keys, make setup cancellation terminal, and add cross-terminal previous-step navigation to setup and model flows.
Refs #92833
The desktop pools per-profile backends and reaps them after ~10 idle minutes; a reaped profile took its cron ticker with it, so its jobs silently stopped until the user next opened that profile. The primary desktop backend (which outlives the pool) now ticks every local profile store, same as a multiplex gateway (#69377 desktop sibling). External cron providers keep single-store semantics (registries are not profile-scoped); enumeration failure fails open to the active profile. Per-store .tick.lock still dedupes against live pool backends.
macOS sleep/wake (or a silent network drop) can leave the renderer's
WebSocket half-open: no close event fires, so connectionState stays
'open' while every RPC hangs until its per-call timeout. prompt.submit's
timeout is 30 minutes, so the user's next message reads as "enter does
nothing until I restart the app".
- Add a minimal ping RPC (tui_gateway/server.py) answered synchronously
on the WS reader thread.
- On wake signals, reconnectNow now probes the open-looking socket with a
5s-bounded ping and force-closes it on failure, letting the existing
reconnect machinery (backoff, tile rebinding, session refresh) take
over. A pre-ping backend answering -32601 is treated as healthy.
- Tests: half-open socket force-reconnects; healthy socket untouched;
method-not-found backend untouched; backend ping envelope contract.
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
Widen the exception guard from OSError to Exception (re-raising
KeyboardInterrupt/EOFError first) so any prompt_toolkit runtime
failure degrades to input() — matching the established pattern in
masked_secret_prompt. ValueError and RuntimeError can arise from
exotic stream wrappers or event-loop issues with the same root cause:
prompt_toolkit cannot attach stdin on the terminal.
Add test_line_input_falls_back_to_input_on_any_prompt_toolkit_failure
covering the ValueError case.
Cover the prompt_toolkit runtime-failure path added in the fix commit: a
tty-reporting stdin where prompt_toolkit raises OSError(22) (macOS kqueue
EINVAL on fd 0 under curl|bash installs) must degrade to input() instead
of aborting the setup wizard.
Replace hardcoded /opt/hermes check with dynamic install-tree detection
using Path(__file__).resolve().parent. This catches ALL install paths
(Docker /opt/hermes, apt /usr/local/lib/hermes-agent, git clone, custom)
instead of just the Docker image path. Also covers subdirectories of the
install tree, not just the top-level dir.
Add regression test test_install_tree_skipped to verify both the install
root and subdirectories are excluded from chmod.
Add contributor email mapping for bradmarshall987.
Follow-up to PR #93050 by @bradmarshall987.
Deterministic, LLM-free conformance cells against the real SessionDB with
real SIGKILL mid-write, per the tracking issue's spot-probe method:
- cell 1: acknowledged-append durability + recovery determinism (adapted
from the issue's 29.5K probe, scaled kill window, identical assertions)
- cell 2: consume-once under 8-process concurrent claim_handoff
- cell 3 (new): compression-rotation atomicity — never a compression-ended
parent without a continuation (#80337 contract; #80487 recovery context)
- cells 4-5: documented stubs interlocked with #82956-#82959 and
#83197/#83557
Journal-mode matrix (resolver default / DELETE / WAL-with-skip-gate) per
cell; every wait deadline-bounded; writers asserted alive at kill time.
Review v2 of #93084: the readable-cmdline path used substring matching
as destructive authority. That fails open on prefix collisions
(--profile timothy vs our profile tim) and same-name profiles under
different roots, so a poisoned record could still reach SIGTERM.
Ownership is now decided by the persisted identity record ALONE — exact
_same_hermes_home equality, bound to the live target by exact pid +
start-time. Missing/legacy/unbound/foreign records all refuse. A
readable argv feeds only a token-exact consistency check
(_looks_like_profile_conflict_from_cmdline via shlex tokens) that
refuses explicit contradictions like --profile timothy under tim; bare
or matching argv adds nothing.
Also fixes the review's source-of-record concern: the guard validates
the record that authorizes get_running_pid()'s answer rather than
assuming {HERMES_HOME}/gateway.pid is always the source.
Adds the requested signal-boundary regression: start_gateway(replace=
True) with unprovable ownership returns False without calling
terminate_pid or writing a takeover marker; the bound same-home
counterpart still reaches the replace flow. Legacy replace-flow tests
updated to stage a valid bound record for their legitimate-replace
fixtures.
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.
Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes the TOCTOU window flagged in review on #93369 (merged via
#93430): the divergence guard compared EXPORT-TIME message counts, but
another backend can append donor messages between the export snapshot
and the retire loop — that growth would be stamped behind the
non-recoverable adopted_by_profile archive, the exact H2 class the
guard exists to prevent, just via a narrower race.
The retire loop now re-reads live donor vs local counts immediately
before end_session and leaves the donor unretired (donor_retired=False,
warn-logged) on any donor-ahead signal; the next resume's export-time
guard then handles the divergence normally. Equal-count CONTENT
divergence (donor rewind+rewrite) remains invisible to count comparison
— documented as accepted: bytes stay in the donor store either way.
New red-first-verified regression simulates the exact race by appending
to the donor from inside an export_session_lineage wrapper.
adoption+ownership suites: 25 passed; ruff clean.
Follow-ups on top of the salvaged #92785 commit:
- Pre-gate now matches base URLs via normalize_route_base_url and
provider ids via custom_provider_aliases, mirroring the semantics of
get_custom_provider_model_capability. The raw string comparison
silently dropped declarations whose config spelling differed only by
host case or trailing slash (proven empirically: …/v1/ vs …/v1 with a
non-matching provider name returned (False, False) despite an explicit
prompt_caching: true).
- get_provider(..., allow_network=False) in the early-init/stub branch:
the policy runs per request destination (MoA aggregator, auxiliary
replans via blank_cache_policy_stub, early agent init) and a cold
models.dev cache triggered a measured ~450 ms foreground registry
fetch from the send path. A catalog miss degrades to the conservative
side.
- Debug-log the previously silent provider-lookup exception fallback.
- Tests: _make_agent defaults _custom_providers=[] (post-init reality;
keeps built-in-route tests off the catalog/config fallback), the two
early-init tests delete the attr explicitly, and three regression
tests pin the URL-drift, spaced-legacy-name, and no-network contracts
(all three fail on the unfixed commit).
Apply explicit per-model prompt_caching capabilities to custom
chat-completions routes, rather than limiting them to recognized providers,
hosts, or model families.
Keep undeclared routes conservative, derive the marker layout from the wire
transport, and leave Responses and Bedrock caching paths unchanged.
The hosted-provider misfire catch-up (fire_overdue_jobs) fired any runnable
overdue job with no one-shot grace check, so a stored past-due one-shot
bypassed ONESHOT_GRACE_SECONDS and executed arbitrarily late after downtime.
Sibling site of the due-scan gate from #89571; pins both directions with
tests.
create_job / update_job / resume_job all reject a one-shot whose run time is
more than ONESHOT_GRACE_SECONDS in the past ("will never fire"), and
_recoverable_oneshot_run_at never recovers such a schedule — but
_get_due_jobs_locked dispatched ANY one-shot whose *persisted* next_run_at was
in the past, even hours later (gateway down past the window, host asleep,
hand-edited jobs.json). A wall-clock one-shot then ran hours late, violating
the "will never fire" contract enforced everywhere else.
- Grace gate: a once-kind job whose next_run_dt is more than
ONESHOT_GRACE_SECONDS in the past is never appended to the due list.
- If no run_claim/fire_claim exists (nothing was ever dispatched), retire the
record with a diagnostic file so it stops being scanned and the miss is
operator-visible.
- If a (possibly stale) claim exists, a run may still be in flight in another
process: skip this scan but KEEP the record so its mark_job_run can land
(avoids re-introducing mid-flight record deletion).
- Manual re-trigger still works: trigger_job sets next_run_at=now (inside
grace) so an explicitly re-run stale one-shot fires.
Tests (tests/cron/test_oneshot_grace_due_scan.py): stale-not-due+retired,
within-grace-still-due, stale+claim-skipped-but-kept, retriggered-is-due, and
recurring-jobs-unaffected.
Behavior-contract tests for sanitize_task_id_for_path (colon/separator
removal, verbatim pass-through for existing safe ids, determinism,
collision-freedom incl. the a:b vs a_b digest case, traversal and
oversized-id bounds) and for the singularity persistent overlay path
(sanitized, verbatim for safe ids, distinct dirs for colon-vs-underscore
ids).
Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>