Two tests in TestReapUnsupervisedGatewayOrphansMacOS:
- test_macos_excludes_launchd_pid_from_kill: verifies a launchd-managed
PID is not SIGTERM'd while a real orphan is
- test_macos_no_orphans_when_only_launchd_gateway_running: verifies the
reaper returns False when the only gateway PID is launchd-managed
Both tests patch is_macos() to True and supports_systemd_services() to
False to simulate the macOS code path where the short-circuit does not
fire.
Background process completions on messaging platforms now default to a
one-line status message (✅/❌ + command + duration; failures append a
short output tail) instead of dumping the raw output buffer into the
chat. New display.background_process_notifications mode 'concise' is
the default; 'all' keeps the old raw-dump behavior for anyone who wants
it. Config migration v35 moves users still on the old implicit default
'all' to 'concise' on their next update; explicit result/error/off
choices are preserved.
A provider-confirmed rearm (#85846) resets the shared attempt budget, but
an earlier insufficient-progress verdict left _preflight_compression_blocked
armed, keeping the pre-API gate dark for the rest of the turn — a later
pressure spike could still grow unchecked until the provider overflow
handler fired. Clear the blocker and the stale pressure reading inside the
provider-confirmed rearm branch: the prompt is proven back below the
threshold, so the old verdict describes a request shape that no longer
exists.
Builds on @h-mascot's #84995, whose commit is preserved on this branch;
his rearm condition was superseded by #85846's latch-verified variant, but
the blocker-clear half was correct and is kept.
* fix(cli): --in accepts Git Bash / MSYS paths on Windows
Under Git Bash, 'hermes chat --in ~' reaches the CLI as /c/Users/<user>
(the shell expands ~ to an MSYS POSIX path; MSYS2 argument conversion
is disabled for native executables), and the isdir check failed with
'--in directory not found: /c/Users/...'. Route the value through the
existing _msys_to_windows_path translator (MSYS + Cygwin + WSL drive
spellings; no-op elsewhere) before expanduser/abspath.
Hit live: Bot Mode's agent-messaging protocol delivers with --in ~, so
every bot-to-bot send from a Git Bash-driven agent failed on Windows.
Tests pin both the translation cases and (source-level) the call site
actually using it.
* test: assert the MSYS translation, not platform abspath
The prior assertion ran os.path.abspath on the translated Windows path,
which on the Linux CI runner (posixpath) treats 'C:\Users\alice' as
relative and prepends the runner cwd. Pin the translation output and
ntpath absoluteness instead — same contract, platform-independent.
Windows contributors' tools default to CRLF; without repo-level
normalization an edit becomes a whole-file phantom diff (583-line
'change' observed today from one 2-line edit), string-match patch
tooling breaks on invisible \r, and review is polluted. Extend the
existing LF rules (shell/Docker) to every source/text extension:
normalize at check-in AND check out as LF so working trees match the
index on all platforms. *.ps1 stays CRLF (PowerShell 5.1 tooling).
git add --renormalize: exactly one tracked file had mixed endings
(tests/tools/test_windows_agent_loop_papercuts.py) — normalized here,
so no phantom diffs land on anyone's next commit.
A marathon tool turn burned all compression_attempts on *successful*
pre-API compactions; the gate then went permanently dark and the context
grew unchecked until the provider rejected the request terminally
("Context length exceeded: max compression attempts (3) reached", session
f087963205f9, 2026-08-01). The budget now refunds at loop top when the
assembled request is back under threshold * 0.8 AND the compressor's own
should_compress() agrees the pressure is gone.
Anti-thrash intent of the cap is preserved (#11529): no-progress passes
never reach the refund margin, divergent-signal cases (should_compress
still True) keep the budget burnt, and the insufficient-progress blocker
is untouched.
Validation: new behavioral suite (7) + all 130 compression/context tests
via scripts/run_tests_hermetic.py.
Rückbau: Commit revertieren; kein Zustand, keine Migration.
(cherry picked from commit 041b489d566bfb1d6816d53d9acd5ac50b8d5af0)
Stop freezing the xAI/xAI-OAuth catalog at import so /model and setup
pick up new Grok IDs after the models.dev cache refreshes. Put xai and
xai-oauth on the shared picker-time models.dev merge path and pin
grok-4.6 as the default headline model.
Box cloud content management via the official @box/cli through the
terminal tool: files, folders, sharing, search, metadata, Box AI,
Hubs, bulk operations, webhooks, and a REST fallback via box request.
OAuth-only auth; SKILL.md routes to ten scoped reference files.
Salvaged from PR #52107 by @iskysun96.
Regression tests salvaged from PR #70522 by @JoaoMarcos44. The Qwen flat
cached_tokens behavior is provided by the shared top-level fallback from
PR #66105 (@mehmetkr-31); the codex cache_write_tokens read landed in the
previous commit.
Salvaged from PR #85702 by @JoaoMarcos44, composed onto the mapping-safe
_usage_get reads (PR #74591 by @RelaxJonh) and the flat cached_tokens /
Anthropic-name fallbacks (PRs #66105, #52571):
- cache-write precedence in the chat_completions branch:
details.cache_write_tokens > details.cache_creation_input_tokens >
usage.cache_creation_input_tokens > usage.cache_write_tokens
- codex_responses branch reads details.cache_write_tokens (GPT-5.6+
documented name) with cache_creation_tokens fallback (from PR #70522)
- _usage_count(): clamp malformed negative counters to 0
- all reads in every branch are mapping-safe via _usage_get
When the Responses API returns usage as a plain dict (e.g. from a
middleware or proxy that deserialises JSON to dict instead of a typed
SDK object), normalize_usage() used getattr() exclusively, which
silently returned 0 for every field on a dict.
Add _usage_get() helper that reads via .get() for dicts and getattr()
for attribute-style objects. All accessor sites in normalize_usage()
now use this helper, so token counts and cost are correct regardless
of the usage object's type.
Regression tests: two new tests feed the same payload as both a dict
and a SimpleNamespace through the codex_responses and
chat_completions branches, asserting identical output and non-zero
values.
Kimi/Moonshot's native API (api.moonshot.cn / .ai) reports context-cache hits
as a top-level ``usage.cached_tokens``. The chat-completions branch of
normalize_usage() walks a fallback chain of
prompt_tokens_details.cached_tokens -> cache_read_input_tokens ->
prompt_cache_hit_tokens; none of those names match, so direct Kimi sessions
normalized to cache_read_tokens=0. The hits were invisible in accounting and
the cached prefix was billed at the full input rate.
Appended as the last link in that chain, so it only fills a genuine zero and
cannot override a provider that reports the nested OpenAI shape or DeepSeek's
prompt_cache_hit_tokens.
Rebuilt on current main rather than rebased — the branch was ~3400 commits
behind. The DeepSeek half of the original branch is dropped: 03c0b00f4
(#65678) landed prompt_cache_hit_tokens on main, so this is Kimi-only as the
review asked. The scripts/release.py addition to the frozen LEGACY_AUTHOR_MAP
is dropped too; contributors/emails/mehmet.kar@std.yildiz.edu.tr already
exists on main.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The command wrapper prints the cwd marker after the command returns. A
killed or timed-out command emits no marker, so ``env.cwd`` still holds the
directory of the last command to FINISH. One local environment serves every
session, because ``_resolve_container_task_id`` collapses cwd-only overrides
to ``"default"``. That leftover directory is therefore routinely another
session's.
The post-command dual-write copied ``env.cwd`` into the interrupted session's
durable record. Every later command in that session then ran in the foreign
directory, and the cwd echo told the model it had moved there. A desktop chat
silently re-homed into a worktree that another chat had opened.
Report the observation instead of inferring it. The marker parse now sets
``result["cwd_observed"]``, and both the record write and the echo read that
flag. The local override clears the flag when it rolls back a path that does
not exist, because the restored value is also unobserved. When a command
reports no cwd, the session keeps the directory it already had.
This needs no second session to be wrong: a lone session that interrupts a
command re-adopts a stale value too. A second session only makes the wrong
directory belong to somebody else.
The same class of write exists in the file-tools rescue for a reaped
environment (#26211). That rescue copied the cached snapshot of the shared
``env.cwd`` into the session record. The rescue is now fill-only: it writes
the snapshot when the session has no record, and it never overwrites a
record that the session wrote for itself.
The tests drive ``terminal_tool`` itself through an interrupt, not a copy of
its gate. Review found that a revert of either call site passed the first
version of the tests. Each gate now has a test that fails when the gate is
removed (verified by mutation).
Two exact-dict assertions in the Vercel sandbox tests now assert the two
fields they care about, so a new result key does not fail them.
Extends the NS-656 memory-pressure surface to cover disk exhaustion
(OOF-2 / OOF-107 lineage: agents fill their data volume — SQLite writes
fail, sessions stop persisting — while every dashboard looks healthy).
- gateway/disk_status.py (new): collect_disk_status() samples
shutil.disk_usage(HERMES_HOME) and classifies pressure
(critical: <256 MB free or >=95% used; elevated: <512 MB free).
Never raises — degrades to pressure="unknown" with null telemetry,
same contract as collect_memory_status().
- /api/status: sibling `disk` block next to `memory`, advisory only —
not folded into component/overall health.
- web: DiskPressureStatus type; MemoryPressureBanner generalized to a
resource banner with worst-first triggers (disk critical > memory
critical > OOM restart > disk elevated > memory elevated) and
cascading dismissals — hiding the top trigger surfaces the next one
instead of silencing everything. All dismissals stay boot_id-scoped.
- i18n: diskCriticalBanner / diskElevatedBanner (en, optional fields
with English fallback per existing pattern).
Tests: gateway/test_disk_status.py (14), web_server disk-block
presence/degradation, banner disk trigger/priority/dismissal-cascade
suite (21 total).
Addresses the human review findings on the memory-pressure feature:
* [P2] Dismissal hid later incidents of the same kind. The gateway now
publishes `boot_id` (the lifecycle sentinel's started_at — changes on
every gateway life) in the /api/status memory block, and the dashboard
keys OOM-restart dismissal on it: acknowledging one restart no longer
mutes the NEXT one (the OOM-loop case this banner exists for). Live
pressure dismissals now also reset once pressure is demonstrably back
to "ok" — "unknown" (stale heartbeat) is absence of evidence and
clears nothing. Dismissal storage moved to a JSON list; old bare-string
entries fail JSON.parse and degrade to a clean reset.
* [P2] suspected_oom is a heuristic (unclean exit + low-memory final
heartbeat), not proof the OOM killer acted — banner copy now says
"restarted unexpectedly, most likely because it ran out of memory"
instead of stating OOM as fact.
* [P3] Mobile header clearance was applied per-banner (mt-14 on both
MemoryPressureBanner and ProfileScopeBanner) AND on the content
(pt-14), double/triple-stacking 56px gaps when banners were visible.
Replaced with a single h-14 spacer above the banner stack.
Hosted agents can be OOM-killed hourly while the dashboard and the NAS
agent card both look perfectly healthy — every memory signal the gateway
already produces (heartbeat mem samples, lifecycle-ledger unclean-exit
verdicts, cache-pressure evictions) dies in server-side log files. The
BlueAtlas incident (NS-608) ran for three days like this.
This is the read-side fix:
* New gateway/memory_status.py distills the existing 30s loop heartbeat
(gateway RSS + system MemAvailable/MemTotal + swap) and the lifecycle
sentinel into a compact `memory` block: pressure ok/elevated/critical/
unknown, coarse MB numbers, and last-boot unclean/suspected-OOM flags.
Pure file reads, no new sampling, no gateway IPC. Stale (>150s) or
future-dated heartbeats degrade pressure to "unknown" so a dead
gateway's final gasp can't render a live "critical" banner forever.
Critical thresholds mirror the ledger's OOM-suspicion heuristics: if a
level would make a later unclean death "suspected OOM", warn at that
level while the process is still alive.
* lifecycle_ledger.record_startup now carries prior_unclean_exit /
prior_suspected_oom onto the reclaimed sentinel — previously the
verdict survived only in append-only diag prose. Flags age out on the
next sentinel rewrite (scoped to the life after the crash).
* /api/status serves the block (profile-aware, executor-offloaded,
fail-safe to pressure=unknown). Deliberately NOT folded into
components/overall: memory pressure is advisory, and flipping overall
to "degraded" on it would page NAS's availability sweep for a
condition the eviction valve is already handling. Public-safety:
coarse numbers/enums/booleans only — same disclosure class as the
existing nous_session_valid field, added for the same NAS-sweep
audience.
* Dashboard: new MemoryPressureBanner (app-shell, next to
ProfileScopeBanner) with worst-first trigger precedence
(critical > suspected-OOM restart > elevated), per-trigger
session-scoped dismissal, and escalation re-opening past a dismissal.
i18n keys optional with English fallbacks, matching the
managingProfileBanner convention.
Tests: gateway/test_memory_status.py (classification bands, staleness,
clock skew, corrupt files, bool-is-not-int), lifecycle sentinel
carry-forward, /api/status contract (block always present, collector
crash degrades instead of 500), and 7 banner component tests.
NAS-side ingestion (agent-card notice + memory-tier upsell) ships
separately.
Refs NS-656; context: NS-608, NS-657, OOF-77.
Azure Foundry's OpenAI-compatible Responses surface rejects the post-tool
follow-up payload with HTTP 400 `invalid_payload` when a replayed encrypted
`reasoning` item is sent alongside `function_call` / `function_call_output`.
The initial function-call request and ordinary multi-turn continuity are both
accepted, so the failure only appears after the first tool executes.
Detect the Foundry endpoint in `ResponsesApiTransport.build_kwargs` and drop
only the encrypted reasoning replay on that follow-up turn, leaving
function_call / function_call_output continuity intact.
Salvage of #59981, rebuilt on current main. Same root cause and fix direction
as the original, which was correct; this version resolves three defects:
- No `chat_completion_helpers.py` change. main already forwards `provider`
and `base_url` to the Responses transport, so the original's re-added
arguments produced `SyntaxError: keyword argument repeated: provider` on
merge. Dropping the hunk removed the syntax error and the conflict.
- Host matching uses `utils.base_url_host_matches`, not a substring test.
`".services.ai.azure.com" in base_url` also matches URLs carrying the
domain in a path or query segment, which would silently disable reasoning
replay on an unrelated provider.
- The post-tool predicate tests the trailing messages, not the whole history.
Scanning for any tool call plus any tool result made it sticky: one tool
call early in a conversation suppressed reasoning on every later turn.
- Tool calls pair on `call_id` as well as `id`. Responses histories carry the
function call id in `call_id` while `id` holds the response item id
(`fc_...`). Identity is resolved via the converter's own
`_split_responses_tool_id`, covering composite `"call_x|fc_y"` ids and bare
`fc_` ids on both sides of the pairing.
Tests: 27 cases across the transport and the live `build_api_kwargs` bridge,
including six parametrized tool-call id shapes, non-Foundry host lookalikes,
the sticky-history guard, parallel tool results, and an unpaired tool result.
Each guard was confirmed to catch its defect by reverting the fix.
Verified with `scripts/run_tests.sh tests/agent/ tests/run_agent/`:
532 files, 5602 tests passed, 0 failed.
Not verified against a live Azure Foundry endpoint — no credentials. The
original HTTP 400 reproduction and post-fix Foundry Project / Azure Container
Apps harness runs are @AshuJoshi's, from #59981. This change is verified at
the payload-construction layer only.
Closes#59981.
Co-authored-by: Ashu Joshi <AshuJoshi@users.noreply.github.com>
- Cold force_refresh (fresh CLI process, e.g. hermes config refresh)
now hydrates the memory cache from disk before fetching, so the
conditional GET actually fires on the flow the feature was built for
instead of silently re-downloading the full ~2 MB registry
(empirically probed: If-None-Match sent, 304 serves disk data).
- Conditional-GET decision is passed in explicitly
(_fetch_models_dev_from_network(conditional=...)) by callers holding
the fetch lock, removing the hidden read of module globals inside
the fetch; the background worker now fetches INSIDE the lock,
symmetric with foreground (true singleflight — no concurrent
double-download, no fetching against mid-commit etag state).
- Corrupt disk cache is QUARANTINED (renamed to .json.corrupt) rather
than left in place: rejection becomes a one-time event instead of a
re-read + re-parse + warning + unlink on every hot-path call while
offline (probed: 1 warning across 5 calls, was 5).
- Dropped the dead _DEFAULT_MODELS_DEV_URL constant; module and
function docstrings updated to match the servable-cache conditional
semantics.
- Conditional GET now requires a servable in-memory registry: an
If-None-Match sent while holding no cache invited a 304 against
nothing, permanently serving {} with a blocking foreground fetch on
every call (the exact #35838 class this PR fixes). Empirically
repro'd and verified fixed (corrupt cache + stale sidecar: was 3
calls -> {} forever; now 1 unconditional fetch -> real data).
- ETag persists atomically WITH the cache body via
_commit_registry -> _save_disk_cache(data, etag), wiring up the
previously-dead etag param; the sidecar can no longer get ahead of
the registry it vouches for. _save_etag now uses
utils.atomic_write_text (unique tempnames + fsync) instead of a
hand-rolled fixed-name .tmp replace.
- Corrupt/unreadable disk cache clears the ETag sidecar so the
refetch is unconditional; _confirm_cache_not_modified keeps a
defense-in-depth guard (clear sidecar + arm backoff) should a 304
ever land on an empty registry.
- allow_network=True paths use the zero-arg fetch_models_dev() call
shape at all sites (was 1 of 5) — ~46 test sites monkeypatch it
with zero-arg lambdas; the unconditional kwarg broke
test_xiaomi_provider (verified fail->pass).
- _get_models_dev_url falls back to the MODELS_DEV_URL module global
(not the constant) so existing patch sites keep working.
- Tests: replaced two mock-riddled corrupt-cache tests with real
tmp_path file tests; added regression tests for the 304/empty-cache
loop, sidecar clearing, and conditional-GET gating.
Harden the models.dev catalog refresh path (#35838) with three missing
pieces:
1. ETag conditional GET — every network request sends If-None-Match
with the last-known ETag (persisted alongside the cache file). A 304
Not Modified re-confirms the existing cache without re-downloading
the full ~2 MB registry. This makes the 4-hour TTL effectively free
to maintain.
2. No-network-on-hot-paths invariant — allow_network=False is now the
default for every query function called on the conversation hot path:
get_model_capabilities, get_model_info, lookup_models_dev_context,
_get_provider_models. These are called during vision routing, image
routing, cost-guard checks, and context-length resolution on every
turn — they must never block on the network. Interactive flows
(model picker, model switch) explicitly pass allow_network=True.
3. Mirror URL override — models_dev.url in config.yaml lets deployments
point at a self-hosted mirror without code changes. Follows the same
pattern as model_catalog.url.
Additional hardening:
- Cache TTL bumped from 1h to 4h (ETag makes refresh cheap)
- Corrupt/empty disk cache is rejected with a warning instead of being
served as {} and silently breaking provider/model resolution
- _validate_registry() guards against non-dict and empty-dict payloads
Fixes#35838
The desktop settings page saves approvals.mode through REST PUT /api/config
(and the raw editor through PUT /api/config/raw). Enforcement follows the
file immediately, because the approval gate re-reads config per command, but
every live session's YOLO/approval indicator repaints only on a session.info
event, and the REST save emitted nothing. The indicator showed bypass OFF
while approvals.mode=off silently auto-approved every dangerous command, and
switching sessions repainted the stale cached per-session state, making the
toggle look like it flipped itself back. The gateway /approvals slash
command had the same gap.
The config.set RPC handler already re-emits session.info to all live
sessions after a mode flip; give the other writers the same contract:
- tui_gateway/server.py: add broadcast_session_info(), which snapshots
_sessions under _sessions_lock and re-emits via
_emit_session_info_for_session. Also call it from the /approvals slash
mirror when a mode argument was persisted (bare /approvals is read-only).
- hermes_cli/web_server.py: after a REST save that actually changed the
normalized approvals.mode, call the broadcast through a sys.modules guard
(no gateway imported means no sessions to notify). The comparison runs on
the in-memory documents (existing vs merged, parsed vs raw): the settings
page PUTs the defaulted GET record while disk holds sparse YAML, so a
block-level compare would broadcast on every autosave, and re-reading
through the config cache after the save could serve the pre-save document
on an (mtime_ns, size) key collision. Own-profile saves only: a
profile-scoped save targets a different HERMES_HOME than this process's
gateway sessions.
No broadcast on saves that leave the effective mode unchanged, so settings
autosave churn (skin, font, TTS) can't spam session.info.
Scope: reaches sessions of the in-process gateway (hermes serve / hermes
dashboard, the topologies the desktop app talks to). A spawned
tui_gateway.entry child gateway has its own process and _sessions; its TUI
statusbar reconciles each turn via the existing session.info emissions.
When a prompt_toolkit run_in_terminal cooked->raw restore is lost (cancelled
coroutine, racing chained cross-thread windows from background-review
summaries / process-notification prints), the tty stays in cooked mode while
the Application still expects raw. The kernel line-buffers keystrokes and the
CLI appears to stop taking input even though the event loop is healthy.
Observed live 2026-08-13: interactive session left in 'icanon echo' after a
background skill-review fork + notify_on_complete turn; only an external
stty rescue restored input.
Fix: _heal_cooked_mode_drift() re-applies prompt_toolkit's own raw-mode flag
surgery when stdin's lflag has drifted cooked, and process_loop's idle branch
runs a rate-limited _check_termios_drift() watchdog that skips legitimate
cooked windows (app._running_in_terminal), agent-running phases, non-tty
stdin, and Windows.
patch_session_model_config merges key-level and only deletes on explicit
None. Dropping falsy values from the top-level patch let a previous
switch's api_mode/base_url survive the next switch — TUI/desktop resume
then restored e.g. openrouter with anthropic_messages wire mode, and a
failed bare-custom heal produced a stale-provider/new-endpoint route.
Write absent top-level values as explicit None so each switch fully
replaces the persisted route. Regression test against a real SessionDB;
mutation-checked. Also correct the heal comment (CLI is deliberately
stricter than the TUI recovery, which keeps bare custom with a base_url).
- Bare 'custom' from ModelSwitchResult.target_provider is the resolved
billing class, not a routable identity; persisting it verbatim made a
later --resume hard-fail once the config default moved off the custom
endpoint. Heal to custom:<name> via canonical_custom_identity at
persist time, and again on restore for rows written by older builds
(mirrors tui_gateway's _stored_session_runtime_overrides recovery).
- --global switches now also update the session row: the row records
what THIS session runs, otherwise resume restored the stale
creation-time model over the user's new global choice.
- Only adopt resolved credential_pool alongside its api_key (don't null
the ambient pool when resolution returns no credentials).
- 3 new tests; healing path mutation-checked.
- Extract the two duplicated /model session-persist blocks into
_persist_model_switch_to_session; persist the route BOTH nested
(gateway_runtime, CLI reader) and top-level (TUI gateway's
_stored_session_runtime_overrides reader) so a CLI switch also
survives a desktop/TUI session.resume.
- Add SessionDB.session_gateway_runtime as the canonical tolerant
row-level route reader (session_yolo_enabled precedent); use it in
_restore_session_model instead of hand-rolled JSON parsing.
- Clear stale launch-time _explicit_api_key/_explicit_base_url when
resume restores a different provider (same leak guard
_apply_model_switch_result already has).
- 12 new tests incl. a real-SessionDB round trip; mutation-checked.
- Single-source the included note as _INCLUDED_NOTE and attach it at
BOTH status='included' sites (the zero-amount pricing-entry branch
previously returned the same status with no note).
- Docstring/comment precision on format_cost_label: the fallback
triggers on 4dp ROUNDING to 0.0000 (banker's rounding includes the
exact $0.00005 boundary), not truncation; note why the rendered-label
guard beats a naive Decimal threshold.
- Tests: replaced a dead assertion with the exact-boundary case
($0.00005), fixed an overclaiming comment, aligned the terminal
cost column.
- Insights formatters now route aggregate estimated cost through the
shared format_cost_label() instead of hardcoded 2dp — a sub-cent
aggregate (one cheap DeepSeek session, ~$0.0046) no longer renders
'Estimated: ~$0.00', the exact bug class this PR fixes (#79220).
- format_cost_label: positive amounts below $0.00005 render '~$<0.0001'
instead of the zero-looking '~$0.0000' 4dp truncation artifact.
- Renamed _format_cost_label -> format_cost_label (now a cross-module
shared helper).
- Tests: renamed test_gateway_format_hides_cost ->
test_gateway_format_hides_cache_details and
test_no_cost_section_when_all_zero ->
test_unknown_bucket_shown_for_costless_session (names contradicted
behavior); restored a real assertion in the custom-models test that
had been weakened to a comment; added sub-cent-aggregate and 4dp-floor
contract tests (mutation-checked).
Three cost-display honesty fixes:
1. Sub-cent cost label rendering (#79220) — _format_cost_label() scales
precision to magnitude: zero renders as '$0.00', sub-cent (< $0.01)
renders at 4 decimal places (e.g. '~$0.0046'), normal costs keep 2dp.
This fixes the bug where DeepSeek per-turn costs of $0.004640 rendered
as '~$0.00' despite amount_usd carrying full Decimal precision.
2. Cost bucket surfacing (#77223) — insights format_terminal and
format_gateway now display three cost buckets: estimated (with dollar
figure), included (session count, labeled 'subscription — no provider
invoice'), and unknown (session count, labeled 'no pricing data').
Previously, included and unknown sessions silently collapsed to $0 in
the aggregate view, hiding 315 of 473 sessions in the reporter's DB.
3. Subscription-included cost notes — estimate_usage_cost now attaches a
'subscription-included; no provider invoice for usage' note to
CostResult for subscription-included routes (openai-codex), so
consumers can distinguish 'free because subscription' from 'free
because $0 pricing'.
Fixes#79220Fixes#77223
- Deleted the id(cfg)-keyed _OVERRIDE_CACHE layer: id() is unique only
among live objects, so a config reload could serve stale overrides
forever when CPython reuses the freed dict's address. The upstream
load_config_readonly is already (mtime,size)-cached (~1 stat/hit), so
the local layer was redundant state with a correctness risk.
- _override_to_catalog_shape returns (patch, vision) instead of
smuggling an in-band _vision_override sentinel key through the merged
dict; removed the two dead call-site pops.
- _find_model_entry gains the :cloud/-cloud suffix fallback that
lookup_models_dev_context already had, so 'catalog hit' means the
same thing to every consumer — a suffix-keyed model (kimi-k2.6:cloud)
now counts as KNOWN and keeps its catalog capabilities instead of
being displaced by a fill-gap _default (mutation-checked contract
test added).
- get_model_info's unknown-model override path seeds the same safe
defaults as get_model_capabilities (200K ctx, tools on, 8192 out),
so a partial override no longer yields ctx=0/tools-off on that path
(contract test added); the DEFAULT_CONFIG defaults claim is now true
for both paths.
- Activated the previously-dead _MODELS_DEV_TO_PROVIDER reverse map
(lazily built, many-to-one aware) and used it in
_provider_override_section instead of a per-call linear scan.
Review follow-ups on the model_overrides feature:
- ONE canonical override schema everywhere. get_model_info previously
merged the override dict raw into the models.dev catalog shape
({**raw, **override}), so the documented context_window/supports_*
keys silently did nothing on that path (cost guard, inventory) while
working in capabilities/context paths — same config key, two
incompatible schemas. Overrides are now translated into the catalog
shape at the get_model_info boundary (_override_to_catalog_shape),
and sub-dicts (limit, modalities) are MERGED, not clobbered — an
override setting only context_window no longer wipes the catalog's
limit.output.
- _default is now a FILL-GAP default, not an override: it applies only
to models the catalog does not know (the #8731/#84482 self-unblock
path) and never displaces catalog data. A
_default: {context_window: 128000} can no longer clamp every model
of a provider. Explicit per-provider+model entries keep their
win-over-catalog semantics.
- Early-chain _override_context_window (model_metadata step 0b) is
explicit-only, so a _default can never preempt custom_providers
per-model settings or live probes; fill-gap defaults apply at the
lookup_models_dev_context catalog-miss boundary (step 5f) instead.
This fixes the precedence inversion where a provider/global _default
silently overrode an explicit per-endpoint per-model context_length.
- Provider keys accept BOTH id spaces (Hermes id and models.dev id:
copilot/github-copilot both work) and model ids match
case-insensitively, mirroring catalog lookup.
- Malformed override values (context_window: '512k') log a one-shot
warning instead of being silently swallowed.
- DEFAULT_CONFIG comment: removed the false family/dated-snapshot
inheritance claim, documented the recognized field list, fill-gap
semantics, and the id-space rule.
- Tests: rewritten for the new contracts (fill-gap invariants,
dual-id-space keys, sub-dict merge preservation, one-shot warning);
added a real-config-yaml e2e plumbing test (mutation-checked: fails
when the config key wiring is broken).
Add a unified model_overrides config section that lets users manually
declare context_window, max_output_tokens, capabilities, cost, and
family for any provider+model — winning over models.dev, OpenRouter, and
hardcoded defaults.
Resolution order (first hit wins):
1. model_overrides.<provider>.<model_id> (per-provider+model)
2. model_overrides.<provider>._default (per-provider default)
3. model_overrides._default (global default)
4. Normal catalog resolution
Key subtlety: an unknown model id (not in the
catalog) derives base metadata from sensible defaults before patching,
so overriding a model the catalog doesn't know yet is the supported
self-unblock path. This is exactly the #84482 scenario (Upstage
solar-pro4/syn-pro wrong context) and the #8731 scenario (custom/local
models with manual capability declaration).
Wired into:
- get_model_capabilities() — patches capability fields; unknown models
get safe defaults (tools on, vision/reasoning off) before patching
- lookup_models_dev_context() — context_window override, checked before
catalog lookup so it works even for providers not in PROVIDER_TO_MODELS_DEV
- get_model_info() — merges override dict onto catalog entry (shallow
merge); for unknown models, the override is the sole source of metadata
- get_model_context_length() — step 0b in the resolution pipeline,
before custom_providers (0c) and before any network probe
Config example:
model_overrides:
upstage:
solar-pro4:
context_window: 524288
syn-pro:
context_window: 65536
custom:my-local-vllm:
my-llava-model:
context_window: 8192
supports_vision: true
supports_reasoning: false
supports_tools: true
_default:
context_window: 128000
Fixes#8731Fixes#84482
Refs #47247
BARE_BILLING_PROVIDERS incorrectly included "openrouter" alongside
"auto" and "custom". OpenRouter is a fully routable provider with
its own API key and base_url — sessions that used OpenRouter store
billing_provider="openrouter", and dropping it forces resume to the
current global model (e.g. a custom endpoint), which is the wrong
provider for the stored model.
Remove "openrouter" from the bare-bucket set so OpenRouter sessions
correctly restore their provider identity on resume.
Fixes#57588
Maintainer fixup on the #79787 salvage:
- An explicit fb.api_mode of "chat_completions" was silently overridden
by the codex_responses / bedrock re-detection pass (which only skipped
re-detection when the pre-computed mode was non-default). Track
explicitness in fb_api_mode_explicit and gate the whole re-detection
block on it.
- Replace the locals().get('fb_api_mode') dead-code hack with clean code
(fb_api_mode is always bound at that point).
- Restore the post-resolve /anthropic + api.anthropic.com host check for
named custom providers whose base_url comes from config rather than
the fallback entry (#32243, #49247), which the PR's restructure dropped.
- Add regression tests: explicit api_mode honored (incl. explicit
chat_completions not overridden), /anthropic-hint fallback detected
pre-rewrite, api_mode forwarded to resolve_provider_client, plain
fallback unchanged.
- /model switch now refreshes agent._custom_providers from the config
loaded during the switch before re-evaluating cache policy — a
prompt_caching flag added to config.yaml after session start was
invisible to a mid-session switch (policy read the stale init-time
snapshot while context_length resolution used the live list).
- Production-path test: real config.yaml in the modern providers: dict
shape through the real loader chain, exercising the init-order fallback
(no _custom_providers attr) for both the fable opt-in and the opus
explicit opt-out.
- Pin operator kill-switch precedence: _cache_disabled (prompt_caching.
cache_ttl falsy) beats an explicit per-model prompt_caching: true.
- Log (debug) instead of silently swallowing capability-lookup failures in
anthropic_prompt_cache_policy — a swallowed failure would otherwise
downgrade an explicit prompt_caching: true to (False, False) with zero
trace. Matches the sibling MoA branch's logger.debug style.
- Use load_config_readonly() for the None-fallback in
get_custom_provider_model_capability: the helper only reads, and the
fallback fires on the blank-stub paths (agent init before
_custom_providers is assigned, MoA/auxiliary destination planning), so
skip the ~135us defensive deepcopy per call.
- Add route-isolation regression tests at both levels (config helper +
agent policy): a prompt_caching declaration for one provider route must
never apply to another route with the same model name. Mutation-checked:
both tests fail when the URL match is disabled.
Now that subscriptions survive `done` (completion is reversible —
on every 5s notifier tick forever. Add
kanban_db.purge_stale_done_notify_subs(): one DELETE removing subs
whose task has been done with no new events past a retention window
(age = latest task event, falling back to completed_at/created_at, so
any activity exempts the task; a reopened task is exempt by status
alone). The notifier watcher runs it per board once at startup and at
most hourly, re-reading kanban.done_sub_retention_days (config.yaml,
default 30; 0 disables) at each sweep.