The consumed-but-uncommitted rotation verdict was process-local
(_SPENT_ROTATION_FINGERPRINTS), while the credential it protects is
explicitly cross-process: ~/.claude/.credentials.json is shared by every
Hermes profile and process. A fresh interpreter could lease the stale
access token or re-POST the already-spent single-use refresh token and
burn the credential family into invalid_grant.
- Persist non-secret one-way fingerprints to a sidecar registry next to
the shared singleton source (claude_code / hermes_pkce), written under
the same path-keyed cross-process lock that serializes refreshes.
- Consult the sidecar in the pool resolver, the pool refresh path, and
the direct claude_code resolver/refresh before leasing or POSTing.
- Two-process regression: A rotates and loses the commit; B (fresh
interpreter, empty local registry) must neither lease the stale pair
nor POST the spent refresh token. Plus a no-verdict control.
Closes the remaining P1 from the exact-head review of f228439b on
PR #87891.
The adapter godfile split moved credential resolution into
agent/anthropic_credentials.py, which silently disarmed two repository
guards still pointed at the old seam:
- tests/conftest.py::_neutralize_macos_keychain_creds patched only the
adapter re-export, so the default suite lost its protection against
reading the operator's real macOS Keychain. Patch the implementation
owner AND the adapter alias.
- test_oauth_setup_token_keeps_inherited_stdin read only the old source
file; it now scans both seams and fails loudly if the call moves again.
- test_hermetic_side_effect_guards isolates the owner module directly.
Two runtime blockers from the exact-head review of c057ef5.
1. A sanitized `claude_code` pool row was treated as token authority.
`claude_code` is a borrowed source: it is absent from the owned-source
allowlist, so `sanitize_borrowed_credential_payload` strips `access_token`
and `refresh_token` before the row reaches `auth.json`. `load_pool()`
re-hydrates the live pair from the singleton on every load, which is what
makes `~/.claude/.credentials.json` — not the pool store — authoritative
for this source.
`_sync_anthropic_entry_from_pool_store()` re-read that persisted row during
refresh. Being token-less, it "differed" from the live entry, so it was
adopted as a rotation performed by another process: `_refresh_entry()`
replaced a usable credential with an empty one and returned it before
`_claude_code_credentials_lock()` and the authoritative re-read were ever
entered. The empty OAuth entry then stayed selectable, because the
empty-runtime-key guard in `_available_entries()` covered API-key rows only.
Repairs: the pool-store sync refuses borrowed sources outright (plus a
defensive refusal of any token-less row, for future sources that sanitize on
write); the `claude_code` branch of `_refresh_entry()` now runs before the
generic adopt-and-return shortcut, so the path-keyed lock and the
authoritative re-read are always entered before deciding to POST or adopt;
and an OAuth entry with no access token is never leased.
2. A failed commit still fell through to the same spent credential.
`_refresh_oauth_token()` correctly returns None when the refresh POST
rotated the single-use token but the replacement could not be committed.
That verdict did not survive the caller: `resolve_anthropic_token()`
continued to `_resolve_anthropic_pool_token()`, which enumerates read-only
(`clear_expired=False, refresh=False`) over a pool that `load_pool()` had
just re-seeded from the unchanged singleton — so the pair whose refresh half
was already spent came back as a healthy token, and
`_refresh_provider_credentials("anthropic")` reported success and evicted
its cached clients.
Repair: every commit-failure path records the consumed pre-rotation pair as
non-reversible fingerprints (bounded, process-local), and both the Claude
Code file resolver and the pool resolver refuse a credential whose
fingerprint is on that list. `_refresh_provider_credentials("anthropic")`
consequently returns False when the spent family is the only credential,
while genuinely independent pool credentials stay eligible.
Coverage: `test_anthropic_borrowed_row_authority.py` starts from
`load_pool()` reading an actually persisted, actually sanitized row, forces
a refresh, and asserts the full pair survives with exactly one POST and one
commit, that the shared-file lock is entered, and that no empty OAuth entry
can be leased. `test_anthropic_spent_rotation_verdict.py` takes the full
resolver path: successful POST plus failed commit must make
`resolve_anthropic_token()` return None, make
`_refresh_provider_credentials("anthropic")` return False, and keep the
spent fingerprint out of every lease — with a control proving a successful
commit quarantines nothing and an independent credential still resolving.
Five of the seven new borrowed-row tests fail on the previous head, and the
three resolution tests fail with the verdict disabled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
`agent/anthropic_adapter.py` was 3,423 lines and this PR adds another auth
boundary to it. Split along the seams that were already there, so the
credential surface this PR changes has a single owner instead of being
interleaved with request building:
- `agent/anthropic_endpoints.py` (258) — base-URL/endpoint-family predicates.
Pure functions over a URL string, which is what lets both of the modules
below depend on it without a cycle.
- `agent/anthropic_message_convert.py` (1,225) — OpenAI-style to Anthropic
Messages payload conversion: model ids, tool schemas, content/thinking
blocks, tool_use pairing, cache_control, screenshot eviction, blank-block
scrubbing.
- `agent/anthropic_credentials.py` (910) — credential sources, the OAuth
flows, and the refresh commit (`CredentialPersistError` and both singleton
writers).
- `agent/anthropic_adapter.py` (1,215) — client construction and the Messages
API call, re-exporting every name from the three modules above so existing
`from agent.anthropic_adapter import ...` imports keep resolving. The
re-export surface was diffed against the pre-split module: nothing dropped.
Call sites that read a moved name through the adapter's namespace at runtime
(`credential_pool._refresh_entry_impl`, `auxiliary_client`) now import it from
the defining module, so there is one patchable seam rather than two bindings
that can disagree. The tests that monkeypatched those seams were retargeted to
match; no assertion was changed.
No behavior change.
Anthropic OAuth refresh tokens are single-use: the POST that returns a new
pair invalidates the one that was sent. The replacement therefore only
becomes real once it reaches its authoritative store -
~/.claude/.credentials.json for claude_code entries,
~/.hermes/.anthropic_oauth.json for hermes_pkce ones. Both writers caught
OSError/IOError, logged at debug level and returned nothing, so no caller
could tell a durable commit from a failed one.
That let a refresh spend the only refresh token, report success, and leave
the consumed pre-rotation pair on disk. _seed_from_singletons() re-reads
those files on every load_pool(), so the next process seeded the spent pair
back over the fresh pool row and the following refresh replayed a consumed
token (invalid_grant / refresh_token_reused) - exactly the failure this PR
set out to remove.
- _write_claude_code_credentials() and _write_hermes_oauth_credentials()
now raise CredentialPersistError instead of swallowing the write error.
- _refresh_oauth_token() treats a failed commit as a failed refresh and
returns None rather than handing back an access token whose refresh half
was lost.
- _refresh_entry_impl() fails closed on both the primary and the recovery
path: the rotated pair is never marked, persisted or returned, and the
entry is quarantined DEAD with a credential_persist_failed reason so it
leaves rotation and surfaces as an explicit re-auth instead of a silent
fallback to another provider. The retry path now commits to the singleton
before persisting the pool row.
- _upsert_entry() no longer treats re-seeding a borrowed source as a
rotation. Borrowed rows (claude_code, env-backed) are written to auth.json
without their secret, so comparing the re-seeded token against the empty
stored value reported a rotation on every load and cleared the DEAD state
the previous process had just written - resurrecting the quarantined,
already-consumed credential on restart. It now compares the incoming
token against the row's secret_fingerprint.
Adds failure-injection coverage for both writers, the direct resolver, the
claude_code and hermes_pkce pool paths and the retry path, each asserting
that a reload cannot bring the pre-refresh pair back as a usable credential.
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).
Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.
A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.
Fixes#87887, #87888, #87889.
The self-improvement review fork advertises the parent's full tool schema
(deliberate — tools[] must stay byte-identical for prompt-cache parity)
but denied everything except memory/skill tools at dispatch. Models
naturally reach for read_file to inspect a SKILL.md before patching, got
denied, then attempted a blind skill_manage patch which the
read-before-write guard correctly refused. One deployment logged ~142
denials + ~204 refusals over 2 days: the self-improvement loop ran
continuously but almost never landed a skill patch.
Fix is dispatch-side ONLY — zero request-body change, cache untouched:
- Whitelist read_file + search_files on the review fork (reads are
side-effect-free). Write tools (write_file/patch/terminal) stay denied:
autonomous maintenance must go through skill_manage's validation.
- read_file now registers full reads with the review fork's
read-before-write guard (same as skill_view), so the natural
read_file -> skill_manage(patch) sequence lands. Partial reads
(offset>1 / truncated) don't count. No-op outside review forks.
- Self-correcting deny message: names skill_view/skill_manage/memory as
substitutes so one denial redirects the model instead of a storm
(the actionable half of #61521's proposal 2).
Rejects #39997's alternative (narrow the advertised schema on local
endpoints): local backends have KV/prefix caches too, and re-prefilling
a large snapshot is most expensive exactly there.
Live A/B (real dispatch path, isolated HERMES_HOME): on main,
read_file DENIED -> patch REFUSED (read-before-write); on this branch,
read_file OK -> patch LANDED. tools[] identical in both.
The rename sweep in the base commit missed the sibling-test blast radius
(18 red files on CI). Three classes, all fixed:
1. Stale old names in tests (todo/cronjob/process/tour/tip) — updated to
todo_list/cronjob_manage/process_manage/gui_tour/show_tip at every
registry.get_entry/dispatch/coerce/preview/allowlist call site, plus
the coding-brief sentence in agent/coding_context.py now names
todo_list (and its gating test).
2. Missed rename in production: AGENT_RUNTIME_POST_HOOK_TOOL_NAMES still
held 'tour' — post-hook ownership would have double-emitted for
gui_tour via the bridge path.
3. Tests pinning pre-deferral assembly (blank-slate surface, modal
sandbox resolution, desktop diet, HUD note) now pin their ACTUAL
contract under the legacy defer:[] override, or assert on granted
tool names instead of visible schemas.
Also fixes a pre-existing ordering flake surfaced by the sweep:
test_holds_exactly_the_gui_affordances depended on whether an earlier
test had imported apply_layout_tool (registry-registered, not in the
static desktop_ui list) — now forces discovery and pins the full set.
649 tests green locally across all touched files, both orderings.
Extends the real-profile machinery (PR #95620) to Brave Origin — Brave's
standalone paid build with a fully separate install identity:
- new canonical key 'brave-origin' in _CHROMIUM_BROWSERS
- Windows: BraveOHTML ProgId -> brave-origin; channel ProgIds BraveOBHTML/
BraveODHTML/BraveOSHTM fail closed (identifiers from brave-core
install_static)
- macOS: com.brave.Browser.origin bundle id (exact match); .beta/.dev/
.nightly channel bundles fail closed; /Applications/Brave Origin.app
- Linux: brave-origin.desktop matched BEFORE the bare 'brave' fragment
(substring scan would otherwise resolve an Origin default to stable
Brave and drive the wrong profile — #95549 wrong-principal invariant);
brave-origin-{beta,nightly,dev} fail closed
- profile dirs: BraveSoftware/Brave-Origin on all three OSes (per
brave-core kProductPathName + Homebrew cask zap paths)
- /browser connect launch tables: Brave Origin split into its OWN group
so a 'brave' executable lookup can never resolve to the Origin binary
- user-facing strings/docs/desktop tooltip updated
Tests: progid/bundle/desktop map params + data-dir resolution for all
three OSes; 125 passed in the three browser test files.
DEFAULT_AGENT_IDENTITY was rewritten in agent/prompt_builder.py (behavior
spec, exploration-thrift line deliberately removed) but the actual seed
written to disk on first run, hermes_cli/default_soul.py's
DEFAULT_SOUL_MD, was never updated. ensure_hermes_home() writes
DEFAULT_SOUL_MD into SOUL.md on every fresh install before the agent's
first turn, so virtually all real users end up as "SOUL.md users" seeded
with the pre-rewrite text -- including the exact "targeted and efficient
exploration" line the rewrite explicitly banned -- while the new
DEFAULT_AGENT_IDENTITY fallback essentially never serves the "fresh
install" audience its own PR body named as the target.
- DEFAULT_SOUL_MD now matches DEFAULT_AGENT_IDENTITY exactly.
- The pre-rewrite text is added to _LEGACY_TEMPLATE_SOULS so installs
already seeded with it self-heal via the existing upgrade-in-place
mechanism (same guarantee as the comment-only scaffold entries: the
string carries zero user intent, so it's safe to replace).
- Synced the other places install.sh's own comment says "MUST match
DEFAULT_SOUL_MD": scripts/install.sh, scripts/install.ps1,
docker/SOUL.md, and the docs/i18n pages that quote the fallback text
verbatim.
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
Sweeper review, all three points:
- Placement: no new plugins/model-providers/ directory. The Token Plan
profiles register from the existing alibaba plugin module — one module
per vendor, matching how the kimi module carries both of its endpoint
variants. Token Plan is the same vendor/service (Model Studio), same
OpenAI-compatible protocol, its own key + endpoints; splitting to a
standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
resolve_runtime_provider() for all four variants — provider, api_key,
api_mode, base_url — alongside the existing zai/minimax/kilocode
runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
with the bundled variants and their env keys/base-url overrides.
The models.dev catalog advertises alibaba-cn, alibaba-coding-plan-cn, and
alibaba-token-plan(-cn), and resolve_provider_full()'s catalog chain lets
the CLI --provider path resolve them — but auth.resolve_provider() (the
credential/runtime path used by 'hermes chat') consults only
PROVIDER_REGISTRY and raised "Unknown provider 'alibaba-coding-plan-cn'"
(hermes_cli/auth.py:1937). PROVIDER_REGISTRY auto-extends from provider
profiles (auth.py:461-490), so the fix registers the missing profiles at
that chokepoint: alibaba-cn joins the alibaba plugin, alibaba-coding-plan-cn
joins alibaba-coding-plan, and a new alibaba-token-plan plugin registers
both regional token-plan tiers. Names match the catalog keys exactly.
No core edits — plugins/model-providers is the designed extension path.
9d9f44d638 removed the desktop platform hint's "never hand-edit
mcp_servers config for them" sentence, reasoning it was a "word-for-word
duplicate of the setup_mcp tool schema... taught on every call." The
schema has never contained that instruction — only "never re-ask after
a decline." setup_mcp is desktop_ui-toolset-only and no runtime guard
in agent/file_safety.py covers mcp_servers config, so removing the only
place teaching this left a real gap: a model asked to add/configure an
MCP server could just write_file into mcp_servers config directly,
bypassing the consent-card/OAuth flow the tool exists to enforce.
Restored the instruction directly in SETUP_MCP_SCHEMA's description —
completing the original commit's stated intent (move it to the schema)
rather than reverting to the platform hint, since the schema reaches
every setup_mcp call regardless of platform hint wording changes.
Added a regression test asserting the schema description forbids
hand-editing mcp_servers config, so a future prompt-diet pass can't
silently drop it again without a test failing.
74a95a3ddf promoted /bg and /btw to independent canonical commands and
retired /background entirely (hermes_cli/commands.py's COMMAND_REGISTRY
has no "background" command or alias). Three places still taught the
old name:
- skills/autonomous-ai-agents/hermes-agent/references/slash-commands.md:
the bundled hermes-agent skill's own slash-command reference — the
skill's SKILL.md explicitly routes the model here for in-session
command questions, so a model following it would emit the dead
`/background <prompt>` and never learn /btw exists.
- ui-tui/README.md: listed /btw as an alias of /background, which is
simply wrong post-split (both are independent, alias-free commands).
- tests/cli/test_cli_background_status_indicator.py: docstring/comments
described the ▶ indicator by the retired command name.
No behavior change; corrects documentation only.
Maintainer A/B (288 live runs, 3 model tiers, results in the PR body):
with the clarify schema visible, models used structured ask-the-user
18/18 on ambiguous tasks (score 1.00 all models). Deferred, usage
collapsed to 7/18 (gpt-terra 0/6) — models still asked, but as
plain-text turn-ending questions: no structured choices, no recommended
option, an extra user round-trip. The ask-the-user affordance has to be
ambient to fire; a catalog stub is not enough (~250 tok to keep eager).
- _DEFAULT_DEFERRED_TOOLS: remove clarify (19 -> 18 deferred)
- regression test pins clarify ∉ default defer set AND assembles direct
while the bridge is active (sabotage-verified: fails with clarify
in the set)
Flip the salvaged --start-now behavior (PR #97958) into the unconditional
default: /loop's first iteration is due the moment the loop is set, then
recurs on the normal cadence. The flag is dropped — it was never released,
so there is nothing to deprecate.
- LoopManager.set(): next_due_at = now for both cadence modes
- drop --start-now parsing, the persisted LoopState.start_now field, and
the flag from help text; confirmation now always says the first wakeup
fires now
- tests updated to pin the new default (incl. the TUI not-due test, which
now has to push next_due_at out explicitly)
- docs: quick-start and command table describe the immediate first run
/loop [interval] <prompt> currently schedules the first wakeup one full
interval after the command runs (next_due_at = now + interval). When the
user just told Hermes what to check, waiting the whole interval before
any output feels like the command was ignored.
Add an opt-in --start-now flag that keeps Claude Code parity as the
default but lets the user run the first iteration immediately, then
continue on the cadence:
/loop 1h check the deploy status # first run in 1h (unchanged)
/loop 1h --start-now check the deploy # first run now, then hourly
- parse_loop_args(): parse and strip --start-now (leading or trailing)
- LoopState: new persisted start_now field (default False, survives
serialization round-trip and old rows missing the field)
- LoopManager.set(): next_due_at = now when start_now, for both fixed
interval and self-paced modes
- dispatch_loop_command(): wire start_now through, update help text, and
report "First wakeup fires now" in the confirmation
- website/docs: document the flag in the /loop guide
- tests: parse (trailing/leading/absent/self-paced/combo/prompt-word),
tick lifecycle (due immediately vs after interval), serde round-trip,
and dispatch-level confirmation
The inline set-comprehension re-implemented modelName extraction that
_extract_model_name() already provides (and that both
union_with_portal_free/paid_recommendations already use). Beyond the
duplication, the inline str(entry.get("modelName", "")) stringified
non-string values — a malformed Portal entry with modelName 5 would have
produced a garbage "5" match that discard("") does not filter. The helper
isinstance-checks and returns None for those, so routing through it makes
the validation tier semantically identical to the union helpers.
Adds test_non_string_model_name_entries_ignored locking the behavior
(mutation-checked: fails on the raw-stringify form, passes on the helper).
Fixes#71312 (duplicate #71313).
When selecting a model via the Telegram /model picker (or any other
messaging-platform slash command, since they all share
validate_requested_model() through gateway/slash_commands.py ->
model_switch.switch_model()), a model available via Nous Portal's live
/api/nous/recommended-models endpoint but not yet in the hardcoded
curated catalog (_PROVIDER_MODELS["nous"]) was rejected with "was not
found in this provider's model listing" -- even though the exact same
model works fine via `hermes chat -m <model> --provider nous`.
Root cause: `hermes chat` merges Portal recommendations into its model
list via union_with_portal_free_recommendations() /
union_with_portal_paid_recommendations() at model-list build time
(hermes_cli/auth.py, web_server.py, model_setup_flows.py,
model_switch.py), so the model already appears "known" by the time
validation runs for that path. validate_requested_model() itself,
which every per-message /model command goes through, only checked the
live /v1/models listing and the curated catalog (_model_in_provider_catalog) --
never the Portal recommendations feed -- so a model that exists only
in Portal Recommendations was rejected on that path specifically.
Fix: add a Nous-specific fallback tier in validate_requested_model(),
checked after the curated-catalog fallback and before the final
rejection, reading the same fetch_nous_recommended_models() feed
(free + paid tiers) the CLI union helpers already use. Scoped to
provider == "nous" only; short-circuits before the network call when
an earlier tier already accepted the model; fails closed (rejects,
doesn't crash) if the Portal feed is unreachable.
Reported two issues filed 3 minutes apart with identical content by
the same author (#71312, #71313) -- commented on #71313 marking it a
duplicate of #71312 and pointing to this fix (could not close it
directly, no admin rights on the repo from this token).
6/6 new tests pass in TestValidateRequestedModelNousPortalRecommendations;
95/95 in the full tests/hermes_cli/test_model_validation.py file;
87/87 in tests/hermes_cli/test_models.py (unaffected, confirmed).
Closes#75364.
`_compress_context_via_codex_app_server` returns the transcript unchanged
when the codex thread reports `interrupted` or `error`. The session is
therefore still above threshold, and nothing records that the attempt
failed — so the next turn retries immediately, and keeps retrying for as
long as the condition persists.
Every other compression path arms the shared failure cooldown, records an
ineffective-compression strike, or both. This path records neither:
* `_hygiene_compression_failure_cooldowns` is set only on
`asyncio.TimeoutError`, or behind `_last_compress_aborted`, which is
assigned exclusively in `context_compressor.py` on the Hermes summarizer
path.
* `compression_ineffective_count` lives in `ContextCompressor`, and this
path returns before any compressor bookkeeping runs.
`compress_context` already documents the rule this path was missing —
"Every automatic entrypoint must honor compressor-owned cooldown and
breaker state" — but the codex branch dispatches above that block and
returns from inside it.
`result.interrupted` needs no unusual configuration to occur: an ordinary
user message arriving mid-compaction sets it (see
`codex_app_server_session.py`, which produces the "compact turn
interrupted" string). Observed in production on a Discord gateway session
at ~315k tokens against a 258k window, where compaction was attempted on
essentially every turn for ~70 minutes; the session's
`compression_ineffective_count` was still 0 afterwards.
This reuses the existing cooldown rather than adding a new mechanism:
* arm `_record_compression_failure_cooldown` with the existing
`_SUMMARY_FAILURE_COOLDOWN_SECONDS` when compaction returns
interrupted/error;
* honor an active cooldown on entry, matching the Hermes path.
`force=True` bypasses both, so an explicit /compress is never braked by a
failure it did not cause, and a successful compaction arms nothing.
The chat_completions and transport-parity Mistral tests pinned the
same branch with different reasoning_config shapes ({effort:none} vs
{enabled:False, effort:none}) — behaviorally identical since effort
short-circuits first, but the drift reads as a semantic difference.
Align both to the explicit form. Also scope the port-guard comment to
the try/except shape it actually shares with hermes_cli/models.py.
urlparse raises ValueError on non-integer / out-of-range ports, and
http://myhost:99999/v1 passes OpenAI-client construction (only httpx
rejects it later), so the crash was reachable from build_kwargs on
every request for such a URL. Wrap the parsed.port check in the same
try/except ValueError guard hermes_cli.models already uses around its
11434 check, and pin it with parametrized tests.
Companion to the connector's delete op (gateway-gateway 119a228). The
fresh-final unfurl route re-posts the completed reply and previously left
the sealed streamed preview behind (double delivery). delete_message now
emits op=delete when the negotiated descriptor advertises it; without the
advertisement it returns False with zero wire traffic, degrading to the
old leave-the-preview behavior against older connectors.
Consumer-level test drives placeholder -> stamped fresh final -> delete
of the original preview id.
Slack evaluates link previews exactly once, at chat.postMessage (live
probe 2026-08-28: URL at post + stamps unfurls; a chat.update that
INTRODUCES the URL never does, stamped or not). Edit-based streaming
posts its first frame before the model produces any URL — on flat DMs
with tool_progress=accumulate that frame is the task card — so a
configured unfurl_links/media: true could never surface a preview:
the only post Slack evaluates carries no link.
RelayAdapter now implements prefers_fresh_final_streaming(): True only
when the Slack unfurl hints contain an explicit True AND the final text
carries a link. The stream consumer then delivers the completed reply
as one fresh send — URL and stamps present at the single moment Slack
looks. False-only hints (enterprise fail-closed posture) keep the edit
lane untouched: suppression rides the placeholder post and edits can
never add a preview, so false inherits with zero streaming-UX cost.
Consumer-level contract test drives the exact regression shape
(placeholder frame -> URL-bearing final) and asserts op=send + stamps;
verified RED against the unfixed adapter, GREEN with the hook.
The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.
- agent/background_review.py: extract the review-fork construction into
build_cache_parity_fork() — same runtime/credentials as the parent,
byte-identical system prompt / tools[] / reasoning config on the
same-model path, shared session_id for prefix warmth, full persistence
detachment (no state.db writes, no rotation, no external memory,
in-place-only compaction). The review thread now calls the helper;
behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
AIAgent exists — replays the untruncated snapshot as warm cache reads,
denies every tool at dispatch via an empty thread whitelist (tools[]
stays byte-identical for cache parity), attributes usage to the parent,
and trims a mid-turn snapshot tail so role alternation holds. The
one-shot digest remains as fallback (no live agent = cold cache anyway,
and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
the chat's cached agent (parity with how turns reuse it).
Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
Follow-up for salvaged PR #28253: current main's generic profile fetch
passes base_url= to fetch_models and merges curated fallback_models
first (658ac1d86 / #46309), so the mocked-signature and exact-equality
assertions from the PR's era no longer match the contract.
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Addresses the automated review on this PR:
- _profile_declared_efforts falls back from provider name to the
endpoint's host (via model_metadata's URL->provider map), so a named
custom provider pointed at api.router.com — which the host mandate
already routes onto this transport — gets the catalog clamp instead
of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
ingest, logging and dropping unrecognized tiers; a model whose whole
vocabulary is unrecognized stays out of the map (transport defaults)
instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
order.
- plugin.yaml credits the human contributor per repo convention.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
Review finding (quality pass): on a single-profile gateway,
_handle_busy_command set _busy_input_mode but left _busy_text_mode
stale, so the adapter refresh a line later re-read the old value —
/busy queue persisted to config but live text messages kept
interrupting until restart. The profile path already re-derives both
from the fresh config; the non-profile path now does the same via
_load_busy_text_mode() (busy_input_mode is the source of truth,
run.py:9877). Regression assertion added to test_set_mode_persists;
verified red without the production fix.
test_update_is_known_command grepped _handle_message's source for the
literal '"update"' — a banned source-reading test (AGENTS.md), broken
by the if-chain -> _gateway_plain_command_handlers() refactor. Assert
the actual dispatch contract instead: the shared handler table maps
'update' to _handle_update_command.