Commit Graph

26217 Commits

Author SHA1 Message Date
Gille cd1c3211ec chore: map contributor email 2026-08-29 18:10:03 -07:00
AideYu 5cd9c4563c fix(telegram): recover exhausted request pool 2026-08-29 18:10:03 -07:00
Teknium e05c91ac71 fix(cli): slow /handoff transfers no longer misreported as "gateway not running"
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.

- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
  rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
  there really does mean no gateway — then up to 15 min for the claimed
  dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
  returns {failed: false, state: running} instead of stomping the claim.

Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
2026-08-29 18:08:55 -07:00
GodsBoy e38cca50d6 fix(memory): keep Mem0 OSS OpenAI requests direct 2026-08-29 18:02:19 -07:00
Ayush Nangia 94aad6dcd2 fix(providers): surface Alibaba China in desktop parity 2026-08-29 18:01:25 -07:00
Ayush Nangia ea129cac54 test(providers): keep Alibaba regression coverage focused 2026-08-29 18:01:25 -07:00
Ayush Nangia 695d86f518 fix(providers): fold Token Plan into the alibaba plugin, add runtime-path regressions, document all variants
Sweeper review, all three points:

- Placement: no new plugins/model-providers/ directory. The Token Plan
  profiles register from the existing alibaba plugin module — one module
  per vendor, matching how the kimi module carries both of its endpoint
  variants. Token Plan is the same vendor/service (Model Studio), same
  OpenAI-compatible protocol, its own key + endpoints; splitting to a
  standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
  resolve_runtime_provider() for all four variants — provider, api_key,
  api_mode, base_url — alongside the existing zai/minimax/kilocode
  runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
  with the bundled variants and their env keys/base-url overrides.
2026-08-29 18:01:25 -07:00
Ayush Nangia 7cf7df0297 fix(providers): register Alibaba China + Token Plan provider profiles (#73265)
The models.dev catalog advertises alibaba-cn, alibaba-coding-plan-cn, and
alibaba-token-plan(-cn), and resolve_provider_full()'s catalog chain lets
the CLI --provider path resolve them — but auth.resolve_provider() (the
credential/runtime path used by 'hermes chat') consults only
PROVIDER_REGISTRY and raised "Unknown provider 'alibaba-coding-plan-cn'"
(hermes_cli/auth.py:1937). PROVIDER_REGISTRY auto-extends from provider
profiles (auth.py:461-490), so the fix registers the missing profiles at
that chokepoint: alibaba-cn joins the alibaba plugin, alibaba-coding-plan-cn
joins alibaba-coding-plan, and a new alibaba-token-plan plugin registers
both regional token-plan tiers. Names match the catalog keys exactly.

No core edits — plugins/model-providers is the designed extension path.
2026-08-29 18:01:25 -07:00
nftpoetrist d6a6d87c4a fix(tools): restore setup_mcp's never-hand-edit instruction
9d9f44d638 removed the desktop platform hint's "never hand-edit
mcp_servers config for them" sentence, reasoning it was a "word-for-word
duplicate of the setup_mcp tool schema... taught on every call." The
schema has never contained that instruction — only "never re-ask after
a decline." setup_mcp is desktop_ui-toolset-only and no runtime guard
in agent/file_safety.py covers mcp_servers config, so removing the only
place teaching this left a real gap: a model asked to add/configure an
MCP server could just write_file into mcp_servers config directly,
bypassing the consent-card/OAuth flow the tool exists to enforce.

Restored the instruction directly in SETUP_MCP_SCHEMA's description —
completing the original commit's stated intent (move it to the schema)
rather than reverting to the platform hint, since the schema reaches
every setup_mcp call regardless of platform hint wording changes.

Added a regression test asserting the schema description forbids
hand-editing mcp_servers config, so a future prompt-diet pass can't
silently drop it again without a test failing.
2026-08-29 17:58:51 -07:00
nftpoetrist bc81d0a665 docs: sync stale /background references with the /bg + /btw split
74a95a3ddf promoted /bg and /btw to independent canonical commands and
retired /background entirely (hermes_cli/commands.py's COMMAND_REGISTRY
has no "background" command or alias). Three places still taught the
old name:

- skills/autonomous-ai-agents/hermes-agent/references/slash-commands.md:
  the bundled hermes-agent skill's own slash-command reference — the
  skill's SKILL.md explicitly routes the model here for in-session
  command questions, so a model following it would emit the dead
  `/background <prompt>` and never learn /btw exists.
- ui-tui/README.md: listed /btw as an alias of /background, which is
  simply wrong post-split (both are independent, alias-free commands).
- tests/cli/test_cli_background_status_indicator.py: docstring/comments
  described the ▶ indicator by the retired command name.

No behavior change; corrects documentation only.
2026-08-29 17:58:46 -07:00
nftpoetrist 8142494401 fix(tui): render nested todo subtasks via the parent field
CLI, ACP, and the desktop app all got nested-subtask rendering (the
optional `parent` field on a todo item), but the TUI never did. Its
TodoItem type had no `parent` field, parseTodos() in turnController.ts
dropped it even if the tool payload sent it, and TodoPanel rendered the
list with a flat map() and a single fixed indent — a session using
nested subtasks showed every subtask at the same visual level as its
parent, with no hierarchy cue, in the terminal UI.

- types.ts: add the optional `parent` field to TodoItem, matching
  apps/desktop/src/lib/todos.ts's TodoItem exactly.
- turnController.ts: parseTodos() now preserves parent (trimmed,
  dropped if empty or self-referential), the same normalization
  desktop's parseArray() applies.
- lib/todo.ts: port todoTree() from apps/desktop/src/lib/todos.ts
  verbatim — same DFS-with-depth algorithm, same dangling/cycle
  handling, so both surfaces render identical hierarchy from the same
  `parent` field.
- todoPanel.tsx: render todoTree(todos) instead of a flat map(), with
  per-row indentation scaled by depth (capped at 4 levels, mirroring
  desktop's status-row.tsx cap).
2026-08-29 17:58:41 -07:00
Teknium 1a47a36422 feat(loop): first wakeup fires immediately by default
Flip the salvaged --start-now behavior (PR #97958) into the unconditional
default: /loop's first iteration is due the moment the loop is set, then
recurs on the normal cadence. The flag is dropped — it was never released,
so there is nothing to deprecate.

- LoopManager.set(): next_due_at = now for both cadence modes
- drop --start-now parsing, the persisted LoopState.start_now field, and
  the flag from help text; confirmation now always says the first wakeup
  fires now
- tests updated to pin the new default (incl. the TUI not-due test, which
  now has to push next_due_at out explicitly)
- docs: quick-start and command table describe the immediate first run
2026-08-29 17:43:38 -07:00
oxngon 796babaaed feat(loop): add --start-now to fire the first wakeup immediately
/loop [interval] <prompt> currently schedules the first wakeup one full
interval after the command runs (next_due_at = now + interval). When the
user just told Hermes what to check, waiting the whole interval before
any output feels like the command was ignored.

Add an opt-in --start-now flag that keeps Claude Code parity as the
default but lets the user run the first iteration immediately, then
continue on the cadence:

  /loop 1h check the deploy status            # first run in 1h (unchanged)
  /loop 1h --start-now check the deploy       # first run now, then hourly

- parse_loop_args(): parse and strip --start-now (leading or trailing)
- LoopState: new persisted start_now field (default False, survives
  serialization round-trip and old rows missing the field)
- LoopManager.set(): next_due_at = now when start_now, for both fixed
  interval and self-paced modes
- dispatch_loop_command(): wire start_now through, update help text, and
  report "First wakeup fires now" in the confirmation
- website/docs: document the flag in the /loop guide
- tests: parse (trailing/leading/absent/self-paced/combo/prompt-word),
  tick lifecycle (due immediately vs after interval), serde round-trip,
  and dispatch-level confirmation
2026-08-29 17:43:38 -07:00
kshitijk4poor 4209d371aa refactor(models): reuse _extract_model_name in the Portal-recommendation validation tier
The inline set-comprehension re-implemented modelName extraction that
_extract_model_name() already provides (and that both
union_with_portal_free/paid_recommendations already use). Beyond the
duplication, the inline str(entry.get("modelName", "")) stringified
non-string values — a malformed Portal entry with modelName 5 would have
produced a garbage "5" match that discard("") does not filter. The helper
isinstance-checks and returns None for those, so routing through it makes
the validation tier semantically identical to the union helpers.

Adds test_non_string_model_name_entries_ignored locking the behavior
(mutation-checked: fails on the raw-stringify form, passes on the helper).
2026-08-29 22:41:55 +05:30
ygd58 0ffad55e09 fix(models): accept live Nous Portal recommendations in /model validation
Fixes #71312 (duplicate #71313).

When selecting a model via the Telegram /model picker (or any other
messaging-platform slash command, since they all share
validate_requested_model() through gateway/slash_commands.py ->
model_switch.switch_model()), a model available via Nous Portal's live
/api/nous/recommended-models endpoint but not yet in the hardcoded
curated catalog (_PROVIDER_MODELS["nous"]) was rejected with "was not
found in this provider's model listing" -- even though the exact same
model works fine via `hermes chat -m <model> --provider nous`.

Root cause: `hermes chat` merges Portal recommendations into its model
list via union_with_portal_free_recommendations() /
union_with_portal_paid_recommendations() at model-list build time
(hermes_cli/auth.py, web_server.py, model_setup_flows.py,
model_switch.py), so the model already appears "known" by the time
validation runs for that path. validate_requested_model() itself,
which every per-message /model command goes through, only checked the
live /v1/models listing and the curated catalog (_model_in_provider_catalog) --
never the Portal recommendations feed -- so a model that exists only
in Portal Recommendations was rejected on that path specifically.

Fix: add a Nous-specific fallback tier in validate_requested_model(),
checked after the curated-catalog fallback and before the final
rejection, reading the same fetch_nous_recommended_models() feed
(free + paid tiers) the CLI union helpers already use. Scoped to
provider == "nous" only; short-circuits before the network call when
an earlier tier already accepted the model; fails closed (rejects,
doesn't crash) if the Portal feed is unreachable.

Reported two issues filed 3 minutes apart with identical content by
the same author (#71312, #71313) -- commented on #71313 marking it a
duplicate of #71312 and pointing to this fix (could not close it
directly, no admin rights on the repo from this token).

6/6 new tests pass in TestValidateRequestedModelNousPortalRecommendations;
95/95 in the full tests/hermes_cli/test_model_validation.py file;
87/87 in tests/hermes_cli/test_models.py (unaffected, confirmed).
2026-08-29 22:41:55 +05:30
Hyusein Leshov 835a913ffd fix(compression): arm the failure cooldown when codex compaction fails
Closes #75364.

`_compress_context_via_codex_app_server` returns the transcript unchanged
when the codex thread reports `interrupted` or `error`. The session is
therefore still above threshold, and nothing records that the attempt
failed — so the next turn retries immediately, and keeps retrying for as
long as the condition persists.

Every other compression path arms the shared failure cooldown, records an
ineffective-compression strike, or both. This path records neither:

* `_hygiene_compression_failure_cooldowns` is set only on
  `asyncio.TimeoutError`, or behind `_last_compress_aborted`, which is
  assigned exclusively in `context_compressor.py` on the Hermes summarizer
  path.
* `compression_ineffective_count` lives in `ContextCompressor`, and this
  path returns before any compressor bookkeeping runs.

`compress_context` already documents the rule this path was missing —
"Every automatic entrypoint must honor compressor-owned cooldown and
breaker state" — but the codex branch dispatches above that block and
returns from inside it.

`result.interrupted` needs no unusual configuration to occur: an ordinary
user message arriving mid-compaction sets it (see
`codex_app_server_session.py`, which produces the "compact turn
interrupted" string). Observed in production on a Discord gateway session
at ~315k tokens against a 258k window, where compaction was attempted on
essentially every turn for ~70 minutes; the session's
`compression_ineffective_count` was still 0 afterwards.

This reuses the existing cooldown rather than adding a new mechanism:

* arm `_record_compression_failure_cooldown` with the existing
  `_SUMMARY_FAILURE_COOLDOWN_SECONDS` when compaction returns
  interrupted/error;
* honor an active cooldown on entry, matching the Hermes path.

`force=True` bypasses both, so an explicit /compress is never braked by a
failure it did not cause, and a successful compaction arms nothing.
2026-08-29 22:29:28 +05:30
kshitijk4poor 1e21fe8624 test(custom): align Mistral-omission test inputs; soften models.py comment
The chat_completions and transport-parity Mistral tests pinned the
same branch with different reasoning_config shapes ({effort:none} vs
{enabled:False, effort:none}) — behaviorally identical since effort
short-circuits first, but the drift reads as a semantic difference.
Align both to the explicit form. Also scope the port-guard comment to
the try/except shape it actually shares with hermes_cli/models.py.
2026-08-29 21:58:00 +05:30
kshitijk4poor c870589831 fix(custom): tolerate malformed ports in the Ollama URL heuristic
urlparse raises ValueError on non-integer / out-of-range ports, and
http://myhost:99999/v1 passes OpenAI-client construction (only httpx
rejects it later), so the crash was reachable from build_kwargs on
every request for such a URL. Wrap the parsed.port check in the same
try/except ValueError guard hermes_cli.models already uses around its
11434 check, and pin it with parametrized tests.
2026-08-29 21:58:00 +05:30
xxxigm 6ba8308309 test(custom): pin think=false to Ollama URLs, omit it for Mistral
Cover the Mistral extra_forbidden case and keep the Ollama dual-emission
contract (think=false + reasoning_effort=none) on port 11434 / ollama hosts.
2026-08-29 21:58:00 +05:30
xxxigm 31f0336da7 fix(custom): omit Ollama-only think=false on strict OpenAI-compat endpoints
reasoning_effort: none was injecting extra_body.think=false for every
custom provider. Mistral (and other extra=forbid hosts) reject that
field with HTTP 422. Keep think=false on Ollama URLs only; still send
top-level reasoning_effort=none so /v1 thinking-off keeps working.
2026-08-29 21:58:00 +05:30
Victor Kyriazakos 3f36c87e1e feat(relay): delete_message over the additive delete op — fresh-final preview cleanup
Companion to the connector's delete op (gateway-gateway 119a228). The
fresh-final unfurl route re-posts the completed reply and previously left
the sealed streamed preview behind (double delivery). delete_message now
emits op=delete when the negotiated descriptor advertises it; without the
advertisement it returns False with zero wire traffic, degrading to the
old leave-the-preview behavior against older connectors.

Consumer-level test drives placeholder -> stamped fresh final -> delete
of the original preview id.
2026-08-29 08:29:15 -07:00
Victor Kyriazakos cef4c88f71 fix(relay): route force-on-unfurl streamed finals through fresh chat.postMessage
Slack evaluates link previews exactly once, at chat.postMessage (live
probe 2026-08-28: URL at post + stamps unfurls; a chat.update that
INTRODUCES the URL never does, stamped or not). Edit-based streaming
posts its first frame before the model produces any URL — on flat DMs
with tool_progress=accumulate that frame is the task card — so a
configured unfurl_links/media: true could never surface a preview:
the only post Slack evaluates carries no link.

RelayAdapter now implements prefers_fresh_final_streaming(): True only
when the Slack unfurl hints contain an explicit True AND the final text
carries a link. The stream consumer then delivers the completed reply
as one fresh send — URL and stamps present at the single moment Slack
looks. False-only hints (enterprise fail-closed posture) keep the edit
lane untouched: suppression rides the placeholder post and edits can
never add a preview, so false inherits with zero streaming-UX cost.

Consumer-level contract test drives the exact regression shape
(placeholder frame -> URL-bearing final) and asserts op=send + stamps;
verified RED against the unfixed adapter, GREEN with the hook.
2026-08-29 08:29:15 -07:00
Teknium 578f85cfb0 feat: /btw rides the background-review cache-parity fork for full-context answers
The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.

- agent/background_review.py: extract the review-fork construction into
  build_cache_parity_fork() — same runtime/credentials as the parent,
  byte-identical system prompt / tools[] / reasoning config on the
  same-model path, shared session_id for prefix warmth, full persistence
  detachment (no state.db writes, no rotation, no external memory,
  in-place-only compaction). The review thread now calls the helper;
  behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
  AIAgent exists — replays the untruncated snapshot as warm cache reads,
  denies every tool at dispatch via an empty thread whitelist (tools[]
  stays byte-identical for cache parity), attributes usage to the parent,
  and trims a mid-turn snapshot tail so role alternation holds. The
  one-shot digest remains as fallback (no live agent = cold cache anyway,
  and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
  the chat's cached agent (parity with how turns reuse it).

Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
2026-08-29 08:23:47 -07:00
kshitijk4poor 360761c8cf docs: surface Tencent TokenPlan + hy4-preview across provider docs
Follow-up for salvaged PR #96939 (docs completeness audit): setup
table, quick commands (TokenHub example refreshed to hy4-preview),
fallback tables, env-var reference, quickstart, --provider choices,
and .env.example (TokenHub block was missing there too).
2026-08-29 20:51:17 +05:30
kshitijk4poor ac5c8f58db fix: drop duplicate hy4-preview context entry — main's 1_048_576 wins
Follow-up for salvaged PR #96939: main already added hy4-preview at
1_048_576 (f7c79efbac); the cherry-picked duplicate key later in the
dict silently overrode it with 1024000.
2026-08-29 20:51:17 +05:30
simonweng e74e594a86 fix:update test info 2026-08-29 20:51:17 +05:30
simonweng 0fb5cab0d4 feat:add hy4-preview model and tokenplan provider 2026-08-29 20:51:17 +05:30
kshitijk4poor b954547e72 fix(nebius): route effort through canonical clamp_effort — hand-rolled map inverted the ladder
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
2026-08-29 20:39:44 +05:30
kshitijk4poor a8e11d21c0 docs: surface Nebius Token Factory across user-facing provider docs
Follow-up for salvaged PR #28253 (docs completeness audit): setup
table, quick command, fallback tables, env-var reference, quickstart,
--provider choices, and .env.example.
2026-08-29 20:39:44 +05:30
kshitijk4poor 0d016231a8 test(nebius): align catalog tests with current fetch_models contract
Follow-up for salvaged PR #28253: current main's generic profile fetch
passes base_url= to fetch_models and merges curated fallback_models
first (658ac1d86 / #46309), so the mocked-signature and exact-equality
assertions from the PR's era no longer match the contract.
2026-08-29 20:39:44 +05:30
amrrs 2c2293dcf4 test(nebius): expect verbose model catalog URL 2026-08-29 20:39:44 +05:30
amrrs 49f5b6a9b4 fix(nebius): request verbose model metadata 2026-08-29 20:39:44 +05:30
amrrs 13bad590f3 feat(providers): add Nebius Token Factory provider 2026-08-29 20:39:44 +05:30
kshitijk4poor 1c5ee5815f fix(router): pytest guard on caps warmer + debug log in fail-open efforts lookup
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
2026-08-29 20:29:04 +05:30
kshitijk4poor fbef1ca805 docs: surface Ramp Router across user-facing provider docs
Follow-up for salvaged PR #93548 (docs completeness audit):
- integrations/providers.md: setup table row, quick command example,
  fallback supported-providers list
- reference/environment-variables.md: RAMP_ROUTER_API_KEY / _BASE_URL
- reference/cli-commands.md: --provider choices
- getting-started/quickstart.md: provider table
- user-guide/features/fallback-providers.md: fallback table
- .env.example: commented key block
2026-08-29 20:29:04 +05:30
Neel Patel ca3060b70f docs: chat-completions is now a compat shim on Router, not a 404
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
2026-08-29 20:29:04 +05:30
Neel Patel eeb7916000 review: host-resolved efforts, ladder-validated ingest, deduped catalog
Addresses the automated review on this PR:

- _profile_declared_efforts falls back from provider name to the
  endpoint's host (via model_metadata's URL->provider map), so a named
  custom provider pointed at api.router.com — which the host mandate
  already routes onto this transport — gets the catalog clamp instead
  of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
  ingest, logging and dropping unrecognized tiers; a model whose whole
  vocabulary is unrecognized stays out of the map (transport defaults)
  instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
  order.
- plugin.yaml credits the human contributor per repo convention.
2026-08-29 20:29:04 +05:30
Neel Patel ceb54f5a52 feat(providers): send Hermes-Agent User-Agent on Router requests
Router attributes coding-agent clients by User-Agent prefix (the way it
already recognizes OpenCode's versioned UA), and its WAF rejects
default/blank client UAs. Mirror the xai profile: declare
User-Agent: Hermes-Agent/<version> in default_headers, which
agent_init's profile-headers fallback applies at client construction.
Re-verified live: one-shot chat through api.router.com still works.
2026-08-29 20:29:04 +05:30
Neel Patel 804f8b4732 feat(providers): add Ramp Router (router.com) provider plugin
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:

- plugins/model-providers/router/: RouterProfile plugin —
  api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
  RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
  GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
  Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
  runtime_provider._detect_api_mode_for_url: api.router.com ->
  codex_responses. The host is Responses-only — POST /v1/chat/completions
  does not exist and 404s — so this is a genuine host mandate (exact
  hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
  hook (tri-state: None=defer, ()=model takes no reasoning params,
  tuple=clamp set). Router validates reasoning.effort per model and
  returns HTTP 400 invalid-argument on levels outside the model's
  published vocabulary, and 400 unsupported_parameter when a
  non-reasoning model receives any reasoning field (both verified live).
  The profile answers from a cached copy of the catalog's
  router.capabilities.reasoning block: cache-only on the hot path,
  seeded for free by fetch_models(), disk-mirrored across processes
  (/cache/router_catalog.json), background-warmed when cold
  — same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
  the generic effort-clamp branch (xai/actual/github branches untouched;
  profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
  document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
  rejection, profile registration + auth auto-registry wiring, catalog
  parsing, and transport clamp/suppression/fallback paths.

Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
2026-08-29 20:29:04 +05:30
kshitij b4b7727ea0 Merge pull request #97912 from kshitijk4poor/chore/author-map-provider-salvages
chore: map provider-salvage contributor emails (Neel49, amrrs)
2026-08-29 20:14:05 +05:30
kshitijk4poor 0d02f0d1dc fix: re-derive the live busy text mode after a non-profile /busy change
Review finding (quality pass): on a single-profile gateway,
_handle_busy_command set _busy_input_mode but left _busy_text_mode
stale, so the adapter refresh a line later re-read the old value —
/busy queue persisted to config but live text messages kept
interrupting until restart. The profile path already re-derives both
from the fresh config; the non-profile path now does the same via
_load_busy_text_mode() (busy_input_mode is the source of truth,
run.py:9877). Regression assertion added to test_set_mode_persists;
verified red without the production fix.
2026-08-29 20:03:48 +05:30
kshitijk4poor 740994a903 test: assert /update dispatch via the handler table, not getsource
test_update_is_known_command grepped _handle_message's source for the
literal '"update"' — a banned source-reading test (AGENTS.md), broken
by the if-chain -> _gateway_plain_command_handlers() refactor. Assert
the actual dispatch contract instead: the shared handler table maps
'update' to _handle_update_command.
2026-08-29 20:03:48 +05:30
Samuel Odio f75f24e92e docs(commands): clarify busy gateway behavior 2026-08-29 20:03:48 +05:30
Samuel Odio 8c23824609 docs: align busy command references 2026-08-29 20:03:48 +05:30
Samuel Odio 597d2230ea refactor(gateway): share plain command dispatch 2026-08-29 20:03:48 +05:30
Samuel Odio fe828188d2 docs(test): clarify busy scope coverage 2026-08-29 20:03:48 +05:30
Samuel Odio e8591ae509 test(gateway): verify routed busy persistence 2026-08-29 20:03:48 +05:30
Samuel Odio 8d1d193f11 fix(gateway): apply busy mode per profile 2026-08-29 20:03:48 +05:30
Samuel Odio ce22b689c4 fix: route /insights through /hermes on Slack 2026-08-29 20:03:48 +05:30
Abdias Joel 5b8074dad4 fix: use event.get_command_args() and add persistence tests
Teknium review items:
1. Parse with event.get_command_args() instead of raw event.text
   (matching _handle_fast_command pattern at line 2842)
2. Add mocked persistence tests for queue/steer/interrupt setter
   success, save-failure, and exception branches (5 new tests)
2026-08-29 20:03:48 +05:30