Commit Graph

1690 Commits

Author SHA1 Message Date
blunkjamie-dev 1885a40ad3 fix(buzz): preserve literal mentions and exact UUID targets 2026-08-31 09:05:41 -07:00
Cameron Aragon b302ae3f3e docs(buzz): clarify reaction-only user precedence 2026-08-31 09:05:41 -07:00
Cameron Aragon 743f86ed36 test(buzz): clarify reaction-only precedence 2026-08-31 09:05:41 -07:00
Cameron Aragon c10c77d577 fix(buzz): acknowledge trusted agent tags without dispatch 2026-08-31 09:05:41 -07:00
Elmar Conradie f40edee459 fix(buzz): trust explicit DM metadata fallback 2026-08-31 09:05:41 -07:00
Elmar Conradie 722209bb51 fix(buzz): require explicit group addressing 2026-08-31 09:05:41 -07:00
arimu1 ef2be55025 fix(buzz): treat NIP-10 replies to own messages as mentions
require_mention gated only on visible text, so Desktop thread replies
(e.g. /approve session) to the agent's own prompts were dropped with no
log. Cache event_id→(author, snippet) from seed/poll/WS/send, resolve
the direct e-tag parent, and dispatch when that parent is ours; also
populate reply_to_* on MessageEvent for gateway context injection.

Fixes #75826
2026-08-31 09:05:41 -07:00
liuhao1024 306dc874c1 fix(buzz): dispatch forum-channel kinds instead of chat kind 9 only
The inbound path hardcoded Nostr kind 9 at both the WebSocket
subscription filter and the dispatch gate (which runs before mention
gating), so Buzz forum channels — kind 45001 thread roots and 45003
comment replies — were silently never dispatched to the agent; chat and
stream channels worked, making the gap invisible (#90309). Block's own
ACP harness documents the forum kinds explicitly.

Introduce _DISPATCH_KINDS = {9, 45001, 45003} for the subscription
filter and dispatch gate. The stream kinds (46010/40007/45002) stay out
of scope until their dispatch semantics are confirmed.
_is_direct_message_event deliberately keeps its kind-9-only check:
widening it would let a p-tagged forum post be reclassified as a DM and
bypass mention gating. The send path already works unchanged (send()
omits --kind and threads via --reply-to).

Fixes #90309
2026-08-31 09:05:41 -07:00
Teknium 972f0314de feat(buzz): compose thread-topology cluster — reply_in_thread opt-out, NIP-10 root anchoring on all send paths, _PLATFORM_DEFAULTS tier
Compose/fix-up on top of the cherry-picked cluster commits:

- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
  field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
  users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
  through _apply_yaml_config and honored by send(), send_image(), and
  _standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
  _progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
  so the synthetic-thread fallback and the progress reply anchor are both
  suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
  _extract_thread_root (marked root > reply > legacy positional e-tag)
  instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
  edit_message now implemented, accumulate-style progress works, but without
  the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
  update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
2026-08-31 07:30:44 -07:00
NanakoIce 66fa6e41c4 fix(buzz): reply in-thread instead of flat channel posts
Buzz has no native thread_id; channel threading is entirely --reply-to on
the triggering event. Interim commentary and progress bubbles only passed
the anchor via metadata.reply_to_message_id (or not at all), so most Kathy
posts landed as new top-level messages and cluttered channels.

- Honor metadata.reply_to_message_id in BuzzAdapter.send
- Pass reply_to on stream commentary sends
- Treat buzz like slack/mattermost for progress thread resolution
- Set _progress_reply_to to the trigger event for buzz
- Add unit tests for adapter metadata and progress routing
2026-08-31 07:30:44 -07:00
Tom Watts cbeb925f0a fix(buzz): reply into the existing thread instead of nesting a new one
Every Buzz reply opened a fresh thread, including when the user was
already replying inside one. A threaded client fills up with an endless
ladder of one-message threads and the conversation becomes unreadable.

The adapter itself never had threading logic; the behaviour comes from
the generic gateway default. `_reply_anchor_for_event()` in
gateway/platforms/base.py returns `event.message_id`, which is right for
reply-style platforms (Telegram/Discord "reply to this message") but
wrong for a thread-style one: anchoring to the message you are answering
nests a new sub-thread under every single turn.

Buzz threads are NIP-10, so the information needed is already on the
inbound event. The adapter now records each inbound message's thread
root from its `e` tags and resolves the outbound anchor to that root, so
a reply joins the thread the user is typing in. When the trigger was
itself top-level there is no root and the anchor passes through
unchanged, preserving the existing behaviour of opening exactly one
thread from a top-level message.

Fixed in the adapter rather than in `_reply_anchor_for_event()`: Buzz is
a plugin-supplied platform, and its NIP-10 tag semantics do not belong in
core. Root extraction prefers an explicit `root` marker, falls back to a
lone `reply` marker (a message bearing only `reply` started the thread,
so that parent is the root for everything after it), and treats a legacy
unmarked `e` tag as the parent. A mention-only `p` tag is not a reply.

The root cache is an OrderedDict bounded at 512 entries with FIFO
eviction so a long-lived gateway cannot leak, and the resolver is applied
to the image send path as well as `send()`. Both helpers tolerate a
missing `_thread_roots` attribute, since the standalone/cron send path
constructs an adapter without running `__init__`.

Tests cover root extraction (top-level, thread opener, nested, legacy
unmarked tag), the top-level passthrough that guards the existing
behaviour, unknown/None anchors, cache bounding and eviction, and an
end-to-end assertion through `send()` that `--reply-to` carries the root.
Verified against the real event shapes returned by a live hosted relay.
2026-08-31 07:30:44 -07:00
al9000-max 5d73a11a97 Buzz adapter: honor reply_to_mode instead of always threading replies
The Buzz adapter appended --reply-to unconditionally, so every agent reply
threaded onto its parent event id with no way to turn it off.

reply_to_mode is already a generic PlatformConfig field, parsed for any
platform from gateway.platforms.<name>.reply_to_mode, and the Discord and
Telegram adapters both honor it. Buzz never read it, so setting it was a
silent no-op.

Read it in __init__ (BUZZ_REPLY_TO_MODE overrides config.yaml, matching how
require_mention and transport already work in this adapter) and skip the
--reply-to append when it is "off", at all three send paths: send(),
send_image(), and the out-of-process _standalone_send() used for
deliver=buzz cron delivery.

Default is unchanged ("first"), so existing installs keep threading.
2026-08-31 07:30:44 -07:00
yuvalfis 09cbce43e0 fix(buzz): preserve stable thread roots 2026-08-31 07:30:44 -07:00
Han Ngo 34c10f83c3 feat(buzz): implement edit_message and delete_message so replies can stream
The gateway already streams by sending a first partial message and re-editing
it as tokens arrive, falling back to that path when an adapter does not support
native drafts. The Buzz adapter never implemented edit_message, so it inherited
the base stub that returns success=False and every reply was delivered in one
block when the turn finished, however long the turn took.

buzz-cli already exposes `messages edit` and `messages delete`, so no new
mechanism is needed.

One detail worth calling out for review: buzz-cli reports a NEW event id for
each edit, but the edit TARGET stays the original id, and the stream consumer
holds a single message_id for the whole stream. edit_message therefore returns
the id it was given rather than the one the CLI reports. Returning the CLI's id
would make every edit after the first address a message that was never sent.

delete_message is included because the consumer's fresh-final cleanup path
calls it when it replaces a preview rather than editing in place.

Tested: 10 new cases in tests/gateway/test_buzz_adapter.py covering the edit
target, stdin content, the returned id, echo suppression, finalize being inert,
both no-op guards, retryable vs non-retryable CLI failures, and delete. The
file goes from 33 passing to 6 failing if the adapter change is reverted while
the tests stay.
2026-08-31 07:30:44 -07:00
Teknium 9113cf24a9 fix(buzz): reconcile scoped auth-tag resolution across salvaged fixes
The WS NIP-42 auth path now prefers the connect()-resolved _auth_tag
(credentials-file aware, #79514) and falls back to a lazy scope-aware
_resolve_auth_tag() so a bare adapter re-auth stays profile-correct
(#98738): scoped multiplex profiles fail closed instead of borrowing
the default profile's tag from os.environ. _exec_buzz fakes updated
for the auth_tag kwarg introduced by the #83155 salvage.
2026-08-31 07:28:30 -07:00
Alex P. Günsberg 29ee8230be fix(buzz): fail closed on multiplex credential discovery 2026-08-31 07:28:30 -07:00
Alex P. Günsberg 8c09c39530 fix(buzz): scope owner credentials per profile 2026-08-31 07:28:30 -07:00
Alex P. Günsberg 51ffeb059c fix(buzz): load owner auth tag from credentials 2026-08-31 07:28:30 -07:00
liuhao1024 56fb2b3b02 fix(buzz): drop unreachable line orphaned by the scoped-secrets refactor
Remove the stray 'return val if val is not None else default' tail left
in _unscoped_profile_secrets() when the new return was added, and note
in the docstring that the process-global cache is startup-gate-only
(review feedback on #95224).
2026-08-31 07:28:30 -07:00
liuhao1024 a684d154bc fix(buzz): let the requirement gate see externally managed secrets
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.
2026-08-31 07:28:30 -07:00
Kosta Gorod a5e7355766 fix(buzz): resolve BUZZ_AUTH_TAG through the profile secret scope
One BUZZ_* read survived the #98738 sweep unscoped: the NIP-42 WebSocket
auth path read BUZZ_AUTH_TAG with a bare os.getenv. Under
gateway.multiplex_profiles the process env holds the default profile's
bridge/.env output, so a scoped secondary profile without its own tag
signed its relay auth event with the default profile's NIP-OA
owner-attestation tag. Reproduced on f3845a72af before the fix; the same
repro now attaches no tag (fail-closed).

The read goes through _get_scoped_secret: scoped multiplex profiles fail
closed to "", while single-profile and unscoped default-profile reads
keep the legacy env behavior. Adversarial coverage added for the leak
itself, the scoped positive control, unscoped precedence, partial-extra
adapter config, scoped validate_config, scoped standalone-send target
resolution, central-authz wildcard/blank-entry/normalization semantics,
and adapter-intake vs central-authz agreement on the same allowlist.

Fixes #98738

Signed-off-by: Kosta Gorod <35299380+KostaGorod@users.noreply.github.com>
2026-08-31 07:28:30 -07:00
liuhao1024 aaa5f27d0f fix(buzz): secondary multiplex profiles must not inherit the default profile's env
Under gateway.multiplex_profiles the default profile's YAML-to-env bridge
writes BUZZ_* values into os.environ, and every Buzz read gave that env
precedence over the secondary profile's PlatformConfig — so each secondary
adapter connected as the default identity, watched its channels, and
resolved its credentials file (#98738).

- Add _profile_scoped()/_scoped_platform_setting(): inside a secondary
  profile scope extra is authoritative and env is not consulted (a missing
  key fails closed to its default instead of borrowing the default
  profile's value); single-profile and unscoped/default-profile reads keep
  the legacy env-over-config precedence.
- Apply the scoped read to BuzzAdapter.__init__ (relay, CLI path, channels,
  home channel, poll interval, require_mention, transport, allowed users),
  _resolve_private_key (BUZZ_CREDENTIALS_FILE), validate_config,
  _standalone_send, and check_requirements (which now consults the
  profile's own config.yaml via the scoped home override).
- _env_enablement() returns None inside a profile scope and
  _apply_yaml_config() skips the env bridge there, so the default profile's
  env cannot fabricate Buzz for a profile that never configured it and a
  secondary profile's YAML cannot be pinned into the process env
  (first-writer-wins, #72348 Telegram/Discord mirror).
- Central authorization now consults a plugin platform's live-adapter
  config.extra.allowed_users (gated on the registry entry declaring
  allowed_users_env, with an optional normalize_user_id hook so Buzz npub
  entries match hex-pubkey user ids) — under multiplex only the default
  profile's list ever reached the env var, so listed secondary-profile
  users were default-denied (#82871). Empty/absent lists change nothing;
  default-deny is preserved.
2026-08-31 07:28:30 -07:00
Teknium a0a63a1bc2 fix(gateway): username-based DISCORD_ALLOWED_USERS no longer locks out the operator after one turn
The Discord adapter resolves username allowlist entries to numeric IDs at
connect and mirrors them into os.environ — but the gateway's per-turn .env
hot-reload (load_hermes_dotenv(override=True)) restores the raw usernames
from the file. From the second agent turn onward, _is_user_authorized
compared numeric user_ids against username strings and dropped every
message from the operator as 'Unauthorized user' while the adapter layer
still admitted them (bot reacted, never replied).

Fix: gateway authz unions the adapter's resolved numeric IDs
(DiscordAdapter.resolved_allowlist_user_ids()) into the env-derived
allowlist. Union only fires when an env allowlist is configured (never a
widening; fail-closed branch unchanged), is duck-typed + isinstance-guarded
against mock adapters, and filters non-numeric entries so unresolved
usernames and '*' can't leak through adapter memory.

Live repro: symptom fired on origin/main (authorized=False after reload),
passes with fix; stranger + empty-allowlist + raising-resolver negatives
hold. Sabotage run: incident test fails on unfixed authz_mixin.
2026-08-31 06:28:25 -07:00
ehz0ah 64b96bb5d2 fix(openviking): synchronize setup connection state 2026-08-31 17:20:45 +05:30
ehz0ah 823bcc887a feat(openviking): use user memory by default
Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.

Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
2026-08-31 17:20:45 +05:30
Teknium d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
kshitijk4poor 6681f9ebc3 refactor(telegram): share exception-graph walk across classifiers
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.

Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
2026-08-31 11:35:24 +05:30
kshitijk4poor 8288129475 fix(telegram): bound polling drain with wall-clock deadline
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.

Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
2026-08-31 11:34:16 +05:30
Teknium d63f996a75 feat(photon): read-receipt toggle, receipt-type alias, docs
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
  messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
  coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
2026-08-30 18:37:50 -07:00
Zihan Huang 9744fc0c99 feat(photon): support iMessage read receipts 2026-08-30 18:37:50 -07:00
Teknium 5bdaea64ed feat(telegram): inline command picker — search every command and skill, no menu cap
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).

- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
  pagination logic (unit-testable without python-telegram-bot). First
  query token filters; the remainder is carried into the sent command as
  its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
  enables inline mode via BotFather /setinline) + _handle_inline_query
  with the same auth path as inline-button callbacks — unauthorized users
  get an empty list, so the installed-skill catalog is not leaked to
  arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
  message starts with /, which reaches the bot even under default privacy
  mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
  /setinline setup.
2026-08-29 20:57:55 -07:00
Teknium 83f4524b42 feat(discord): expose /plan in the native slash-command picker
Text-message /plan already works on Discord via the gateway fall-through;
this makes it discoverable in the / picker alongside /steer and /compress.
2026-08-29 19:14:15 -07:00
AideYu 5cd9c4563c fix(telegram): recover exhausted request pool 2026-08-29 18:10:03 -07:00
GodsBoy e38cca50d6 fix(memory): keep Mem0 OSS OpenAI requests direct 2026-08-29 18:02:19 -07:00
Ayush Nangia 94aad6dcd2 fix(providers): surface Alibaba China in desktop parity 2026-08-29 18:01:25 -07:00
Ayush Nangia 695d86f518 fix(providers): fold Token Plan into the alibaba plugin, add runtime-path regressions, document all variants
Sweeper review, all three points:

- Placement: no new plugins/model-providers/ directory. The Token Plan
  profiles register from the existing alibaba plugin module — one module
  per vendor, matching how the kimi module carries both of its endpoint
  variants. Token Plan is the same vendor/service (Model Studio), same
  OpenAI-compatible protocol, its own key + endpoints; splitting to a
  standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
  resolve_runtime_provider() for all four variants — provider, api_key,
  api_mode, base_url — alongside the existing zai/minimax/kilocode
  runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
  with the bundled variants and their env keys/base-url overrides.
2026-08-29 18:01:25 -07:00
Ayush Nangia 7cf7df0297 fix(providers): register Alibaba China + Token Plan provider profiles (#73265)
The models.dev catalog advertises alibaba-cn, alibaba-coding-plan-cn, and
alibaba-token-plan(-cn), and resolve_provider_full()'s catalog chain lets
the CLI --provider path resolve them — but auth.resolve_provider() (the
credential/runtime path used by 'hermes chat') consults only
PROVIDER_REGISTRY and raised "Unknown provider 'alibaba-coding-plan-cn'"
(hermes_cli/auth.py:1937). PROVIDER_REGISTRY auto-extends from provider
profiles (auth.py:461-490), so the fix registers the missing profiles at
that chokepoint: alibaba-cn joins the alibaba plugin, alibaba-coding-plan-cn
joins alibaba-coding-plan, and a new alibaba-token-plan plugin registers
both regional token-plan tiers. Names match the catalog keys exactly.

No core edits — plugins/model-providers is the designed extension path.
2026-08-29 18:01:25 -07:00
kshitijk4poor 1e21fe8624 test(custom): align Mistral-omission test inputs; soften models.py comment
The chat_completions and transport-parity Mistral tests pinned the
same branch with different reasoning_config shapes ({effort:none} vs
{enabled:False, effort:none}) — behaviorally identical since effort
short-circuits first, but the drift reads as a semantic difference.
Align both to the explicit form. Also scope the port-guard comment to
the try/except shape it actually shares with hermes_cli/models.py.
2026-08-29 21:58:00 +05:30
kshitijk4poor c870589831 fix(custom): tolerate malformed ports in the Ollama URL heuristic
urlparse raises ValueError on non-integer / out-of-range ports, and
http://myhost:99999/v1 passes OpenAI-client construction (only httpx
rejects it later), so the crash was reachable from build_kwargs on
every request for such a URL. Wrap the parsed.port check in the same
try/except ValueError guard hermes_cli.models already uses around its
11434 check, and pin it with parametrized tests.
2026-08-29 21:58:00 +05:30
xxxigm 31f0336da7 fix(custom): omit Ollama-only think=false on strict OpenAI-compat endpoints
reasoning_effort: none was injecting extra_body.think=false for every
custom provider. Mistral (and other extra=forbid hosts) reject that
field with HTTP 422. Keep think=false on Ollama URLs only; still send
top-level reasoning_effort=none so /v1 thinking-off keeps working.
2026-08-29 21:58:00 +05:30
kshitijk4poor b954547e72 fix(nebius): route effort through canonical clamp_effort — hand-rolled map inverted the ladder
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
2026-08-29 20:39:44 +05:30
amrrs 49f5b6a9b4 fix(nebius): request verbose model metadata 2026-08-29 20:39:44 +05:30
amrrs 13bad590f3 feat(providers): add Nebius Token Factory provider 2026-08-29 20:39:44 +05:30
kshitijk4poor 1c5ee5815f fix(router): pytest guard on caps warmer + debug log in fail-open efforts lookup
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
2026-08-29 20:29:04 +05:30
Neel Patel ca3060b70f docs: chat-completions is now a compat shim on Router, not a 404
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
2026-08-29 20:29:04 +05:30
Neel Patel eeb7916000 review: host-resolved efforts, ladder-validated ingest, deduped catalog
Addresses the automated review on this PR:

- _profile_declared_efforts falls back from provider name to the
  endpoint's host (via model_metadata's URL->provider map), so a named
  custom provider pointed at api.router.com — which the host mandate
  already routes onto this transport — gets the catalog clamp instead
  of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
  ingest, logging and dropping unrecognized tiers; a model whose whole
  vocabulary is unrecognized stays out of the map (transport defaults)
  instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
  order.
- plugin.yaml credits the human contributor per repo convention.
2026-08-29 20:29:04 +05:30
Neel Patel ceb54f5a52 feat(providers): send Hermes-Agent User-Agent on Router requests
Router attributes coding-agent clients by User-Agent prefix (the way it
already recognizes OpenCode's versioned UA), and its WAF rejects
default/blank client UAs. Mirror the xai profile: declare
User-Agent: Hermes-Agent/<version> in default_headers, which
agent_init's profile-headers fallback applies at client construction.
Re-verified live: one-shot chat through api.router.com still works.
2026-08-29 20:29:04 +05:30
Neel Patel 804f8b4732 feat(providers): add Ramp Router (router.com) provider plugin
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:

- plugins/model-providers/router/: RouterProfile plugin —
  api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
  RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
  GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
  Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
  runtime_provider._detect_api_mode_for_url: api.router.com ->
  codex_responses. The host is Responses-only — POST /v1/chat/completions
  does not exist and 404s — so this is a genuine host mandate (exact
  hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
  hook (tri-state: None=defer, ()=model takes no reasoning params,
  tuple=clamp set). Router validates reasoning.effort per model and
  returns HTTP 400 invalid-argument on levels outside the model's
  published vocabulary, and 400 unsupported_parameter when a
  non-reasoning model receives any reasoning field (both verified live).
  The profile answers from a cached copy of the catalog's
  router.capabilities.reasoning block: cache-only on the hot path,
  seeded for free by fetch_models(), disk-mirrored across processes
  (/cache/router_catalog.json), background-warmed when cold
  — same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
  the generic effort-clamp branch (xai/actual/github branches untouched;
  profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
  document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
  rejection, profile registration + auth auto-registry wiring, catalog
  parsing, and transport clamp/suppression/fallback paths.

Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
2026-08-29 20:29:04 +05:30
Teknium 74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Teknium ccc367dce0 fix(prompt)+feat(gateway): platform-hint truth pass + universal voice-bubble transcode (all 22 hints source-verified) (#97873)
* fix(prompt): platform-hint truth pass — CLI/TUI file-delivery reality (paths/URLs only, MEDIA: prints literally), CLI no-markdown verified live, Slack/Discord markdown+tables truth, shared local-cron constant

* feat(gateway): universal voice-bubble delivery — shared transcode_to_ogg_opus; telegram [[audio_as_voice]] any-format; feishu native voice; hints to new truth

* chore: delete the webui ghost hint (tombstone comment, audit-verified); sync send_voice signature pin in tts routing test
2026-08-29 05:57:13 -07:00