Commit Graph

1559 Commits

Author SHA1 Message Date
Teknium 83f4524b42 feat(discord): expose /plan in the native slash-command picker
Text-message /plan already works on Discord via the gateway fall-through;
this makes it discoverable in the / picker alongside /steer and /compress.
2026-08-29 19:14:15 -07:00
AideYu 5cd9c4563c fix(telegram): recover exhausted request pool 2026-08-29 18:10:03 -07:00
GodsBoy e38cca50d6 fix(memory): keep Mem0 OSS OpenAI requests direct 2026-08-29 18:02:19 -07:00
Ayush Nangia 94aad6dcd2 fix(providers): surface Alibaba China in desktop parity 2026-08-29 18:01:25 -07:00
Ayush Nangia 695d86f518 fix(providers): fold Token Plan into the alibaba plugin, add runtime-path regressions, document all variants
Sweeper review, all three points:

- Placement: no new plugins/model-providers/ directory. The Token Plan
  profiles register from the existing alibaba plugin module — one module
  per vendor, matching how the kimi module carries both of its endpoint
  variants. Token Plan is the same vendor/service (Model Studio), same
  OpenAI-compatible protocol, its own key + endpoints; splitting to a
  standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
  resolve_runtime_provider() for all four variants — provider, api_key,
  api_mode, base_url — alongside the existing zai/minimax/kilocode
  runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
  with the bundled variants and their env keys/base-url overrides.
2026-08-29 18:01:25 -07:00
Ayush Nangia 7cf7df0297 fix(providers): register Alibaba China + Token Plan provider profiles (#73265)
The models.dev catalog advertises alibaba-cn, alibaba-coding-plan-cn, and
alibaba-token-plan(-cn), and resolve_provider_full()'s catalog chain lets
the CLI --provider path resolve them — but auth.resolve_provider() (the
credential/runtime path used by 'hermes chat') consults only
PROVIDER_REGISTRY and raised "Unknown provider 'alibaba-coding-plan-cn'"
(hermes_cli/auth.py:1937). PROVIDER_REGISTRY auto-extends from provider
profiles (auth.py:461-490), so the fix registers the missing profiles at
that chokepoint: alibaba-cn joins the alibaba plugin, alibaba-coding-plan-cn
joins alibaba-coding-plan, and a new alibaba-token-plan plugin registers
both regional token-plan tiers. Names match the catalog keys exactly.

No core edits — plugins/model-providers is the designed extension path.
2026-08-29 18:01:25 -07:00
kshitijk4poor 1e21fe8624 test(custom): align Mistral-omission test inputs; soften models.py comment
The chat_completions and transport-parity Mistral tests pinned the
same branch with different reasoning_config shapes ({effort:none} vs
{enabled:False, effort:none}) — behaviorally identical since effort
short-circuits first, but the drift reads as a semantic difference.
Align both to the explicit form. Also scope the port-guard comment to
the try/except shape it actually shares with hermes_cli/models.py.
2026-08-29 21:58:00 +05:30
kshitijk4poor c870589831 fix(custom): tolerate malformed ports in the Ollama URL heuristic
urlparse raises ValueError on non-integer / out-of-range ports, and
http://myhost:99999/v1 passes OpenAI-client construction (only httpx
rejects it later), so the crash was reachable from build_kwargs on
every request for such a URL. Wrap the parsed.port check in the same
try/except ValueError guard hermes_cli.models already uses around its
11434 check, and pin it with parametrized tests.
2026-08-29 21:58:00 +05:30
xxxigm 31f0336da7 fix(custom): omit Ollama-only think=false on strict OpenAI-compat endpoints
reasoning_effort: none was injecting extra_body.think=false for every
custom provider. Mistral (and other extra=forbid hosts) reject that
field with HTTP 422. Keep think=false on Ollama URLs only; still send
top-level reasoning_effort=none so /v1 thinking-off keeps working.
2026-08-29 21:58:00 +05:30
kshitijk4poor b954547e72 fix(nebius): route effort through canonical clamp_effort — hand-rolled map inverted the ladder
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
2026-08-29 20:39:44 +05:30
amrrs 49f5b6a9b4 fix(nebius): request verbose model metadata 2026-08-29 20:39:44 +05:30
amrrs 13bad590f3 feat(providers): add Nebius Token Factory provider 2026-08-29 20:39:44 +05:30
kshitijk4poor 1c5ee5815f fix(router): pytest guard on caps warmer + debug log in fail-open efforts lookup
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
2026-08-29 20:29:04 +05:30
Neel Patel ca3060b70f docs: chat-completions is now a compat shim on Router, not a 404
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
2026-08-29 20:29:04 +05:30
Neel Patel eeb7916000 review: host-resolved efforts, ladder-validated ingest, deduped catalog
Addresses the automated review on this PR:

- _profile_declared_efforts falls back from provider name to the
  endpoint's host (via model_metadata's URL->provider map), so a named
  custom provider pointed at api.router.com — which the host mandate
  already routes onto this transport — gets the catalog clamp instead
  of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
  ingest, logging and dropping unrecognized tiers; a model whose whole
  vocabulary is unrecognized stays out of the map (transport defaults)
  instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
  order.
- plugin.yaml credits the human contributor per repo convention.
2026-08-29 20:29:04 +05:30
Neel Patel ceb54f5a52 feat(providers): send Hermes-Agent User-Agent on Router requests
Router attributes coding-agent clients by User-Agent prefix (the way it
already recognizes OpenCode's versioned UA), and its WAF rejects
default/blank client UAs. Mirror the xai profile: declare
User-Agent: Hermes-Agent/<version> in default_headers, which
agent_init's profile-headers fallback applies at client construction.
Re-verified live: one-shot chat through api.router.com still works.
2026-08-29 20:29:04 +05:30
Neel Patel 804f8b4732 feat(providers): add Ramp Router (router.com) provider plugin
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:

- plugins/model-providers/router/: RouterProfile plugin —
  api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
  RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
  GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
  Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
  runtime_provider._detect_api_mode_for_url: api.router.com ->
  codex_responses. The host is Responses-only — POST /v1/chat/completions
  does not exist and 404s — so this is a genuine host mandate (exact
  hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
  hook (tri-state: None=defer, ()=model takes no reasoning params,
  tuple=clamp set). Router validates reasoning.effort per model and
  returns HTTP 400 invalid-argument on levels outside the model's
  published vocabulary, and 400 unsupported_parameter when a
  non-reasoning model receives any reasoning field (both verified live).
  The profile answers from a cached copy of the catalog's
  router.capabilities.reasoning block: cache-only on the hot path,
  seeded for free by fetch_models(), disk-mirrored across processes
  (/cache/router_catalog.json), background-warmed when cold
  — same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
  the generic effort-clamp branch (xai/actual/github branches untouched;
  profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
  document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
  rejection, profile registration + auth auto-registry wiring, catalog
  parsing, and transport clamp/suppression/fallback paths.

Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
2026-08-29 20:29:04 +05:30
Teknium 74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Teknium ccc367dce0 fix(prompt)+feat(gateway): platform-hint truth pass + universal voice-bubble transcode (all 22 hints source-verified) (#97873)
* fix(prompt): platform-hint truth pass — CLI/TUI file-delivery reality (paths/URLs only, MEDIA: prints literally), CLI no-markdown verified live, Slack/Discord markdown+tables truth, shared local-cron constant

* feat(gateway): universal voice-bubble delivery — shared transcode_to_ogg_opus; telegram [[audio_as_voice]] any-format; feishu native voice; hints to new truth

* chore: delete the webui ghost hint (tombstone comment, audit-verified); sync send_voice signature pin in tts routing test
2026-08-29 05:57:13 -07:00
kshitijk4poor 2f01ec9fa4 fix(teams): allowlist-gate BF attachment auth, stream downloads under media cap, lock token refresh
Second follow-up for salvaged PR #94547, folding in review findings from the
duplicate-PR cluster (#47015, #55054, #58476, #72977, #73685 all fix the
same 401) and the sweeper review of #73685:

- Replace the dot-anchored suffix predicate with exact-match against the
  existing _ALLOWED_TEAMS_SERVICE_HOSTS allowlist (two of the five
  duplicate PRs converged on this independently). Any Azure customer can
  register <name>.trafficmanager.net profiles, so suffix matching was not
  safe. Also requires https on the default port — :444 on an allowlisted
  host no longer receives the bearer (sweeper finding on #73685).
- Stream _fetch_attachment_bytes through _read_httpx_body_with_limit
  instead of buffering response.content — the shared inbound media cap now
  applies to authenticated downloads too (sweeper finding: a lying
  Content-Length must not OOM the gateway).
- Serialize token refresh with a lazily-bound asyncio.Lock so concurrent
  attachments share one STS POST (review finding on #94547).
- Token expiry now uses time.monotonic() (from #55054) — wall-clock jumps
  can't extend a stale token.
- Tests updated: exact-allowlist predicate (lookalike/subdomain/port/scheme
  negatives), streaming fake client, and a concurrent-cold-cache lock test.
  Mutation-checked: suffix match, silent drop, no-lock, and unbounded
  buffer each fail a test.
2026-08-29 11:04:38 +05:30
kshitijk4poor c23d40af17 fix(teams): dot-anchor Bot Framework host check, log dropped BF images, tests
Follow-up for salvaged PR #94547 (Sibbern's Bot Framework attachment auth):

- The host check used bare endswith('trafficmanager.net') /
  endswith('botframework.com'), which matched attacker lookalike hosts
  (evil-trafficmanager.net) — sending the bot's bearer token off-platform.
  Same threat model _ALLOWED_TEAMS_SERVICE_HOSTS already documents. Now a
  single dot-anchored predicate, _is_botframework_attachment_host(),
  used by both _fetch_attachment_bytes and the _on_message image branch
  (was copy-pasted in two places).
- BF images whose bytes fail image validation were silently dropped with
  no log (cache_media_bytes returns None; old path logged). Add the
  missing else-warning, mirroring the document branch.
- Record cached_m.media_type instead of the raw content_type so the
  MessageEvent MIME matches what was actually cached.
- Init _bf_token_cache in __init__ (was masked by getattr).
- Tests: host predicate dot-anchoring (attacker lookalikes blocked),
  BF routing vs generic helper, bearer attach + attacker-host block end
  to end, token acquire + cache reuse, token-failure degradation, and
  the silent-drop regression guard. Mutation-checked: reverting the
  dot-anchor, the else-warning, or the token cache each fails a test.
2026-08-29 11:04:38 +05:30
Jacob S. 0eff6bc200 fix(teams): authenticate Bot Framework connector attachment downloads
Inline/pasted images in Teams arrive with a contentUrl on
smba.trafficmanager.net (/v3/attachments/...). Unlike file uploads
(pre-authenticated SharePoint downloadUrls), these connector URLs
require the bot's own bearer token; fetching them anonymously fails
with 401 Unauthorized and the image is silently dropped.

- add _get_botframework_token(): client-credentials token for
  https://api.botframework.com/.default, cached until ~5 min before
  expiry
- _fetch_attachment_bytes(): attach the token when the attachment
  host is *.trafficmanager.net / *.botframework.com; SharePoint and
  other URLs remain auth-free as before
- route image/* attachments with Bot Framework contentUrls through
  the authenticated fetch instead of cache_image_from_url (which
  sends no Authorization header)

SSRF guards unchanged. Verified on a live Teams personal-scope bot:
pasted images previously logged '[teams] Failed to cache image
attachment: 401 Unauthorized' and now cache and deliver correctly.
2026-08-29 11:04:38 +05:30
Brooklyn Nicholson 5e550838f7 feat(kanban): board export/import REST endpoints
POST /boards/{slug}/export and POST /boards/import, so the desktop and
dashboard can drive board transfer. Both exchange filesystem paths
rather than bytes, the same contract profile export/import uses and for
the same reason: the client runs its native save/open dialog on the
machine that hosts the backend, so a path is all either side needs, and
a board carrying a few hundred megabytes of attachments never has to
cross the renderer heap.

Also covers the pre-existing rename (PATCH) and delete endpoints, which
had no tests — including that delete archives to a restorable directory
and refuses to touch `default`.
2026-08-28 22:38:01 -05:00
Teknium 3340bbbdad feat(a2a): client tools config-gated — disabled unless enabled (−561 tok/call on unconfigured installs) (#97421)
* feat(a2a): outbound client tools are config-gated — served only when a2a_agents configured, inbound platform enabled, or A2A_PORT set (-561 tok/call on unconfigured installs)

* ci: retrigger after runner startup_failure on rerun attempt
2026-08-28 15:08:08 -07:00
Teknium a641644f12 fix(paths): display_hermes_home renders POSIX separators on Windows — kills ~/AppData\Local\hermes chimeras in tool schemas and user-facing messages (#97137) 2026-08-28 05:45:00 -07:00
Teknium 536adb35c4 refactor(video_generate): capability-gated dynamic schema (~814 → 458/377 tok/call) (#97095)
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests

* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)

* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
2026-08-28 05:10:33 -07:00
Teknium a619db6633 refactor(image_generate): capability-gated dynamic schema (554 → 317 tok/call, −43%) (#97057)
* refactor(image_generate): capability-gated dynamic schema — args render only when the active model honors them (554 -> 317 tok/call, -43%)

* fix(image_gen): fleet-wide supports_upscale declarations — krea (Enhance) + fal plugin (Clarity passthrough) declare it; declaration<->implementation contract-tested across all 7 in-tree providers
2026-08-28 03:58:52 -07:00
Teknium 34393c32aa feat(plugins): wire plugin platform handlers into a2a, buzz, and qqbot adapters
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
2026-08-27 07:51:37 -07:00
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
teknium1 c96f830252 feat(plugins): let plugins register Telegram PTB handlers via ctx.register_telegram_handler
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.

Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
2026-08-27 07:51:37 -07:00
wansui 9faa953cd1 fix(wecom): eliminate duplicate + split bubbles in native streaming
Fixes two related native-streaming bubble defects surfaced in production:

1. Duplicate bubble on long turns — when keep-alive already refreshed the
   6-min reply window, the Layer-2 clock fallback still declined the finalize
   frame and forced a proactive send(), duplicating the message. Skip the
   clock fallback while keep-alive is active; intermediate-frame failures are
   now fully fire-and-forget (only a failed FINAL frame falls back to send()).

2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
   a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
      notice) called _reset_segment_state(), clearing the cumulative
      _accumulated so the next delta + finalize frame carried only the few
      chars accumulated after the reset. Native streaming now skips that
      reset (commentary still posts as its own message via send()).
   b) The adapter-side _BlockChunker.update() 'only grow' guard silently
      dropped any cumulative snapshot shorter than its high-water mark, so
      after a baseline reset the leading characters were stranded before
      _emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
      layer entirely; intermediate frames are pure identity-dedup, matching
      the fire-and-forget model.

Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.

Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).
2026-08-27 07:33:36 -07:00
wansui 2ecb544551 feat(wecom): native reply streaming (per-turn isolation, dedup-safe delivery, interaction boundaries)
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.

Includes the machinery intrinsic to native streaming on WeCom:

- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
  frames update it, a finalize frame closes it; native-streaming adapters are
  let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
  long-connection mode has no documented edit-rate limit); an adapter-level
  frame cap is retained. (Early builds gated frames behind a char throttle;
  removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
  isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
  (control vs normal) plus a per-chat token bucket to stay under WeCom's rate
  limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
  delivered the moment it is emitted; failures logged, not re-sent; delivery
  marked once per turn), per-req_id reply queue with ack tracking, and the
  timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
  846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
  the prompt is the last thing on screen and never traps a lingering bubble;
  eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
  stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
  turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
  messages; image+text double-callback merged into one turn.

Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.
2026-08-27 07:33:36 -07:00
Ben Barclay 5521265def fix(slack): coerce string unfurl knobs on the native plane
Relay-plane parity: hermes config set / Railway persist YAML booleans
as strings, and _slack_unfurl_kwargs silently dropped them — so
'unfurl_links: "false"' was a no-op on native while working on relay.
Coerce recognized string booleans exactly as _slack_unfurl_hints does;
unrecognized values still drop so junk config keeps Slack's default
instead of accidentally suppressing previews.

Replaces test_send_ignores_non_boolean_unfurl_options (which froze the
dropped-string behavior) with coercion + junk-drop tests.
2026-08-27 15:33:32 +10:00
Will Lynas 91dbd7a6f8 fix(slack): preserve unfurl controls during streaming 2026-08-27 15:33:32 +10:00
Will Lynas b91845c2a0 fix(slack): honor unfurl controls for media captions 2026-08-27 15:33:32 +10:00
Andrew Bennett aeaa0c784e feat(slack): add link unfurl controls 2026-08-27 15:33:32 +10:00
Teknium 7a7a371c59 fix: harden claim-release guard for bare test doubles; repoint source-pinning test at the impl
The wrapper now getattr-defaults _processed_message_ts (object.__new__
adapters in sibling suites lack it), and the reaction-guard source pin
reads _handle_slack_message_impl where the production expression lives.
2026-08-26 15:54:53 -07:00
Teknium 39a5838f07 fix(slack): release a failed handler's fresh ts claim so the turn isn't swallowed
Follow-up for the #95417 salvage, addressing the review finding: the entry
claim closes the unfurl race but a handler that raises mid-enrichment would
hold the claim forever — neither a Slack retry nor a user edit could ever
re-drive the message. _handle_slack_message is now a thin guard around the
impl that releases only claims taken by the failed invocation itself, with a
warning log so swallowed turns are traceable. Pre-existing claims from a
successful turn are never released. Two failure-path tests pin both sides.
2026-08-26 15:54:53 -07:00
Richard Hojun Jang 708f84c477 fix(slack): claim message ts before enrichment so link unfurls can't duplicate a turn
Slack emits `message_changed` for a link unfurl carrying a DIFFERENT event ts
than the original message. That ts legitimately misses the `_dedup` check, so
`_processed_message_ts` is the only guard against it becoming a second user
turn -- but it was only populated at the END of `_handle_slack_message`, after
thread context, permalink resolution and file downloads had all awaited.

An unfurl landing inside that window found the guard empty and was promoted to
a duplicate turn: a spurious "Interrupting current task" banner plus the same
answer posted twice.

Production capture (adminbot, 2026-08-22 02:23:30-31Z, channel C0BF1EYUA9H):

  02:23:30.718  message      ts=1787365409.908499  dedup_hit=False
  02:23:30.737  app_mention  ts=1787365409.908499  dedup_hit=True
  02:23:31.675  message      ts=1787365411.012100  dedup_hit=False   <- leaked
                subtype=message_changed

The original copy was still resolving two Slack permalinks when the unfurl
arrived 957ms later.

Claim the message ts once every filter has passed and the event is certain to
be delivered, before the slow enrichment awaits. Claiming any earlier (right
after the dedup check) also claims messages the handler then discards, which
breaks summoning the bot by editing "@bot" into a previously ignored message
(tests/gateway/test_slack.py::TestMessageRouting::
test_message_edit_with_new_mention_processed).

Eviction logic is extracted to `_remember_processed_message_ts` so both call
sites share one bounded implementation.
2026-08-26 15:54:53 -07:00
Teknium b0563d5ad3 feat: MiniMax H3 Max joins the FAL video picker (t2v + i2v)
fal's post-trained H3 variant — #1-ranked quality/prompt adherence/
aesthetics, 5s 768p video in under 3 seconds, $0.04/s launch pricing.

- New minimax-h3-max family: minimax/h3-max/{text,image}-to-video
- Inherits base-H3 wire quirks (integer duration, i2v drops
  aspect_ratio) but caps at 768P (480P/768P enums, no 2K/4K) and
  declares seed on both endpoints
- New generic static_payload family flag: constant keys the endpoint
  requires on every request (H3 Max lists prompt_expansion_mode in its
  required array; sent as 'balanced')

Payload asserted against the endpoint OpenAPI schema; 73/73 targeted
tests green (surface matrix auto-covers the new family).
2026-08-26 15:48:43 -07:00
pierrenode 6766732620 fix(memory-setup): route .env writer through save_env_value's validation gate
hermes_cli/memory_setup.py::_write_env_vars() wrote provider-controlled
.env entries with a direct Path.write_text() + post-hoc chmod, bypassing
the denylist/regex/CRLF-stripping/atomic-replace validation that
hermes_cli/config.py::save_env_value() already provides for every other
.env writer in the codebase. A malicious or buggy memory-provider plugin
declaring a crafted env-var name/value in its setup schema could inject
arbitrary lines into .env.

Routes memory-provider env writes through save_env_value(), and fixes a
regression this surfaced in plugins/memory/supermemory/__init__.py::
post_setup(), which called the old two-parameter _write_env_vars(env_path,
values) signature — restores the caller via context-local
hermes_constants.set_hermes_home_override()/reset_hermes_home_override()
instead of a removed env_path parameter, so explicit HERMES_HOME overrides
during setup still resolve correctly.

Adds test_env_file_created_with_secure_permissions, guarded on Windows
(POSIX mode bits aren't enforced there, mirroring the existing skip in
test_openviking_provider.py / test_supermemory_provider.py) since
save_env_value's atomic-replace path creates the temp file at 0o600 before
writing content, closing the TOCTOU window the old direct-write + chmod
implementation had.
2026-08-26 15:48:39 -07:00
Nikita Barkov 2e80d7fa05 fix(slack): keep the resolved proxy on bolt's per-request client
slack_bolt builds a fresh AsyncWebClient for every inbound request and
copies proxy=app.client.proxy into its constructor, where slack_sdk reads
a None/blank proxy *argument* as "unspecified" and reloads HTTP(S)_PROXY
from the environment. aiohttp then treats that env value as an explicit
proxy and skips its own NO_PROXY check, so the adapter's resolved decision
to go direct - a NO_PROXY bypass, or a proxy scheme aiohttp cannot use -
holds on every client except the one authorization spends on auth.test.

The failure looks like a healthy bot: Socket Mode connects, outbound sends
keep working, and every inbound event is rejected with "Failed to authorize
with the given token" - forever, since a failed auth_test_result is not
cached and never retried differently.

Re-apply the resolved proxy through AsyncApp(before_authorize=...), which
bolt inserts before the authorization middleware: the request-scoped client
already exists there and has not been used yet. Assigning the attribute
post-construction is the only way to express "no proxy" to slack_sdk.

Co-authored-by: Junie <junie@jetbrains.com>
2026-08-26 10:35:05 -07:00
Nikita Barkov bac960e23d fix(slack): stop injecting thread roots as reply context 2026-08-26 10:26:00 -07:00
Nikita Barkov 5538bd1f93 fix(slack): prevent duplicate rich-text message content
Slack sends an authored message twice: flat in `event.text` and structurally
in `event.blocks`. The blocks are rendered so quoted and forwarded content is
not lost, and whatever the render carries beyond the flat text is appended to
the message. That comparison had several ways to fail on the *same* sentence,
each of which showed the author their own words a second time:

1. HTML entities — the flat copy escapes `&`/`<`/`>` while `blocks[].link.url`
   stays raw, so any link with query parameters (every "Copy link" on a
   thread) mismatched.
2. Permalink unfurls — the live inbound path skips `is_msg_unfurl`
   attachments, thread/parent hydration did not, so the linked message's body
   was appended again.
3. The Block Kit dump — it serialized the authored `rich_text` alongside the
   UI blocks it exists for, and its allowlist drops `url`, so the sentence
   reappeared with every link removed.
4. Unknown inline elements — the renderer knew eight types and silently
   dropped the rest. A pasted message permalink arrives as `message_mention`,
   so the link vanished from the render and the sides stopped comparing equal.
5. `message_mention` without a url — `url` is optional on that element while
   `channel_id` and `message_ts` are not, so the element rendered as nothing
   and the sentence came back with a blank in the link's place.
6. `date` elements — `fallback` and `url` are both optional, and the flat
   `<!date^…>` form was never read down to what the rich text renders.
7. Labelled mentions — Slack may attach a label (`<@U…|name>`,
   `<#C…|general>`, `<!subteam^S…|@marketing>`, `<!here|@here>`) in the flat
   text while the blocks carry the bare id. The bot's own mention is one of
   these, and stripping only its bare form left it in the flat copy.
8. Autolink schemes — only `https` and `mailto` were matched, so a `tel:` link
   kept its angle brackets and mismatched too.

Unknown inline types are now read by their `url`/`text`/`fallback` so a type
Slack adds later still renders, and `team`, `color` and a fallback-less `date`
render into the flat form Slack sends. Every field is read as a string or
not at all: Block Kit carries text as an object in many places, and a
non-string one reaches the renderer's `str.join` and raises there, which
costs the whole message. `channel_id` and `message_ts` are the
permalink's own components, so a url-less `message_mention` renders the
permalink's tail; the workspace host and the thread query cannot be rebuilt
from the element, so a permalink on either side is reduced to that same tail.
Canonicalization is used for matching only -- the authored text still reaches
the agent verbatim, so a mistake here can cost an unrendered element, never an
altered or missing message.

An element carrying neither a url nor a label still renders as nothing, and a
message containing one is still appended twice. Suppressing such a render was
tried and is worse: an app message whose body lives only in the blocks
disappears, and a forwarded quote is dropped. Genuinely additional content --
quotes, lists, code blocks, attachments, interactive bot blocks -- is
unaffected throughout.

Tests cover both merge sites (live inbound and thread hydration) and the
negative cases.
2026-08-26 10:10:51 -07:00
kshitijk4poor fab534b503 fix: omit User-Agent from anonymous OpenViking identity probes
Anonymous probes (_anonymous_json) are designed to probe server identity
before disclosing credentials. Sending the Hermes version on these probes
would fingerprint the exact version to an untrusted/MITM endpoint.

Keep User-Agent on authenticated requests (_headers) and multipart uploads
(_multipart_headers), which already send credentials.
2026-08-26 12:43:17 +05:30
ehz0ah 3db5267008 feat(openviking): identify Hermes requests 2026-08-26 12:43:17 +05:30
Gille 7c5c994397 fix(teams): request supported transcript content format 2026-08-26 02:54:35 +05:30
Jeffrey Quesnelle 5b82658b3c Merge pull request #90129 from rroverin/fix/nvidia-nim-tool-message-name-field
fix(providers): strip name/tool_name from NVIDIA NIM tool messages
2026-08-25 16:39:36 -04:00
kshitijk4poor b0cf2597c2 fix: follow-up for salvaged PR #93985 — cache key, snapshot, dead code
- Key _user_space_cache on _conn_snapshot instead of client object identity,
  so _new_client() results from the same connection share the cached user
  (previously every on_memory_write triggered an uncached /api/v1/system/status
  probe with a 30s default timeout)
- Thread a short timeout (0.05s) through the write-path identity probe
- Harden _tool_remember to snapshot the client before URI construction + POST,
  matching the pattern already established in on_memory_write
- Remove dead instance method _user_scoped_uri (zero callers; all call sites
  use the module-level function directly)

Co-authored-by: ehz0ah <haozhe4547@gmail.com>
2026-08-25 12:27:21 +05:30
ehz0ah 4387e03960 fix(memory): keep OpenViking identity operations consistent 2026-08-25 12:27:21 +05:30