Commit Graph

1477 Commits

Author SHA1 Message Date
Teknium ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal 28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
xxxigm b577353636 fix(teams-pipeline): parse Graph users('id') meeting paths
getAllTranscripts @odata.id often uses users('{organizer}')/onlineMeetings('{id}'),
which the slash-only parser missed, so new jobs still stored the transcript id
and hit the refuse-GET guard. Accept the quoted-user form and re-parse the
stored notification on run so replay picks the meeting id.
2026-08-21 12:15:36 +05:30
Gille 790c850144 fix(telegram): preserve rich finals after DM drafts 2026-08-20 21:58:18 -07:00
Teknium 19c6a11924 feat(opencode): sync Zen/Go catalogs — Ox Alpha stealth model, Grok routing, new free tier
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
  listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
  flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
  hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
  muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
  glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
  endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.

Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
2026-08-20 19:41:03 -07:00
Abdulkadir Ateş 5969ea1558 fix(opencode): route Muse Spark through the Responses API
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.

- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
2026-08-20 19:41:03 -07:00
Teknium 4ea69d9d2c feat: keyless web tier becomes a 5-vendor round-robin ring (adds Tavily, Firecrawl, Keenable)
Fresh installs with zero web credentials now rotate web_search/
web_extract across FIVE vendors' public free tiers — Exa, Parallel,
Tavily, Firecrawl, Keenable — instead of a 2-vendor 50/50 split, with
next-in-line ring failover on rate limits (multi-hop until a vendor
serves or the ring is exhausted; served_by marks the actual vendor).

- plugins/web/keenable/: new bundled provider (search via /v1/search,
  fetch via /v1/fetch; keyed Bearer or keyless with the mandatory
  X-Keenable-Title app header). Credit: integration proposed by
  Ilya Gusev (Keenable) in #49758; Free/Paid picker rows included.
- keyless_mcp: tavily/firecrawl/keenable keyless search+extract
  wrappers, _KEYLESS_RING + per-process round-robin cursor (seeded by
  the random session id, advances per unpinned request), pinned-vendor
  entry (pin = start there; rotation off), paid-pinned vendors excluded
  from the ring entirely.
- Tavily/Firecrawl providers route keyless traffic through the ring;
  both are now default-on ring members (no longer selection-gated).
- web_tools/registry: keenable in backend sets, auto-detect, availability
  probes; _keyless_preference() delegates to the ring cursor.
- KEENABLE_API_KEY in OPTIONAL_ENV_VARS; docs updated (ring semantics).

Live E2E: all 10 vendorXcapability paths (5 search + 5 extract) served
real results keyless; rotation cycled all five vendors over 5 dispatch
calls; double-throttle failover walked exa->parallel->tavily.
2026-08-20 00:17:25 -07:00
Teknium 797bc4bf9b Merge remote-tracking branch 'origin/main' into feat/keyless-tavily-firecrawl-failover 2026-08-19 23:04:36 -07:00
Slobaka 6ff341c4d6 fix(doctor): web readiness reflects the selected provider's real state (#78412)
Salvaged from #78434 by @Slobaka (also the issue reporter; earlier than
the competing #78436). hermes doctor no longer paints a green web check
when the explicitly selected provider cannot initialize — web splits
into per-capability rows (web search / web extract) resolved through
the same registry resolvers the dispatchers use, with readiness from a
true availability probe (_provider_is_ready).

Keyless-tier integration on top of the salvage:
- _provider_is_ready counts is_keyless_available() as ready — keyless
  mode is a working state, not a misconfiguration (zero-config installs
  and selected-keyless Tavily/Firecrawl show ok, not warn)
- Tavily/Firecrawl gain is_keyless_available() (True only when
  explicitly selected — they stay out of the zero-config fallback)
- doctor triggers plugin discovery before reading the registry (fresh
  doctor processes saw an empty registry and warned on everything)

E2E: searxng-selected-without-URL warns (the #78412 repro);
zero-config, tavily-keyless, firecrawl-keyless all read ok;
parallel pinned paid without a key warns.
2026-08-19 23:03:58 -07:00
liuhao1024 fbca706789 fix(telegram): log the first confirmed getUpdates progress per generation
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).

Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.

Fixes #90504
2026-08-20 11:28:26 +05:30
Teknium 2eb5217af9 feat: keyless free-tier failover + Tavily/Firecrawl salvage integration
- Cross-vendor failover: when Exa's or Parallel's keyless free tier
  returns a rate-limit-shaped error, the request retries once on the
  other vendor's free endpoint (search + whole-batch extract). Result
  notes served_by; a peer pinned to its paid tier is never used;
  non-throttle errors never fail over.
- Docs: failover note + Tavily/Firecrawl keyless-when-selected rows.
- Firecrawl keyless test expectations aligned with the keyless tier.
2026-08-19 22:58:05 -07:00
LeonSGP43 f51e61136a fix(web): explicit Firecrawl selection works keyless against the public cloud API
Salvaged from #50659 by @LeonSGP43 onto current main (the client
resolver was rewritten for strict-selection semantics since the PR;
reapplied the keyless mode as a third client_mode inside the new
resolver). An explicit firecrawl selection with no FIRECRAWL_API_KEY /
FIRECRAWL_API_URL now routes through a minimal REST client (v2 search +
scrape, no Authorization header) instead of erroring. Unconfigured
installs never route here — the keyless path requires the explicit
selection. Fixes #49912.
2026-08-19 22:54:28 -07:00
Lakshya Agarwal 6bf4375755 feat(onboarding): enhance Tavily backend support for keyless access
- Added support for keyless Tavily integration in the onboarding flow, allowing it to be recognized as available without an API key.
2026-08-19 22:44:11 -07:00
Lakshya Agarwal ee37f3d897 feat(tavily): update Tavily integration to support keyless access
- Updated the Tavily API key description to clarify that it is optional and keyless access is supported.
- Modified the Tavily plugin and provider to handle requests with or without an API key, using Bearer authentication when the key is provided.
- Enhanced documentation to reflect the new keyless functionality and updated environment variable descriptions.
- Added tests to ensure correct behavior for both keyed and keyless requests.
2026-08-19 22:42:58 -07:00
Gille cefeed4ca8 fix(a2a): expose schemas through tool describe 2026-08-20 10:41:12 +05:30
Brooklyn Nicholson b4f978d983 fix(nous): treat "takes no reasoning parameter" as a definitive no
Both the wire path and the picker only consulted the catalog's
`mandatory` flag, so a route the Portal lists as accepting no reasoning
parameter at all still got sent a disable, and still offered a Thinking
toggle in the model picker.

For a route it serves, the aggregator's own catalog outranks the
models.dev inference: `supports_reasoning: false` now suppresses the
disable on the wire and drops reasoning controls from the picker
entirely, so there is no disable left to describe.
2026-08-19 23:28:14 -05:00
Brooklyn Nicholson d39a031329 fix(nous): stop dropping "thinking off" on Portal models that can honor it
reasoning: {enabled: false} is the only shape the Portal honors, and the
profile refused to send it for every model. Sending nothing means the
upstream default instead, which on a thinking-first route like
deepseek/deepseek-v4-pro (catalog: default_effort high) is thinking ON — so
turning thinking off kept billing reasoning tokens on every turn.

The blanket omission was over-broad. The Portal only rejects a disable on
reasoning-mandatory routes ("Reasoning is mandatory for this model"), which
its catalog flags per model, so that flag now gates the omission. Models the
catalog can't speak to keep the old behavior rather than risk the 400.

extra_body.thinking, DeepSeek's own disable shape, is not forwarded upstream
by the Portal and is not an option here.
2026-08-19 22:14:56 -05:00
Teknium aebab05f9e Merge pull request #90313 from NousResearch/feat/keyless-web-search-fallback
feat: web search works keyless on fresh installs (Parallel + Exa free tiers)
2026-08-19 19:47:12 -07:00
Teknium 1fa66f2577 Merge remote-tracking branch 'origin/main' into feat/keyless-web-search-fallback
# Conflicts:
#	website/docs/user-guide/configuration.md
2026-08-19 19:36:12 -07:00
Axmr1 4511ba49dd fix(image_gen/openai-codex): do not save progressive partial frames as finals
Codex Responses streams can emit partial_image_b64 previews without a final
image_generation_call.result. The provider treated any b64 as success and could
let a partial overwrite a coexisting final in the same payload, delivering
smeared intermediates as finished GPT Image 2 outputs.

Request partial_images=0, prefer final over partial in extraction, fail closed
(with one content-agnostic retry) unless source=final, and surface image_source
plus pixel_size for QA.
2026-08-19 19:35:07 -07:00
Teknium f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00
Jeffrey Quesnelle 612b3633d2 Merge pull request #77915 from bbednarski9/feat/relay-native-plugin-init
feat(relay)!: initialize static/dynamic plugins via native integration, remove opt-in plugin
2026-08-19 22:11:13 -04:00
Teknium 095f003377 Merge remote-tracking branch 'origin/main' into feat/keyless-web-search-fallback
# Conflicts:
#	hermes_cli/tools_config.py
2026-08-19 16:48:40 -07:00
Teknium 7f83d3808c fix(browser): strict cloud-provider selection; camofox becomes a selection
An explicitly stored browser.cloud_provider that names no registered
plugin now raises the honest selection-naming error instead of warning
and silently auto-detecting; the auto-detect walk (including the managed
gateway entitlement probe) runs only when no cloud_provider key was ever
written. The 'nous' selection routes to the Browser Use provider, whose
config resolver is now a strict switch: 'nous' => managed only, stored
vendor => direct BROWSER_USE_API_KEY only with a selection-naming error
when missing. Camofox is selected via browser.cloud_provider: camofox;
CAMOFOX_URL stays the server ADDRESS only and can no longer override an
explicit different selection (never-configured installs keep the legacy
env-var activation).
2026-08-19 16:10:01 -07:00
Teknium d7119ea2a6 fix(web): honor the stored web backend selection; no silent backend swaps
_get_backend returns the stored web.backend verbatim (mapping the managed
'nous' selection to the firecrawl provider) — unknown names surface the
honest selection-naming error at dispatch instead of silently rerouting
through the credential ladder, which now runs only on never-configured
installs. _get_capability_backend no longer discards an explicit
search/extract backend when its availability probe fails. The firecrawl
client resolves strictly: 'nous' => managed gateway only (unavailable =>
selection-naming error), stored vendor => direct only (no FIRECRAWL key
=> error, never a silent managed fallback billed to Nous).
2026-08-19 16:10:01 -07:00
Teknium 2dea073a1c fix(tools): dispatch image/video FAL strictly on the stored hermes tools selection
Add read_selection()/selection_exists()/selection_error() to
tool_backend_helpers: one provider string per category ('nous' = managed
Nous Tool Gateway, vendor name = direct with the user's own credentials,
no key ever written = legacy credential autodetect). Legacy configs are
interpreted at read time only (use_gateway: true => nous); nothing is
migrated on disk, and the DEFAULT_CONFIG-seeded stt.provider: local is
treated as never-configured.

_resolve_managed_fal_gateway / _resolve_managed_fal_video_gateway now
switch on that string: 'nous' routes managed only (unentitled => error
naming the selection), a stored vendor routes direct only (missing
FAL_KEY => error naming FAL_KEY and the selection, no silent managed
reroute), and FAL_KEY presence no longer selects the route. Krea's
model-driven managed interception now requires no stored provider (or
the managed selection) instead of merely provider != krea, and the
image/video registries map the 'nous' selection to the FAL plugin.
2026-08-19 16:10:01 -07:00
Teknium 9ec5750aca fix: widen reasoning-effort wire translation to sibling sites (#89503 class)
The chat_completions chokepoint fix (ultra->max for every model,
cherry-picked from #89509) has siblings with the same bug shape:

- codex.py: ultra->max was gated on gpt-5.6 only; now baseline for all
  Responses-API models (backend-specific branches still override).
- Kimi top-level reasoning_effort: K3 accepts low/high/max only —
  'medium' and upper-ladder levels were dropped to the medium default
  (400s on K3, ladder inversion on K2). Full ladder mapped per family,
  mirroring the kimi-coding plugin's K3 map.
- TokenHub: 'minimal' fell through to the 'high' default (asked least,
  got most); full ladder now mapped onto low/medium/high.
- auxiliary_client Responses path: ultra->max alongside the existing
  minimal->low clamp.
- custom provider plugin: ultra capped at max instead of forwarded
  verbatim to GLM/vLLM/SGLang backends that reject it.
- copilot plugin: ad-hoc downgrade rules replaced with the shared
  clamp_reasoning_effort_to_supported ladder walk so ultra/max resolve
  to the strongest supported level instead of medium (#74295).

Sabotage-verified: new sibling-site tests fail 6/10 without the fixes.
2026-08-19 16:04:22 -07:00
Teknium f08d3e400f feat: hermes tools lets Exa/Parallel users pick the free keyless or paid keyed endpoint
Exa and Parallel now each render as two picker rows in hermes tools —
'Free (keyless)' and 'Paid (API key)'. Selection persists to
web.provider_tier.<name>:
- free: always the anonymous public endpoint, even with a key set
- paid: always the keyed SDK path; missing key errors instead of
  silently downgrading to the free tier (is_keyless_available also
  returns False so the auto-fallback walk can't route there)
- unset: auto (key present -> paid, else keyless)

Mechanism: get_setup_schema() gains a 'variants' list the picker
flattens into sibling rows sharing one web_backend; selection writes
the tier via both _write_provider_config sites; active-row detection
matches the tier (auto mirrors use_keyless). Routing goes through a
single use_keyless() chokepoint shared by search+extract in both
providers.

Live E2E: tier=free with a fake key present searched keyless OK (a
keyed call would have 401'd); tier=paid without key errored naming
PARALLEL_API_KEY; picker rows verified for both vendors x both tiers.
2026-08-19 15:36:20 -07:00
Teknium 2d9dad0bae docs: correct Exa keyless rate-limit characterization
A 12-request sequential burst from the same IP that earlier saw the
free-tier rate-limit error went 12/12 OK — the limit is a transient
burst/load control, not a tight standing per-IP quota. Soften the docs
and setup-schema wording accordingly (opencode users hit Exa keyless
as their default path in practice without throttling).
2026-08-19 15:24:27 -07:00
Teknium 96c2fd3c04 feat: web search/extract now work keyless on fresh installs via Parallel + Exa free tiers
With zero web credentials configured, web_search/web_extract previously
resolved to the nonfunctional firecrawl sentinel and errored. Now the
backend resolution walks a strictly-last keyless tier: Parallel's and
Exa's public anonymous MCP endpoints (the same free tiers opencode ships
as its default search path).

- plugins/web/keyless_mcp.py: minimal JSON-RPC tools/call client for
  mcp.exa.ai + search.parallel.ai (SSE + plain JSON parsing, typed
  errors, per-process random session id, no user identifiers)
- WebSearchProvider.is_keyless_available(): separate weaker tier that
  never leaks into is_available(), so keyed setups are never pre-empted
- Exa/Parallel providers: route to keyless endpoints when their key is
  absent; keyed SDK path unchanged
- registry + _get_backend(): keyless walk (parallel -> exa) strictly
  after every keyed/importable candidate; check_web_api_key() lights
  the tools up on zero-credential installs
- web.keyless_fallback config key (default true) to disable the tier
- docs: web-search.md + configuration.md

E2E-verified against both live endpoints from an isolated HERMES_HOME
(search + extract via the real dispatchers, disable-flag negative path).
2026-08-19 15:15:27 -07:00
Teknium aba96d5251 feat(image-gen): route live-catalog models to the Image API; merge picker catalogs; docs
Follow-ups on top of the salvaged #82631 surface:

- _select_surface: an unknown model id found in the live /images/models
  catalog now ROUTES to the dedicated Image API instead of only logging a
  hint — without this, a model picked from the live picker that postdates
  the curated snapshot would fall onto chat-completions and fail. Curated
  defaults stay pinned to chat (no behaviour change for existing setups);
  offline probes still fall back to chat. _HINTED_MODELS removed.
- list_models (OpenRouter): union of the live GET /images/models catalog
  (43 models today) and the chat-completions image models, deduped,
  defaults first; curated metadata wins for known ids, API names for the
  rest. Nous Portal (no /images route) keeps its chat-only catalog.
  Offline fallback: static chain + curated Image API snapshot.
- Tests updated/added: unknown-id routing (flipped from the hint-only
  pinning test), non-catalog id stays on chat, merged-picker union/dedupe/
  order, Nous exclusion.
- Docs: image-generation.md gains the OpenRouter Image API section and an
  editing-support row.

Live-verified: picker lists 43 models; generation succeeded through the
dedicated API on google/gemini-3.1-flash-lite-image and on the previously
unreachable black-forest-labs/flux.2-klein-4b (config-selected, no kwarg).
2026-08-19 14:44:36 -07:00
AI Staff d6e6e8b602 feat(plugins): add OpenRouter Image API surface to openrouter image_gen backend 2026-08-19 14:44:36 -07:00
Bryan Bednarski 8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski 0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Teknium 9c6ecf2ca7 feat(image-gen): OpenRouter image picker lists every live image-output model; xAI edits honor dispatched model
- plugins/image_gen/openrouter: list_models() now queries the endpoint's
  /models catalog filtered to output_modalities containing "image"
  (per-backend 5-min cache, 10s timeout, static 2-model chain as offline
  fallback; openrouter/auto* router pseudo-models excluded). Every image
  model OpenRouter serves — including future releases — is selectable in
  `hermes tools` with no code change. Applies to Nous Portal too via the
  shared provider class.
- plugins/image_gen/xai: forward the dispatched model kwarg into
  _resolve_edit_model() so an explicitly selected edit-capable model is
  honored on /images/edits (extends the salvaged #55893 fix to the edit
  path; text-only models still fall back to quality).
- Tests: OpenRouter live-catalog filtering/exclusions/order, offline
  fallback, cache single-fetch; xAI edit-kwarg forwarding incl. the
  text-only-hijack negative case.

Live-verified against openrouter.ai: 9 image-output models returned and
rendered, matching the public models?output_modalities=image listing.
2026-08-19 02:04:29 -07:00
srojk34 008d469991 fix(xai): forward image_gen.model kwarg to _resolve_model in generate() 2026-08-19 02:04:29 -07:00
69k4xmdfm2-blip 21260c3281 fix(gateway): carry the profile in adapter-derived session keys (#88404)
Adapter ingress derives a session key BEFORE the runner stamps
source.profile in _make_profile_message_handler, so the namespace fell
back to the active profile and every bot in a multiplexed gateway
produced agent:main:<platform>:<chat>. A Telegram private chat reports
the user's own id as chat.id, identical for every bot, so two profiles
sharing one human collapsed onto a single lane: _pending_text_batches,
_active_sessions, the busy-session guard and _post_delivery_callbacks are
all keyed on that string. A day of production logs across two bots shows
60 flushes, none carrying the secondary profile's namespace.

set_owner_profile records credential ownership on the adapter and
_session_key_profile resolves the namespace as source.profile ->
_owner_profile -> the session store's resolver, so a secondary adapter
keys into its own namespace even before the source is stamped. Stamped
sources keep priority, so relay/connector ingress, which routes per event
rather than per credential, is unchanged. _configure_profile_adapter
installs the owner alongside the other handlers, covering startup and
reconnect.

Every candidate is type-checked as a non-blank str, and every attribute
read goes through getattr: adapters are routinely built without
BasePlatformAdapter.__init__, and a duck-typed session store returns a
truthy non-string that would otherwise be interpolated into the key as
agent:<MagicMock ...>:.

Also routes the four call sites that passed no profile at all (feishu
media batches, raft, slack _session_key_for_source, telegram photo
batches) through the same resolver.

test_multiplex_busy_input_mode's secondary-adapter busy case seeded
_active_sessions with the unstamped agent:main: key, asserting the
pre-fix collapse. It now seeds the lane the profile-owned adapter
actually derives.

A primary adapter has no owner and an unstamped source, so it resolves
exactly as before; with multiplex_profiles off the resolver returns None
and every key is byte-identical to today's.
2026-08-19 02:00:16 -07:00
Teknium ac0a8cd281 feat(image-gen): xAI Grok image catalog goes live-driven; grok-imagine-image-2.0 selectable
- plugins/image_gen/xai: merge the live /v1/image-generation-models catalog
  (5-min cache, 10s timeout, static-table fallback when offline/unauth)
  into the picker so new xAI Imagine models appear automatically the day
  they launch, with generic metadata until curated text is added.
- Add grok-imagine-image-2.0 to the curated static table (typography/
  layout-aware model, API-available since Aug 8 2026).
- Edits honor an explicitly selected image-input-capable model
  (e.g. grok-imagine-image-2.0) instead of always forcing
  grok-imagine-image-quality; quality remains the default edit baseline.
- Tests: hermetic autouse fixture keeps unit runs offline; new coverage
  for live-merge, unknown-future-model selection, offline fallback, and
  edit-model resolution. Docs model table updated (en + zh-Hans).

Live-verified: /image-generation-models returns grok-imagine-image,
grok-imagine-image-2.0, grok-imagine-image-quality; real generation with
2.0 succeeded end to end.
2026-08-19 01:19:37 -07:00
zhuermu c6b680c445 fix(image-gen): handle stale OpenRouter model defaults 2026-08-19 01:19:37 -07:00
Teknium c7d0f6c35f fix(providers): honor a custom base_url over models_url in fetch_models
Follow-up to the salvaged CommandCode signature fix: accepting base_url
but ignoring it left custom endpoints (user-configured model.base_url /
COMMANDCODE_BASE_URL proxies) fetching the public catalog instead of the
configured one. Reviewer dansigma flagged this on PR #88851.

Class-wide fix, not a CommandCode patch:

- providers/base.py: a caller base_url that DIFFERS from the profile's
  default now wins over models_url. Equality with the default means "not
  customised" (callers pass base_url unconditionally, defaulting to the
  profile's own URL) and keeps models_url as the endpoint, preserving the
  OpenRouter-style split-catalog behavior.
- commandcode: _fetch_commandcode_models() takes the endpoint override;
  both profile overrides forward base_url.
- Tests: base-class precedence (custom beats models_url, default does
  not), CommandCode redirect via live local HTTP server incl. claude-*
  filter, and default-echo hitting the canonical endpoint. All verified
  to fail against the pre-fix implementation (sabotage run).
2026-08-18 14:27:36 -07:00
greyvito f5ea3fa9cb fix(commandcode): accept base_url kwarg in fetch_models overrides
The model picker's generic live-fetch path (hermes_cli/models.py
provider_model_ids) calls profile.fetch_models(api_key=..., base_url=...).
Both CommandCode overrides only accepted api_key/timeout, so every picker
open raised TypeError, which was silently swallowed, leaving the provider
with zero models.

Match the base ProviderProfile.fetch_models signature (base_url kwarg) and
add a regression test asserting both profiles accept it.
2026-08-18 14:27:36 -07:00
Teknium d03fe2adf4 fix: detect base64-encoded transcript markers in Graph ids
The getAllTranscripts resourceData.id from the field report is a base64url
blob whose DECODED payload ends in "-TranscriptV2" while the encoded form
contains no readable marker, so the substring heuristic in
looks_like_transcript_id missed it. Add a best-effort base64 decode hint so
degraded notifications (no @odata.id) are still refused with the clear
guidance error instead of a cryptic Graph 400.
2026-08-18 12:37:40 -07:00
kyssta-exe db004d1801 fix(teams-pipeline): support organizer-scoped meeting lookup (#83422) 2026-08-18 12:37:40 -07:00
xxxigm 182a96645a fix(teams-pipeline): resolve Graph transcript notifications to the meeting id
getAllTranscripts webhooks put the callTranscript id in resourceData.id. Using that as an onlineMeeting id makes Graph v1.0 return 400 Unexpected id format. Parse meeting and organizer ids from @odata.id and use the organizer-scoped users path.
2026-08-18 12:37:40 -07:00
Jeffrey Quesnelle 75bbc055a7 Merge pull request #85579 from bbednarski9/codex/fix-relay-canonical-operation-names
fix(relay): use canonical managed operation names
2026-08-17 21:29:18 -04:00
Teknium aa8ceed4b6 fix(memory): keep a stale holder's late close() from evicting a fresh registry entry
Follow-up to the #88347 salvage: after release_all_under() force-closes a
profile's shared connection, a store re-created on the same path registers
a fresh entry under the same key. A stale holder that later calls close()
would pop that fresh entry (its refs were transferred nowhere), letting a
third store open a second connection to the same database — exactly the
multi-writer contention the shared registry exists to prevent. close()
now evicts the registry entry only when it is still its own.
2026-08-17 16:33:52 -07:00
liuhao1024 4f354c27b7 fix(profiles): release memory-store handles before rmtree on profile delete
The desktop's main serve process opens memory_store.db for every known
profile and nothing closed those connections before delete_profile's
rmtree — on Windows the open SQLite handles make the removal fail with
WinError 32 for both the CLI and the DELETE /api/profiles/<name> route
(#88347). POSIX unlinking of open files hid the same leak.

MemoryStore.close() is refcount-driven, so a live holder keeps the
handle forever; add MemoryStore.release_all_under(directory) to
force-close every shared connection under a directory, and call it in
delete_profile after stopping the profile backends. Inside serve the
handles live in that very process and get released; from the CLI it is
a no-op.

Fixes #88347
2026-08-17 16:33:52 -07:00
Brooklyn Nicholson 20ec564684 docs: align the ::preview description and storage example with shipped behavior
The SDK doc still described the v1 frame (fixed height attribute, rail card
under the frame) — chrome that no longer exists. And plugin_storage's usage
example used `with plugin_db(...)`, which reads as auto-close but sqlite3's
context manager only scopes transactions; the example now closes explicitly.
2026-08-17 15:12:53 -05:00
Brooklyn Nicholson 8f2ddc9676 feat(plugins): per-plugin durable data directory that survives plugin update and removal
Plugins that persist state have been writing into their own install tree
(<hermes home>/plugins/<name>/), which `hermes plugins update` git-pulls and
`hermes plugins remove` deletes — user data dies with the code that wrote it.

plugins/plugin_storage.py is the sanctioned home: plugin_data_dir(name) gives
one data root per plugin under <hermes home>/plugin-data/<name>/ (profile-
aware, created on first use, names validated against traversal), and
plugin_db(name) opens a WAL-mode SQLite database inside it. Secrets stay on
the existing secret-scope path — this is state, not credentials.

hermes-achievements, the in-tree offender, converts with a legacy-file
migration on first read.
2026-08-17 15:12:53 -05:00