661 Commits

Author SHA1 Message Date
teknium1 019103bfbc docs(mem0): comment and #99121 test state the one fail-loud scope contract
The retained comment still claimed a scope-less multiplex caller could load an
OSS config; identity reads above it raise UnscopedSecretError first. Both the
comment and the regression test now say the same thing: callers are scoped, an
OSS profile whose scope lacks MEM0_API_KEY initializes, and a scope-less caller
is a spawn-site bug that raises.
2026-09-13 05:19:48 -07:00
teknium1 43ec2036ec fix(hindsight): corrupt config.json falls through to the legacy file and env, as before
read_json_or_empty returns {} for malformed JSON; returning that directly from
_load_config made a corrupt profile config.json yield an empty (silently
unconfigured) mapping where the pre-dedup loop kept walking to the legacy
file and then the env branch. Only a non-empty parsed object is now returned.
2026-09-13 05:19:48 -07:00
teknium1 4e74beb064 fix(langfuse): atexit finalizer flushes settled clients without reading credentials
_finalize_all_traces called _get_langfuse(), which lazily BUILDS a client when the
launch-profile slot is empty. Under a multiplex gateway every turn runs scoped, so
only the per-home slots are populated; at exit (no scope) the credential read
raised UnscopedSecretError and the flush loop was skipped for every profile,
losing all pending traces. The finalizer now iterates the already-settled
clients (launch slot + per-home map) and never initializes. Test drives two
profiles under real home-override + secret scopes, then finalizes unscoped.
2026-09-13 05:19:48 -07:00
teknium1 d364473620 refactor(model-providers): fetch_models stubs drop to the ABC; commandcode and AI Gateway catalog GETs go through open_credentialed_url
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
2026-09-13 05:19:48 -07:00
teknium1 f678ed8299 refactor(model-providers): thinking-toggle XOR effort translation lives in agent.reasoning_effort; opencode-free imports it instead of borrowing via sys.modules
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
2026-09-13 05:19:48 -07:00
teknium1 e7d9e07dbc refactor(plugins): HERMES_HOME resolves through hermes_constants.get_hermes_home everywhere; no ~/.hermes fallbacks
Four bundled plugins wrapped get_hermes_home() in a try/except that fell back
to ~/.hermes on ImportError (a2a/protocol._hermes_home, photon/auth
._auth_json_path, google_chat adapter inline, openviking done in the previous
commit). A bundled plugin cannot lose hermes_constants -- each already imports
gateway.* / agent.* from the same tree -- so the fallback was dead code that
was also wrong on Windows (%LOCALAPPDATA%/hermes) and under a profile override.
telegram's gmail-triage verb path used Path.home()/".hermes" outright, ignoring
profiles. mem0/_oss_providers baked os.path.expanduser("~/.hermes/mem0_qdrant")
into VECTOR_PROVIDERS at import time, so the Qdrant default landed in the
user's ~/.hermes for every profile; the default is now a lazy callable resolved
by vector_default_config(provider_id) at setup time. a2a/security.py only
changes its import (it borrowed protocol._hermes_home).

Behavior change: on Windows and under profiles these paths now follow the
active HERMES_HOME (they were previously anchored to ~/.hermes in the impossible
fallback / at import time); the default-profile POSIX layout is unchanged.

Test: tests/plugins/test_plugin_paths_follow_profile.py asserts each resolver
(a2a conversations, photon auth.json, mem0 qdrant default, openviking log)
lands inside a HERMES_HOME ContextVar override; sabotage red for the mem0
import-time constant and the photon ~/.hermes fallback.
2026-09-13 05:19:48 -07:00
teknium1 ca2ca5b9e1 fix(secret-scope): mem0, langfuse and azure credential readers stop swallowing UnscopedSecretError
Three shims caught UnscopedSecretError and degraded to "" (mem0._scoped_env)
or to os.environ (langfuse._secret, azure_identity_adapter._scoped_env). Under
multiplex os.environ holds the DEFAULT profile's .env, so the langfuse/azure
fallback could ship another profile's keys, and the mem0 fallback silently
routed a mis-spawned turn's memories into the default profile's account. The
exception exists to surface exactly that spawn-site bug (agent/AGENTS.md:
never add environ fallthrough, never swallow it). All three now call
agent.secret_scope.get_secret directly: with a scope installed a miss returns
the default; single-profile deployments (multiplex off) still read the process
env inside get_secret; a scope-less multiplex caller raises.

Behavior change: a mis-spawned child under gateway.multiplex_profiles now
fails loud with UnscopedSecretError instead of running silently unauthenticated
/ on the default profile's identity. The #99121 contract (OSS mode needs no
MEM0_API_KEY in scope) is unchanged and its test now installs an empty profile
scope, which is the situation the issue described; a genuinely scope-less
caller is asserted to raise in a new test.

Tests: tests/plugins/test_scoped_secret_readers_fail_closed.py (scope wins over
environ; scope-less multiplex raises) for langfuse + azure, sabotage red when
the langfuse fallthrough is restored; tests/plugins/memory/test_mem0_v3.py::
test_load_config_fails_closed_without_scope_even_for_identity_settings,
sabotage red with a swallowing wrapper reinstated.
2026-09-13 05:19:48 -07:00
teknium1 9ba5850e95 refactor(memory-plugins): background threads inherit the profile context; JSON sidecar reads and the holographic config.yaml write use core primitives
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.

Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.

holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.

openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.

Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
2026-09-13 05:19:48 -07:00
teknium1 9401cc1643 test(redact): a2a superset test asserts every credential class; drop dead mask branch
The a2a invariant test skipped any synthesized token the canonical
redactor itself let through, so a boundary/shape regression would have
passed silently. Every class scrubs today, so the escape hatch goes and
the soft ">= 50" count becomes the exact registry size.

The gateway body test still asserted a "[REDACTED]" fallback marker that
no longer exists; redact_for_egress masks via _mask_token ("***"), so
assert that alone. honcho oauth.py's `re` import became unused when its
private pattern list moved to the registry.
2026-09-13 05:07:50 -07:00
teknium1 226df89f74 refactor(redact): one secret-pattern source; a2a, gateway chat and monitoring egress scrub through redact_for_egress
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.

Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.

Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).

Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
2026-09-13 05:07:50 -07:00
Teknium 63584da036 feat(video): add MiniMax H3 Max Turbo family (fal post-train, 480P-1080P)
fal launched H3 Max Turbo on Sep 3 (minimax/h3-max-turbo/{text,image}-to-video):
a throughput-tuned post-train of H3 Max with a 1080P tier Max lacks, at
$0.025/s 480p / $0.04/s 768p / $0.08/s 1080p list ($0.00625-0.02/s promo until
Sep 14). Schema matches Max's shape — required prompt_expansion_mode static
key, int duration 5-15, seed on both endpoints, i2v drops aspect_ratio — plus
the new 1080P resolution enum, so it reuses the existing family capability
flags with a Turbo-specific resolution alias map.

Schema verified against the FAL queue OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403 "Exhausted balance"); portal
allowlist/pricing needed for managed users on the 2 new endpoints.
2026-09-12 21:58:25 -07:00
teknium1 baf1200e0a refactor(commandcode): drop the deepseek ImportError guard, trim tests to two invariants
The deepseek plugin is bundled and always loads, so the try/except around
the delegation import was defense-in-depth for a path that cannot fail.
Tests reduced to the two contracts that matter: /reasoning none reaches
the wire as thinking.disabled, and DeepSeek ids produce exactly the native
DeepSeek profile's output while non-DeepSeek families stay a no-op.
2026-09-12 20:41:32 -07:00
liuhao1024 fbb3a244b7 fix(commandcode): degrade to no-op when the deepseek plugin shim is missing
The bundled-plugin loader pops half-registered modules when a plugin
fails to load, so the lazy 'from plugins.model_providers.deepseek import
deepseek' could raise ImportError on every DeepSeek-routed CommandCode
turn — turning the soft 'thinking uncontrollable' bug into a hard
crash. Catch ImportError, log, and return the pre-fix no-op (review
feedback on #95241).
2026-09-12 20:41:32 -07:00
liuhao1024 6f88fb030a fix(commandcode): forward DeepSeek reasoning controls through the wire
CommandCode fronts DeepSeek with vendor-prefixed ids
(deepseek/deepseek-v4-flash). DeepSeek V4+ defaults to thinking mode
when the thinking field is omitted, so /reasoning none changed the
Hermes session state but not the actual request -- the turn sat in
reflecting.../brainstorming... for minutes (#95232). Strip the vendor
prefix for DeepSeek-family ids and delegate to the native DeepSeek
profile's build_api_kwargs_extras (extra_body.thinking +
reasoning_effort mapping); other CommandCode model families keep the
base no-op behavior. The prior no-op tests codified the bug and are
rewritten to pin the new contract.
2026-09-12 20:41:32 -07:00
teknium1 c6f87deb2c feat(video): OpenRouter backend covers every model on the live video catalog
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):

- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
  no duration are outside the unified video_generate surface and are dropped)
  with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
  advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
  live limits (nearest by value/height/ratio) and drops generate_audio/seed
  for models that lack them (the API 400s otherwise); reference images ride
  in input_references; local file inputs are refused (OpenRouter fetches
  URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
  never to a provider-supplied unsigned_urls host (kept from #103267)

Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.

Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
2026-09-12 13:44:52 -07:00
cedanoagent 387ac50d85 feat(video): add OpenRouter Hailuo 3 Max provider 2026-09-12 13:44:52 -07:00
Ada 7c734c7838 fix(opencode-go): send reasoning for the canonical deepseek-flash id
The version-less canonical Flash id is not matched by _is_deepseek_thinking_model, so agent.reasoning_effort and every auxiliary reasoning_effort were silently dropped on the OpenCode Go relay while the direct provider was fixed (8435a3ae00/aeecb110f8). Match it through a version-less id set, mirroring plugins/model-providers/deepseek. Verified against the Go relay: named levels are graded (low 906 / high 1371 / max >=2500 reasoning tokens on a multi-step prompt) and integer efforts are rejected, so the named level is the knob to send.
2026-09-12 08:06:48 -07:00
Teknium 39ab25a6ac test(mem0): drop the inside-the-cap no-op case (salvage bar: 2 invariant tests) 2026-09-12 05:09:43 -07:00
John Paul Soliva aca4dc86a3 fix(memory/mem0): trim a synced message at the last sentence boundary, not the first separator kind
_truncate_for_sync documents "the last sentence boundary within max_len", but it
looped over separator KINDS and returned on the first kind that qualified. An early
"。" therefore outranked a "." 240 characters later, and in pure ASCII "." outranked
a later "!" or "?" purely because it comes first in the tuple.

With the 450-char default, "a"*200 + "。" + "b"*240 + "." + "c"*100 kept 201 of the
442 characters available: 241 characters the embedder would have accepted were
discarded, so any fact in the second half of the turn never reached extraction. The
add() call succeeds, so unlike #106235 nothing is logged — the turn is simply
remembered from its first sentence. Raising sync_max_chars widens the gap rather
than closing it.

Take the max over every separator instead, from a named tuple so the set is not
buried in the loop. ".\n" is dropped: its index can never exceed the bare "." it
starts with, so under a max it is unreachable. The first-third guard and the hard-cut
fallback for unsegmented input are unchanged.

Fixes #108868
2026-09-12 05:09:43 -07:00
Teknium 208bd0b65a fix(multiplex): per-profile plugin caches and the Yuanbao active adapter
Ramp Router efforts cache + warm/disk flags, xAI and OpenRouter image catalogs,
Hindsight append-capability verdict, memory-provider skill registry, OpenViking
atexit provider, Honcho loopback flow status, Langfuse client (os.environ-only
credentials) and disk-cleanup's protected cron paths held one profile's
credential- or home-derived value process-wide; YuanbaoAdapter._active_instance
was last-connected-wins across profiles.

Keyed by home key / credential fingerprint under an override, credentials read
through the secret scope, warm threads run under copy_context(); unscoped module
slots stay for the single-profile path and the existing monkeypatch tests.
2026-09-12 01:35:05 -07:00
Teknium 5e04961072 test(hindsight): trim salvaged #103957 tests to the three invariants
Keep: scopeless worker resolves the on-disk key (durability core), a keyed
build still rewrites (rotation), and the API-prefixed vault name resolves.
Drop the mock-patched gate predicate test, the inlined copy of the same
predicate, and the empty-build/empty-disk arm — they re-assert the helper's
shape rather than a behaviour contract.
2026-09-11 15:26:46 -07:00
Teknium a9838c2100 fix(multiplex): tool and memory-provider env reads stay inside the routed profile
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.

Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.

Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.

Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.

Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.

Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.

Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.

session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.

Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.

Fixes #82903
Fixes #65941
Fixes #99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
2026-09-11 15:26:46 -07:00
이민재 a1b689f12a fix(mem0): skip platform key lookup in OSS mode
Conflict resolution on top of main also routes MEM0_MODE / MEM0_HOST / MEM0_AGENT_ID /
MEM0_USER_ID through get_secret: identity and host are .env values like the key, and a
raw environ read hands a secondary profile the default profile's mem0 account.

(cherry picked from commit 7a5863ebafdb193d8ec2f89a7aca309774d6d69a)
2026-09-11 15:26:46 -07:00
red4711 778ad61017 fix(hindsight): accept HINDSIGHT_API_LLM_API_KEY vault name in key resolution
The vault item is named HINDSIGHT_API_LLM_API_KEY (matching the daemon env var), while the setup wizard uses HINDSIGHT_LLM_API_KEY. Vault-fed scopes carrying only the API-prefixed name resolved empty, so the scopeless path — and the whole fail-closed machinery — never engaged on real deployments. Accept both names.

Proven end-to-end: hydrated scope now resolves the 73-char key and the rewrite gate passes.

(cherry picked from commit 2ce8c2980188ae8e4a3ed398e14fa8ed89a40644)
2026-09-11 15:26:46 -07:00
red4711 d0d5c7745a test(hindsight): run pure gate tests on Windows too
test_gate_refuses_keyless_build_against_keyed_disk and test_rewrite_allowed_when_build_carries_key touch no filesystem modes, so the nt skip guard was needlessly excluding them. Addresses review feedback on PR #103957.

(cherry picked from commit 196cf723571c87829be61e17df27869948db00a0)
2026-09-11 15:26:46 -07:00
red4711 2c24ed4b6e fix(hindsight): keep scopeless daemon worker from clobbering profile env key
The embedded daemon-start worker usually runs with no secret scope, so _embedded_llm_api_key resolved empty and both the profile-env compare and the upstream manager merge overwrote the file's good key with emptiness.

- _embedded_llm_api_key: explicit config, then secret scope, then on-disk profile env fallback; UnscopedSecretError treated as empty without touching os.environ (multiplex safety).
- Add _may_rewrite_profile_env gate: refuse rewrite on keyless-build plus keyed-disk, and skip the daemon stop in that case.
- Document key resolution order in plugins/memory/hindsight/README.md.
- Add 5 tests pinning disk fallback, the gate's refuse branch, rotation, keyless providers, and file survival.

(cherry picked from commit 48a4bcd536fa799b666a8d5eaae73c1816ca9e33)
2026-09-11 15:26:46 -07:00
Erosika f7d5ac3230 feat(honcho): write bot dms into their own a2a session
A DM relayed from another Hermes profile ran as a turn in the recipient's
Bot Chat session. sync_turn wrote the bot's words and the recipient's reply
into that session, and before per-author writes they landed under the
human's peer. The human's representation absorbed conversations the human
never had.

The turn context now marks such turns with scope a2a:<bot id>.
sync_turn routes a bot-authored turn into a separate Honcho session keyed
<session>:a2a:<sanitized bot id>, created with the sender bot as its user
peer, and never writes it into the human's session. The key is deterministic
so every turn from the same bot reaches the same session, and it stays
inside Honcho's 100 character session id limit. Recall still reads the
human's session only.

a2aSessions (host block, then root, default true) turns the routing on.
With it off, bot-authored turns are skipped. A bot turn that names no
author id is skipped as well, because nothing can key its session. Human
turns are unchanged.

get_or_create takes a user_peer_id override so the a2a session's roster is
the bot and the assistant, not the runtime human.
2026-09-10 10:45:57 -07:00
Teknium aeecb110f8 fix(deepseek): deepseek-flash is the canonical Flash id; retired names fold onto it
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.

Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.

Builds on YipTszkwan's #107126 (earliest fix in the cluster).
2026-09-10 02:44:26 -07:00
YipTszkwan 8435a3ae00 fix(deepseek): recognise the version-less deepseek-flash model id
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:

* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
  omitted `extra_body.thinking`. The server then defaults to thinking-on, so
  the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
  V-series regex), so the id a user picked never reached the wire and the
  config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
  instead of the real 1M window, capping the model at an eighth of its
  context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
  detector at its 180s default instead of the 600s reasoning-model floor.

Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".

Adds the id to all four gates plus regression coverage for each site.
2026-09-10 02:44:26 -07:00
teknium1 acf9177c70 test(fal): trim the billing-409 salvage to two invariant tests
Keep the two tests that fail on main without the fix:
- test_fal_common: a keyed managed submit makes exactly one POST (plus the
  negative arm: an unkeyed submit still goes through the SDK retry ladder)
- test_image_generation: the 409 BILLING_ERROR body surfaces
  `unsupported_pricing_meter` instead of the generic "not yet enabled" text

Dropped from #106484: the duplicate video-plugin billing test (same helper,
same assertion), the `_fal_client = fake` / `import_fal_client` stub churn
and the `tools.lazy_deps` stub — fal-client is installed in CI (`--extra fal`)
so those fixtures were not needed; the `_load_fal_client` no-op fixture on
TestManagedGatewayErrorTranslation for the same reason.

Also drop the redundant `retry_request is None` re-check in
`_ManagedFalSyncClient.submit` — `__init__` already raises when the helper
is missing.
2026-09-09 11:46:36 -07:00
Matt Earls 0c6b94e499 fix(image-gen): preserve managed FAL billing errors
Avoid retrying idempotent managed FAL submissions because the retry can mask the initial billing failure. Surface structured Nous billing diagnostics consistently for image and video paths, with hermetic regression coverage.

(cherry picked from commit 289ce039e9a522dc8016ae4a512214c05d0a8bc0)
2026-09-09 11:46:36 -07:00
teknium1 322905e91b feat(memory): make the mem0 sync char cap configurable via mem0.json
A flat 450-char cap fits 512-token embedders (bge-small-zh-v1.5,
all-minilm) but stores only ~5% of the window on 8192-token models
(text-embedding-3-small, jina-embeddings-v3, bge-m3), degrading memory
quality for users those models served fine before truncation existed.

Read `sync_max_chars` from mem0.json once in initialize() (450 default)
and pass it to _truncate_for_sync(). Config over auto-detection: the
Ollama /api/show probe + known-model table proposed in #37427 adds a
network call and a curated list for a number the operator already knows
from their embedder choice; the setup wizard's mem0.json is the plugin's
behavioral-settings surface (no new HERMES_* env var). Documented in the
plugin README and the memory-providers docs page.

Dynamic-cap requirement and measurements (450 OK / 600 -> HTTP 500 on
bge-small-zh-v1.5:f16) by @szicely in #106235.

Refs #37421 #106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-09 10:55:02 -07:00
liuhao1024 040742d06b fix(memory): truncate oversized mem0 sync messages at the source
sync_turn() sent the whole turn to backend.add() untruncated. OSS
embedding models with small context windows (Ollama bge-small-zh-v1.5:
512 tokens) reject the request with HTTP 500, and hosted APIs answer
INPUT_TOKEN_LIMIT_EXCEEDED — in both cases _try() only logs, silently
dropping the turn's memory extraction after long conversations.

Cap each synced message at its last sentence boundary within 450 chars
before ingestion: short turns pass through unchanged, long turns keep a
coherent statement for fact extraction. This replaces the previous
retry-on-error approach, which could not match Ollama's HTTP 500 shape.

Salvaged from #37427 with the test suite trimmed to the one invariant
(oversized turn still reaches a small-context backend, short text
untouched, no breaker failure).

Refs #37421 #106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
2026-09-09 10:55:02 -07:00
Gianpietro Dal Zio 653418d842 fix(dashboard): coalesce expensive reads before worker admission 2026-09-09 21:16:40 +05:30
Teknium e1838c5b5a fix: trim opencode-go 422 salvage to the invariant set
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
2026-09-09 03:52:47 -07:00
ericmaddox bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
Teknium 7777f8c350 feat: add GPT Image 2.5 generation and editing to OpenAI provider 2026-09-08 14:57:15 -07:00
kshitijk4poor ab2f4602de refactor: MessageEvent to gateway/platforms/event.py; ElicitationHandler takes a call_context thunk
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.

gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).

tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.

ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)

Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
2026-09-07 22:47:33 +05:30
Teknium 567d53db23 test: use an isolated real ledger for Chronos claim rearming 2026-09-07 05:57:26 -07:00
Teknium 5bd439d3ed refactor(plugins): own dual-kind hook fallback in the ledger mixin
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
2026-09-06 13:36:12 -07:00
Joey 684a2cfbd7 fix(plugins): give dual-kind memory hooks a single owner 2026-09-06 13:36:12 -07:00
Teknium c8cbc07030 test(mem0): setup module keeps post_setup when discovery imports the package first
The plugin loader execs sibling modules before the package __init__, so a
module-level `from . import _read_mem0_json` in _setup.py failed against the
empty parent shell and the whole module silently dropped out — every later
`hermes memory setup mem0` died with "cannot import name 'post_setup'".
The test loads mem0 through the real discovery path from a cold sys.modules
and asserts the cached _setup module exposes post_setup (red on base).

Also maps the contributor email for #103078 credit.

Campaign tracker: https://github.com/NousResearch/hermes-agent/issues/104154
2026-09-06 05:33:25 -07:00
Teknium f1ccf436a2 feat(web): Perplexity Search API as a web_search + web_extract backend
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:

- search: POST https://api.perplexity.ai/search (documented Search API),
  search_context_size=low so `snippet` stays description-sized;
  results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
  route behind `pplx content snippets` (the CLI's `content fetch` is
  deprecated upstream). web_extract has no query, so the URLs' path words
  serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
  backend set + credential ladder + availability probe (web_tools),
  registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
  dump key lists, nous_subscription direct-credential detection, setup
  summary, test conftests, docs.

Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
2026-09-04 07:17:00 -07:00
Teknium da7ee6353e simplify(compat): tools/browser_tool tests — repoint 56 test files from tools.browser_tool.<name> to the defining browser_tool_* sibling (patch where the name is looked up) 2026-09-03 14:18:25 -07:00
Teknium 7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium e3ab65fe80 simplify(compat): kanban_db — drop 73 re-exports/aliases, repoint 794 callers 2026-09-03 13:48:14 -07:00
Teknium fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium d179f28307 simplify(compat): anthropic_adapter — drop 30 re-exports + 1 alias, repoint 22 caller files (32 sites), 38 test files (~125 sites) 2026-09-03 13:16:47 -07:00
Teknium 67ccfaed39 simplify(compat): image_generation_tool — drop 2 re-export noqa blocks (13 names), repoint 2 callers + 2 tests to image_generation_catalog 2026-09-03 13:10:38 -07:00
Teknium b610e603db simplify(compat): plugins/platforms+web — drop 7 re-exports + 3 aliases, repoint 1 caller + 11 tests, re-remove credential_summary
dingtalk: drop DINGTALK_TYPE_MAPPING/EXT_MAP re-exports. google_chat: card_spec_to_cards_v2 test -> .cards.
matrix: drop module-level MAX_MESSAGE_LENGTH alias (no importers). teams: drop TeamsSummaryWriter re-export
(teams_pipeline/runtime + tests -> summary_writer). wecom: drop WeComStreamExpiredError/STREAM_EXPIRED_ERRCODE/
MAX_INTERMEDIATE_FRAMES re-exports (tests -> .streaming). parallel: drop _get_parallel_client/_get_async_parallel_client
aliases (tests -> _get_sync_client). email: drop stale 'alias' comment (_esecret_int is the only name).
photon: re-remove credential_summary() (shim-only, cb9b7c36f3); its no-leak test now drives print_credential_summary.
2026-09-03 13:04:17 -07:00