Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
Fourteen platform plugins hand-rolled the "already configured? Reconfigure? [y/N]"
gate at the top of interactive_setup (env check + info line + prompt_yes_no(..., False)),
with drifting wording ("X: already configured" vs "X is already configured." vs
"already enabled") and, for LINE and SimpleX, raw input() loops with their own
EOF/KeyboardInterrupt handling and no gate at all. Fixes to the gate (non-interactive
handling, wording, default) therefore reached only the core Telegram/BlueBubbles/webhook
wizards.
- hermes_cli/setup_platforms.py: `_declines_reconfigure` becomes the public
`declines_reconfigure(label, question, *env_vars)` (any-of env check, so Matrix's
token-or-password gate fits); `_save_prompted` becomes `save_prompted` alongside it.
No alias kept; the three core callers are updated.
- buzz, dingtalk, discord, feishu, google_chat, irc, matrix, mattermost, raft, slack,
teams, wecom: the hand-rolled gate is replaced by one `declines_reconfigure(...)` call;
post-decline extras (Discord allowlist nudge, Slack manifest refresh, Raft "Keeping"
line) stay local and unchanged.
- line, simplex: the raw input() loops move onto hermes_cli.cli_output.prompt (masked
for secrets, "" on Ctrl-C/EOF) and gain the shared gate on their primary env var.
Behavior change: the gate's info line is now uniformly "<Label>: already configured"
(DingTalk/Feishu/WeCom lose the trailing period + inline ID; Buzz/IRC/Google Chat/Raft/
Teams no longer echo the current value in that line). Feishu and WeCom now gate on the
app/bot ID alone instead of ID AND secret. LINE and SimpleX gain a "Reconfigure?" [y/N]
prompt when already configured; their prompts now honour HERMES_NONINTERACTIVE and print
via the CLI helpers instead of bare print(). Prompt defaults (No) are unchanged everywhere.
Not touched: WhatsApp's gate keys on WHATSAPP_ENABLED being truthy (a "false" value must
not count as configured), which the shared any-set gate cannot express — left hand-rolled.
Test: tests/plugins/platforms/test_interactive_setup_reconfigure_gate.py parametrized over
the 14 wizards — with the primary env var set and the user declining, each wizard must have
called declines_reconfigure with that var and returned without prompting or saving.
Sabotage: reverting mattermost's gate fails that row.
Six adapters defined their own `_cancel_task` and nine more inlined the same
cancel + suppress(CancelledError) + await block; five kept a hand-rolled TTL-dict
`_is_duplicate` next to the existing `helpers.MessageDeduplicator`; three carried a
`_bounded_put`. Each copy fixed the same bugs on its own schedule (self-cancel deadlock,
done-task re-await, eviction under load).
- `helpers.cancel_task`: None/done no-op, never awaits the current task, swallows the
task's own exception at teardown. Replaces qqbot/signal/yuanbao/buzz/photon/simplex
definitions and the inline copies in weixin, discord, email, irc, line, mattermost,
whatsapp and telegram.
- `helpers.MessageDeduplicator` replaces `_is_duplicate` in qqbot, ntfy, photon,
wecom_callback and LINE's `_MessageDeduplicator`; every site keeps its own
max_size/TTL (qqbot and ntfy 1000/300s, photon 4000/48h, wecom_callback 2000/300s,
LINE 1000/no TTL).
- `helpers.bounded_put` replaces photon/wecom/whatsapp_cloud copies; a re-put now
refreshes the key to the newest slot at every site.
- telegram gmail-triage scripts resolve under `get_hermes_home()` instead of a hard
`~/.hermes`, so profiles with HERMES_HOME set find them.
Not changed: `get_chat_info` stays `@abstractmethod` because
tests/gateway/test_relay_capability_surface.py locks the abstract set to exactly
{connect, disconnect, send, get_chat_info} as a cross-repo contract, so the ~17 no-op
overrides remain.
Behavior change: whatsapp_cloud `_bounded_put` was a pure FIFO (no refresh on re-put);
it now refreshes like the other two sites. Task cancellation at the migrated sites
swallows a task's terminal exception where a few copies previously only suppressed
CancelledError (all are shutdown/disconnect paths).
Photon forked BasePlatformAdapter._send_with_retry before three fixes landed there: the server's
retry_after is honoured over exponential backoff, a long server penalty (> 60s) returns a typed
failure instead of sleeping inline (#91969), and an exhausted rate-limited send no longer posts the
failure notice inside the flood penalty. The fork got none of them. Photon's two genuine differences
are now hooks on the base — `_send_retry_is_final(result)` (structured auth/target refusals are
returned as-is, no retry, no plain-text resend) and `_send_plain_fallback(...)` (no Markdown banner,
richlink() bypassed) — and the 42-line fork is deleted.
Slack's `_retry_after_from_exc` and Discord's `_extract_discord_retry_after` parsed the header by
hand and only understood the numeric form; both now call agent.retry_utils.parse_retry_after_seconds
(numeric or HTTP-date, either header casing). Discord keeps its `retry_after` attribute path, the
`X-RateLimit-Reset-After` fallback and the 1s floor.
Behavior change: Photon retries now add up to 1s of jitter to the backoff and honour a server
retry_after; Slack/Discord recognise an HTTP-date Retry-After they previously ignored.
The retained comment still claimed a scope-less multiplex caller could load an
OSS config; identity reads above it raise UnscopedSecretError first. Both the
comment and the regression test now say the same thing: callers are scoped, an
OSS profile whose scope lacks MEM0_API_KEY initializes, and a scope-less caller
is a spawn-site bug that raises.
read_json_or_empty returns {} for malformed JSON; returning that directly from
_load_config made a corrupt profile config.json yield an empty (silently
unconfigured) mapping where the pre-dedup loop kept walking to the legacy
file and then the env branch. Only a non-empty parsed object is now returned.
_finalize_all_traces called _get_langfuse(), which lazily BUILDS a client when the
launch-profile slot is empty. Under a multiplex gateway every turn runs scoped, so
only the per-home slots are populated; at exit (no scope) the credential read
raised UnscopedSecretError and the flush loop was skipped for every profile,
losing all pending traces. The finalizer now iterates the already-settled
clients (launch slot + per-home map) and never initializes. Test drives two
profiles under real home-override + secret scopes, then finalizes unscoped.
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
Four bundled plugins wrapped get_hermes_home() in a try/except that fell back
to ~/.hermes on ImportError (a2a/protocol._hermes_home, photon/auth
._auth_json_path, google_chat adapter inline, openviking done in the previous
commit). A bundled plugin cannot lose hermes_constants -- each already imports
gateway.* / agent.* from the same tree -- so the fallback was dead code that
was also wrong on Windows (%LOCALAPPDATA%/hermes) and under a profile override.
telegram's gmail-triage verb path used Path.home()/".hermes" outright, ignoring
profiles. mem0/_oss_providers baked os.path.expanduser("~/.hermes/mem0_qdrant")
into VECTOR_PROVIDERS at import time, so the Qdrant default landed in the
user's ~/.hermes for every profile; the default is now a lazy callable resolved
by vector_default_config(provider_id) at setup time. a2a/security.py only
changes its import (it borrowed protocol._hermes_home).
Behavior change: on Windows and under profiles these paths now follow the
active HERMES_HOME (they were previously anchored to ~/.hermes in the impossible
fallback / at import time); the default-profile POSIX layout is unchanged.
Test: tests/plugins/test_plugin_paths_follow_profile.py asserts each resolver
(a2a conversations, photon auth.json, mem0 qdrant default, openviking log)
lands inside a HERMES_HOME ContextVar override; sabotage red for the mem0
import-time constant and the photon ~/.hermes fallback.
Three shims caught UnscopedSecretError and degraded to "" (mem0._scoped_env)
or to os.environ (langfuse._secret, azure_identity_adapter._scoped_env). Under
multiplex os.environ holds the DEFAULT profile's .env, so the langfuse/azure
fallback could ship another profile's keys, and the mem0 fallback silently
routed a mis-spawned turn's memories into the default profile's account. The
exception exists to surface exactly that spawn-site bug (agent/AGENTS.md:
never add environ fallthrough, never swallow it). All three now call
agent.secret_scope.get_secret directly: with a scope installed a miss returns
the default; single-profile deployments (multiplex off) still read the process
env inside get_secret; a scope-less multiplex caller raises.
Behavior change: a mis-spawned child under gateway.multiplex_profiles now
fails loud with UnscopedSecretError instead of running silently unauthenticated
/ on the default profile's identity. The #99121 contract (OSS mode needs no
MEM0_API_KEY in scope) is unchanged and its test now installs an empty profile
scope, which is the situation the issue described; a genuinely scope-less
caller is asserted to raise in a new test.
Tests: tests/plugins/test_scoped_secret_readers_fail_closed.py (scope wins over
environ; scope-less multiplex raises) for langfuse + azure, sabotage red when
the langfuse fallthrough is restored; tests/plugins/memory/test_mem0_v3.py::
test_load_config_fails_closed_without_scope_even_for_identity_settings,
sabotage red with a swallowing wrapper reinstated.
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.
Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.
holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.
openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.
Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
The a2a invariant test skipped any synthesized token the canonical
redactor itself let through, so a boundary/shape regression would have
passed silently. Every class scrubs today, so the escape hatch goes and
the soft ">= 50" count becomes the exact registry size.
The gateway body test still asserted a "[REDACTED]" fallback marker that
no longer exists; redact_for_egress masks via _mask_token ("***"), so
assert that alone. honcho oauth.py's `re` import became unused when its
private pattern list moved to the registry.
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.
Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.
Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).
Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
fal launched H3 Max Turbo on Sep 3 (minimax/h3-max-turbo/{text,image}-to-video):
a throughput-tuned post-train of H3 Max with a 1080P tier Max lacks, at
$0.025/s 480p / $0.04/s 768p / $0.08/s 1080p list ($0.00625-0.02/s promo until
Sep 14). Schema matches Max's shape — required prompt_expansion_mode static
key, int duration 5-15, seed on both endpoints, i2v drops aspect_ratio — plus
the new 1080P resolution enum, so it reuses the existing family capability
flags with a Turbo-specific resolution alias map.
Schema verified against the FAL queue OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403 "Exhausted balance"); portal
allowlist/pricing needed for managed users on the 2 new endpoints.
The deepseek plugin is bundled and always loads, so the try/except around
the delegation import was defense-in-depth for a path that cannot fail.
Tests reduced to the two contracts that matter: /reasoning none reaches
the wire as thinking.disabled, and DeepSeek ids produce exactly the native
DeepSeek profile's output while non-DeepSeek families stay a no-op.
The bundled-plugin loader pops half-registered modules when a plugin
fails to load, so the lazy 'from plugins.model_providers.deepseek import
deepseek' could raise ImportError on every DeepSeek-routed CommandCode
turn — turning the soft 'thinking uncontrollable' bug into a hard
crash. Catch ImportError, log, and return the pre-fix no-op (review
feedback on #95241).
CommandCode fronts DeepSeek with vendor-prefixed ids
(deepseek/deepseek-v4-flash). DeepSeek V4+ defaults to thinking mode
when the thinking field is omitted, so /reasoning none changed the
Hermes session state but not the actual request -- the turn sat in
reflecting.../brainstorming... for minutes (#95232). Strip the vendor
prefix for DeepSeek-family ids and delegate to the native DeepSeek
profile's build_api_kwargs_extras (extra_body.thinking +
reasoning_effort mapping); other CommandCode model families keep the
base no-op behavior. The prior no-op tests codified the bug and are
rewritten to pin the new contract.
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):
- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
no duration are outside the unified video_generate surface and are dropped)
with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
live limits (nearest by value/height/ratio) and drops generate_audio/seed
for models that lack them (the API 400s otherwise); reference images ride
in input_references; local file inputs are refused (OpenRouter fetches
URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
never to a provider-supplied unsigned_urls host (kept from #103267)
Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.
Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
The version-less canonical Flash id is not matched by _is_deepseek_thinking_model, so agent.reasoning_effort and every auxiliary reasoning_effort were silently dropped on the OpenCode Go relay while the direct provider was fixed (8435a3ae00/aeecb110f8). Match it through a version-less id set, mirroring plugins/model-providers/deepseek. Verified against the Go relay: named levels are graded (low 906 / high 1371 / max >=2500 reasoning tokens on a multi-step prompt) and integer efforts are rejected, so the named level is the knob to send.
_truncate_for_sync documents "the last sentence boundary within max_len", but it
looped over separator KINDS and returned on the first kind that qualified. An early
"。" therefore outranked a "." 240 characters later, and in pure ASCII "." outranked
a later "!" or "?" purely because it comes first in the tuple.
With the 450-char default, "a"*200 + "。" + "b"*240 + "." + "c"*100 kept 201 of the
442 characters available: 241 characters the embedder would have accepted were
discarded, so any fact in the second half of the turn never reached extraction. The
add() call succeeds, so unlike #106235 nothing is logged — the turn is simply
remembered from its first sentence. Raising sync_max_chars widens the gap rather
than closing it.
Take the max over every separator instead, from a named tuple so the set is not
buried in the loop. ".\n" is dropped: its index can never exceed the bare "." it
starts with, so under a max it is unreachable. The first-third guard and the hard-cut
fallback for unsegmented input are unchanged.
Fixes#108868
Ramp Router efforts cache + warm/disk flags, xAI and OpenRouter image catalogs,
Hindsight append-capability verdict, memory-provider skill registry, OpenViking
atexit provider, Honcho loopback flow status, Langfuse client (os.environ-only
credentials) and disk-cleanup's protected cron paths held one profile's
credential- or home-derived value process-wide; YuanbaoAdapter._active_instance
was last-connected-wins across profiles.
Keyed by home key / credential fingerprint under an override, credentials read
through the secret scope, warm threads run under copy_context(); unscoped module
slots stay for the single-profile path and the existing monkeypatch tests.
Keep: scopeless worker resolves the on-disk key (durability core), a keyed
build still rewrites (rotation), and the API-prefixed vault name resolves.
Drop the mock-patched gate predicate test, the inlined copy of the same
predicate, and the empty-build/empty-disk arm — they re-assert the helper's
shape rather than a behaviour contract.
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.
Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.
Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.
Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.
Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.
Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.
Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.
session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.
Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.
Fixes#82903Fixes#65941Fixes#99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
Conflict resolution on top of main also routes MEM0_MODE / MEM0_HOST / MEM0_AGENT_ID /
MEM0_USER_ID through get_secret: identity and host are .env values like the key, and a
raw environ read hands a secondary profile the default profile's mem0 account.
(cherry picked from commit 7a5863ebafdb193d8ec2f89a7aca309774d6d69a)
The vault item is named HINDSIGHT_API_LLM_API_KEY (matching the daemon env var), while the setup wizard uses HINDSIGHT_LLM_API_KEY. Vault-fed scopes carrying only the API-prefixed name resolved empty, so the scopeless path — and the whole fail-closed machinery — never engaged on real deployments. Accept both names.
Proven end-to-end: hydrated scope now resolves the 73-char key and the rewrite gate passes.
(cherry picked from commit 2ce8c2980188ae8e4a3ed398e14fa8ed89a40644)
test_gate_refuses_keyless_build_against_keyed_disk and test_rewrite_allowed_when_build_carries_key touch no filesystem modes, so the nt skip guard was needlessly excluding them. Addresses review feedback on PR #103957.
(cherry picked from commit 196cf723571c87829be61e17df27869948db00a0)
The embedded daemon-start worker usually runs with no secret scope, so _embedded_llm_api_key resolved empty and both the profile-env compare and the upstream manager merge overwrote the file's good key with emptiness.
- _embedded_llm_api_key: explicit config, then secret scope, then on-disk profile env fallback; UnscopedSecretError treated as empty without touching os.environ (multiplex safety).
- Add _may_rewrite_profile_env gate: refuse rewrite on keyless-build plus keyed-disk, and skip the daemon stop in that case.
- Document key resolution order in plugins/memory/hindsight/README.md.
- Add 5 tests pinning disk fallback, the gate's refuse branch, rotation, keyless providers, and file survival.
(cherry picked from commit 48a4bcd536fa799b666a8d5eaae73c1816ca9e33)
A DM relayed from another Hermes profile ran as a turn in the recipient's
Bot Chat session. sync_turn wrote the bot's words and the recipient's reply
into that session, and before per-author writes they landed under the
human's peer. The human's representation absorbed conversations the human
never had.
The turn context now marks such turns with scope a2a:<bot id>.
sync_turn routes a bot-authored turn into a separate Honcho session keyed
<session>:a2a:<sanitized bot id>, created with the sender bot as its user
peer, and never writes it into the human's session. The key is deterministic
so every turn from the same bot reaches the same session, and it stays
inside Honcho's 100 character session id limit. Recall still reads the
human's session only.
a2aSessions (host block, then root, default true) turns the routing on.
With it off, bot-authored turns are skipped. A bot turn that names no
author id is skipped as well, because nothing can key its session. Human
turns are unchanged.
get_or_create takes a user_peer_id override so the a2a session's roster is
the bot and the assistant, not the runtime human.
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.
Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.
Builds on YipTszkwan's #107126 (earliest fix in the cluster).
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:
* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
omitted `extra_body.thinking`. The server then defaults to thinking-on, so
the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
V-series regex), so the id a user picked never reached the wire and the
config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
instead of the real 1M window, capping the model at an eighth of its
context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
detector at its 180s default instead of the 600s reasoning-model floor.
Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".
Adds the id to all four gates plus regression coverage for each site.
Keep the two tests that fail on main without the fix:
- test_fal_common: a keyed managed submit makes exactly one POST (plus the
negative arm: an unkeyed submit still goes through the SDK retry ladder)
- test_image_generation: the 409 BILLING_ERROR body surfaces
`unsupported_pricing_meter` instead of the generic "not yet enabled" text
Dropped from #106484: the duplicate video-plugin billing test (same helper,
same assertion), the `_fal_client = fake` / `import_fal_client` stub churn
and the `tools.lazy_deps` stub — fal-client is installed in CI (`--extra fal`)
so those fixtures were not needed; the `_load_fal_client` no-op fixture on
TestManagedGatewayErrorTranslation for the same reason.
Also drop the redundant `retry_request is None` re-check in
`_ManagedFalSyncClient.submit` — `__init__` already raises when the helper
is missing.
Avoid retrying idempotent managed FAL submissions because the retry can mask the initial billing failure. Surface structured Nous billing diagnostics consistently for image and video paths, with hermetic regression coverage.
(cherry picked from commit 289ce039e9a522dc8016ae4a512214c05d0a8bc0)
A flat 450-char cap fits 512-token embedders (bge-small-zh-v1.5,
all-minilm) but stores only ~5% of the window on 8192-token models
(text-embedding-3-small, jina-embeddings-v3, bge-m3), degrading memory
quality for users those models served fine before truncation existed.
Read `sync_max_chars` from mem0.json once in initialize() (450 default)
and pass it to _truncate_for_sync(). Config over auto-detection: the
Ollama /api/show probe + known-model table proposed in #37427 adds a
network call and a curated list for a number the operator already knows
from their embedder choice; the setup wizard's mem0.json is the plugin's
behavioral-settings surface (no new HERMES_* env var). Documented in the
plugin README and the memory-providers docs page.
Dynamic-cap requirement and measurements (450 OK / 600 -> HTTP 500 on
bge-small-zh-v1.5:f16) by @szicely in #106235.
Refs #37421#106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
sync_turn() sent the whole turn to backend.add() untruncated. OSS
embedding models with small context windows (Ollama bge-small-zh-v1.5:
512 tokens) reject the request with HTTP 500, and hosted APIs answer
INPUT_TOKEN_LIMIT_EXCEEDED — in both cases _try() only logs, silently
dropping the turn's memory extraction after long conversations.
Cap each synced message at its last sentence boundary within 450 chars
before ingestion: short turns pass through unchanged, long turns keep a
coherent statement for fact extraction. This replaces the previous
retry-on-error approach, which could not match Ollama's HTTP 500 shape.
Salvaged from #37427 with the test suite trimmed to the one invariant
(oversized turn still reaches a small-context backend, short text
untouched, no breaker failure).
Refs #37421#106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
The plugin loader execs sibling modules before the package __init__, so a
module-level `from . import _read_mem0_json` in _setup.py failed against the
empty parent shell and the whole module silently dropped out — every later
`hermes memory setup mem0` died with "cannot import name 'post_setup'".
The test loads mem0 through the real discovery path from a cold sys.modules
and asserts the cached _setup module exposes post_setup (red on base).
Also maps the contributor email for #103078 credit.
Campaign tracker: https://github.com/NousResearch/hermes-agent/issues/104154
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:
- search: POST https://api.perplexity.ai/search (documented Search API),
search_context_size=low so `snippet` stays description-sized;
results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
route behind `pplx content snippets` (the CLI's `content fetch` is
deprecated upstream). web_extract has no query, so the URLs' path words
serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
backend set + credential ladder + availability probe (web_tools),
registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
dump key lists, nous_subscription direct-credential detection, setup
summary, test conftests, docs.
Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.