The stack inserted the errcode table directly after _strip_reply_fallback's
return with no blank lines, which reads as if the constant belongs to the
function body and trips E305. Two blank lines restore the module-level
boundary; no behaviour change.
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.
With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).
The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.
A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.
Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.
Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.
(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.
I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.
I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.
While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.
Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.
On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.
(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.
I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.
Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.
(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)
One reader (gateway.platforms._shared.extra_or_secret) now implements the
precedence every per-profile setting follows for the OWNING profile:
explicit scoped env/.env → that profile's config.yaml (PlatformConfig.extra)
→ the adapter's default. A scoped miss returns the default, never the launch
process's os.environ; single-profile / default-profile installs keep the
documented env-over-YAML contract.
Why: 545e74d0ea (#108705) stopped bridging a secondary's YAML into the
process env and moved readers to config.extra, but the shared reader and the
hand-rolled helpers in Discord/Slack/Matrix/Telegram consulted YAML FIRST and
then fell back to a scoped env read. Two bug classes followed (#108440
post-merge review by andrexibiza, #109032):
- an explicit env value could no longer beat YAML for the owning profile
(DISCORD_ALLOW_MENTION_EVERYONE=false lost to allow_mentions.everyone: true;
TELEGRAM_REACTIONS=true lost to the stock reactions: false);
- a secondary that OMITTED a key inherited the launch profile's bridged env
through the fallback (Matrix process_notices/session_scope, Discord
auto_thread/reactions/mentions, Slack reactions/ignored_channels).
Consumers migrated to the shared reader: Discord _build_allowed_mentions and
_extra_or_env_flag; Slack _slack_allow_bots, _reactions_enabled (the
_extra_or_env_* getters already used it); Matrix _extra_truthy, _extra_csv_set,
session_scope, reactions, require_mention parsers, and — new — the
allowed_users / ignore_user_patterns consumers that never read the seeded YAML
lists; Telegram _extra_bool, _extra_str_set, _reactions_enabled; Feishu
allow_bots; WhatsApp dm_policy/group_policy.
Refs #108440, #109032
#69090 scoped MATRIX_RECOVERY_KEY itself (via _scoped_recovery_key())
so a secondary profile resolves its own recovery key under multiplex,
but left its sibling, MATRIX_RECOVERY_KEY_OUTPUT_FILE, on a bare
os.getenv(). _recovery_key_output_path() is called from inside
_verify_or_bootstrap_cross_signing(), which runs fully inside
_profile_runtime_scope for a secondary profile: when that profile
bootstraps a new recovery key, it either doesn't get written to a
file at all, or gets written to the default profile's configured
path, depending on which one has the env var set.
Route it through the same _get_scoped_secret() helper _scoped_recovery_key()
already uses.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit fb765ee49b2f1a1853e52fb901d25f767180a2fc)
545e74d0 (post-0.21.2) made the Matrix YAML bridge seed its values into
PlatformConfig.extra so secondary multiplex profiles read their own config. The
"csv" bridge kind seeds any non-None value, so `free_response_rooms: ''` now
reaches extra as '' — and the readers' `if raw is None` fallback no longer fires,
so MATRIX_FREE_RESPONSE_ROOMS is ignored and require_mention drops every
un-mentioned message. Before that commit the bridge only wrote env and the key
was absent from extra, so the env value applied.
Route the three identity-check readers (_extra_csv_set, _extra_truthy,
_resolve_max_message_length — the last a three-tier chain where '' also
short-circuited the plugin-registry default) through the shared
gateway.platforms._shared.extra_or_secret, whose default already treats a blank
string as unset (the idiom mattermost/dingtalk/slack readers use). Explicit
scalars, bools and lists (including []) stay authoritative.
Two invariant tests replace the salvaged suite (moved to
tests/plugins/platforms/matrix/ to mirror the source path): blank falls through
for all three readers; explicit values still beat env.
Fixes#109358
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
Fourteen platform plugins hand-rolled the "already configured? Reconfigure? [y/N]"
gate at the top of interactive_setup (env check + info line + prompt_yes_no(..., False)),
with drifting wording ("X: already configured" vs "X is already configured." vs
"already enabled") and, for LINE and SimpleX, raw input() loops with their own
EOF/KeyboardInterrupt handling and no gate at all. Fixes to the gate (non-interactive
handling, wording, default) therefore reached only the core Telegram/BlueBubbles/webhook
wizards.
- hermes_cli/setup_platforms.py: `_declines_reconfigure` becomes the public
`declines_reconfigure(label, question, *env_vars)` (any-of env check, so Matrix's
token-or-password gate fits); `_save_prompted` becomes `save_prompted` alongside it.
No alias kept; the three core callers are updated.
- buzz, dingtalk, discord, feishu, google_chat, irc, matrix, mattermost, raft, slack,
teams, wecom: the hand-rolled gate is replaced by one `declines_reconfigure(...)` call;
post-decline extras (Discord allowlist nudge, Slack manifest refresh, Raft "Keeping"
line) stay local and unchanged.
- line, simplex: the raw input() loops move onto hermes_cli.cli_output.prompt (masked
for secrets, "" on Ctrl-C/EOF) and gain the shared gate on their primary env var.
Behavior change: the gate's info line is now uniformly "<Label>: already configured"
(DingTalk/Feishu/WeCom lose the trailing period + inline ID; Buzz/IRC/Google Chat/Raft/
Teams no longer echo the current value in that line). Feishu and WeCom now gate on the
app/bot ID alone instead of ID AND secret. LINE and SimpleX gain a "Reconfigure?" [y/N]
prompt when already configured; their prompts now honour HERMES_NONINTERACTIVE and print
via the CLI helpers instead of bare print(). Prompt defaults (No) are unchanged everywhere.
Not touched: WhatsApp's gate keys on WHATSAPP_ENABLED being truthy (a "false" value must
not count as configured), which the shared any-set gate cannot express — left hand-rolled.
Test: tests/plugins/platforms/test_interactive_setup_reconfigure_gate.py parametrized over
the 14 wizards — with the primary env var set and the user declining, each wizard must have
called declines_reconfigure with that var and returned without prompting or saving.
Sabotage: reverting mattermost's gate fails that row.
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).
`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.
Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.
`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.
Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
Nine surfaces (feishu, teams, slack, telegram, whatsapp_cloud, qqbot, matrix, discord, relay)
each re-derived the approval choice set — [Allow Once]; session + always unless smart-denied;
[Deny] — and four of them (discord, slack, teams, whatsapp_cloud) never adopted
base._format_exec_approval, so header/reason/smart-deny wording and truncation budgets drifted
per adapter. Three separate commits had to touch 5–9 adapters for one semantic fix.
BasePlatformAdapter.send_exec_approval now builds an ExecApprovalPrompt (shared text via
_format_exec_approval, shared `(label, choice, style)` rows via _exec_approval_actions) and
hands it to the `_send_exec_approval_prompt` hook. Each adapter keeps only its widget mapping
(~10–20 LOC); platform wording stays via the existing `_EA_*` class attrs, and a new
`_exec_approval_cmd_budget` hook lets Slack/Discord budget the command against their hard
message caps (3000-char section / 2000-char message) instead of computing it inline.
`_EA_REASON_BUDGET` covers Slack's 500 / Discord's 300 reason caps.
The runner used to detect button support by `hasattr(type(adapter), "send_exec_approval")`;
that is now true for every adapter, so `_renders_exec_approval_buttons` asks
`supports_exec_approval_buttons()` (hook overridden?) and keeps the duck-typed check for
non-BasePlatformAdapter classes.
Visible text changes (button semantics unchanged everywhere):
- Discord: the smart-deny line now follows the reason (was inside the header before the
fence); the truncation marker is "..." not "\n... [truncated]".
- Slack: smart-deny line follows the reason instead of the header.
- Teams: unchanged (same 2000-char preview, same smart-deny block).
- WhatsApp Cloud: identical text; body still capped at 1024.
- QQBot/relay: unchanged.
Eight adapters (discord, telegram, wecom, matrix, whatsapp, simplex, feishu, weixin) each kept a
copy of the delayed text-batch flush that base._enqueue_text_event schedules. Two correctness
fixes had landed in single copies only: Discord's asyncio.shield around the dispatch (#12444 —
a late chunk cancelling the flush task aborted the in-flight agent turn) and WeCom/Weixin's
synchronous task-identity check before the pop (a superseded task waking late popped the event
and the successor found nothing). The other adapters carried both bugs latent.
The base now owns `_flush_text_batch` with both fixes, plus the `_pending_text_batches` /
`_pending_text_batch_tasks` dicts and `_SPLIT_THRESHOLD` / delay attrs (defaults; adapters set
their own). Platform policy goes through three small hooks instead of a copied body:
`_text_batch_delay_for(pending)` (Telegram's fast/short tiers, WeCom's attachment-only wait),
`_pop_text_batch(key)` (Feishu's side count table) and `_dispatch_text_batch(event)` (Feishu's
per-chat lock). Telegram keeps its `_flush_buffered` body because its contract differs on
purpose — a cancel after the pop must hold-and-re-raise so teardown can stop a flush — and gains
the identity check there. Matrix's `_split_threshold` is renamed to the shared `_SPLIT_THRESHOLD`;
SimpleX exposes its single delay through the shared attr names.
`helpers.TextBatchAggregator` (zero users) sits inside the revert-scheduled PLUGIN-COMPAT block
and is left for that revert.
The Matrix docs described six agent-exposed matrix_* tools and three MATRIX_TOOLS_ALLOW_* env gates that were never implemented — the tool names and gates appear nowhere in code. Docs now describe actual behavior: no Matrix-specific agent tools; reactions/redactions are internal to approval prompts and pickers; MATRIX_ALLOWED_ROOMS scopes responses. Fixes#100535.
Under gateway.multiplex_profiles a served secondary profile's adapter is built and
connected inside _profile_runtime_scope while os.environ still holds the DEFAULT
profile's .env. Credentials and allowlists were already read through the profile
scope (get_scoped_secret / _platform_gate_env), but the non-credential SETTINGS the
adapters read with bare os.getenv were not, so a served profile silently ran with the
default profile's values: webhook listener host/port/URL (SMS, Teams, LINE, Feishu,
BlueBubbles), Signal's connect URL/account gate, mention gating and reactions (Slack,
Matrix, Signal, Feishu, BlueBubbles, Discord), Matrix thread/session/E2EE policy and
message-length limits, Discord backfill/command-sync/attachment caps, Buzz reply mode
and env enablement seed, A2A agent name/port/description/toolsets, and the
api_server model alias.
Every such read now goes through the existing scoped reader (get_scoped_secret, or the
adapter's own scope-aware helper): under a secondary's scope the profile's own .env is
authoritative and a miss yields the default -- never another profile's value; the
default profile and single-profile gateways keep reading os.environ exactly as before.
Buzz and A2A previously short-circuited to "extra only / built-in default" under a
scope, which also dropped the profile's OWN .env; they now read the scope so a served
profile matches its standalone gateway.
The parity harness (temp HERMES_HOME, default + 2 secondaries with distinct values for
every env var each adapter reads, real load_gateway_config + adapter factory in both
topologies) went from 70 raw process-env bypass sites across 14 adapters to only the
HERMES_<PLATFORM>_* perf knobs and the api_server listener vars, which are process-
global by design (agent.secret_scope._GLOBAL_ENV_*).
Under gateway.multiplex_profiles every secondary profile's config loads inside
_profile_runtime_scope, yet every apply_yaml_config_fn hook (feishu, matrix,
whatsapp, slack, dingtalk, discord non-gate keys, telegram non-gate keys) and
gateway/config_loader.py::bridge_core_env_settings still wrote os.environ there.
First-writer-wins: the first secondary with a require_mention / allowlist /
allow_bots / reactions block made that policy the DEFAULT profile's (live:
TELEGRAM_REQUIRE_MENTION written from a secondary load), and secondaries read
the default's env for the same keys.
- gateway/platforms/_shared.py::yaml_env_setter: the one env-write shape for
YAML->env bridges — env wins, skipped under an active secondary scope.
- Every hook now seeds its values into the profile's PlatformConfig.extra and
uses yaml_env_setter; bridge_core_env_settings seeds telegram/signal
require_mention into extra and skips the env write under scope.
- Readers that bypassed extra/scope now consult extra first (matrix flags +
session_scope, slack reactions, telegram reactions/mention_patterns/_extra_bool,
discord reactions/auto_thread/history_backfill/approval_mentions/allow_mentions,
feishu allow_bots, dingtalk mention_patterns, signal require_mention).
Salvages the shape of PR #100604 (whatsapp, earliest report #80099), #100435
(discord) and #100448 (telegram) by @nftpoetrist on top of current main.
Fixes#80099
Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env. Several
adapter-owned authorization gates still read GATEWAY_ALLOW_ALL_USERS, GATEWAY_ALLOWED_USERS
or their platform allowlist/allow-all raw from os.environ, so the default profile opting
into open access opened every secondary email/QQ/WhatsApp/Matrix/Teams/Slack/LINE/DingTalk
bot to any sender (email additionally skipped From: authentication), the default's Matrix
allowlist decided who may approve tool calls on a secondary bot, and a secondary that
opted in only in its own .env was silently deny-all.
Every such read now goes through the adapter's existing module-local scoped reader
(gateway.platforms._shared.get_scoped_secret / matrix _startup_env_secret): profile
scope first, scoped miss = default, never os.environ; the unscoped default-profile and
single-profile paths keep the environ read, where it IS the profile's own value.
Sites: email _allow_all_senders/_allowlist_in_effect; qqbot _open_dm_opted_in;
whatsapp_common _open_dm_opted_in/_live_dm_allow_from; teams _card_action_denied;
matrix _is_authorized_user, MATRIX_ALLOWED_USERS, MATRIX_IGNORE_USER_PATTERNS,
_extra_csv_set (allowed/free-response rooms); slack _slack_allow_bots/_slack_api_human_users;
line _truthy_env/allowlist (allow-all, user/group/room allowlists); dingtalk _extra_get
(allowed_users/chats, free-response chats, require_mention).
Live repro (temp HERMES_HOME, multiplex on, default env GATEWAY_ALLOW_ALL_USERS=true,
secondary scope without opt-in): EmailAdapter._allow_all_senders() True -> False,
QQAdapter._open_dm_opted_in() True -> False, Matrix _is_authorized_user('@stranger')
True -> False, Teams card action allowed -> denied.
Co-authored-by: Drexuxux <drexux0@gmail.com>
Co-authored-by: MoonsvnLyn <FirmamentalSpring@users.noreply.github.com>
Co-authored-by: svector-anu <anuoluwakolapo94@gmail.com>
Co-authored-by: babatorik <durgun.ismail@gmail.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's
values and a secondary profile's .env exists only in its secret scope.
Matrix, Home Assistant, Teams, DingTalk, SMS and Telegram already read
their *secret* scope-aware, but read the paired endpoint/identity raw from
os.environ — so a secondary's token/secret/password was paired with the
default profile's host or app identity:
- Matrix: MATRIX_HOMESERVER / MATRIX_USER_ID / MATRIX_DEVICE_ID in
__init__ and MATRIX_HOMESERVER in _standalone_send -> the secondary's
access token was sent to the default's homeserver (and its E2EE device
id reused).
- Home Assistant: HASS_URL in __init__ and _standalone_send -> the
secondary's HASS_TOKEN was posted to the default's HA instance.
- Teams: TEAMS_CLIENT_ID / TEAMS_TENANT_ID (env-first, even over extra),
TEAMS_SERVICE_URL, TEAMS_HOME_CHANNEL(_NAME) in _credentials,
_env_enablement, _standalone_send and __init__ -> Bot Framework token
requested for the default's app with the secondary's secret; cron
home channel was the default's conversation.
- DingTalk: DINGTALK_CLIENT_ID in _credentials, DINGTALK_WEBHOOK_URL in
_standalone_send -> wrong app id; cron output posted to the default's
robot webhook (URL carries its access_token).
- Telegram: TELEGRAM_WEBHOOK_URL in connect() -> the default's public
webhook URL registered on the secondary bot, mixing update traffic.
Each read now goes through the same scoped reader its sibling secret
already uses (_get_scoped_secret / _startup_env_secret / get_secret with
the UnscopedSecretError fallback), so a scoped miss is empty rather than
the default profile's value; the unscoped default-profile path and
single-profile deployments keep reading os.environ, which there IS the
profile's own value. Teams _credentials now lets extra (per-profile
config.yaml) win over env, matching every other adapter and its own
__init__.
Live repro (/tmp/mux_audit/fix-endpoint-identity/repro.py, secondary
scope "bot2" with the default profile in os.environ): 19 of 24 reads
returned the default's endpoint/identity before; 0 after.
Salvage: SMS from-number fix cherry-picked from #100625 (tests trimmed to
two invariants). Teams identity (#100622) and DingTalk (#100615)
reimplemented on the minimal shape — the Teams PR added a _profile_scoped()
gate to _env_enablement and the DingTalk PR only covered the YAML bridge,
not DINGTALK_CLIENT_ID / DINGTALK_WEBHOOK_URL. Matrix, Home Assistant and
Telegram parts have no prior PR.
Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
_standalone_send builds formatted_body through its own markdown call and
never went through _markdown_to_html, so cron-delivered equations still
arrived as raw dollars. Tokenize/expand around that conversion as well.
Keep the two tests that fail without the fix: inline + display math reach
formatted_body as data-mx-maths markup with the TeX HTML-escaped while the
plain body keeps the raw TeX; unpaired dollars and text colliding with the
sentinel format pass through unchanged (no IndexError). The 17 unit tests
from #106452 are dropped per the salvage bar (≤ 2 invariant tests).
Also drop the dead `_tex_store = []` pre-assignment and the `_fb`/`html`
temporaries in `_markdown_to_html` — pure tidy, no behaviour change.
Element (feature_latex_maths) typesets <div|span data-mx-maths="TEX">
elements at display time, but the outbound HTML sanitizer allowlists tags
and attributes, so data-mx-maths markup sent by the gateway never reaches
Element intact - messages containing $...$ render as raw dollars.
Convert $...$ (inline) and $$...$$ (display) to opaque sentinel tokens
before Markdown conversion and expand them to data-mx-maths markup after
sanitization. Tokens are printable text with no HTML/Markdown meaning, so
neither the converter nor the sanitizer touches the TeX. Unpaired dollars
(prices, literals) are untouched, and adversarial text colliding with the
sentinel format passes through verbatim (index-checked expansion).
(cherry picked from commit eb73aafe0e08502c54e63dbd03e99796c3b7fee3)
#106167 widened the base contract to SendResult but left six native-batch
overrides (Discord, Email, Matrix, Mattermost, Slack, Telegram) returning
None, which (a) meant a media-only reply on those platforms still reported
FAILURE because _record_delivery(None) records nothing, and (b) produced new
`ty` invalid-method-override diagnostics against the widened base (#106192).
Make the contract honest instead of annotating it Optional: each override
now rolls its batches (and any per-image fallback) into one SendResult, so
the turn-outcome accounting works on every platform, not just Signal and
the base loop. The "legacy overrides return None" comment in
_send_image_batch goes away with the legacy.
ty on the 8 touched files: origin/main 241 diagnostics / 20 override,
this branch 241 / 20 — byte-identical diagnostic set; the intermediate
`-> SendResult` head without this commit had 246 / 25.
Refs #106192
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
Review findings on #102117 (independent reviewer + itsflownium):
* hermes_cli.kanban_db.connect / connect_closing pointed at hermes_cli.projects_db (different DB, no
board= parameter). The compat generator ranked candidate homes by path proximity when a name is
defined in several modules. Now it requires shape compatibility with the BASE definition (same
literal for constants, superset of parameter names for defs) and prefers the facade's own
<stem>_* sibling. Same class fixed for tools.tts_tool.DEFAULT_XAI_BASE_URL (-> tts_tool_providers),
and 17 constants/defs that had been pointed at same-named strangers (Matrix MAX_MESSAGE_LENGTH ->
Signal's 8000, tts MAX_TEXT_LENGTH -> BlueBubbles', honcho/retaindb/supermemory *_SCHEMA -> another
plugin's schema, ...) are now restored from BASE verbatim instead.
* send_yuanbao_direct (restored-def): body called adapter._outbound.send_direct, which HEAD moved to
the sender; rewritten to adapter._outbound.sender.send_direct.
* COMPAT_MANIFEST.md states the scope explicitly: public top-level names only; private names and
test monkeypatch seams are not preserved.
* scripts/check_subprocess_stdin.py: _splat_carries_stdin looked 30 lines ahead in the file text
and was satisfied by an unrelated later stdin=; it now finds the splatted name's definition via AST
and requires stdin inside that expression/body.
Tests: tests/test_compat_manifest_targets.py (pointer identity vs the facade's sibling; kanban
connect(board=) opens a Kanban DB, not projects.db; both FAIL on the previous layer),
test_subprocess_stdin_guard gains the false-negative probe, and the MoA -Q quiet-output contract
tests are back (tests/agent/test_moa_quiet_reference_output.py) against build_moa_facade.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
dingtalk: drop DINGTALK_TYPE_MAPPING/EXT_MAP re-exports. google_chat: card_spec_to_cards_v2 test -> .cards.
matrix: drop module-level MAX_MESSAGE_LENGTH alias (no importers). teams: drop TeamsSummaryWriter re-export
(teams_pipeline/runtime + tests -> summary_writer). wecom: drop WeComStreamExpiredError/STREAM_EXPIRED_ERRCODE/
MAX_INTERMEDIATE_FRAMES re-exports (tests -> .streaming). parallel: drop _get_parallel_client/_get_async_parallel_client
aliases (tests -> _get_sync_client). email: drop stale 'alias' comment (_esecret_int is the only name).
photon: re-remove credential_summary() (shim-only, cb9b7c36f3); its no-leak test now drives print_credential_summary.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
The multiplex gateway imports plugins/platforms/matrix/adapter.py once, so
the module-level _STORE_DIR/_CRYPTO_DB_PATH resolved against the root
HERMES_HOME for every profile: all bots' Olm identities landed in one
crypto.db and inbound E2EE failed with "no session found" (#89168).
connect() runs inside _profile_runtime_scope, so resolve the store dir
there via get_hermes_dir (honors the context-local HERMES_HOME) and cache
it on the instance -- diagnostics and error-log paths read outside the
scope then still report the store actually in use. Mirrors the
pairing-store fix (a6397c379).
Salvage of #89169 (per-call resolvers collapsed into one cached resolve;
dead `_CRYPTO_DB_PATH = None` alias dropped -- no external importers).
Also routes the last raw MATRIX_HOMESERVER read in check_matrix_requirements
through _startup_env_secret like its token/password neighbours (#69943).
Fixes#89168
Co-authored-by: Michael Short <18595461+mjshorty@users.noreply.github.com>
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.
- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.
Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Adapter ingress derives a session key BEFORE the runner stamps
source.profile in _make_profile_message_handler, so the namespace fell
back to the active profile and every bot in a multiplexed gateway
produced agent:main:<platform>:<chat>. A Telegram private chat reports
the user's own id as chat.id, identical for every bot, so two profiles
sharing one human collapsed onto a single lane: _pending_text_batches,
_active_sessions, the busy-session guard and _post_delivery_callbacks are
all keyed on that string. A day of production logs across two bots shows
60 flushes, none carrying the secondary profile's namespace.
set_owner_profile records credential ownership on the adapter and
_session_key_profile resolves the namespace as source.profile ->
_owner_profile -> the session store's resolver, so a secondary adapter
keys into its own namespace even before the source is stamped. Stamped
sources keep priority, so relay/connector ingress, which routes per event
rather than per credential, is unchanged. _configure_profile_adapter
installs the owner alongside the other handlers, covering startup and
reconnect.
Every candidate is type-checked as a non-blank str, and every attribute
read goes through getattr: adapters are routinely built without
BasePlatformAdapter.__init__, and a duck-typed session store returns a
truthy non-string that would otherwise be interpolated into the key as
agent:<MagicMock ...>:.
Also routes the four call sites that passed no profile at all (feishu
media batches, raft, slack _session_key_for_source, telegram photo
batches) through the same resolver.
test_multiplex_busy_input_mode's secondary-adapter busy case seeded
_active_sessions with the unstamped agent:main: key, asserting the
pre-fix collapse. It now seeds the lane the profile-owned adapter
actually derives.
A primary adapter has no owner and an unstamped source, so it resolves
exactly as before; with multiplex_profiles off the resolver returns None
and every key is byte-identical to today's.
Matrix m.audio/m.file/m.video events populate content.body with the
uploaded filename when the sender adds no caption. The adapter already
blanks that for m.image (PR #16821, issue #13482) but not for audio,
file, or video msgtypes, so the filename survives into event.text and
is appended after the transcript, where the model reads it as the
user message rather than as transport noise.
Extend the existing adapter-level blanking: add
_looks_like_matrix_media_filename() with the same conservative
heuristic (single token, no whitespace, no path separators, known
media suffix or mimetypes audio/video match) and apply it at the
media message handler for m.audio, m.file, and m.video.
Salvage of #87968 by @AiwendilInTheWoods — reworked from shared
gateway path to adapter-level fix for consistency with the existing
m.image blanking.
The reply fallback stripping loop was copy-pasted between _handle_text_message
and _handle_media_message. Extract into the module-level _strip_reply_fallback
helper, matching _extract_reply_fallback's existing pattern.