connect() called _register_slash_commands inline on every (re)connect, and
/reload-skills called refresh_skill_group inline; both run
discord_skill_commands_by_category, the same per-skill path-resolution walk
the Telegram menu paid in #110707. On a 1.5k-skill install that holds the
loop past the liveness watchdog. Registration now hops through
asyncio.to_thread from connect(); refresh_skill_group is a coroutine that
hops the rescan the same way (the reload handler already awaits an awaitable
result). Contextvar-scoped profile overrides propagate through to_thread.
One invariant test: the loop keeps ticking while the scan blocks, from both
sites. Red on the PR head, green here.
Review finding: Discord _register_slash_commands/refresh_skill_group ran the skill catalog disk scan synchronously on the event loop.
The gateway's idle-command path resolved skill slash commands inline on the
loop: a cold skill scan, skill file loads and the unavailable-skill rglob over
every skills dir. On a 1.5k-skill install that held the loop ~2 minutes, the
loop-liveness watchdog fired and the gateway exited mid-session (#111091).
_hm_skill_slash_rewrite now runs through _run_in_executor_with_context so the
profile contextvars the scan is scoped to survive the hop. Known commands still
short-circuit before any I/O (previous commit).
Same class in the Telegram inline picker: build_inline_results rebuilds the
command/skill catalog per keystroke via _collect_gateway_skill_entries, the
same path-resolution pass #110707 traced in the command menu.
One invariant test: a known command triggers no scan; an unknown command's scan
runs while the loop keeps ticking. Red on origin/main and on the reorder-only
tree, green here.
Follow-up to the salvaged #110716 commit. Drops the module-level
_build_telegram_command_menu wrapper (telegram_menu_max_commands is a config
read and stays on the loop; asyncio.to_thread takes kwargs directly), and
applies the same hop to _ensure_forum_commands, which rebuilds the menu on the
inbound-message path for every forum chat until registration succeeds — the
second live fire site named in #110707.
Replaces the contributor's test with one that proves the invariant for both
sites: the loop keeps ticking while telegram_menu_commands blocks. Red on
origin/main, green here.
The reply-pill fix split the body on a leading "> " so the strip would not
rewrite the "> <@bot:srv>" pill. That keyed the exemption on the body shape,
not on the relation: a hand-typed blockquote in a plain (non-reply) message
that mentions the bot inside the quote reached the agent with the raw
"@hermes:example.org" text, which main used to strip. Split around the quote
only when m.relates_to carries m.in_reply_to; otherwise strip the whole body.
Review finding: quote-block exemption keyed on body.startswith('> ') instead of m.in_reply_to.
The Matrix adapter strips the bot's mention from the inbound body before
_extract_reply_context parses the inline reply fallback, and that strip is a
blind whole-body replace. A reply to the bot names the bot in the fallback
pill ("> <@bot:server> quoted"), which is exactly what makes the message
count as a mention under the default MATRIX_REQUIRE_MENTION=true -- so the
strip runs on every reply-to-the-bot and rewrites the pill to "> <>", after
which the pill regex no longer matches. reply_to_author_id is lost,
reply_to_text becomes the mangled "<> quoted" remnant, and the prompt
renders "[Replying to: "<> ..."]" instead of "[Replying to your previous
message: ...]".
Split the body into (quote block, reply text) and strip the mention from the
reply text only. The mention gate still sees the raw body, so a reply to the
bot keeps waking the bot; the visible reply text and the quote-block strip
are unchanged.
Fixes#111233
Drop the unreachable handle_message assertion in the drop tests
(_prefilter_inbound never calls it) and state the deliberate
file_comment drop decision in the allowlist comment.
Housekeeping subtypes (channel_join/leave/topic/name/purpose,
convert_to_private/public, pins, deletions) are not a person speaking,
yet _prefilter_inbound only rejected message_changed/message_deleted,
so each of them started a full agent turn in free-response channels.
Replace the denylist with an allowlist: a message passes when subtype
is absent, file_share, thread_broadcast or me_message; everything else
is dropped. Fixes#110778.
`_clarify_callback_sync` decided "no answer arrived" by testing whether the
response text starts with '[' (the shape of the timeout / undeliverable
sentinels). A real answer can start with '[' too — a "[A] staging" choice
label picked by number, or "[urgent] ..." free text after Other — so the
clarify resolved and the agent got the answer, yet the Slack card was
rewritten to "This prompt expired" and typing was never re-armed.
`_clarify_send_then_wait` now returns `(response, answered)` and the runner
branches on that flag only.
The Slack click handler popped the retire entry as soon as Other was
clicked, but Other is not terminal: the clarify stays pending for typed
text, so a later timeout or /new reset found nothing to retire and the card
stayed stuck on "Awaiting typed answer". The entry is now popped only on a
terminal outcome (a choice click, or Other on an already-dead entry).
A typed answer to a native card (numeric pick, or text after Other) never
reaches the click handler, so the card kept its buttons forever; the
TEXT_RESOLVED intercept now retires it with the answer.
Review finding: '[' prefix mistaken for the timeout sentinel; Other click dropped the retire entry; typed answers never rewrote the card.
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:
- TurnRunner._clarify_callback_sync: when the bounded wait returns a
sentinel (timeout, /new or run-end clear_session), schedule the retire
with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
the prose is routed as a follow-up (#111019). Lookup is on the adapter
class so MagicMock doubles cannot fabricate the method; no platform ==
SLACK special-case.
The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.
Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
- Drop the 900s "inbound liveness" INFO line from the salvage: a periodic
log heartbeat is a feature with a separate scoping call (#111211 item 3);
the existing stall watchdog already escalates when getUpdates stops.
- The polling error_callback interpolated the raw exception and left the
redaction call inside the format string, so its lines read
"Telegram network _redact_telegram_error_text(error), scheduling
reconnect: ..." and leaked unredacted text. Redact for real.
- Tests: keep two invariants (recovered wording after network errors;
clean bootstrap stays "confirmed healthy") and the transport test.
httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.
Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.
Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.
A dead background event loop tears the loop's default executor down, and
after that every inbound message was dropped inside the dedup gate with a
RuntimeError out of asyncio.to_thread — the adapter went permanently deaf
while the gateway process, websocket and service all stayed healthy.
#10849 already moved the outbound SDK calls onto an adapter-owned,
self-healing pool; this gives the inbound dedup-state flush the same
treatment, so a default-executor teardown can no longer wedge message
intake.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
Review finding on #90154: only the wildcard ("*") branch of
_iter_missed_message_backfill_candidates skipped obfuscated channels. A
channel the bot previously talked in that later lost VIEW_CHANNEL still
entered the explicit-id branch and produced a failing history read on
every startup. Apply the predicate once, after both branches build the
candidate list, and import it at module level like the other helpers.
(During the rebase the predicate was also tightened, resolved in the original
commit: is_discord_channel_obfuscated catches AttributeError (the precise
expected failure) instead of a bare Exception so a genuine attribute bug
is not hidden behind the name fallback, and its docstring states the
deliberate bias of the sentinel-name fallback (a visible channel literally
named ___hidden___ is also skipped on discord.py builds without the flag).)
Discord's Channel Obfuscation change (announced Aug 12 2026, HTTP
enforcement Nov 16 2026) dispatches channels the bot lacks VIEW_CHANNEL
on with name "___hidden___", flag 1 << 17 (CHANNEL_OBFUSCATED), and
nulled fields. Without filtering, the channel directory lists phantom
"___hidden___" entries the agent can never post to, and wildcard
missed-message backfill wastes history reads on channels that always
403.
Adds is_discord_channel_obfuscated() to gateway/platforms/helpers.py
(checks the flag bit plus the sentinel name for discord.py builds that
don't expose the new flag) and applies it at both enumeration sites.
The profile scope (HERMES_HOME override, secret scope, terminal policy) is a
contextvar bundle bound per turn. A bare threading.Thread / Timer / gRPC
callback starts with an empty context and resolves the LAUNCH profile:
- agent/title_generator.py: the auto-title thread read
auxiliary.title_generation (model, language, provider key) from the default
profile's config and billed the default's key for a secondary's session.
Spawn via agent.memory_provider.spawn_context_thread (copy_context).
- tui_gateway/session_lifecycle.py: every teardown caller is a bare Timer
(ws-orphan reap), the idle-reaper thread, atexit _shutdown_sessions, the
session.close pool RPC, superseded_by_resume or compute_host flush - none
carries a scope, yet on_session_end / commit_memory_session / agent.close ->
shutdown_memory_provider read the provider's config + credentials at call
time. Under multiplex they failed closed (tail never committed, #110622
class); on the Desktop backend a secondary's transcript went to the launch
profile's memory tenant. _finalize_session and _teardown_session now bind
_session_profile_runtime_scope(session) around those blocks, which covers
every spawn site through the single chokepoint.
- plugins/platforms/google_chat/adapter.py: Pub/Sub callbacks run on the gRPC
SubscriberClient's threads and run_coroutine_threadsafe copies THAT empty
context onto the loop task, so _dispatch_message and everything under it
(attachment cache, per-user OAuth token store via _acquire_user_chat_api ->
_load_per_user_chat_api, TTS keys, delivery ledger, bot-id cache) resolved
the launch profile. connect() captures its scope; _on_pubsub_message and
_submit_on_loop run under a per-callback copy of it.
spawn_context_thread gains a kwargs passthrough for the title thread's
callbacks.
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:
- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.
tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.
Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
DiscordAdapter.connect() set self._running = True directly instead of
calling self._mark_connected(), unlike every other platform adapter
(Telegram, WeCom, Matrix, Feishu, Google Chat, IRC, LINE, Mattermost,
ntfy, photon, raft, simplex, a2a, buzz, dingtalk, whatsapp).
_mark_connected() clears _fatal_error_code/_fatal_error_message/
_fatal_error_retryable and rewrites the runtime status file as
"connected". Bypassing it meant a transient connect failure (e.g. a
one-off DNS blip: "Cannot connect to host discord.com:443 ssl:default
[Temporary failure in name resolution]") left the platform reported as
permanently fatal in gateway_state.json / the dashboard, even after the
adapter successfully reconnected and was actively serving messages for
hours.
Reproduced under gateway.multiplex_profiles: true with a secondary
profile's Discord bot (sarathi:discord) — the bot reconnected
repeatedly ("Connected as ..." logged many times over 18+ hours) while
the dashboard kept showing the original fatal error the entire time.
Fixes#102554.
Added a regression test asserting connect() clears a previously
recorded fatal error.
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.
Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.
`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
The stack inserted the errcode table directly after _strip_reply_fallback's
return with no blank lines, which reads as if the constant belongs to the
function body and trips E305. Two blank lines restore the module-level
boundary; no behaviour change.
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.
With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).
The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.
A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.
Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.
Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.
(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.
I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.
I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.
While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.
Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.
On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.
(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.
I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.
Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.
(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)
The zero-knob test only asserted that ack_stale still trips with the knob
at 0. Because _read_websocket_health evaluates ack age before the
event-silence dimension, the test stayed green even with the
`_event_max_silence_seconds > 0` guard deleted: it never observed the
stale stamp being ignored. Assert (True, "healthy") with knob=0, a stale
stamp and a green transport first, then make the ACK stale and keep the
ack_stale assertion. Dropping the guard now fails this test.
Also correct the on_socket_event_type comment: discord.py dispatches
socket_event_type before the op-code switch but only for a non-null `t`;
heartbeat ACK frames carry `t: null`, they do not "return before" it.
The #109521 explanation (socket_event_type fires for every parsed
DISPATCH frame and is not debug-gated unlike on_socket_raw_receive;
op-11 ACKs carry no event type) was written out four times: knob init,
stamp init, the on_socket_event_type handler, and _read_websocket_health.
Keep the authoritative paragraph on the handler that actually stamps, and
point the other three sites at it so a future edit has one place to go.
Drop the math.isfinite(event_silence) half of the silence guard. Both
operands are our own perf_counter floats, so their difference cannot be
non-finite; the ack_age guard is different because _last_ack comes from
discord.py. math stays imported for the remaining finiteness checks.
`_warn_liveness_config_disabled` told operators an unusable value turns
off "the websocket liveness probe". That is true for the interval,
threshold, ack-age and latency knobs, which sit in the probe's startup
guard, but `websocket_event_max_silence_seconds` is only checked inside
`_read_websocket_health`, so ack-age/latency keep guarding. Say so, or
an operator reading the log would believe the whole watchdog is down.
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.
This adds the dispatch-side signal that was requested instead:
- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
every parsed DISPATCH frame with no debug gate (verified live against
the real received_message path: 4/4 frames fired with
enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
incident report's field-proven operator bound); 0 opts out of this
dimension ONLY — the #109782 review failure put the knob in
_start_liveness_probe's all-or-nothing guard, killing the whole
watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out
Fixes#109521
(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
The rebase moved the inbound attachment loop into SlackAdapter._append_link_unfurls,
so the nested-table hunk now lives there and is asserted directly. Drop the source-
provenance references and duplicate cell-level cases; one ragged/malformed-row test
covers raw_text, rich_text, None and unknown cell types.
Port from qwibitai/nanoclaw#3666: Slack represents a pasted table as
'table' blocks — usually nested in attachments[].blocks[], sometimes
top-level. They appear in neither the message text nor the file list,
so the agent received the sentence before the table and nothing else.
- _render_slack_table_block(): projects rows as 'cell | cell' lines,
collecting text leaves from raw_text/rich_text cell subtrees; capped
at 20k chars with a visible '[table truncated]' marker.
- Wired into all three ingestion paths: _extract_text_from_slack_blocks
(thread history + attachment-nested blocks), the live inbound
attachment loop, and _extract_additional_text_from_slack_blocks
(top-level blocks on live messages).
- _serialize_slack_blocks_for_agent skips 'table' blocks — the
allowlist drops 'rows', so it only emitted an empty husk.
send_voice built its own discord.File from the audio bytes and never ran
the size preflight, so an oversized audio attachment still burned the
doomed 413 round-trip that #50846 is about. Route it through the same
_reject_oversized_upload helper as _send_file_attachment.
The limit constant and _discord_upload_limit_bytes lived on the adapter
facade while every consumer is in adapter_media.py; the facade+sibling
layout puts topic code in the sibling, so they move there.
Discord raised the default file upload limit from 10 MiB to 20 MiB for
users, bots, webhooks and interaction responses (developer changelog,
Sep 3 2026). The 25 MiB constant here predates the preflight salvage and
never matched the platform; more importantly discord.py 2.7.1 still
reports 10 MiB via guild.filesize_limit for unboosted guilds, so the
guild-aware path under-reported the cap and rejected 10-20 MiB files
Discord now accepts. Floor the guild value at the platform default so a
stale library constant can only widen, never shrink, the preflight.
The adapter facade is ~6.6k lines; new behaviour belongs in a topical sibling per the
facade+siblings layout. expand_link_entities() now lives in telegram_entities.py and
reuses the encode/decode UTF-16 slicing the adapter already uses for entity spans.
Also: skip inlining when the anchor text already is the URL (no 'url (url)' duplication),
trim the test file to the invariants and point it at the sibling.
Image.open() was never closed; convert()/resize() return new images so the
source file object lingered until GC (a real leak on Windows where the open
handle blocks later deletion of the original). Use the context manager.
- Keep main's media_write_timeout=60s (PR's HERMES_* env var dropped per
.env-is-secrets-only policy; main already fixed the timeout half).
- Replace stdlib imghdr (removed in Python 3.13) with a magic-byte sniff.
- Exclude GIFs: JPEG conversion flattens animations to one frame.
- Fix transparent-PNG handling: RGBA hit the len(getbands())==4 branch
before the white-background composite, rendering transparency black.
- Clean up temp JPEGs after send (both single and media-group paths);
the docstring promised caller cleanup that neither call site did.
- Write temp files via tempfile default dir instead of an undefined
DEFAULT_OUTPUT_DIR (NameError at runtime in the original PR).
- Add real-Pillow regression tests incl. a sabotage-verified
white-background test.
Behind an HTTP proxy (e.g. tgapi.indevs.in) the PTB
media_write_timeout (~20s) is exceeded by raw PNGs > 1-2MB, causing
TimedOut errors on both send_photo and the send_document fallback.
Add TelegramAdapter._compress_image_to_jpeg() which converts large
PNG/raster images (>1MB) to progressive JPEG at 85% quality, with
optional resize above 1600px. Applied in send_image_file() and in the
media-group path of send_multiple_images(). Compression is a no-op for
JPEGs, small files, and non-raster formats, and falls back gracefully
if Pillow is unavailable.
Co-authored-by: user