Commit Graph

744 Commits

Author SHA1 Message Date
kshitijk4poor 6681f9ebc3 refactor(telegram): share exception-graph walk across classifiers
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.

Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
2026-08-31 11:35:24 +05:30
kshitijk4poor 8288129475 fix(telegram): bound polling drain with wall-clock deadline
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.

Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
2026-08-31 11:34:16 +05:30
Teknium d63f996a75 feat(photon): read-receipt toggle, receipt-type alias, docs
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
  messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
  coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
2026-08-30 18:37:50 -07:00
Zihan Huang 9744fc0c99 feat(photon): support iMessage read receipts 2026-08-30 18:37:50 -07:00
Teknium 5bdaea64ed feat(telegram): inline command picker — search every command and skill, no menu cap
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).

- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
  pagination logic (unit-testable without python-telegram-bot). First
  query token filters; the remainder is carried into the sent command as
  its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
  enables inline mode via BotFather /setinline) + _handle_inline_query
  with the same auth path as inline-button callbacks — unauthorized users
  get an empty list, so the installed-skill catalog is not leaked to
  arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
  message starts with /, which reaches the bot even under default privacy
  mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
  /setinline setup.
2026-08-29 20:57:55 -07:00
Teknium 83f4524b42 feat(discord): expose /plan in the native slash-command picker
Text-message /plan already works on Discord via the gateway fall-through;
this makes it discoverable in the / picker alongside /steer and /compress.
2026-08-29 19:14:15 -07:00
AideYu 5cd9c4563c fix(telegram): recover exhausted request pool 2026-08-29 18:10:03 -07:00
Teknium 74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Teknium ccc367dce0 fix(prompt)+feat(gateway): platform-hint truth pass + universal voice-bubble transcode (all 22 hints source-verified) (#97873)
* fix(prompt): platform-hint truth pass — CLI/TUI file-delivery reality (paths/URLs only, MEDIA: prints literally), CLI no-markdown verified live, Slack/Discord markdown+tables truth, shared local-cron constant

* feat(gateway): universal voice-bubble delivery — shared transcode_to_ogg_opus; telegram [[audio_as_voice]] any-format; feishu native voice; hints to new truth

* chore: delete the webui ghost hint (tombstone comment, audit-verified); sync send_voice signature pin in tts routing test
2026-08-29 05:57:13 -07:00
kshitijk4poor 2f01ec9fa4 fix(teams): allowlist-gate BF attachment auth, stream downloads under media cap, lock token refresh
Second follow-up for salvaged PR #94547, folding in review findings from the
duplicate-PR cluster (#47015, #55054, #58476, #72977, #73685 all fix the
same 401) and the sweeper review of #73685:

- Replace the dot-anchored suffix predicate with exact-match against the
  existing _ALLOWED_TEAMS_SERVICE_HOSTS allowlist (two of the five
  duplicate PRs converged on this independently). Any Azure customer can
  register <name>.trafficmanager.net profiles, so suffix matching was not
  safe. Also requires https on the default port — :444 on an allowlisted
  host no longer receives the bearer (sweeper finding on #73685).
- Stream _fetch_attachment_bytes through _read_httpx_body_with_limit
  instead of buffering response.content — the shared inbound media cap now
  applies to authenticated downloads too (sweeper finding: a lying
  Content-Length must not OOM the gateway).
- Serialize token refresh with a lazily-bound asyncio.Lock so concurrent
  attachments share one STS POST (review finding on #94547).
- Token expiry now uses time.monotonic() (from #55054) — wall-clock jumps
  can't extend a stale token.
- Tests updated: exact-allowlist predicate (lookalike/subdomain/port/scheme
  negatives), streaming fake client, and a concurrent-cold-cache lock test.
  Mutation-checked: suffix match, silent drop, no-lock, and unbounded
  buffer each fail a test.
2026-08-29 11:04:38 +05:30
kshitijk4poor c23d40af17 fix(teams): dot-anchor Bot Framework host check, log dropped BF images, tests
Follow-up for salvaged PR #94547 (Sibbern's Bot Framework attachment auth):

- The host check used bare endswith('trafficmanager.net') /
  endswith('botframework.com'), which matched attacker lookalike hosts
  (evil-trafficmanager.net) — sending the bot's bearer token off-platform.
  Same threat model _ALLOWED_TEAMS_SERVICE_HOSTS already documents. Now a
  single dot-anchored predicate, _is_botframework_attachment_host(),
  used by both _fetch_attachment_bytes and the _on_message image branch
  (was copy-pasted in two places).
- BF images whose bytes fail image validation were silently dropped with
  no log (cache_media_bytes returns None; old path logged). Add the
  missing else-warning, mirroring the document branch.
- Record cached_m.media_type instead of the raw content_type so the
  MessageEvent MIME matches what was actually cached.
- Init _bf_token_cache in __init__ (was masked by getattr).
- Tests: host predicate dot-anchoring (attacker lookalikes blocked),
  BF routing vs generic helper, bearer attach + attacker-host block end
  to end, token acquire + cache reuse, token-failure degradation, and
  the silent-drop regression guard. Mutation-checked: reverting the
  dot-anchor, the else-warning, or the token cache each fails a test.
2026-08-29 11:04:38 +05:30
Jacob S. 0eff6bc200 fix(teams): authenticate Bot Framework connector attachment downloads
Inline/pasted images in Teams arrive with a contentUrl on
smba.trafficmanager.net (/v3/attachments/...). Unlike file uploads
(pre-authenticated SharePoint downloadUrls), these connector URLs
require the bot's own bearer token; fetching them anonymously fails
with 401 Unauthorized and the image is silently dropped.

- add _get_botframework_token(): client-credentials token for
  https://api.botframework.com/.default, cached until ~5 min before
  expiry
- _fetch_attachment_bytes(): attach the token when the attachment
  host is *.trafficmanager.net / *.botframework.com; SharePoint and
  other URLs remain auth-free as before
- route image/* attachments with Bot Framework contentUrls through
  the authenticated fetch instead of cache_image_from_url (which
  sends no Authorization header)

SSRF guards unchanged. Verified on a live Teams personal-scope bot:
pasted images previously logged '[teams] Failed to cache image
attachment: 401 Unauthorized' and now cache and deliver correctly.
2026-08-29 11:04:38 +05:30
Teknium 3340bbbdad feat(a2a): client tools config-gated — disabled unless enabled (−561 tok/call on unconfigured installs) (#97421)
* feat(a2a): outbound client tools are config-gated — served only when a2a_agents configured, inbound platform enabled, or A2A_PORT set (-561 tok/call on unconfigured installs)

* ci: retrigger after runner startup_failure on rerun attempt
2026-08-28 15:08:08 -07:00
Teknium a641644f12 fix(paths): display_hermes_home renders POSIX separators on Windows — kills ~/AppData\Local\hermes chimeras in tool schemas and user-facing messages (#97137) 2026-08-28 05:45:00 -07:00
Teknium 34393c32aa feat(plugins): wire plugin platform handlers into a2a, buzz, and qqbot adapters
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
2026-08-27 07:51:37 -07:00
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
teknium1 c96f830252 feat(plugins): let plugins register Telegram PTB handlers via ctx.register_telegram_handler
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.

Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
2026-08-27 07:51:37 -07:00
wansui 9faa953cd1 fix(wecom): eliminate duplicate + split bubbles in native streaming
Fixes two related native-streaming bubble defects surfaced in production:

1. Duplicate bubble on long turns — when keep-alive already refreshed the
   6-min reply window, the Layer-2 clock fallback still declined the finalize
   frame and forced a proactive send(), duplicating the message. Skip the
   clock fallback while keep-alive is active; intermediate-frame failures are
   now fully fire-and-forget (only a failed FINAL frame falls back to send()).

2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
   a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
      notice) called _reset_segment_state(), clearing the cumulative
      _accumulated so the next delta + finalize frame carried only the few
      chars accumulated after the reset. Native streaming now skips that
      reset (commentary still posts as its own message via send()).
   b) The adapter-side _BlockChunker.update() 'only grow' guard silently
      dropped any cumulative snapshot shorter than its high-water mark, so
      after a baseline reset the leading characters were stranded before
      _emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
      layer entirely; intermediate frames are pure identity-dedup, matching
      the fire-and-forget model.

Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.

Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).
2026-08-27 07:33:36 -07:00
wansui 2ecb544551 feat(wecom): native reply streaming (per-turn isolation, dedup-safe delivery, interaction boundaries)
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.

Includes the machinery intrinsic to native streaming on WeCom:

- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
  frames update it, a finalize frame closes it; native-streaming adapters are
  let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
  long-connection mode has no documented edit-rate limit); an adapter-level
  frame cap is retained. (Early builds gated frames behind a char throttle;
  removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
  isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
  (control vs normal) plus a per-chat token bucket to stay under WeCom's rate
  limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
  delivered the moment it is emitted; failures logged, not re-sent; delivery
  marked once per turn), per-req_id reply queue with ack tracking, and the
  timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
  846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
  the prompt is the last thing on screen and never traps a lingering bubble;
  eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
  stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
  turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
  messages; image+text double-callback merged into one turn.

Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.
2026-08-27 07:33:36 -07:00
Ben Barclay 5521265def fix(slack): coerce string unfurl knobs on the native plane
Relay-plane parity: hermes config set / Railway persist YAML booleans
as strings, and _slack_unfurl_kwargs silently dropped them — so
'unfurl_links: "false"' was a no-op on native while working on relay.
Coerce recognized string booleans exactly as _slack_unfurl_hints does;
unrecognized values still drop so junk config keeps Slack's default
instead of accidentally suppressing previews.

Replaces test_send_ignores_non_boolean_unfurl_options (which froze the
dropped-string behavior) with coercion + junk-drop tests.
2026-08-27 15:33:32 +10:00
Will Lynas 91dbd7a6f8 fix(slack): preserve unfurl controls during streaming 2026-08-27 15:33:32 +10:00
Will Lynas b91845c2a0 fix(slack): honor unfurl controls for media captions 2026-08-27 15:33:32 +10:00
Andrew Bennett aeaa0c784e feat(slack): add link unfurl controls 2026-08-27 15:33:32 +10:00
Teknium 7a7a371c59 fix: harden claim-release guard for bare test doubles; repoint source-pinning test at the impl
The wrapper now getattr-defaults _processed_message_ts (object.__new__
adapters in sibling suites lack it), and the reaction-guard source pin
reads _handle_slack_message_impl where the production expression lives.
2026-08-26 15:54:53 -07:00
Teknium 39a5838f07 fix(slack): release a failed handler's fresh ts claim so the turn isn't swallowed
Follow-up for the #95417 salvage, addressing the review finding: the entry
claim closes the unfurl race but a handler that raises mid-enrichment would
hold the claim forever — neither a Slack retry nor a user edit could ever
re-drive the message. _handle_slack_message is now a thin guard around the
impl that releases only claims taken by the failed invocation itself, with a
warning log so swallowed turns are traceable. Pre-existing claims from a
successful turn are never released. Two failure-path tests pin both sides.
2026-08-26 15:54:53 -07:00
Richard Hojun Jang 708f84c477 fix(slack): claim message ts before enrichment so link unfurls can't duplicate a turn
Slack emits `message_changed` for a link unfurl carrying a DIFFERENT event ts
than the original message. That ts legitimately misses the `_dedup` check, so
`_processed_message_ts` is the only guard against it becoming a second user
turn -- but it was only populated at the END of `_handle_slack_message`, after
thread context, permalink resolution and file downloads had all awaited.

An unfurl landing inside that window found the guard empty and was promoted to
a duplicate turn: a spurious "Interrupting current task" banner plus the same
answer posted twice.

Production capture (adminbot, 2026-08-22 02:23:30-31Z, channel C0BF1EYUA9H):

  02:23:30.718  message      ts=1787365409.908499  dedup_hit=False
  02:23:30.737  app_mention  ts=1787365409.908499  dedup_hit=True
  02:23:31.675  message      ts=1787365411.012100  dedup_hit=False   <- leaked
                subtype=message_changed

The original copy was still resolving two Slack permalinks when the unfurl
arrived 957ms later.

Claim the message ts once every filter has passed and the event is certain to
be delivered, before the slow enrichment awaits. Claiming any earlier (right
after the dedup check) also claims messages the handler then discards, which
breaks summoning the bot by editing "@bot" into a previously ignored message
(tests/gateway/test_slack.py::TestMessageRouting::
test_message_edit_with_new_mention_processed).

Eviction logic is extracted to `_remember_processed_message_ts` so both call
sites share one bounded implementation.
2026-08-26 15:54:53 -07:00
Nikita Barkov 2e80d7fa05 fix(slack): keep the resolved proxy on bolt's per-request client
slack_bolt builds a fresh AsyncWebClient for every inbound request and
copies proxy=app.client.proxy into its constructor, where slack_sdk reads
a None/blank proxy *argument* as "unspecified" and reloads HTTP(S)_PROXY
from the environment. aiohttp then treats that env value as an explicit
proxy and skips its own NO_PROXY check, so the adapter's resolved decision
to go direct - a NO_PROXY bypass, or a proxy scheme aiohttp cannot use -
holds on every client except the one authorization spends on auth.test.

The failure looks like a healthy bot: Socket Mode connects, outbound sends
keep working, and every inbound event is rejected with "Failed to authorize
with the given token" - forever, since a failed auth_test_result is not
cached and never retried differently.

Re-apply the resolved proxy through AsyncApp(before_authorize=...), which
bolt inserts before the authorization middleware: the request-scoped client
already exists there and has not been used yet. Assigning the attribute
post-construction is the only way to express "no proxy" to slack_sdk.

Co-authored-by: Junie <junie@jetbrains.com>
2026-08-26 10:35:05 -07:00
Nikita Barkov bac960e23d fix(slack): stop injecting thread roots as reply context 2026-08-26 10:26:00 -07:00
Nikita Barkov 5538bd1f93 fix(slack): prevent duplicate rich-text message content
Slack sends an authored message twice: flat in `event.text` and structurally
in `event.blocks`. The blocks are rendered so quoted and forwarded content is
not lost, and whatever the render carries beyond the flat text is appended to
the message. That comparison had several ways to fail on the *same* sentence,
each of which showed the author their own words a second time:

1. HTML entities — the flat copy escapes `&`/`<`/`>` while `blocks[].link.url`
   stays raw, so any link with query parameters (every "Copy link" on a
   thread) mismatched.
2. Permalink unfurls — the live inbound path skips `is_msg_unfurl`
   attachments, thread/parent hydration did not, so the linked message's body
   was appended again.
3. The Block Kit dump — it serialized the authored `rich_text` alongside the
   UI blocks it exists for, and its allowlist drops `url`, so the sentence
   reappeared with every link removed.
4. Unknown inline elements — the renderer knew eight types and silently
   dropped the rest. A pasted message permalink arrives as `message_mention`,
   so the link vanished from the render and the sides stopped comparing equal.
5. `message_mention` without a url — `url` is optional on that element while
   `channel_id` and `message_ts` are not, so the element rendered as nothing
   and the sentence came back with a blank in the link's place.
6. `date` elements — `fallback` and `url` are both optional, and the flat
   `<!date^…>` form was never read down to what the rich text renders.
7. Labelled mentions — Slack may attach a label (`<@U…|name>`,
   `<#C…|general>`, `<!subteam^S…|@marketing>`, `<!here|@here>`) in the flat
   text while the blocks carry the bare id. The bot's own mention is one of
   these, and stripping only its bare form left it in the flat copy.
8. Autolink schemes — only `https` and `mailto` were matched, so a `tel:` link
   kept its angle brackets and mismatched too.

Unknown inline types are now read by their `url`/`text`/`fallback` so a type
Slack adds later still renders, and `team`, `color` and a fallback-less `date`
render into the flat form Slack sends. Every field is read as a string or
not at all: Block Kit carries text as an object in many places, and a
non-string one reaches the renderer's `str.join` and raises there, which
costs the whole message. `channel_id` and `message_ts` are the
permalink's own components, so a url-less `message_mention` renders the
permalink's tail; the workspace host and the thread query cannot be rebuilt
from the element, so a permalink on either side is reduced to that same tail.
Canonicalization is used for matching only -- the authored text still reaches
the agent verbatim, so a mistake here can cost an unrendered element, never an
altered or missing message.

An element carrying neither a url nor a label still renders as nothing, and a
message containing one is still appended twice. Suppressing such a render was
tried and is worse: an app message whose body lives only in the blocks
disappears, and a forwarded quote is dropped. Genuinely additional content --
quotes, lists, code blocks, attachments, interactive bot blocks -- is
unaffected throughout.

Tests cover both merge sites (live inbound and thread hydration) and the
negative cases.
2026-08-26 10:10:51 -07:00
xxxigm 6c1bfff65c fix(telegram): wait for reconnect before failing send as Not connected
A short Telegram drop used to fail the final reply immediately. The
answer then sat in the delivery ledger until the next gateway boot.
Wait up to 15s for the bot (or a replacement adapter) so a brief
blip delivers now, matching QQBot.
2026-08-25 12:07:36 +05:30
xxxigm f84f94b494 fix(teams): do not call App() when the SDK was never bound
find_spec("microsoft_teams") can be true from sibling namespace packages
while App is still None, so a failed lazy-install crashed connect with
'NoneType' object is not callable instead of a missing-SDK error.
Probe microsoft_teams.apps via the parent first — a dotted find_spec
raises ModuleNotFoundError on 3.11 when the namespace is absent.
2026-08-25 12:07:33 +05:30
kshitijk4poor 111d809562 refactor(telegram): migrate _await_with_thread_deadline onto agent.deadline.run_bounded_async (#85125 2f)
The adapter's private thread-deadline helper was the ancestor of the
unified deadline layer's run_bounded_async (#85147 was extracted from
it, plus the caller-cancellation leak fix the original still lacked).
Consolidate: the helper body becomes a thin wrapper mapping
BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call
sites (the PTB retry ladder) expect. ~90 duplicated lines die, along
with the adapter-local copies of the abandon-cleanup runner and the
blocked-loop faulthandler diagnostics (both live in agent/deadline.py).

Everything the call sites rely on is preserved by the unified layer:
- thread-timer deadline that survives a blocked event loop (#63309)
- abandonment of cancellation-shielded tasks (PTB/httpcore anyio init)
- detached best-effort on_abandon cleanup (no httpx pool leak per retry)
- off-loop stack dump when the loop never processes the expiry
Plus one behavior IMPROVEMENT inherited from the shared copy: a caller
cancelling the wrapper no longer leaks the inner task unobserved (the
telegram original had that leak; the extraction fixed it).

test_telegram_init_deadline.py: the #63309 diagnostics probe now pins
the shared layer's dump hook (label "telegram-init") — same contract,
new seam. Wedge + cleanup-crash tests pass unchanged.
2026-08-24 17:04:55 +05:30
aniruddhaadak80 7befc1d2dd fix(gateway): route platform authorization reads through the profile secret scope
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):

- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
  WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
  GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
  GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
  hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
  _sig_secret helper; empty scoped group list previously meant "drop all
  groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
  WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
  while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
  startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
  though its sibling dm/group reads already used the scoped _getenv

All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.

Fixes #93522
2026-08-24 03:20:06 -07:00
chelsealong d7e4204e77 fix(gateway): scope multiplex-profile authorization reads (weixin/yuanbao/wecom)
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.

Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.

Fixes #93522.
2026-08-24 03:20:06 -07:00
Finn763 bf15b050b1 fix(telegram): watchdog silent long-poll death via last getUpdates progress (#92991) 2026-08-23 19:26:41 -07:00
kshitij 261a4efb90 Merge pull request #92214 from kshitijk4poor/discord-picker-constants-followup
refactor(discord): derive picker capacity constants; drop last bare 25-option literal
2026-08-22 16:35:15 +05:30
Kshitij Kapoor 1db0a7d825 refactor(telegram): share the flood-wait cap between send and edit paths
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
2026-08-22 16:26:51 +05:30
kshitijk4poor 9154421b11 refactor(discord): use _DISCORD_SELECT_MAX_OPTIONS in ChoicePickerView
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
2026-08-22 16:23:39 +05:30
kshitijk4poor c925cc8eb8 refactor(discord): derive model-select capacity from the row/option constants
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
2026-08-22 16:23:39 +05:30
HexLab98 a444b673ad fix(telegram): fail closed on long send-path flood waits
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
2026-08-22 15:25:30 +05:30
kshitijk4poor 209e2ebdda refactor(discord): derive model-select counts, name the 25-option cap
Review follow-ups for the salvaged picker partitioning:
- Drop the _model_chunks stash — it went stale when navigating to a
  provider with an empty model list (early return skipped the
  reassignment), producing a wrong 'N more available' count. Derive
  shown = min(len(models), 75) directly instead.
- Add _DISCORD_SELECT_MAX_OPTIONS / _DISCORD_SELECT_MAX_ROWS constants
  per the file's named-limits convention; replaces 4 bare literals.
- Remove dead total_rows variable.
2026-08-22 14:55:04 +05:30
Epic ab3e2f563b fix(discord): render provider model lists >25 options across multiple select menus
The Discord /model picker built a single discord.ui.Select filled with
models[:25], silently dropping any models beyond the first 25. Discord
caps a single select at 25 options but allows up to 5 component rows, so
partition the list across up to 3 select menus (25 each; Back/Cancel use
the other 2) instead of truncating.

This fixes providers like Nous (curated list + Portal recommendations
exceed 25) whose tail — including free-tier :free Portal picks — was
previously clipped on Discord while showing fine in the Portal UI / CLI.

- _build_model_select: slice into <=25-option chunks, one select per chunk
  (custom_id model_model_select_<i>), all via _on_model_selected. Multi-row
  menus get a (n/total) placeholder suffix.
- _on_provider_selected: 'N more available' count reflects models actually
  rendered across the partitioned menus.
- Add regression test covering the 37-model Nous case (no truncation/dupes,
  per-menu 25 cap holds).
2026-08-22 14:55:04 +05:30
Gille e9a7c7aa4d fix(telegram): omit topic routing from rich edits 2026-08-21 19:07:14 -07:00
kshitijk4poor bad2ed866c fix(telegram): honor the direct-messages-topic alias in the fresh-final gate
prefers_fresh_final_streaming read only the raw direct_messages_topic_id
key; the adapter's canonical accessor _metadata_direct_messages_topic_id
also accepts the documented telegram_direct_messages_topic_id alias
(treated as equivalent in gateway/delivery.py), so an alias-only lane
would still flatten tables. Route the gate through the accessor and pin
the alias with a regression (mutation-checked: raw-key gate fails it).
Also reshape the happy-path endpoint assertion into the actual invariant
(sendRichMessage present, no rich draft frames) instead of a frozen call
list. Surfaced during review of PR #91436.
2026-08-22 03:31:54 +05:30
HexLab98 194729c95f fix(telegram): keep DM-topic tables on sendRichMessage when drafts degrade
#91241 stopped root-DM tables collapsing to bullets by keeping native
draft transport when rich_drafts is off. Private Telegram topics still
reject sendMessageDraft (string thread ids, forum-style thread fields),
so the stream consumer falls back to edit-in-place. Telegram then
rejects a rich edit of that plain MarkdownV2 preview and format_message
permanently rewrites pipe tables into bullet lists — the remaining
report after that merge.

Route drafts through the same integer topic kwargs as send(), and on
that degraded topic path prefer a fresh sendRichMessage (then delete
the preview) instead of the table-to-bullets formatter.
2026-08-22 03:31:54 +05:30
kshitijk4poor 3841910cee fix(telegram): widen cancellation-shielded stop to sibling paths
The network-error reconnect path (PR #91524) was the only site converted
from asyncio.wait_for to _await_with_thread_deadline.  The same
cancellation-shielding vulnerability exists at two more updater.stop()
sites:

- Conflict-retry path: asyncio.wait_for could hang forever if PTB/AnyIO
  cleanup swallowed CancelledError, stalling the conflict-retry ladder.
  Now uses _await_with_thread_deadline and escalates to fatal on timeout
  (same reasoning: cannot safely reuse an Updater whose lifecycle lock
  may still be held).

- Conflict-exhausted fatal path: asyncio.wait_for could hang before the
  fatal notification fired.  Now uses _await_with_thread_deadline; the
  timeout handler already proceeds to fatal notify, so no behavior change
  beyond the deadline mechanism.

All three asyncio.wait_for(updater.stop()) sites now use the
thread-deadline helper consistently.
2026-08-22 03:20:48 +05:30
Good Chang 9e36774d77 fix(telegram): rebuild after cancellation-shielded stop
Use the existing wall-clock deadline helper for updater.stop() during network recovery. If PTB cleanup remains cancellation-shielded past the deadline, escalate to retryable fatal recovery so the runner builds a fresh adapter instead of calling start_polling() while the old Updater may still hold its lifecycle lock.

Add regression coverage with stop() swallowing cancellation while holding the same lock start_polling() needs, and verify the old Updater is never reused.
2026-08-22 03:20:48 +05:30
Gille 790c850144 fix(telegram): preserve rich finals after DM drafts 2026-08-20 21:58:18 -07:00
liuhao1024 fbca706789 fix(telegram): log the first confirmed getUpdates progress per generation
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).

Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.

Fixes #90504
2026-08-20 11:28:26 +05:30
Gille cefeed4ca8 fix(a2a): expose schemas through tool describe 2026-08-20 10:41:12 +05:30