#106167 widened the base contract to SendResult but left six native-batch
overrides (Discord, Email, Matrix, Mattermost, Slack, Telegram) returning
None, which (a) meant a media-only reply on those platforms still reported
FAILURE because _record_delivery(None) records nothing, and (b) produced new
`ty` invalid-method-override diagnostics against the widened base (#106192).
Make the contract honest instead of annotating it Optional: each override
now rolls its batches (and any per-image fallback) into one SendResult, so
the turn-outcome accounting works on every platform, not just Signal and
the base loop. The "legacy overrides return None" comment in
_send_image_batch goes away with the legacy.
ty on the 8 touched files: origin/main 241 diagnostics / 20 override,
this branch 241 / 20 — byte-identical diagnostic set; the intermediate
`-> SendResult` head without this commit had 246 / 25.
Refs #106192
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
The redelivery hook keyed on the canonical flood_control:<seconds> result, so
two real refusals slipped past it and armed no timer, leaving the reply for the
next restart. A short wait that outlived the send retries raised instead of
failing closed, and an edit refused again after its inline wait returned the
platform's raw text. Both now fail closed canonically, the second carrying the
new delay rather than the first refusal's.
The ledger also accepts a row still carrying the platform's own wording, so a
row persisted by an unnormalized path is dated from the delay it states instead
of the generic default. Without that a boot sweep claims it at once and spends
its one attempt inside the penalty. Matching requires the flood wording as well
as a delay, so an unrelated retry suggestion is never read as a flood.
Six new assertions fail without this change. 770 passed across the ledger,
Telegram, send-retry and queued suites.
Two corrections on top of the mention-preservation pick:
- Stripping our own trigger handle from every group message was replaced by never stripping
it in groups, which broke the pre-existing contract that `@bot 2` answers a pending clarify
prompt and `@bot ok` approves (the intercepts match the exact stripped text). Strip only when
no other participant is named; when other bots/users are mentioned the text is kept verbatim
so `@research_bot , @ops_bot …` no longer arrives as `, @ops_bot …`.
- The addressing block carried a per-message fact ("explicitly mentions you: yes/no") inside
`channel_prompt`, which is part of the cached-agent signature: a mention turn followed by a
reply turn rebuilt the AIAgent (prompt-cache miss) every time the shape flipped. Keep only
the username line, byte-identical for the life of the session.
Tests rewritten to pin both contracts (sole-addressee stripping; identical channel_prompt and
agent signature across a mention turn and a reply turn).
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Raw yaml.safe_load of config.yaml is only legal in owner modules; the
write-back round-trip uses the sanctioned raw primitive. Behavior-identical
(same file, same parse, {} on missing).
Follow-ups on the #101406 salvage (#101391):
- The issue's measured incident was a healthy connect followed by silent
polling loss; nothing republished for 11 h. `_schedule_polling_recovery`
(the single entry to the recovery ladder) now publishes "retrying" while
the adapter is running; `_record_polling_progress` already flips it back.
- `BasePlatformAdapter.send_path_degraded` property (default False) replaces
the runner's `getattr(adapter, "_send_path_degraded")` reach into a
plugin-private attribute; Telegram overrides it. `_mark_connected()` reads
the property instead of taking a `degraded=` kwarg, and `_mark_degraded()`
+ `DEGRADED_STATUS_MESSAGE` replace three copies of the same write/string.
- The startup connect stamp (sibling of the reconnect stamp) honours the
same flag.
- Tests: mid-session death publishes retrying; pre-connect death does not;
progress never flips while not running / fatal; runner reconnect stamp
honours the adapter flag.
connect() intentionally returns True when Telegram polling starts in
degraded mode (or a reconnect's require_progress=False skips the strict
readiness gate), so background recovery can retry without failing
gateway startup. But _mark_connected() published platform_state
"connected" unconditionally, and the reconnect watcher in gateway/run.py
stamped "connected" again right after -- so gateway_state.json was
indistinguishable from a healthy adapter for as long as recovery took
(observed ~11h on one seat).
_send_path_degraded already tracks exactly this (set at polling
generation start, cleared on the first confirmed getUpdates round-trip),
so:
- BasePlatformAdapter._mark_connected() takes a `degraded` flag and
publishes "retrying" instead of "connected" when set.
- TelegramAdapter passes its current _send_path_degraded into
_mark_connected().
- The gateway/run.py reconnect watcher checks the same flag before
overwriting the adapter's own status write.
- _record_polling_progress() republishes "connected" the moment
polling actually proves progress, so the state does not stay wedged
at "retrying" until the next disconnect.
Fixes#101391
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.
- `_make_adapter_auth_check`: for the shared primary adapter under
multiplex, mirror the inbound message path exactly — stamp the
`profile_routes` match so the routed profile's pairing store is
consulted, and authorize under the TRANSPORT home via
`_is_user_authorized_for_source` (same split as
`_make_default_profile_message_handler`, 2afed50863). A rejected route
fails closed like the ingress gate. Retain the receiving adapter as
`_transport_adapter_ref` so config.yaml policy reads stay on it.
Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
`thread_id` as keywords only when set, so legacy 3-positional callbacks
keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
introspection class — fall back to the injected `gateway_runner` and
the adapter's owner profile.
Fixes#86296Fixes#92840
Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
_is_callback_user_authorized resolved the gateway's auth chain through
_message_handler.__self__. For a secondary multiplexed adapter the
message handler is a per-profile closure with no __self__, so the
introspection silently fell through to the env-only fallback -- which
knows nothing about config allowlists or the pairing store, denying
every button caller on that profile (fail-closed, but wrong).
Prefer the auth callback GatewayRunner already injects at connection
time via set_authorization_check (registered for primary and multiplexed
adapters alike, delegating to the full _is_user_authorized chain), and
keep the introspection plus env fallback for adapters wired without it.
Same resolution pattern the admin-tier gate uses.
Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
The "Connecting to Telegram (attempt N/8)…" line logs at WARNING and
reaches the gateway's default stderr handler, but the matching
"Connected to Telegram (… mode)" line was INFO and went to the log file
only. A healthy startup therefore looked permanently hung at
"attempt 1/8" on the terminal — the logging-illusion half of #90835.
Promote the success line to WARNING so both sides of the connect
transition share the same console sink; a genuine hang is now the
absence of the success line. Adds an AST-level regression test pinning
the level pairing. Sibling adapters (homeassistant, wecom) log both
sides at INFO, so they don't have this asymmetry.
Fixes#90835
Prevent the dedicated getUpdates pool from reusing server-closed connections and replace a polling HTTP client left open after a timed-out CLOSE-WAIT drain. Keep the general Bot API pool reusable so concurrent sends and edits are unaffected. Add regression coverage for both transport limits and stale-client replacement. Fixes#87057
After updater.stop() times out, HTTPXRequest.initialize() is a no-op unless
the client is already closed, so start_polling reused the wedged socket and
the gateway stayed alive but deaf. Rebuild the polling client after a hung
drain, watch getUpdates I/O independently of get_me(), and enable TCP
keepalive on the fallback transport.
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.
Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.
Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).
- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
pagination logic (unit-testable without python-telegram-bot). First
query token filters; the remainder is carried into the sent command as
its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
enables inline mode via BotFather /setinline) + _handle_inline_query
with the same auth path as inline-button callbacks — unauthorized users
get an empty list, so the installed-skill catalog is not leaked to
arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
message starts with /, which reaches the bot even under default privacy
mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
/setinline setup.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
A short Telegram drop used to fail the final reply immediately. The
answer then sat in the delivery ledger until the next gateway boot.
Wait up to 15s for the bot (or a replacement adapter) so a brief
blip delivers now, matching QQBot.
The adapter's private thread-deadline helper was the ancestor of the
unified deadline layer's run_bounded_async (#85147 was extracted from
it, plus the caller-cancellation leak fix the original still lacked).
Consolidate: the helper body becomes a thin wrapper mapping
BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call
sites (the PTB retry ladder) expect. ~90 duplicated lines die, along
with the adapter-local copies of the abandon-cleanup runner and the
blocked-loop faulthandler diagnostics (both live in agent/deadline.py).
Everything the call sites rely on is preserved by the unified layer:
- thread-timer deadline that survives a blocked event loop (#63309)
- abandonment of cancellation-shielded tasks (PTB/httpcore anyio init)
- detached best-effort on_abandon cleanup (no httpx pool leak per retry)
- off-loop stack dump when the loop never processes the expiry
Plus one behavior IMPROVEMENT inherited from the shared copy: a caller
cancelling the wrapper no longer leaks the inner task unobserved (the
telegram original had that leak; the extraction fixed it).
test_telegram_init_deadline.py: the #63309 diagnostics probe now pins
the shared layer's dump hook (label "telegram-init") — same contract,
new seam. Wedge + cleanup-crash tests pass unchanged.
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
prefers_fresh_final_streaming read only the raw direct_messages_topic_id
key; the adapter's canonical accessor _metadata_direct_messages_topic_id
also accepts the documented telegram_direct_messages_topic_id alias
(treated as equivalent in gateway/delivery.py), so an alias-only lane
would still flatten tables. Route the gate through the accessor and pin
the alias with a regression (mutation-checked: raw-key gate fails it).
Also reshape the happy-path endpoint assertion into the actual invariant
(sendRichMessage present, no rich draft frames) instead of a frozen call
list. Surfaced during review of PR #91436.
#91241 stopped root-DM tables collapsing to bullets by keeping native
draft transport when rich_drafts is off. Private Telegram topics still
reject sendMessageDraft (string thread ids, forum-style thread fields),
so the stream consumer falls back to edit-in-place. Telegram then
rejects a rich edit of that plain MarkdownV2 preview and format_message
permanently rewrites pipe tables into bullet lists — the remaining
report after that merge.
Route drafts through the same integer topic kwargs as send(), and on
that degraded topic path prefer a fresh sendRichMessage (then delete
the preview) instead of the table-to-bullets formatter.
The network-error reconnect path (PR #91524) was the only site converted
from asyncio.wait_for to _await_with_thread_deadline. The same
cancellation-shielding vulnerability exists at two more updater.stop()
sites:
- Conflict-retry path: asyncio.wait_for could hang forever if PTB/AnyIO
cleanup swallowed CancelledError, stalling the conflict-retry ladder.
Now uses _await_with_thread_deadline and escalates to fatal on timeout
(same reasoning: cannot safely reuse an Updater whose lifecycle lock
may still be held).
- Conflict-exhausted fatal path: asyncio.wait_for could hang before the
fatal notification fired. Now uses _await_with_thread_deadline; the
timeout handler already proceeds to fatal notify, so no behavior change
beyond the deadline mechanism.
All three asyncio.wait_for(updater.stop()) sites now use the
thread-deadline helper consistently.
Use the existing wall-clock deadline helper for updater.stop() during network recovery. If PTB cleanup remains cancellation-shielded past the deadline, escalate to retryable fatal recovery so the runner builds a fresh adapter instead of calling start_polling() while the old Updater may still hold its lifecycle lock.
Add regression coverage with stop() swallowing cancellation while holding the same lock start_polling() needs, and verify the old Updater is never reused.
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).
Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.
Fixes#90504