Commit Graph

14 Commits

Author SHA1 Message Date
Faych 9bd7a37a77 fix: _extract_retry_after returns None for non-retryable errors (#394)
* fix: _extract_retry_after returns None for non-retryable errors

* fix: update _extract_retry_after to handle generic transient errors with default retry delay

* fix: extend base _non_retryable_patterns in channel subclasses

* fix(channels): merge structured SDK error check into non-retryable step

* fix(channels): decouple status code and SDK error code extraction in retry logic

- Independently evaluate HTTP status codes and structured SDK error codes
- Fix misleading doc comments for Feishu and DingTalk patterns
- Remove redundant try-except AttributeError on getattr with default
- Expand test coverage for dual-signal matrix and header parsing

* refactor(channels): simplify status code and SDK error extraction via channel overrides

- Handle httpx and aiohttp exceptions in base Channel class
- Override _extract_status_code and _extract_sdk_error_code in SlackChannel and DiscordChannel
- Replace mock exception types in comprehensive test suite with real httpx and aiohttp errors
- Add dedicated Slack and Discord retry error extraction test suites

* fix(channels): clean up Slack and Discord error code extraction

- Remove defensive string checks and attribute guards in SlackChannel
- Directly access exc.response.status_code and exc.response.get('error') in SlackChannel
- Remove unnecessary _extract_sdk_error_code override in DiscordChannel
- Use real SlackApiError, SlackResponse, and discord.HTTPException in unit tests

* refactor: reorder retry logic to prioritize non-retryable checks, remove aiohttp dependency, and clean up exception handling in base and channel modules.

* test(channels): skip Slack/Discord retry tests when the SDK extra is absent

The retry-extraction tests build real SlackApiError / discord.HTTPException
objects, but slack-sdk and discord.py are optional extras that the dev
dependency group does not install. Under CI's `uv sync --dev` all nine
tests failed with ModuleNotFoundError raised from the channel override.

Gate both test classes with skipif(find_spec(...) is None) so the suite is
green without the extras and the tests still run wherever they are installed.

* refactor(channels): replace retry-delay lookup with _extract_retry_delay

_extract_retry_after still read the server-supplied delay by probing
exc.retry_after and exc.response.headers via getattr/hasattr, the last
remnant of the pattern the extractors moved away from. Replace both steps
with one overridable hook, _extract_retry_delay, implemented against the
real exception types:

- base: httpx.HTTPStatusError -> Retry-After header (httpx.Headers is
  case-insensitive; HTTP-date form remains unsupported)
- SlackChannel: SlackApiError -> Retry-After, matched case-insensitively
  because SlackResponse.headers is a plain dict whose casing depends on the
  HTTP client (same approach as slack_sdk's RateLimitErrorRetryHandler)
- TelegramChannel: telegram.error.RetryAfter.retry_after (int, or timedelta
  under PTB_TIMEDELTA)
- DiscordChannel: discord.RateLimited.retry_after, which the old duck-typed
  getattr matched and would otherwise have been lost

Drop the isinstance(retry, bool) and val >= 0 guards; no SDK produces those.
Delete the test that asserted the duck-typed attribute; add real-object tests
for each override, guarded like the existing SDK-dependent classes.

* ci: install the all-channels extra so SDK-dependent channel tests run

The Slack, Discord, and Telegram retry tests build real SDK exception
objects and are skipped when the SDK is absent. CI only ran `uv sync --dev`,
so those tests never executed there. Install the existing all-channels
extra alongside the dev group; the skipif guards remain for lean local runs.

* fix(channels): honor HTTP-date Retry-After and tolerate malformed values

RFC 9110 allows Retry-After as either delay-seconds or an HTTP-date. The
httpx path treated a date as unparseable and fell back to the 1.0 s default,
so a 503 asking for a specific wait was retried too early. Add
Channel._parse_retry_after, which returns delay-seconds as-is and converts
an HTTP-date to the non-negative seconds until it (tz-less dates read as
UTC).

SlackChannel used a bare float() on the header. A non-numeric value raised
inside the retry predicate, which escapes retry_async and drops the chunk
instead of retrying. Route Slack through the same helper so a bad header
falls back to _rate_limit_delay.

Addresses CodeRabbit review comments on base.py:866 and slack/channel.py:229.

* fix(channels): treat HTTP 400 and 404 as non-retryable

Both are permanent for a given request, so retrying burns the attempt
budget for nothing. Add them to _non_retryable_status_codes alongside
401/403.

Deliberately not a 4xx range check: 408 and 425 are retryable by
definition and 429 is handled by the rate-limit path. A test pins 408 as
still retryable so the range shortcut is not reintroduced later.

Partially addresses CodeRabbit's outside-diff comment on base.py:749-750.

* fix(channels): guard Slack retry extractors against raw aiohttp responses

slack_sdk attaches the bare aiohttp.ClientResponse to SlackApiError when a
JSON-declared body fails to parse. That object has neither status_code nor
get(), so _extract_status_code raised AttributeError inside should_retry,
replacing the original error and skipping the remaining attempts. Narrow
both extractors to SlackResponse/AsyncSlackResponse so such errors fall
through to the message patterns and retry as before. Add a wire-level
regression test against a local aiohttp server.

---------

Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-09-10 22:25:12 +00:00
houren Antony da15b70535 fix(channels): reject unsigned webhook POSTs on encryption-configured channels (#401)
Closes #392 (uncontroversial part).

WeChat (`_handle_message`) and Feishu (`_handle_event`) gated their
signature/decryption checks behind a condition the REQUEST controls:

- WeChat: `if encrypt and self._crypto:` -- a POST with no `<Encrypt>`
  element took the false branch and reached `_safe_process_message`
  without any verification, even when `encoding_aes_key` + `token` were
  configured.
- Feishu: `if self.config.encrypt_key and "encrypt" in body:` -- a
  plaintext body skipped decryption entirely and was processed directly.

Since the webhook port is the channel's only inbound boundary, an
attacker could POST forged plaintext and reach the agent, spoofing
`sender_id` / `FromUserName` (and, with an empty allowlist, passing the
sender gate).

Fix: when encryption is configured, an inbound POST MUST carry the
encrypted field (`<Encrypt>` / `encrypt`) -- otherwise it is rejected
with 403 and never reaches the agent. Plaintext mode (no encryption
configured) is unchanged, so existing plaintext deployments are not
affected. The remaining fail-closed question (what to do when
credentials are entirely unset) is left for the maintainers to decide
as the policy part of the issue.

Regression tests (9 new):
- WeChat: plaintext rejected / missing Encrypt rejected / bad signature
  rejected / valid signature decrypts and processes / plaintext still
  accepted when no crypto.
- Feishu: plaintext rejected / non-dict body rejected / encrypted body
  decrypts and processes / plaintext still accepted when no encrypt_key.

93 tests in the two channel files pass; full suite 3045 passed, 13
skipped; ruff clean.

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-08-14 17:25:57 +01:00
Ziheng Zhang d2283397a4 feat(feishu): scan-to-create QR onboarding + silence unsubscribed WS events (#239)
* feat(feishu): scan-to-create QR onboarding flow

Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.

- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
  helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
  domain switch based on the scanning user's tenant_brand, and a
  best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
  in the Feishu branch, ask for region (feishu vs lark), then call
  qr_register and populate feishu_app_id / feishu_app_secret /
  feishu_domain; add qrcode>=7.4 to the feishu pip extras

* fix(feishu): silently absorb unsubscribed WebSocket events

Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.

The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.

Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 22:53:16 +08:00
MuXinCG a813f8afd9 fix(feishu): isolate SDK event loop to prevent cross-thread RuntimeError on Linux
lark_oapi.ws.client captures the main thread's event loop in a
module-level variable at import time. When the WebSocket SDK thread
calls loop.run_until_complete() on that shared loop, nest_asyncio's
global patches cause task-tracking conflicts on Linux/Python 3.12:
    - RuntimeError: Leaving task … does not match the current task
    - AttributeError: 'NoneType' object has no attribute 'select'

Replace the previous Handle._run monkey-patch (which only suppressed
symptoms) with a proper fix: create a fresh event loop in the SDK
thread and swap the module-level loop variable so the SDK operates
on a fully isolated loop with no cross-thread interaction.

Closes #97
2026-03-27 22:23:08 +08:00
Ziheng Zhang 65cec64445 feat(feishu): add WebSocket long connection subscription mode (#87)
* feat(feishu): add WebSocket long connection subscription mode

Add WebSocket (长连接) mode as an alternative to webhook for Feishu
event subscription. This allows running without a public IP, port
forwarding, or tunnel — ideal for local dev and NAT/firewall setups.

- New `feishu_subscription_mode` config: "webhook" (default) or "websocket"
- WebSocket mode uses official `lark-oapi` SDK with thread-safe queue bridge
- Onboard wizard: mode selection, SDK install prompt for websocket
- CLI: `--mode webhook|websocket` for standalone serve
- `pip install evoscientist[feishu]` optional dependency
- 5 new tests covering config, SDK missing error, message bridge, cleanup
- Docs: subscription mode comparison table, prerequisites per mode

* Fix: Ruff

* Fix: small fix
2026-03-22 14:45:28 +00:00
Jan Piotrowski 81316ffb2d chore: resolve all ruff linting and static analysis errors in EvoScientist/
- Fix RUF006: Implement background task tracking in Discord, iMessage, WeChat, and TUI to prevent premature GC of fire-and-forget tasks.
- Fix B904: Add explicit exception chaining (raise ... from) across all exception handlers.
- Fix RUF012: Annotate mutable class attributes with ClassVar for command arguments and media maps.
- Fix B008: Refactor Typer commands in cli/commands.py to use Annotated for argument and option defaults.
- Fix B023/B018: Resolve late-binding issues in lambdas and remove useless expressions.
- Fix syntax errors in retry.py docstrings and models.py lambda parameter ordering.
2026-03-19 17:04:03 +01:00
Jan Piotrowski 4a3d6c0318 chore: add ruff lint rules and turn on formatting 2026-03-19 17:04:02 +01:00
X-iZhang c5a4d559a2 Refactor test cases for improved readability and consistency
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
2026-03-15 21:12:51 +00:00
MuXinCG 5fd8e1a8ad fix feishu bugs 2026-02-17 16:08:41 +08:00
MuXinCG bec3fc0b67 pass linter 2026-02-17 15:40:57 +08:00
MuXinCG b85faf7516 Add Dingtalk Feishu 2026-02-17 10:54:15 +08:00
X-iZhang eb10135314 Revert "Merge pull request #8 from EvoScientist/feature/channel-unification"
This reverts commit 84da81256a, reversing
changes made to c2c8e27b46.
2026-02-15 16:32:44 +00:00
MuXinCG 9b93cea12b fix linter bug 2026-02-15 21:07:47 +08:00
MuXinCG 61cf79e525 for merge 2026-02-15 02:11:38 +08:00