* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* test: add autouse fixture for watcher cleanup
* refactor: remove redundant hasattr calls
* refactor: extract shared HITL/ask_user interaction grammar
Extract prompt/question formatting, the reply grammar (approval letters,
ask_user choice letters + the 'Other' sub-flow, stop-commands), the
ApprovalPolicy (config auto-approve rule + session registry +
session-key derivation), per-flow timeout constants, and the bilingual
feedback strings into channels/interaction.py. Both drivers now point at
the shared functions: this reverses cli/channel.py's imports of consumer
privates and closes the /stop drift at the parsing layer (serve-mode
ask_user now checks stop-commands before parsing an answer, matching the
CLI path).
* refactor: add interaction engine + registry; port InboundConsumer
Introduce InteractionIO (transport adapter Protocol),
PendingReplyRegistry (one asyncio-based reply router per process), and
the engine coroutines resolve_ask_user / resolve_approval in
channels/interaction.py. Port InboundConsumer onto them: a _ConsumerIO
adapter over bus.publish_outbound + the registry, one ApprovalPolicy
replacing the config/session auto-approve checks, and a single
reply-interception point (registry.try_resolve) replacing the parallel
ask_user/HITL pending dicts. _resolve_ask_user and the approval section
of _stream_with_hitl are now thin engine calls.
Behavior unification (serve mode): an unrecognized HITL reply now
declines with the shared 'Unrecognized reply' notice instead of
rejecting-and-refeeding as a fresh turn, and /stop mid-approval cancels
cleanly — both via the shared parser.
* refactor: port CLI channel bridge onto the interaction engine
Replace the ~250-line parallel bodies of channel_ask_user_prompt /
channel_hitl_prompt with thin bridges that run resolve_ask_user /
resolve_approval on the bus loop via
run_coroutine_threadsafe(...).result() (outer = engine per-flow timeout
+ slack, so the engine's own timeout fires first). The 15s send timeout
moves into the _BridgeIO adapter.
Delete the _pending_hitl / _hitl_lock / _hitl_auto_approve module
globals and the _register_hitl_wait / _try_set_hitl_reply /
_pop_hitl_reply helpers, absorbed by one bus-loop PendingReplyRegistry +
one ApprovalPolicy. The bus consumer feeds the registry via try_resolve
ahead of normal enqueue.
* refactor: restore serve-mode refeed for unrecognized HITL replies
Gate-review fix: the engine no longer decides transport policy for
unparseable approval replies. resolve_approval now returns an
ApprovalOutcome carrying unrecognized_reply (raw text) when parsing
fails, sending no feedback itself; recognized reject keeps the sharedi
rejection message.
Consumer driver (serve mode) restores the pre-engine semantics: an
unrecognized reply rejects the pending action, confirms with the
rejection message, and the text is re-dispatched as a NEW agent turn —
_stream_with_hitl returns the captured text and _handle_message starts
the refeed turn only after the current one has released the chat lock
(old fall-through ordering). CLI bridge keeps its old no-refeed path
byte-for-byte: 'Unrecognized reply. Action rejected.' and decline.
Tests: serve refeed pinned end-to-end (prompt → unrecognized text →
rejection feedback → text reaches the stream path as a new turn), CLI
no-refeed pinned (notice sent, nothing enqueued), engine test updated to
assert the outcome struct with no engine-side feedback.
* refactor: polish the interaction engine surface
- English feedback strings (Approved / Rejected / auto-approving)
- drop the consumer's backwards-compatible re-exports and both modules'
private timeout aliases; callers use the canonical interaction names
- replace byte-for-byte prompt goldens with structural format tests and
assert feedback via the shared constants instead of string literals
- strip audit/design shorthand (R1/R2/G3, stage numbers) from comments
* fix: propagate pending reply task cancellation
* fix: preserve reply context when refeeding HITL replies
* fix: honor HITL session grants without bus loop
* fix: bound bridge waits by send latency
* fix: handle empty ask_user replies explicitly
* chore: remove stale interaction helpers
* fix: harden interaction engine reply edge cases
Review follow-ups on the interaction engine:
- normalize ask_user choices before .get(): the tool args come from model
JSON and only presence is validated, so plain-string choices must render
and parse instead of crashing the turn
- treat only None as an approval timeout, so an empty/media-only reply
flows through the unrecognized path and serve mode refeeds it with its
preserved context
- intercept prompt replies before _get_thread_id so a consumed reply
cannot create an orphan graph thread or touch the sender-session LRU
- close engine coroutines the bridge failed to schedule (no bus loop /
scheduling error) to avoid never-awaited warnings
- clear pending-response and channel-request state in the bridge test
fixture
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(gateway): graph gateway protocol
* refactor(cli): wire gateway in cli/tui
* refactor(gateway): centralize runtime gateway init
* chore(gateway): restrict RunRequest message type
* feat(gateway): add langgraph server gateway
* chore(cli): tighten serve runtime state typing
* refactor(cli): route async task state reads through graph gateway
* refactor(gateway): support graph targets in server gateway
* refactor(cli): route session commands through graph gateway
* refactor(cli): fold thread store under graph gateway
* refactor(gateway): route graph state access through gateway
* refactor(channels): wire graph gateway
* refactor(memory): preserve graph threads for cloning
* feat(gateway): add thread cloning
* fix(tui): pass effective workspace for thread creation
* chore(memory): add workspare dir to memory worker metadata
* fix(sessions): filter preloaded UUID registy entries by the current scope
* test(fakes): use https
* refactor(consumer): consolidate imports
* fix(stream): optional summarization event
* fix(gateway): resolve abbreviated thread IDs by search
* fix(gateway): page server thread listings
* fix(gateway): emit pending interrupt events
* style: fmt
* feat(gateway): persist workspace_dir & model in thread metadata
* fix(gateway): page server thread prefix resolution
* fix(gateway): expose server thread list metadata
* refactor: add back type def
* refactor: tighten types
* revert: add back worker thread deletion
The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.
* fix(gateway): apply compaction to server thread history
* refactor(stream): restore direct summary replay suppression
* fix(gateway): preserve compaction state and server stream output
* fix(gateway): close local stream generator on cancellation
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: implement configurable sandbox execute timeout and enhance recovery instructions
* feat: add background process management tools and middleware for sandbox execution
* feat: enhance background process management with completion notifications and deduplication
* feat: enhance sandbox execution timeout validation and update related messages
* feat: enhance background process management with thread-specific completion notifications and HITL approval handling
* test: assert completion notification waits for process finish timestamp
* feat(qq): add inline keyboard buttons for C2C HITL approval
QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).
Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
generic `{text, value, type}` shape to QQ's `{render_data, action}`
with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
attached the fallback content gets a textual `Reply: 1=Approve, …`
hint built from the button list. `_parse_approval_reply` accepts
the same values typed manually, so the user is never stuck.
Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
InboundMessage, runs it through inbound middleware (Dedup suppresses
retry callbacks), and publishes directly to the bus — bypassing the
per-sender debounce buffer so the click value isn't merged with any
text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
QQ doesn't show the button as "expired", even if middleware drops
the click or something throws downstream.
`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.
* fix(qq): button-value coercion, ACK timing, HITL consumer wiring
Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.
- Plain-text fallback no longer crashes on non-string `value` (e.g.
`{"text": "OK", "value": 42}`). Extracted `_normalize_button` is now
the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
so the QQ `inline_buttons=True` capability is actually used end-to-end
(HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
not in `finally` after a `return`).
Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.
* refactor(qq): slim button helpers and explicit has_buttons flag
Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).
Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.
Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.
* feat(qq): send post-decision confirmation after HITL approval
Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.
Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.
CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings. Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
* fix: subagent summerize
* fix: group subagent text by agent name for parallel fallback
The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.
Add comprehensive tests for the new behavior (24 tests).
* fix: group subagent text by agent name for parallel fallback
The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.
Add comprehensive tests for the new behavior (24 tests).
* fix: group subagent text by agent name for parallel fallback
The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.
Add comprehensive tests for the new behavior (24 tests).
* chore: fix multiple agent
* test: add test
* fix linter
* remove redundant
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
- Added HITL approval lifecycle to ToolCallWidget, including a new "rejected" status.
- Introduced configuration options for auto-approval and shell command allow list in EvoScientistConfig.
- Developed approval handling functions in display module to manage HITL interrupts and user decisions.
- Created ApprovalWidget for user interaction during approval prompts.
- Enhanced StreamEventEmitter to support interrupt events.
- Updated StreamState to track pending interrupts.
- Implemented tests for HITL functionality, including event structure, state handling, and approval logic.
Introduce the new channel unification architecture:
- Core framework: base channel class, bus, consumer, channel manager,
middleware, mixins, capabilities, formatter, retry, plugin system
- Telegram channel implementation with bot token validation
- Discord channel implementation with bot token validation
- Refactored iMessage to use new base channel architecture
- Updated CLI: /channel commands, serve mode, channel setup wizard
- Updated onboard wizard to support multi-channel selection
- Config settings for all channel types
- Stream events: media attachment support, done event content field
- Comprehensive test coverage for channels, bus, and manager