* feat(gateway): graph gateway protocol
* refactor(cli): wire gateway in cli/tui
* refactor(gateway): centralize runtime gateway init
* chore(gateway): restrict RunRequest message type
* feat(gateway): add langgraph server gateway
* chore(cli): tighten serve runtime state typing
* refactor(cli): route async task state reads through graph gateway
* refactor(gateway): support graph targets in server gateway
* refactor(cli): route session commands through graph gateway
* refactor(cli): fold thread store under graph gateway
* refactor(gateway): route graph state access through gateway
* refactor(channels): wire graph gateway
* refactor(memory): preserve graph threads for cloning
* feat(gateway): add thread cloning
* fix(tui): pass effective workspace for thread creation
* chore(memory): add workspare dir to memory worker metadata
* fix(sessions): filter preloaded UUID registy entries by the current scope
* test(fakes): use https
* refactor(consumer): consolidate imports
* fix(stream): optional summarization event
* fix(gateway): resolve abbreviated thread IDs by search
* fix(gateway): page server thread listings
* fix(gateway): emit pending interrupt events
* style: fmt
* feat(gateway): persist workspace_dir & model in thread metadata
* fix(gateway): page server thread prefix resolution
* fix(gateway): expose server thread list metadata
* refactor: add back type def
* refactor: tighten types
* revert: add back worker thread deletion
The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.
* fix(gateway): apply compaction to server thread history
* refactor(stream): restore direct summary replay suppression
* fix(gateway): preserve compaction state and server stream output
* fix(gateway): close local stream generator on cancellation
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(memory): add observation memory lifecycle
Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.
Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.
Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.
* fix(cli): sync background agent server on resume
Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.
Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.
Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.
Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.
* fix(cli): prepare serve resume workspace before adopting
Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.
* fix(memory): untrack abandoned worker status watches
Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.
* fix(cli): report channel command failures accurately
Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.
* fix(stream): clear memory counters for resume streams
Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.
* docs(tools): make observation recording guidance conditional
Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.
* feat(config): add controls for profile and observation memory
Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.
Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.
Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.
* test(cli): include memory defaults in serve config stubs
* fix(memory): offload async worker launch blocking calls
Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.
* chore(memory): harden turn worker subagent guardrail
* chore(memory): refresh profile context per request
* fix(memory): offload async profile file reads
* fix(memory): offload async worker completion accounting
* Add status bar and compact summary widgets with context window resolution
- Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows.
- Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format.
- Introduced a `CompactingWidget` to indicate ongoing compacting processes.
- Added a base class `TimedStatusWidget` for widgets that require a timer.
- Developed context window resolution helpers to retrieve context window sizes from various model attributes.
- Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios.
- Updated existing tests to cover new features and maintain code quality.
* refactor(Channel): simplify lambda function in _send_with_retry method
* feat: enhance context editing logic and improve error handling in StreamState
* refactor(Channel): streamline lambda function in _send_with_retry method
* feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic
* feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution
* fix: correct formatting of console message for MCP server configuration status
* feat: add check for None summary_message in _apply_summarization_event to prevent errors
* feat: enhance _load_checkpoint_messages to validate message format and apply summarization event
* feat: enable reasoning for OpenRouter via extra_body to prevent multi-turn errors
* feat: implement OpenRouter native reasoning support and patch langchain-openrouter bug
* feat: add OpenRouter reasoning effort configuration and update related tests
* feat: add langchain-openrouter dependency for enhanced reasoning support
* fix: correct spacing in reasoning effort choice label
* feat: implement patch for OpenRouter reasoning details to prevent Pydantic errors
* feat: add patches for OpenRouter reasoning and content handling utilities
* feat: prevent multiple patches of OpenRouter reasoning details by using a global flag
* feat: update OpenRouter reasoning patch to ensure single application with global flag
* feat: refine OpenAI responses API handling to apply only for OpenAI provider
* feat: Enhance TUI interaction by updating todo widget positioning and skipping empty tool call chunks
* feat: Update tool selector threshold and adjust logging level for selector failures
* feat: Temporarily disable timestamp toast in tool call widget for UX review
* feat: Re-enable timestamp toast in tool call widget on click
* feat: Add context management middleware for improved error handling and context editing
* feat: Implement LLMToolSelectorMiddleware for enhanced tool selection and tracking
* feat(tests): update test functions to include mock timestamp parameter
* refactor: simplify tool selection state storage and update comments in middleware
* feat: Enhance tool selection handling and suppress structured output for improved event streaming
* refactor: simplify patching in test_create_tool_selector functions
* feat: add model parameter to create_tool_selector_middleware for enhanced flexibility
* feat: enhance tool selection suppression with JSON buffering for improved accuracy
* feat: Upgrade ccproxy to version 0.2.7 and remove deprecated thinking tag handling
* feat: Enhance ccproxy compatibility and strip legacy thinking tags
* fix: subagent summerize
* fix: group subagent text by agent name for parallel fallback
The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.
Add comprehensive tests for the new behavior (24 tests).
* fix: group subagent text by agent name for parallel fallback
The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.
Add comprehensive tests for the new behavior (24 tests).
* fix: group subagent text by agent name for parallel fallback
The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.
Add comprehensive tests for the new behavior (24 tests).
* chore: fix multiple agent
* test: add test
* fix linter
* remove redundant
- Fix RUF006: Implement background task tracking in Discord, iMessage, WeChat, and TUI to prevent premature GC of fire-and-forget tasks.
- Fix B904: Add explicit exception chaining (raise ... from) across all exception handlers.
- Fix RUF012: Annotate mutable class attributes with ClassVar for command arguments and media maps.
- Fix B008: Refactor Typer commands in cli/commands.py to use Annotated for argument and option defaults.
- Fix B023/B018: Resolve late-binding issues in lambdas and remove useless expressions.
- Fix syntax errors in retry.py docstrings and models.py lambda parameter ordering.
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
- Implemented `ask_user` middleware to facilitate agent-initiated questions during research workflows.
- Created `AskUserWidget` for interactive user prompts, supporting both text and multiple choice questions.
- Enhanced `StreamEventEmitter` to handle `ask_user` interrupts and updated event handling in `stream_agent_events`.
- Added state management for pending `ask_user` events in `StreamState`.
- Developed validation for question structures and parsing for responses.
- Introduced unit tests covering middleware functionality, event handling, and widget behavior.
- Added HITL approval lifecycle to ToolCallWidget, including a new "rejected" status.
- Introduced configuration options for auto-approval and shell command allow list in EvoScientistConfig.
- Developed approval handling functions in display module to manage HITL interrupts and user decisions.
- Created ApprovalWidget for user interaction during approval prompts.
- Enhanced StreamEventEmitter to support interrupt events.
- Updated StreamState to track pending interrupts.
- Implemented tests for HITL functionality, including event structure, state handling, and approval logic.
- Introduced a new test file `test_rich_escape.py` to validate the safety of Rich markup escaping in the ToolResultFormatter.
- Added tests to ensure that tool names and error messages containing brackets do not cause crashes during formatting.
Introduce the new channel unification architecture:
- Core framework: base channel class, bus, consumer, channel manager,
middleware, mixins, capabilities, formatter, retry, plugin system
- Telegram channel implementation with bot token validation
- Discord channel implementation with bot token validation
- Refactored iMessage to use new base channel architecture
- Updated CLI: /channel commands, serve mode, channel setup wizard
- Updated onboard wizard to support multi-channel selection
- Config settings for all channel types
- Stream events: media attachment support, done event content field
- Comprehensive test coverage for channels, bus, and manager
- Introduced a new `sessions.py` module for handling session persistence using SQLite.
- Added CRUD operations for threads, including listing, checking existence, finding similar threads, and deleting threads.
- Enhanced the interactive CLI to support commands for managing sessions: `/current`, `/threads`, `/resume`, and `/delete`.
- Updated the `cmd_interactive` function to handle session metadata and improve user experience with session history rendering.
- Modified the streaming functions to include metadata for checkpoint persistence.
- Added unit tests for session management functionalities to ensure reliability and correctness.
- Updated iMessage channel to read send_thinking preference from config.
- Modified _auto_start_channel to accept send_thinking parameter.
- Removed view_image tool and adjusted related functionality.
- Enhanced README to reflect changes in media handling.
- Updated tests to cover new behavior and removed tests for view_image.