* feat: add STT voice transcription for all channels
Automatically transcribes audio/voice messages (Telegram, WeChat, Slack,
etc.) into text before the agent sees them. Enabled via config, off by default.
Changes:
- EvoScientist/stt.py: new STT engine using faster-whisper with lazy
model loading and per-language model selection (zh/en/auto)
- EvoScientist/channels/base.py: hook in _enqueue_raw() to transcribe
audio files and prepend transcript to message text; removes the raw
[voice: ...] annotation after successful transcription so the agent
does not attempt further audio processing
- EvoScientist/config/settings.py: stt_enabled (default False),
stt_language (default "auto")
- pyproject.toml: optional [stt] dependency group (faster-whisper>=1.0)
- tests/test_stt.py: unit tests covering all backends and channel integration
Usage:
pip install 'EvoScientist[stt]'
EvoSci config set stt_enabled true
EvoSci config set stt_language zh # zh / en / auto
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: remove unused imports (ruff F401)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address PR #28 reviewer feedback
Changes per SemiGlassFace review (CHANGES_REQUESTED):
1. Cache config at channel __init__ — no longer calls load_config() on
every incoming message; STT settings stored as instance attributes
(_stt_enabled, _stt_language, _stt_model, _stt_device,
_stt_compute_type) set once during Channel.__init__().
2. Replace deprecated asyncio.get_event_loop() with get_running_loop()
to avoid DeprecationWarning on Python 3.12+.
3. Annotation removal now uses exact path matching instead of substring
search — checks fp == a or a.endswith(f": {fp}]") so only the
correct annotation is removed after transcription.
4. Expose stt_model, stt_device, stt_compute_type as config fields so
users can override the HuggingFace model id, inference device, and
quantisation without touching code. transcribe_file() forwards all
three to the engine.
Also: _engines dict replaced with single _engine + _engine_key tuple
(model_id, device, compute_type) — reuses cached model unless settings
change, simpler than a dict.
Tests: 19 STT-specific tests all pass; total 1105 tests green, ruff clean.
* fix: resolve ruff lint errors (UP037, I001, PT006)
* style: apply ruff format
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add priority binding for TAB to intercept before Textual's focus_next
- Remove duplicate up/down handling in on_key (now handled by priority
bindings from PR #76)
- Update tests to use cmd_manager.list_commands() instead of removed
_TUI_SLASH_COMMANDS
Closes#57
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add DeepSeek as a recognized third-party provider
Register DeepSeek API (https://api.deepseek.com) with DEEPSEEK_API_KEY
env var and add model short names: deepseek-r1 → deepseek-reasoner,
deepseek-v3 → deepseek-chat.
* feat: add _flatten_message_content utility for list-to-string conversion
Extract text from content block lists while skipping thinking/reasoning
blocks. This handles the case where LangChain stores assistant messages
with content as a list of content blocks instead of a plain string.
* fix: flatten list content to strings for OpenAI-compatible providers
Add _patch_openai_compat_content() that wraps _generate/_agenerate to
sanitize message content before API calls. Apply it for all third-party
OpenAI-compat providers and native OpenAI proxies.
This fixes "invalid type: sequence, expected a string" errors from
strict APIs like DeepSeek that reject list-format content in assistant
messages during multi-turn conversations.
* feat: add DeepSeek API key validation and integrate into onboarding process
test: implement unit tests for content flattening utility in OpenAI-compatible providers
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Comprehensive step-by-step guide covering:
- ccproxy installation from source (patched fork required for Claude 4+)
- ccproxy OAuth login via `ccproxy auth login claude-api`
- EvoScientist install with telegram + stt extras using uv
- STT model pre-download to avoid first-message delay
- launchd plist setup for auto-start on login with KeepAlive
- Full troubleshooting section based on real deployment experience:
- OAuth token not found (.credentials.json location)
- Packages installed in wrong Python environment (conda vs .venv)
- Shell glob eating brackets in pip install 'pkg[extra]'
- Whisper hallucination / VAD filter
- ffmpeg approval prompts → --auto-approve
- heredoc variable expansion gotcha
- External drive mount timing with launchd
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
- Fix RUF006: Implement background task tracking in Discord, iMessage, WeChat, and TUI to prevent premature GC of fire-and-forget tasks.
- Fix B904: Add explicit exception chaining (raise ... from) across all exception handlers.
- Fix RUF012: Annotate mutable class attributes with ClassVar for command arguments and media maps.
- Fix B008: Refactor Typer commands in cli/commands.py to use Annotated for argument and option defaults.
- Fix B023/B018: Resolve late-binding issues in lambdas and remove useless expressions.
- Fix syntax errors in retry.py docstrings and models.py lambda parameter ordering.
* feat: update MiniMax integration to use Anthropic-compatible endpoint and enhance routing logic
* refactor: streamline OAuth install hint and update ccproxy health check timeout
Add MiniMax (api.minimax.io/v1) as a first-class third-party provider,
enabling direct API access without routing through NVIDIA/SiliconFlow/
OpenRouter intermediaries. Includes M2.5 and M2.5-highspeed models
with 204K context window.
Changes:
- Register "minimax" in _THIRD_PARTY_PROVIDERS with MINIMAX_API_KEY
- Add MiniMax-M2.5 and MiniMax-M2.5-highspeed model entries
- Add minimax_api_key to config, env mappings, and env export
- Add MiniMax to onboarding wizard with API key validation
- Update .env.example, README.md, README.zh-CN.md
- Add 9 unit tests and 3 integration tests (all passing)
Co-authored-by: PR Bot <pr-bot@minimaxi.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(mcp): eliminate duplicate config loading during startup
load_mcp_config() was called twice on every startup — once in
_mcp_config_signature() to compute the cache key, and again inside
load_mcp_tools(). This caused warnings to appear twice.
Merge the two calls into _load_mcp_config_once() which returns both the
signature and the parsed config, then pass the config through to
load_mcp_tools() via a new optional parameter.
* test(mcp): fix existing cache tests and add coverage for single-load guarantee
- Update fake_load_mcp_tools to accept optional config kwarg
- Add test_load_mcp_config_called_once_per_cache_miss: verifies
load_mcp_config is called exactly once per cache miss (the bug)
- Add test_cached_config_passed_to_load_mcp_tools: verifies the
pre-loaded config dict is forwarded to load_mcp_tools
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* refactor(mcp): extract MCP server registry from onboard into shared module
Move _RECOMMENDED_MCP_SERVERS, _install_pip_package, and
_pip_install_hint from config/onboard.py into mcp/registry.py as a
shared MCPServerEntry dataclass and registry functions. This enables
reuse by the new /install-mcp command and marketplace integration.
* feat(mcp): add /install-mcp command for browsing and installing MCP servers
Interactive browser for MCP servers (built-in registry + EvoSkills
marketplace), supporting three modes:
- /install-mcp — interactive tag filter + checkbox selection
- /install-mcp <name> — direct install by name or tag pre-filter
- /install-mcp file.yaml — import servers from arbitrary YAML file
Also available as /mcp install and EvoSci mcp install. Includes TUI
browser widget (MCPBrowserWidget) mirroring the skill browser UX.
* chore(tavily): conditionally pass tavily_search when TAVILY_API_KEY is set
* fix(tui): Enter key detection for /install-mcp
- Fix message handler names: Textual converts MCPBrowserWidget to
mcpbrowser_widget (not mcp_browser_widget), so Confirmed/Cancelled
messages were never received by the app
- Distinguish empty selection from cancel in result handling
* chore(widgets): stop auto-advancing cursor on Space toggle in browser widgets
* style: linter
* refactor(mcp): simplify MCP registry to marketplace-only
Remove built-in server list and arbitrary YAML import — all server
definitions now come from the EvoSkills marketplace (mcp/*.yaml).
Onboarding filters by the `onboarding` tag instead of a hardcoded list.
* refactor(mcp): consolidate /install-mcp into /mcp install
Remove standalone /install-mcp command — use /mcp install as the
single entry point. CLI adapter now delegates logic to the shared
InstallMCPCommand class, keeping only the questionary UI layer.
* chore(cli): rm reference to yaml import
- adjust discord max message length to 2000 (the actual limit)
- /install-skill will fallback to searching EvoSkills repository if path not found to allow for installing skills from channels
- remove discord config tests (it asserted @dataclass behaviour)
* fix: add config option for ccproxy port number
* feat: add user prompt for ccproxy port configuration and validation
fix: update is_ccproxy_running to use health check endpoint
test: enhance tests for ccproxy port handling and validation
* fix: streamline ccproxy installation process using _install_pip_package
* fix: auto-patch ccproxy adapter for correct OAuth beta header
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
load_dotenv was called in three scattered places: commands.py (main and
serve entry points) and search.py (at module import time). This could
re-override env vars after config resolution.
I moved the single load_dotenv call into get_effective_config() so it
participates in the config priority chain, and removed it from all
other call sites.
* feat: add tag support to skill metadata parser
- Add `tags` field to `SkillInfo` dataclass
- Extend `_parse_skill_md` to return `SkillInfo` directly (instead of
dict), extracting tags from top-level `tags` or `metadata.tags`
fallback
- Accept `source` as keyword argument in `_parse_skill_md` to avoid
post-hoc mutation
- Add `_normalize_tags` helper (handles list, comma-string, missing)
- Add `list_skills_by_tag()` for filtering installed skills by tag
- Add `get_all_tags()` returning tags sorted by count then
alphabetically
- Add `fetch_remote_skill_index()` with shallow-clone and 10-min cache
- Add 12 new tests covering tag parsing, filtering, and remote index
* feat: add /install-skills command and browse action to skill_manager tool
CLI:
- Add `/install-skills` slash command with interactive tag picker and
skill checkbox (questionary-based, for CLI mode)
- Accepts optional tag argument for pre-filtering: `/install-skills
core`
- Update `/skills` listing to show tags per skill
- Register command in interactive.py dispatch
LangChain tool:
- Add `browse` action to `skill_manager` tool with optional `tag` filter
- Update `list` and `info` actions to include tags in output
* feat: add interactive skill browser widget for TUI
New widget (skill_browser.py):
- Two-phase keyboard-driven widget mounted inline in chat
- Phase 1: tag picker (arrow keys + Enter, Esc to cancel)
- Phase 2: skill checkbox (Space to toggle, Enter to install, Esc back)
- Width-aware description truncation with ellipsis
- Installed skills shown as non-toggleable with checkmark
TUI integration (tui_interactive.py):
- Register /install-skills command with async widget flow
- Echo executed commands in cyan before output
- Add SkillBrowserWidget keyboard delegation (up/down/esc)
- Refocus prompt input after any widget dismissal (also fixes
pre-existing /resume and /delete focus bug)
- Show tags as bulleted newlines in /skills table
- Dynamic autocomplete padding based on longest item
* "fix(onboard): guide ccproxy install and auth in OAuth flow
- Always show API Key / OAuth choice; prompt to install evoscientist[oauth]
when ccproxy missing (mirrors iMessage imsg install UX)
- Add _ccproxy_exe() helper: checks PATH then env bin dir (fixes conda envs
where shutil.which may not find newly installed binaries)
- Fix check_ccproxy_auth() false positive: ccproxy auth status exits 0 even
when not authenticated; detect via output content + filter structlog noise
- Silent install/login subprocesses; show browser URL as fallback
- Reset anthropic/openai auth_mode to api_key when switching to non-Anthropic/
OpenAI provider, preventing stale oauth config from triggering ccproxy error
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>"
* fix(onboard): update OAuth support details in README and zh-CN translation
* feat(cli): add --version / -V flag
Uses importlib.metadata to read the version from the installed package.
* chore: improve bug report and feature request issue templates
- Bug report: replace web-app steps with CLI-oriented examples, add
error output section, add Python version and LLM provider fields
- Feature request: add note directing niche features to EvoSkills,
set default label
* chore: add documentation issue template and issue chooser config
- Add documentation template for reporting missing or unclear docs
- Add config.yml to disable blank issues and link to EvoSkills and
Discord as contact options
* chore: add PR template and improve CONTRIBUTING.md
- Add PR template with type-of-change checkboxes, issue linking for
new features, and CI checklist
- CONTRIBUTING.md: add development setup, PR workflow, and code style
sections; fix wording; make Discord link clickable
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.