main
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1f3e8f57a3 |
fix: resolve subagent tools at execution decision (#381)
* fix: defer subagent tool resolution Signed-off-by: Aditya Datta <crazyme07071996@gmail.com> * fix: preserve inherited subagent tools * test: cover injected subagent tools * style: format subagent regression --------- Signed-off-by: Aditya Datta <crazyme07071996@gmail.com> |
||
|
|
3c5cc831c0 |
Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400) WebUI mode was only reachable from the machine running it: the front-end got no bind interface, and `start_langgraph_dev(...)` was called without a host, so both servers stayed on loopback with no way to widen them. Adds two config fields with deliberately different defaults: webui_host = 0.0.0.0 front-end serves the app shell, no secrets langgraph_dev_host = 127.0.0.1 unauthenticated API, agent can run shell The design hinges on separating bind address from client address. Only bind() uses the configured interface; every consumer that *connects* (health probes, occupancy checks, async sub-agent self-dispatch) goes through the new `_probe_host`, which maps a wildcard bind back to loopback and honors a pinned interface verbatim. `_can_bind_port` is the one exception and binds the literal host, since it must replicate the bind the server itself will attempt. - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`; host kwarg threaded through the probes and `start_langgraph_dev`, which now emits `--host` and propagates EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess - sdk.py: `langgraph_dev_url` tracks host as well as port; EvoScientist.py reuses it instead of an inline f-string - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND banner whenever the bind is not provably loopback - webui.py: forwards both hosts; the front-end is widened via HOSTNAME because @evoscientist/webui ships no --host flag — its bin launcher does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is gated on the backend host only, so the shipped front-end default doesn't print a banner on every launch Verified end to end against a live server: requesting 0.0.0.0 yields a socket listening on 0.0.0.0 with the health probe correctly resolved to 127.0.0.1, while the default still binds 127.0.0.1 only. Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading users will find the front-end reachable from the LAN. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400) Completes the remaining items from #400. - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`. Remote WebUI use needs both anyway (the UI reaches the backend from the browser, not server-side), so a loopback backend default just meant every remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where these servers listen. SECURITY: this exposes an unauthenticated API whose agent can run shell commands. The red PUBLIC BIND banner consequently fires on every launch while exposed — kept deliberately, since the exposure is real and the escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is only discoverable if we say so. READMEs now lead with the warning and document the SSH-tunnel alternative. - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In WebUI mode they are two halves of one surface; moving only one leaves the UI loading but unable to reach the agent. Blank values are dropped rather than written as an empty override that would beat the config file. - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` / `http://localhost:{port}` (steps.py:160, :223) — both render the configured bind through `_base_url` / `_format_hostport`, so a pinned interface is reported honestly and a wildcard still shows loopback. Verified against a live server: with no host argument at all, resolution through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL of http://127.0.0.1, and the warning gate returning True. Still open and tracked separately: the front-end takes its backend URL from browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change in that repo. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced onto Node.js 24. v7.0.0 is the release that made that switch, so anything >= v7 clears the warning; v9.0.0 is current. Pinned to the full tag deliberately: setup-uv stopped publishing major and minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return 404 and would fail the job outright. Releases are immutable from v8 on, so the full tag is as tamper-proof as a SHA. Comment left in lint.yml because "simplifying" this back to `@v9` is an easy and CI-breaking mistake. actions/checkout@v5 is already node24 and needs no change. Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to ease load on PyPI infrastructure). None of these workflows set it, so they follow the new default and Actions cache usage may grow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(cli): correct --host help text and warn on public bind in non-WebUI modes The --host help claimed "WebUI mode only", which is wrong in a way that matters for security. `--host` writes `langgraph_dev_host` unconditionally, and `_ensure_async_subagent_server` auto-starts that backend for tui / cli / serve as well — the langgraph dev server is shared across UI modes. So the flag narrows or widens the agent API in every mode, and only `webui_host` is actually WebUI-specific. Reported against cli/commands.py. The documentation error hid a real gap: the PUBLIC BIND banner lived only in deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound 0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so fixing the text alone would not surface it. Added the same banner to the shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev fails soft (async degrades to in-process delegation), and warning about a bind that never happened would be worse than staying quiet. READMEs (EN + zh-CN) get the same correction — the warning block sat inside the Desktop WebUI section and read as WebUI-scoped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(deploy): strip the config-derived bind host, not just the CLI one `deploy()` only stripped the `--host` branch. When the flag was omitted, `getattr(config, "langgraph_dev_host", ...)` flowed unstripped into `_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and the banner. `run_webui` already strips unconditionally; this aligns the two. Reachable because `deploy()` reads through `getattr` and is routinely handed duck-typed config objects (tests, embedders) that never run `EvoScientistConfig.__post_init__`, which is what normally normalizes these fields. Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is False, so a padded loopback value would print a false PUBLIC BIND warning while binding a string socket.bind() rejects outright — a security banner saying the opposite of the truth. Three regression tests added, each verified to fail against the old code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: apply ruff format to the bind-host changes The Lint workflow runs both `ruff check` and `ruff format --check`; I had only been running the former locally, so five files landed unformatted and failed CI. Whitespace and line-wrapping only — no semantic change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(security): keep the langgraph dev backend on loopback by default The backend is an unauthenticated API whose agent can run shell commands, and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a 0.0.0.0 default put it on the network for users who never asked. Restore 127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in. webui_host keeps its 0.0.0.0 default: the front-end serves the app shell only and holds no credentials. run_webui already prints a remote-backend hint when the front-end is exposed and the backend is not. Help text and both READMEs are reframed around widening rather than narrowing; the escape-hatch tests are inverted to assert the public-bind opt-in survives into argv. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
80f1f4fa0f |
feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system - Added async notifier functionality to handle notifications for sub-agents reaching terminal states. - Introduced `AsyncTaskNotification` dataclass for structured notification data. - Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications. - Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers. - Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM. - Added tests for notification handling, including draining, deduplication, and formatting. - Patched deepagents to integrate the new watcher functionality into start and update tools. * Enhance async notifier with per-thread notification routing and error handling - Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session. - Implemented per-thread notification queues to handle notifications based on the originating thread. - Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues. - Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits. - Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads. - Added a fixture to restore the async watcher patch state in tests to prevent state leakage. - Updated deepagents patching to capture the main agent's CLI thread ID for notification routing. * feat: Enhance async notifier with thread-specific watcher management and notification filtering * test: Enhance notification draining logic for cleaner test setup * refactor: Remove summary field from AsyncTaskNotification and update related tests * feat: Enhance async notification handling with target thread ID support * Refactor async notifier and middleware for improved task management - Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed. - Updated watch_run_and_notify to clarify notification handling and race conditions. - Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls. - Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach. - Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls. - Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation. - Enhanced test coverage for async watcher middleware, including edge cases and error handling. * feat(tests): add fixture to reset notifier state before each test |
||
|
|
f41584e10b |
Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support - Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory. - Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations. - Added new langgraph_dev module for managing async sub-agent lifecycle and deployment. - Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment. - Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations. - Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files. * Refactor code for improved readability by consolidating conditional statements and formatting * feat: enhance async sub-agent support with workspace synchronization and user feedback - Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience. - Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations. - Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents. - Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states. * feat: add async sub-agent configuration and server management functions * feat: improve port occupation handling and log file management in start_langgraph_dev * feat: enhance async sub-agent handling and introduce comprehensive tests - Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff. - Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service. - Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations. - Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness. - Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`. * fix(docs): clarify sub-agent configuration in README * test(manager): isolate _PID_DIR + tighten reuse-path assertion Addresses CodeRabbit review on tests/test_langgraph_manager.py: - Patch _PID_DIR to tmp_path so the FileLock setup in ensure_langgraph_dev doesn't mkdir the user's real ~/.config/evoscientist/ dir as a test side-effect. - Tighten "result is None or hasattr(result, 'poll')" to a strict "result is None" — the reuse path returns None unconditionally, so the OR clause was hiding potential regressions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(manager): clean up stale PID file when unrelated process reuses PID * feat(tests): add validation tests for async flag in load_subagents * fix(load_subagents): restrict to .yaml files and clarify configuration handling * fix(load_subagents): improve error handling for non-dict specifications in YAML * feat(onboard): add "LangGraph Port" step to onboarding process * feat(langgraph): add concurrency configuration for langgraph dev workers * feat(async-subagents): enhance MCP tool routing for async sub-agents * fix(manager): update exception handling for connection errors and prevent zombie processes --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |