* feat: configurable bind host for WebUI and langgraph dev (refs #400)
WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.
Adds two config fields with deliberately different defaults:
webui_host = 0.0.0.0 front-end serves the app shell, no secrets
langgraph_dev_host = 127.0.0.1 unauthenticated API, agent can run shell
The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.
- manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
host kwarg threaded through the probes and `start_langgraph_dev`,
which now emits `--host` and propagates
EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
- sdk.py: `langgraph_dev_url` tracks host as well as port;
EvoScientist.py reuses it instead of an inline f-string
- server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
banner whenever the bind is not provably loopback
- webui.py: forwards both hosts; the front-end is widened via HOSTNAME
because @evoscientist/webui ships no --host flag — its bin launcher
does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
gated on the backend host only, so the shipped front-end default
doesn't print a banner on every launch
Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.
Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes#400)
Completes the remaining items from #400.
- `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
Remote WebUI use needs both anyway (the UI reaches the backend from the
browser, not server-side), so a loopback backend default just meant every
remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
`DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
these servers listen.
SECURITY: this exposes an unauthenticated API whose agent can run shell
commands. The red PUBLIC BIND banner consequently fires on every launch
while exposed — kept deliberately, since the exposure is real and the
escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
only discoverable if we say so. READMEs now lead with the warning and
document the SSH-tunnel alternative.
- `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
WebUI mode they are two halves of one surface; moving only one leaves the
UI loading but unable to reach the agent. Blank values are dropped rather
than written as an empty override that would beat the config file.
- Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
`http://localhost:{port}` (steps.py:160, :223) — both render the
configured bind through `_base_url` / `_format_hostport`, so a pinned
interface is reported honestly and a wildcard still shows loopback.
Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.
Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime
GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.
Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.
actions/checkout@v5 is already node24 and needs no change.
Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(cli): correct --host help text and warn on public bind in non-WebUI modes
The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.
The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.
READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(deploy): strip the config-derived bind host, not just the CLI one
`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.
Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.
Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.
Three regression tests added, each verified to fail against the old code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* style: apply ruff format to the bind-host changes
The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): keep the langgraph dev backend on loopback by default
The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.
webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.
Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL
* fix: keep parent env authoritative over workspace .env for mapped keys
* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins
* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS
* fix: merge .env via dotenv_values to close empty-value and RMW-race edges
* chore: align docstrings after .env-merge rework
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(llm): respect reasoning_effort setting on native OpenAI path
The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.
Adds a regression test and isolates the existing xhigh test from the
env var.
* fix(llm): preserve model reasoning defaults
* fix(llm): preserve GPT-5.6 reasoning default
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True),
override=True), so any test that loads config injected the repo's real
.env into os.environ for the rest of the pytest process. An
empty-valued line like MINIMAX_BASE_URL= then made
os.environ.get(key, default) return '' instead of the default,
failing the MiniMax routing tests in full-suite runs while they
passed in isolation.
Generalizes the find_dotenv redirect that test_config.py's
temp_config_dir fixture already applied locally into a suite-wide
autouse fixture, pointing at a never-created path so tests writing
their own tmp_path/.env cannot collide with it. Adds a regression
test reproducing the leak.
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(llm): add OpenRouter app attribution headers (#339)
Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.
Closes#339
* refactor(llm): centralize OpenRouter attribution defaults + cap categories
Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
(the config fields and llm/models.py both use them) instead of duplicating
the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
sent list to OpenRouter's 2-per-request limit, warning when a configured list
exceeds it, so extras are dropped predictably (and surfaced) here rather than
being silently truncated server-side.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add scheduler functionality with cron-style task management
- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.
* fix(async-notifier): ensure fallback hint is used for unknown notification kinds
* feat: enhance scheduling functionality and improve system message handling
* feat(dangerous-mode): implement real-filesystem access with safety checks
- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.
* feat(dangerous-mode): enhance logging and environment management for dangerous mode
* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
* feat(middleware): reposition code interpreter middleware in the stack
* feat(models): add qwen3.7-plus model entry and update context window comment
* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope
* feat(auxiliary): implement auxiliary model support for background tasks and tool selection
- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.
* feat(steps): update UI backend selection options and descriptions
* Refactor code structure for improved readability and maintainability
* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors
* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py
* feat(config): add auxiliary model and provider environment variables to test setup
* feat(memory): add observation memory lifecycle
Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.
Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.
Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.
* fix(cli): sync background agent server on resume
Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.
Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.
Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.
Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.
* fix(cli): prepare serve resume workspace before adopting
Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.
* fix(memory): untrack abandoned worker status watches
Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.
* fix(cli): report channel command failures accurately
Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.
* fix(stream): clear memory counters for resume streams
Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.
* docs(tools): make observation recording guidance conditional
Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.
* feat(config): add controls for profile and observation memory
Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.
Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.
Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.
* test(cli): include memory defaults in serve config stubs
* fix(memory): offload async worker launch blocking calls
Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.
* chore(memory): harden turn worker subagent guardrail
* chore(memory): refresh profile context per request
* fix(memory): offload async profile file reads
* fix(memory): offload async worker completion accounting
* feat: implement configurable sandbox execute timeout and enhance recovery instructions
* feat: add background process management tools and middleware for sandbox execution
* feat: enhance background process management with completion notifications and deduplication
* feat: enhance sandbox execution timeout validation and update related messages
* feat: enhance background process management with thread-specific completion notifications and HITL approval handling
* test: assert completion notification waits for process finish timestamp
* feat(cli): add --debug flag for verbose logging in serve mode
* feat(cli): add log_level config field with priority over env var
Replace dead `debug` parameter in `main()` with a proper `log_level`
config field in EvoScientistConfig. Enables `EvoSci config set log_level
debug` with priority: config file > EVOSCIENTIST_LOG_LEVEL env var.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
feat(prompts): update system prompt to eliminate numeric limits for sub-agents and delegation rounds
test(tests): adjust tests to reflect changes in configuration and onboarding logic
- Introduced UserMessage widget for displaying user input with a styled prompt.
- Updated onboarding steps to include UI backend selection (Rich CLI or Textual TUI).
- Modified EvoScientistConfig to store selected UI backend.
- Enhanced configuration handling to support UI backend environment variable.
- Updated README with new UI backend options and commands.
- Added tests for new UI backend functionality and UserMessage widget.
- Removed obsolete test files and ensured existing tests are updated accordingly.