Files
EvoScientist-Multi/tests/test_cli_tui_dispatch.py
T
Xiaohui Yan 3c5cc831c0 Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400)

WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.

Adds two config fields with deliberately different defaults:

  webui_host        = 0.0.0.0    front-end serves the app shell, no secrets
  langgraph_dev_host = 127.0.0.1  unauthenticated API, agent can run shell

The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.

  - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
    host kwarg threaded through the probes and `start_langgraph_dev`,
    which now emits `--host` and propagates
    EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
  - sdk.py: `langgraph_dev_url` tracks host as well as port;
    EvoScientist.py reuses it instead of an inline f-string
  - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
    banner whenever the bind is not provably loopback
  - webui.py: forwards both hosts; the front-end is widened via HOSTNAME
    because @evoscientist/webui ships no --host flag — its bin launcher
    does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
    gated on the backend host only, so the shipped front-end default
    doesn't print a banner on every launch

Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.

Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400)

Completes the remaining items from #400.

  - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
    Remote WebUI use needs both anyway (the UI reaches the backend from the
    browser, not server-side), so a loopback backend default just meant every
    remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
    `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
    these servers listen.

    SECURITY: this exposes an unauthenticated API whose agent can run shell
    commands. The red PUBLIC BIND banner consequently fires on every launch
    while exposed — kept deliberately, since the exposure is real and the
    escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
    only discoverable if we say so. READMEs now lead with the warning and
    document the SSH-tunnel alternative.

  - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
    WebUI mode they are two halves of one surface; moving only one leaves the
    UI loading but unable to reach the agent. Blank values are dropped rather
    than written as an empty override that would beat the config file.

  - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
    `http://localhost:{port}` (steps.py:160, :223) — both render the
    configured bind through `_base_url` / `_format_hostport`, so a pinned
    interface is reported honestly and a wildcard still shows loopback.

Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.

Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime

GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.

Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.

actions/checkout@v5 is already node24 and needs no change.

Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): correct --host help text and warn on public bind in non-WebUI modes

The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.

The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.

READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(deploy): strip the config-derived bind host, not just the CLI one

`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.

Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.

Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.

Three regression tests added, each verified to fail against the old code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style: apply ruff format to the bind-host changes

The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): keep the langgraph dev backend on loopback by default

The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.

webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.

Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 16:15:49 +08:00

281 lines
10 KiB
Python

"""Tests for CLI interactive UI backend dispatch."""
from types import SimpleNamespace
import pytest
from EvoScientist.cli.commands import _is_fresh_interactive_session
from EvoScientist.cli.interactive import cmd_interactive
@pytest.mark.parametrize(
("prompt", "thread_id", "expected"),
[
(None, None, True), # bare `EvoSci` → fresh → WebUI launches
("what is 1+1", None, False), # `-p` one-shot → terminal (Rich CLI)
(None, "47bcffcd", False), # `--resume <id>` → terminal (Rich CLI)
("hi", "47bcffcd", False), # both → terminal
("", None, True), # empty `-p` is falsy → treated as fresh
],
)
def test_is_fresh_interactive_session(prompt, thread_id, expected):
"""WebUI only launches for a fresh interactive session; `-p` / `--resume`
fall back to the terminal."""
assert _is_fresh_interactive_session(prompt, thread_id) is expected
def _invoke_main(monkeypatch, argv):
"""Invoke the EvoSci main callback with ui_backend=webui and all heavy setup
mocked. Returns (calls, result): calls["dispatch"] is "webui" if run_webui
ran, or ("cli", <ui_backend>) if cmd_interactive ran."""
from typer.testing import CliRunner
import EvoScientist.cli.commands as cmds
import EvoScientist.cli.interactive as interactive_mod
import EvoScientist.config as cfg_mod
import EvoScientist.deploy.webui as webui_mod
from EvoScientist.cli._app import app
from EvoScientist.config.settings import EvoScientistConfig
calls: dict[str, object] = {}
def _fake_config(overrides):
cfg = EvoScientistConfig()
calls["overrides"] = dict(overrides or {})
# Apply every override the real merge would, so tests can assert on
# flags (like --host) that reach the config rather than the callback.
for key, value in (overrides or {}).items():
setattr(cfg, key, value)
# Mirror the real --ui override; default to webui for this test.
cfg.ui_backend = overrides.get("ui_backend") or "webui"
return cfg
def _fake_run_webui(config, **_kw):
calls["dispatch"] = "webui"
calls["webui_config"] = config
monkeypatch.setattr(cfg_mod, "get_effective_config", _fake_config)
monkeypatch.setattr(cfg_mod, "apply_config_to_env", lambda cfg: None)
monkeypatch.setattr(cmds, "ensure_dirs", lambda: None)
monkeypatch.setattr(cmds, "_ensure_async_subagent_server", lambda *a, **k: None)
monkeypatch.setattr(webui_mod, "run_webui", _fake_run_webui)
monkeypatch.setattr(
interactive_mod,
"cmd_interactive",
lambda **kw: calls.__setitem__("dispatch", ("cli", kw.get("ui_backend"))),
)
result = CliRunner().invoke(app, argv, catch_exceptions=False)
return calls, result
def test_main_callback_launches_webui_for_fresh_session(monkeypatch):
"""Bare `EvoSci` with ui_backend=webui opens the browser app."""
calls, result = _invoke_main(monkeypatch, [])
assert result.exit_code == 0
assert calls.get("dispatch") == "webui"
def test_main_callback_resume_falls_back_to_cli(monkeypatch):
"""`EvoSci --resume <id>` with ui_backend=webui does NOT open the browser;
it resumes the conversation in the Rich CLI (ui_backend forced to 'cli')."""
calls, result = _invoke_main(monkeypatch, ["--resume", "abc123"])
assert result.exit_code == 0
assert calls.get("dispatch") == ("cli", "cli")
# =============================================================================
# --host override
# =============================================================================
def test_host_flag_drives_both_servers(monkeypatch):
"""One flag, both halves. In WebUI mode the front-end and backend are two
halves of one surface, so `--host` has to move them together — widening
only one leaves the UI loading but unable to reach the agent."""
calls, result = _invoke_main(monkeypatch, ["--host", "0.0.0.0"])
assert result.exit_code == 0
cfg = calls["webui_config"]
assert cfg.webui_host == "0.0.0.0"
assert cfg.langgraph_dev_host == "0.0.0.0"
def test_host_flag_is_stripped(monkeypatch):
calls, result = _invoke_main(monkeypatch, ["--host", " 192.168.1.5 "])
assert result.exit_code == 0
assert calls["webui_config"].langgraph_dev_host == "192.168.1.5"
def test_blank_host_flag_leaves_config_defaults(monkeypatch):
"""An all-whitespace value must not write an unusable empty host into the
override dict, where it would beat the config file."""
calls, result = _invoke_main(monkeypatch, ["--host", " "])
assert result.exit_code == 0
assert "langgraph_dev_host" not in calls["overrides"]
assert calls["webui_config"].langgraph_dev_host == "127.0.0.1"
def test_no_host_flag_leaves_config_defaults(monkeypatch):
calls, result = _invoke_main(monkeypatch, [])
assert result.exit_code == 0
assert "webui_host" not in calls["overrides"]
assert calls["webui_config"].webui_host == "0.0.0.0"
def _run_ensure_backend(monkeypatch, config, *, server_up=True):
"""Drive ``_ensure_async_subagent_server`` and capture console output."""
import EvoScientist.cli.commands as cmds
printed: list[str] = []
monkeypatch.setattr(
"EvoScientist.langgraph_dev.manager.ensure_langgraph_dev",
lambda config, *, workspace_dir: None,
)
monkeypatch.setattr(
"EvoScientist.langgraph_dev.manager.is_async_subagents_available",
lambda: server_up,
)
monkeypatch.setattr(cmds, "_reconcile_autoskill_schedule", lambda *a, **k: None)
monkeypatch.setattr(
cmds.console, "print", lambda *a, **k: printed.append(str(a[0]) if a else "")
)
monkeypatch.setattr(
cmds.console,
"status",
lambda *a, **k: __import__("contextlib").nullcontext(),
)
cmds._ensure_async_subagent_server(config, workspace_dir="/tmp/workspace")
return printed
@pytest.mark.parametrize("exposed", ["0.0.0.0", "192.168.1.5", "::"])
def test_cli_mode_warns_on_public_backend_bind(monkeypatch, exposed):
"""The langgraph dev backend is shared across UI modes, so a plain
`EvoSci` session must warn too — otherwise `--host 0.0.0.0` (or a config
file with it) puts an unauthenticated shell-capable API on the network in
every mode with no signal."""
config = SimpleNamespace(langgraph_dev_host=exposed)
printed = _run_ensure_backend(monkeypatch, config)
assert any("PUBLIC BIND" in line for line in printed)
@pytest.mark.parametrize("loopback", ["127.0.0.1", "::1", "localhost"])
def test_cli_mode_silent_on_loopback_backend_bind(monkeypatch, loopback):
config = SimpleNamespace(langgraph_dev_host=loopback)
printed = _run_ensure_backend(monkeypatch, config)
assert not any("PUBLIC BIND" in line for line in printed)
def test_no_warning_when_backend_failed_to_start(monkeypatch):
"""ensure_langgraph_dev fails soft (async degrades to in-process). Warning
about a bind that never happened is worse than saying nothing."""
config = SimpleNamespace(langgraph_dev_host="0.0.0.0")
printed = _run_ensure_backend(monkeypatch, config, server_up=False)
assert not any("PUBLIC BIND" in line for line in printed)
def test_background_agent_server_starts_even_when_async_subagents_disabled(
monkeypatch,
):
import EvoScientist.cli.commands as cmds
calls = []
def fake_ensure(config, *, workspace_dir):
calls.append((config, workspace_dir))
monkeypatch.setattr(
"EvoScientist.langgraph_dev.manager.ensure_langgraph_dev",
fake_ensure,
)
config = SimpleNamespace(enable_async_subagents=False)
cmds._ensure_async_subagent_server(config, workspace_dir="/tmp/workspace")
assert calls == [(config, "/tmp/workspace")]
async def test_resume_workspace_sync_runs_even_when_async_subagents_disabled(
monkeypatch,
):
import EvoScientist.cli.commands as cmds
calls = []
def fake_ensure(config, *, workspace_dir):
calls.append((config, workspace_dir))
monkeypatch.setattr(
"EvoScientist.langgraph_dev.manager.ensure_langgraph_dev",
fake_ensure,
)
config = SimpleNamespace(enable_async_subagents=False)
await cmds._sync_background_agent_server_workspace(
config,
workspace_dir="/tmp/resumed-workspace",
)
assert calls == [(config, "/tmp/resumed-workspace")]
def test_cmd_interactive_dispatches_to_textual(monkeypatch):
captured: dict[str, object] = {}
captured_kwargs: list[dict[str, object]] = []
effective_config = SimpleNamespace(langgraph_dev_port=9999)
def _fake_resolve_ui_backend(value, *, warn_fallback=False):
captured["resolved_input"] = value
captured["warn_fallback"] = warn_fallback
return "tui"
def _fake_run_textual_interactive(**kwargs: object):
captured_kwargs.append(kwargs)
monkeypatch.setattr(
"EvoScientist.cli.interactive.resolve_ui_backend",
_fake_resolve_ui_backend,
)
monkeypatch.setattr(
"EvoScientist.cli.interactive.run_textual_interactive",
_fake_run_textual_interactive,
)
cmd_interactive(
show_thinking=True,
channel_send_thinking=True,
workspace_dir="/tmp/workspace",
workspace_fixed=True,
mode="daemon",
model="demo-model",
provider="demo-provider",
run_name="demo-run",
thread_id="thread-1",
ui_backend="tui",
config=effective_config,
)
assert captured["resolved_input"] == "tui"
assert captured["warn_fallback"] is True
assert len(captured_kwargs) == 1
kwargs = captured_kwargs[0]
assert kwargs["workspace_dir"] == "/tmp/workspace"
assert kwargs["workspace_fixed"] is True
assert kwargs["mode"] == "daemon"
assert kwargs["model"] == "demo-model"
assert kwargs["provider"] == "demo-provider"
assert kwargs["run_name"] == "demo-run"
assert kwargs["thread_id"] == "thread-1"
assert kwargs["config"] is effective_config
assert kwargs["channel_send_thinking"] is True
assert callable(kwargs["load_agent"])
assert callable(kwargs["create_session_workspace"])