Commit Graph

31 Commits

Author SHA1 Message Date
Xiaohui Yan 3c5cc831c0 Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400)

WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.

Adds two config fields with deliberately different defaults:

  webui_host        = 0.0.0.0    front-end serves the app shell, no secrets
  langgraph_dev_host = 127.0.0.1  unauthenticated API, agent can run shell

The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.

  - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
    host kwarg threaded through the probes and `start_langgraph_dev`,
    which now emits `--host` and propagates
    EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
  - sdk.py: `langgraph_dev_url` tracks host as well as port;
    EvoScientist.py reuses it instead of an inline f-string
  - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
    banner whenever the bind is not provably loopback
  - webui.py: forwards both hosts; the front-end is widened via HOSTNAME
    because @evoscientist/webui ships no --host flag — its bin launcher
    does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
    gated on the backend host only, so the shipped front-end default
    doesn't print a banner on every launch

Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.

Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400)

Completes the remaining items from #400.

  - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
    Remote WebUI use needs both anyway (the UI reaches the backend from the
    browser, not server-side), so a loopback backend default just meant every
    remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
    `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
    these servers listen.

    SECURITY: this exposes an unauthenticated API whose agent can run shell
    commands. The red PUBLIC BIND banner consequently fires on every launch
    while exposed — kept deliberately, since the exposure is real and the
    escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
    only discoverable if we say so. READMEs now lead with the warning and
    document the SSH-tunnel alternative.

  - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
    WebUI mode they are two halves of one surface; moving only one leaves the
    UI loading but unable to reach the agent. Blank values are dropped rather
    than written as an empty override that would beat the config file.

  - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
    `http://localhost:{port}` (steps.py:160, :223) — both render the
    configured bind through `_base_url` / `_format_hostport`, so a pinned
    interface is reported honestly and a wildcard still shows loopback.

Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.

Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime

GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.

Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.

actions/checkout@v5 is already node24 and needs no change.

Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): correct --host help text and warn on public bind in non-WebUI modes

The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.

The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.

READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(deploy): strip the config-derived bind host, not just the CLI one

`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.

Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.

Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.

Three regression tests added, each verified to fail against the old code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style: apply ruff format to the bind-host changes

The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): keep the langgraph dev backend on loopback by default

The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.

webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.

Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 16:15:49 +08:00
nb213 4ddf7ebe52 Add Atlas Cloud LLM provider (#388)
* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-29 14:09:44 +01:00
jfilipiuk 10c032450e fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-28 12:26:22 +00:00
Mani Saint-Victor da6ca38d53 fix(llm): respect reasoning_effort setting on native OpenAI path (#321)
* fix(llm): respect reasoning_effort setting on native OpenAI path

The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.

Adds a regression test and isolates the existing xhigh test from the
env var.

* fix(llm): preserve model reasoning defaults

* fix(llm): preserve GPT-5.6 reasoning default

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:42:55 +00:00
Mani Saint-Victor f3e65a446f fix(tests): isolate the developer's real .env from the test suite (#329)
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True),
override=True), so any test that loads config injected the repo's real
.env into os.environ for the rest of the pytest process. An
empty-valued line like MINIMAX_BASE_URL= then made
os.environ.get(key, default) return '' instead of the default,
failing the MiniMax routing tests in full-suite runs while they
passed in isolation.

Generalizes the find_dotenv redirect that test_config.py's
temp_config_dir fixture already applied locally into a suite-wide
autouse fixture, pointing at a never-created path so tests writing
their own tmp_path/.env cannot collide with it. Adds a regression
test reproducing the leak.

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:27:40 +00:00
Zixin Dong f72f7b93d5 feat(llm): add OpenRouter app attribution headers (#339) (#344)
* feat(llm): add OpenRouter app attribution headers (#339)

Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.

Closes #339

* refactor(llm): centralize OpenRouter attribution defaults + cap categories

Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
  (the config fields and llm/models.py both use them) instead of duplicating
  the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
  sent list to OpenRouter's 2-per-request limit, warning when a configured list
  exceeds it, so extras are dropped predictably (and surfaced) here rather than
  being silently truncated server-side.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 11:08:59 +01:00
dinos a7b9e175c1 fix(config): set config.yaml permissions to 0x600 (#336) 2026-07-08 19:25:01 +01:00
renaissancefieldlite f086d77756 Fix UTF-8 config reads on Windows (#318)
* Fix UTF-8 config reads on Windows

* test: cover utf8 production loaders

* fix: read and write settings as utf8

* Apply ruff formatting
2026-07-05 05:10:24 +00:00
dinos 2b244888ec feat(memory): autoskills (#319) 2026-07-03 11:16:33 +02:00
Xi Zhang 7ccfe68f3f feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management

- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.

* fix(async-notifier): ensure fallback hint is used for unknown notification kinds

* feat: enhance scheduling functionality and improve system message handling
2026-06-25 17:31:23 +01:00
dinos 6eb467e70b feat(llm): make OpenRouter Anthropic prompt cache opt-out (#299)
* feat(llm): make OpenRouter Anthropic prompt cache opt-out

* fix(llm): restore truthy env flag helper
2026-06-18 15:09:22 +02:00
dinos f356de36a6 feat: memory retrieval (#281) 2026-06-16 09:14:29 +02:00
Xi Zhang c02be519f6 feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks

- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.

* feat(dangerous-mode): enhance logging and environment management for dangerous mode

* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
2026-06-10 18:19:13 +01:00
dinos cd2baa9588 feat(models): add opt-in prompt caching support for anthropic via openrouter (#272)
* feat(models): add opt-in prompt caching support for anthropic via openrouter

* chore: don't coerce model_kwargs to dict
2026-06-09 15:13:49 +01:00
Xi Zhang 63969b596d Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack

* feat(models): add qwen3.7-plus model entry and update context window comment

* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope

* feat(auxiliary): implement auxiliary model support for background tasks and tool selection

- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.

* feat(steps): update UI backend selection options and descriptions

* Refactor code structure for improved readability and maintainability

* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors

* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py

* feat(config): add auxiliary model and provider environment variables to test setup
2026-06-07 00:52:59 +01:00
dinos 92d95dee68 feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle

Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.

Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.

Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.

* fix(cli): sync background agent server on resume

Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.

Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.

Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.

Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.

* fix(cli): prepare serve resume workspace before adopting

Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.

* fix(memory): untrack abandoned worker status watches

Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.

* fix(cli): report channel command failures accurately

Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.

* fix(stream): clear memory counters for resume streams

Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.

* docs(tools): make observation recording guidance conditional

Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.

* feat(config): add controls for profile and observation memory

Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.

Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.

Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.

* test(cli): include memory defaults in serve config stubs

* fix(memory): offload async worker launch blocking calls

Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.

* chore(memory): harden turn worker subagent guardrail

* chore(memory): refresh profile context per request

* fix(memory): offload async profile file reads

* fix(memory): offload async worker completion accounting
2026-06-05 15:11:20 +01:00
Xi Zhang a13904185d Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions

* feat: add background process management tools and middleware for sandbox execution

* feat: enhance background process management with completion notifications and deduplication

* feat: enhance sandbox execution timeout validation and update related messages

* feat: enhance background process management with thread-specific completion notifications and HITL approval handling

* test: assert completion notification waits for process finish timestamp
2026-05-31 15:11:25 +01:00
X-iZhang 721a03c25b feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments 2026-05-29 00:40:46 +01:00
Ziheng Zhang 3e493233cf feat(channels): simplify debug tracing and add serve debug mode (#143)
* feat(channels): simplify debug tracing and add serve debug mode

* Fix

* fix(channels): remove serve loop patch
2026-04-09 12:47:48 +02:00
Xiaohui Yan e89b71aa60 feat(cli): add --debug flag for verbose logging in serve mode (#141)
* feat(cli): add --debug flag for verbose logging in serve mode

* feat(cli): add log_level config field with priority over env var

Replace dead `debug` parameter in `main()` with a proper `log_level`
config field in EvoScientistConfig. Enables `EvoSci config set log_level
debug` with priority: config file > EVOSCIENTIST_LOG_LEVEL env var.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-08 13:23:10 +08:00
Jan Piotrowski 4a3d6c0318 chore: add ruff lint rules and turn on formatting 2026-03-19 17:04:02 +01:00
X-iZhang 62fe970476 feat: add OpenAI support with OAuth and environment configuration 2026-03-16 00:25:19 +00:00
X-iZhang c5a4d559a2 Refactor test cases for improved readability and consistency
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
2026-03-15 21:12:51 +00:00
X-iZhang 1ce2d059ca feat: implement OAuth support for Anthropic using ccproxy and add related configuration options 2026-03-15 16:52:11 +00:00
X-iZhang 07e57e110b feat(ui): update UI backend options to use 'cli' and 'tui' 2026-03-08 15:34:22 +00:00
X-iZhang 2f6cf824f2 refactor(config): remove max_concurrent and max_iterations from config and onboarding steps
feat(prompts): update system prompt to eliminate numeric limits for sub-agents and delegation rounds
test(tests): adjust tests to reflect changes in configuration and onboarding logic
2026-02-27 18:57:47 +00:00
X-iZhang ea00578cee feat(onboarding): add Ollama provider support with connection validation 2026-02-24 19:45:02 +00:00
X-iZhang 03f4134190 feat: add UserMessage widget and UI backend selection during onboarding
- Introduced UserMessage widget for displaying user input with a styled prompt.
- Updated onboarding steps to include UI backend selection (Rich CLI or Textual TUI).
- Modified EvoScientistConfig to store selected UI backend.
- Enhanced configuration handling to support UI backend environment variable.
- Updated README with new UI backend options and commands.
- Added tests for new UI backend functionality and UserMessage widget.
- Removed obsolete test files and ensured existing tests are updated accordingly.
2026-02-21 01:18:33 +00:00
X-iZhang 80baf85a2c update 2026-02-05 23:58:55 +00:00
X-iZhang 97f104015b update 2026-02-04 23:53:24 +00:00
X-iZhang f46e328545 update 2026-02-04 23:45:09 +00:00