Files
kalisgd0h bd54eaa0a4 feat(cli): add --output-format stream-json for headless clients (#309)
* feat(cli): add --output-format stream-json for headless clients

Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.

- stream/json_sink.py: write_events_as_json + stream_json sink, plus
  redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
  the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
  requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cli): honor explicit --no-auto-mode over config in stream-json

Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 04:53:52 +00:00

82 lines
5.2 KiB
Markdown

# `stream-json` output protocol
`EvoSci --output-format stream-json` runs a single-shot (`-p`) session and emits
EvoScientist's **native event stream** as line-delimited JSON (JSONL) on stdout —
one self-describing JSON object per line. This is the integration surface for
programmatic clients that drive EvoScientist headlessly (for example, an agent
runtime that assigns work and renders live progress).
```bash
EvoSci -p "Summarize the attention mechanism into notes.md" \
--output-format stream-json \
--auto-mode \
--workdir /path/to/task
```
## Stream contract
- **stdout is pure JSONL.** Every line is one complete JSON object. Parse it with
a line reader + `json.loads` per line. Nothing else is written to stdout.
- **stderr carries everything human** — status lines ("Loading agent…"),
separators, the resume hint, and error panels. A consumer should treat stderr
as logs, not protocol.
- **Each object has a `type` field** used for dispatch. A consumer that does not
recognize a `type` should **ignore that line** rather than fail — new event
types may be added over time, and the protocol is forward-compatible by design.
- **The stream ends with a `done` event** carrying the final response text. On an
unhandled failure an `error` event is emitted instead/also.
- **Unattended by default.** `stream-json` is headless, so `--auto-mode` is
enabled automatically: approval and `ask_user` gates are auto-handled and the
run proceeds straight to its `done` event. Pass `--no-auto-mode` to opt out —
a human-in-the-loop `interrupt` / `ask_user` is then emitted as a normal event
and the single-shot run ends right after it (`… → interrupt → done → EOF`). It
does **not** block waiting for input, but it also stops before finishing the
task; answering the event requires re-invoking with `--resume` (experimental).
## Event types
Each line is a self-contained JSON object with a `type` field. Most fields are
scalars, but some events carry nested payloads (e.g. `args`, `action_requests`,
`questions`) — parse each line as a full object, not a flat key/value map. The
fields beyond `type` are listed below.
| `type` | Fields | Meaning |
|------------------------|------------------------------------------------------------|---------|
| `thinking` | `content`, `id` | Model reasoning text |
| `text` | `content` | Assistant output text |
| `tool_call` | `name`, `args`, `id` | Tool invocation |
| `tool_result` | `name`, `content`, `success`, `id` | Tool result (`id` matches the `tool_call`) |
| `subagent_start` | `name`, `description` | Sub-agent delegation begins |
| `subagent_tool_call` | `subagent`, `name`, `args`, `id` | Tool call inside a sub-agent |
| `subagent_tool_result` | `subagent`, `name`, `content`, `success`, `id` | Tool result inside a sub-agent |
| `subagent_text` | `subagent`, `content`, `instance_id` | Text from a sub-agent |
| `subagent_end` | `name` | Sub-agent delegation completes |
| `tool_selection` | `tools` | Tool-selector middleware picked tools |
| `summarization_start` | — | Context summarization begins |
| `summarization` | `content` | Context summarization output |
| `usage_stats` | `input_tokens`, `output_tokens` | Token usage |
| `interrupt` | `interrupt_id`, `action_requests`, `review_configs` | HITL approval interrupt |
| `ask_user` | `interrupt_id`, `questions`, `tool_call_id` | Agent-initiated clarifying question |
| `error` | `message` | Error during the run |
| `done` | `content`, `response` | Final response; end of stream |
## Example transcript
```jsonl
{"type": "thinking", "content": "I should write the notes file.", "id": 0}
{"type": "tool_call", "name": "write_file", "args": {"path": "notes.md", "content": "..."}, "id": "call_1"}
{"type": "tool_result", "name": "write_file", "content": "wrote 412 bytes", "success": true, "id": "call_1"}
{"type": "text", "content": "Done — notes.md now summarizes attention."}
{"type": "usage_stats", "input_tokens": 5123, "output_tokens": 388}
{"type": "done", "content": "Done — notes.md now summarizes attention.", "response": "Done — notes.md now summarizes attention."}
```
## Notes for client implementers
- Read stdout line by line; do not assume a single JSON document.
- Accumulate `text` events for the running assistant message; the terminal
`done.response` is the authoritative final text.
- `tool_call` / `tool_result` correlate by `id`.
- Token usage may arrive across multiple `usage_stats` events; sum them.
- Treat unknown `type` values and unknown fields as non-fatal.