b8dd1bf3a5b94de3eb705d6d94f289b47e4b2715
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a124d16764 |
perf: cut first-turn time-to-first-token by ~80% (all platforms) (#59332)
Four independent pre-request stalls sat on the critical path between prompt submission and the first streamed token, measured with cProfile against a live process: 1. Discord capability detection (~2.0s, worst 5s): get_tool_definitions -> _get_dynamic_schema made a BLOCKING https call to discord.com inside AIAgent.__init__ for any user with DISCORD_BOT_TOKEN set, on every platform, every cold process. Now non-blocking: memory cache -> 24h disk cache -> permissive default + one background detection that seeds the disk cache for the next process. The permissive default is pinned per-process so tool schemas never flip mid-conversation (prompt-cache safety); it mirrors the existing detection-failure fallback (all actions exposed, 403s enriched at call time). 2. Ollama /api/show probe (~0.3s): get_model_context_length step 5e POSTed to <base_url>/api/show for KNOWN providers (openrouter etc.), got a 404, and never cached the miss - so every fresh process paid a full HTTP round-trip. Known non-Ollama providers now skip the probe; local/custom/unknown endpoints keep the exact previous behavior. 3. env_probe subprocess sweep (~0.5s): the Python-toolchain probe ran 4-8 subprocess calls inside the FIRST system prompt build. Now warmed off-thread during agent init; the prompt build hits the cache (same lock, so a mid-flight warm just joins instead of recomputing). 4. tools.mcp_tool import (~0.4s): the between-turns MCP refresh in build_turn_context imported the whole mcp package even with zero MCP servers configured. MCP tools can only exist if tools.mcp_tool was already imported (discovery/reload paths), so gate the import on sys.modules membership - no behavior change for MCP users. CLI additionally pre-imports run_agent + openai off-thread during the idle banner window (same pattern as the /model picker prewarm), hiding the remaining ~1.5s of module imports while the user types. Fixes 1-4 apply to every interaction layer (CLI, gateway, TUI, desktop, cron). Measured cold first turn (submit -> request dispatched, openrouter, discord token set): 4.3s before -> 0.9s after CLI prewarm (~80%); the agent-side non-import cost drops 2.9s -> 0.36s (init) + 0.27s (turn prologue). |
||
|
|
d1f23bb2d5 |
fix: prevent TUI gateway stdin EOF crash across all TUI-context subprocess calls
When Hermes runs in TUI mode, the gateway child process communicates with the Node.js parent over a JSON-RPC protocol on stdin. Subprocess calls that inherit this stdin fd can trigger a race condition where the child's stdin read returns EOF, causing the gateway to exit cleanly (exit code 0) mid-tool- execution. This is the same root cause as issue #14036 (byterover plugin) and PR #39257 (SSH environment backend). This commit applies the fix — stdin=subprocess.DEVNULL — to all 85 subprocess.run() and subprocess.Popen() calls that execute inside the TUI gateway child process. Scope: TUI-context code only (agent/, tools/, plugins/, tui_gateway/server.py). CLI code (cli.py, hermes_cli/), tests, scripts, and gateway process management are excluded — they don't run inside the TUI child and inherit the terminal's stdin, not the JSON-RPC pipe. 85 call sites across 28 files. All files pass syntax check. |
||
|
|
a4d8f0f62a |
feat(prompt): universal task-completion guidance + local Python toolchain probe (#34340)
* fix(codex): surface error code in Responses 'failed' status errors
When a Codex Responses turn ends with status=failed, the response carries
the failure details under `response.error` as
`{code, message, param, ...}`. The previous extractor pulled only
`message`, so users seeing a rate-limit failure got a bare "Slow down"
string indistinguishable from a generic stream truncation; an
internal_error with empty message degraded to a dict dump
("{'code': 'internal_error', 'message': ''}").
Extract a `_format_responses_error()` helper that:
- prefixes `code` when both code and message are present
(e.g. 'rate_limit_exceeded: Slow down')
- falls back to the bare `code` when message is empty
- accepts both dict and attribute-style payloads (SDK and JSON-RPC paths)
- preserves the prior status-only fallback when no error payload exists
Apply the same helper at the sibling site in
`codex_app_server_session.run_turn()` so codex-CLI subprocess turn
failures get the same treatment.
Tests:
- 8 new unit tests for `_format_responses_error` covering both shapes,
empty/missing fields, non-string fields, and the status-only fallback.
- 2 regression tests on `_normalize_codex_response` for failed status
with and without a code, asserting the exact RuntimeError message.
- All 3603 tests in tests/agent/ pass.
Adapted from anomalyco/opencode#28757.
* feat(prompt): universal task-completion guidance + local Python toolchain probe
Two cross-model failure modes get a single-line answer in the cached
system prompt. Both gated by config (default on), both add zero overhead
when not needed, both verified via real AIAgent prompt builds.
## What changed
`TASK_COMPLETION_GUIDANCE` — short prompt block applied to ALL models.
Targets two failure modes observed on a real Sarasota real-estate build
task: (1) Opus stopped after writing an 85-byte stub and gave a prose
response with finish_reason=stop on call #3 of 90; (2) DeepSeek pushed
through a PEP-668 wall, then returned fabricated listings instead of
admitting the blocker. Both behaviors are model-family-agnostic, so the
guidance lives outside the existing tool_use_enforcement gate (~192
tokens, paid once per session via prefix cache).
`tools/env_probe.py` — local Python toolchain probe. Detects
python3/pip/uv/PEP-668 state and emits ONE short line in the system
prompt when something is non-default. Emits NOTHING when the env is
clean (zero token cost for normal users). Skipped entirely for remote
terminal backends (docker/modal/ssh) — they have their own probe.
Example output on a broken environment (the actual case):
Python toolchain: python3=3.11.15 (no pip module),
python=missing (use python3), pip→python3.12 (mismatch),
PEP 668=yes (use venv or uv).
## Config
Both flags live under `agent.` in config.yaml, default True:
agent:
task_completion_guidance: true # universal "finish the job" block
environment_probe: true # local Python toolchain hints
Neither addition required a `_config_version` bump — deep-merge fills
defaults in for existing user configs.
## Validation
| Test surface | Result |
|---|---|
| tests/tools/test_env_probe.py | 10/10 pass (probe unit) |
| tests/run_agent/test_run_agent.py — new classes | 8/8 pass (integration) |
| TestToolUseEnforcementConfig | 17/17 pass (no regression) |
| TestBuildSystemPrompt | 9/9 pass (no regression) |
| TestInvalidateSystemPrompt | 2/2 pass (no regression) |
| tests/agent/test_prompt_builder.py | 124/124 pass (no regression) |
| tests/hermes_cli/ | 5662/5662 pass (config defaults) |
| E2E AIAgent build (broken env) | Both blocks present, 2,178 chars |
| E2E AIAgent build (clean env) | 771-char net overhead, env probe silent |
|