Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.
The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.
Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.
- agent/review_engine.py: shared engine (snapshot, briefing,
auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
(never model-facing) resolved through the same credential system as
delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
slash_commands handler (binds the approval session key so the
completion routes back), TUI/Desktop live dispatch in
tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
dispatch tests fail without the fix), 4 gateway handler tests
through the real async rail
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.
Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.
Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.
- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
`unproven_worktree_payload()`. Both producers of this schema now call them,
so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
hand-building the dict (-16 lines). The re-import is guarded: the outer
`except` can be entered because the `from tools import subagent_worktree`
itself failed, in which case the name is unbound -- an inline fallback keeps
the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
compares its key set against live `finalize_subagent_worktree()` output. Same
contract, no source reading, refactor-proof, and it actually executes the
code.
Also folded from the same review:
- Fail-closed on an unmeasurable commit count. With no `base_commit` the
rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
tree still reached `git worktree remove --force` + `git branch -D` -- the
exact bug class #88113 is about, on a public function that takes a
caller-supplied dict. Now returns un-inspected instead, with a test driving a
real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
names WHICH probe failed, its exit code, and a bounded git stderr tail, so
the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
removed the duplicated index-corruption block in favor of the existing
`_break_git_index()` helper.
Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.
- The distinguishability test asserted the failure payload was byte-identical
to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
That freezes the AMBIGUITY as a required property: emitting `commits: None`
for "unknown" would improve exactly what #88113 is about and fail the test.
Now asserts what the parent actually depends on -- both keep the worktree,
and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
happy path, forbidding an always-present-but-False flag (a legitimately
better JSON contract: stable key set for serializers). Now
`assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
was not even a cross-producer contract: delegate_tool's note said "state
unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
assert the note names the worktree AND branch -- the actionable part for a
human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
before any git call would keep it green while proving nothing). Now checks
`call_count` and mirrors the branch-survival + note-names-path legs its
sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
parent reads, but delegate_tool's fallback -- the second producer -- was
verified only by reading. It now AST-parses the real fallback dict literal
and compares against live `finalize_subagent_worktree()` output, so the two
producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
raising, handled in delegate_tool), and the module docstring listed
`inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
`_break_git_index()` beside the file's other module-level helpers.
Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
The preserved worktree is invisible to the only consumer that can act on it.
Completes the #88113 fix. That change correctly stops the destructive prune
when a git probe fails, but still returns commits=0 / dirty=False -- values
that were never measured. Those are the defaults the prune used to delete on,
so the failure payload is byte-identical to "inspected fine, child left
nothing":
inspection FAILED, uncommitted work kept -> {commits: 0, dirty: False, pruned: False}
inspected OK, child produced nothing -> {commits: 0, dirty: False, pruned: False}
The only failure signal was a logger.warning, and the sole consumer of this
payload is the parent agent reading the serialized delegate_task entry -- it
cannot read logs (no in-repo code reads the key back). So the parent's rational
reading of the failure case is "the child produced no work", which is the exact
wrong conclusion: a worktree possibly full of uncommitted work is preserved and
then never looked at. The data survives but nobody is told to recover it.
Changes:
- subagent_worktree: one _unproven() helper stamps inspection_failed + a note
naming the worktree/branch, warns, and returns the payload. Both unproven
exits route through it, so they cannot drift apart again.
- subagent_worktree: the pre-existing exception path (timeout, OSError, a
non-numeric rev-list stdout) produced the same unproven payload but logged at
DEBUG -- effectively silent. It now takes the same flagged path as a non-zero
exit; identical outcomes get identical reporting.
- delegate_tool: the caller's finalize-raised fallback assigned the
creation-side metadata dict (path/branch/repo_root/base_commit) -- a disjoint
schema missing commits/dirty/pruned. It now emits the same flagged shape, and
logs at WARNING.
- Docs + docstring + module contract now state that pruning requires
affirmative proof, so a future cleanup doesn't "fix" the preserved worktree
by restoring the unconditional prune and reintroducing this P1.
Purely additive: the happy-path payload shape is unchanged, so no existing
reader can break.
Validation:
- 18/18 tests/tools/test_subagent_worktree.py; 127 passed across the delegation
suites (test_delegate, batch_validation, control_actions, timeout_diagnostic).
- 3 new guards mutation-checked: neutering the flag fails all three; reverting
the production file to pre-fix main fails all three. Restores checksum-verified.
- E2E on real git: inspection-failure now returns inspection_failed=true with
work intact on disk; proven-clean still prunes (pruned=true).
When delegation.provider/model is explicitly pinned, the child no longer
inherits the parent's fallback chain: a mid-run auth/429 failure on the
pin previously rerouted the quiet-mode child onto parent fallback models
with no surfaced signal. Same treatment as the existing override_provider
OpenRouter filter-clearing — explicit pins are honored or fail loudly.
Also upgrades the pinned delegation.command-missing-from-PATH case from
warning + silent transport fallback to a loud spawn refusal, both at
credential preflight and in _build_child_agent.
Fixes#80450 (tracker #79686 audit item).
delegation.max_concurrent_children caps how many delegated children run in
parallel per batch (and concurrent background delegation units). The old default
of 3 needlessly serialized independent fan-outs (e.g. reviewing/investigating N
PRs or issues at once), so large batches ran in slow chunks of 3.
Raise the shipped default to 10, which sits at/below the existing high-cost
advisory threshold (>10), so the default never trips the warning. Each child
still consumes API tokens independently, so this is a throughput/latency win the
user pays for in parallel token spend — the floor stays 1 and there is no
ceiling, so anyone can tune it down or up.
- config_defaults.py: default 3 -> 10; _config_version 36 -> 37.
- delegate_tool.py: _DEFAULT_MAX_CONCURRENT_CHILDREN 3 -> 10 (+ docstring).
- config_migrations.py: _migrate_to_37 lifts configs pinned at exactly the old
default 3 to 10 (deliberate non-3 overrides preserved; unset inherits 10).
- cli-config.yaml.example: documented default updated.
Verified: default/fallback read 10, version 37, and the migration lifts 3->10,
preserves an explicit 5, and leaves unset untouched.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
A bare SessionDB() resolves the launch profile's default state.db, but
parents can hold non-default per-profile handles (tui_gateway opens
SessionDB(db_path=<profile_home>/state.db) for non-launch profiles and
hands them to agents via _transfer_db_to_agent). A child of such a
parent would write its transcript into the WRONG database — cross-
profile leakage that breaks parent_session_id lineage and
session_search. Open the dedicated handle at the parent handle's
db_path instead (AsyncSessionDB forwards .db_path via __getattr__, so
the gateway wrapper path works too). Regression test verified RED on
the pre-fix code.
Review follow-up (cc3f18197): if AIAgent() raises inside _build_child_agent
the freshly-opened dedicated handle has no owner and no child close() will
ever run — release it on the exception path so the sqlite fds don't
outlive the failed spawn. Also pin the degradation contract with a test:
a parent without a SessionDB still yields session_db=None children.
Cron run_job closes its per-job SessionDB in its finally block while a
fire-and-forget background delegation subagent is still flushing on a
daemon thread. The child shared the parent's SessionDB object, so every
subsequent flush hit the closed handle ('NoneType' object has no
attribute 'execute') and the child's whole transcript was silently
dropped. The same teardown-while-child-alive shape exists on gateway
session end and /new mid-delegation.
Each child now opens its own SessionDB connection (owned flag set at
construction so child.close() releases it), so no parent teardown can
close the child's handle out from under it.
Regression test proves the child gets a distinct live handle that
survives the parent's close().
A delegated subagent that exhausts its per-child iteration budget
(delegation.max_iterations) still returns a summary, so the result carries
status='completed' even though the child's exit_reason is 'max_iterations' and
its work was cut off mid-task. The parent then reads 'completed', trusts the
partial summary, and only discovers the truncation by parsing the prose (where
the child happens to mention 'hit the iteration limit'). That wastes parent
turns and risks acting on incomplete work.
exit_reason is already computed authoritatively and threaded to every
parent-visible surface; it just wasn't reflected anywhere the parent reads at a
glance. This surfaces it:
- delegate_tool.py: add a parent-visible boolean 'truncated' (= exit_reason ==
'max_iterations') to each task entry, alongside the existing exit_reason.
- process_registry._format_async_delegation: for both the batch and single-task
paths, when truncated -> use a warning icon, append
'TRUNCATED: hit max_iterations — work may be incomplete' to the header/Status
line, and prefix the summary with an unmissable truncation notice. status
semantics are left unchanged (stays 'completed') so existing icon/summary
branch logic and ~10 tests asserting status=='completed' stay valid.
Tests: single-task truncated -> banner; single-task clean -> no banner; batch
marks only the truncated task, not its clean sibling. 23/23 in the async-
delegation suite.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
delegation.max_iterations is the per-subagent tool-call budget. The old
default of 50 truncated substantial delegated work: leaf agents spend
~15-20 turns on reconnaissance before producing output, then ran out of
budget mid-task and returned 'completed but unfinished' summaries. 250
gives real delegated work room to finish.
Changes:
- config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36
- tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in
sync with the shipped default to prevent drift)
- config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly
the OLD default 50 -> 250 on update, so existing installs inherit the new
headroom. Any other explicit value (deliberate override) is preserved;
unset inherits 250 at read time.
- cli-config.yaml.example: doc the new default
The cap is per-child and children run concurrently (max_concurrent_children
default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds
(default 0 = off) remains available as a wall-clock guardrail, and users can
still pin a lower max_iterations explicitly.
Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset
untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
Spill files (terminal overflow, hook context, web_extract full text,
subagent summaries) were written with plain open()/write_text into
predictable directories. A pre-planted symlink at any of those paths
redirected the write onto an arbitrary user-owned file, and raw
pre-redaction terminal/hook spills landed world-readable under the
default umask.
New tools/spill_safety.py helpers create files with
O_CREAT|O_EXCL|O_NOFOLLOW (a link-shaped path fails the write instead of
following it) and overwrite via lstat-checked unlink + exclusive
re-create, so even the redaction rewrite cannot be diverted. Private
tier (0o700 dir / 0o600 file) covers raw terminal and hook spills;
cache/web and cache/delegation keep umask perms because those dirs are
bind-mounted into remote backends that must read them.
Pattern borrowed from DeepSeek Harness dsh-spill-local (MIT):
private root + exclusive owner-only opens for spill artifacts.
delegate_task gains a control plane: action='list' / 'steer' / 'stop'
let the parent agent see, redirect, and early-stop its own running
subagents mid-flight — the model-facing counterpart of the TUI's
delegation.pause / subagent.interrupt / subagent.steer RPCs.
- action='list': live children of this conversation's spawn tree
(ids, goal, status, running_seconds, accepting_steer, live
transcript path). Ownership is enforced via a _delegate_parent_ref
weakref chain stamped at child build time, so a conversation can
only control its own descendants, never a sibling tree.
- action='steer': queues text into a running child via the existing
steer_subagent() registry path (delivered at the child's next tool
boundary; missed steers surface as missed_steer in the completion).
- action='stop': interrupt_subagent() — child stops at its next
iteration boundary, partial result still re-enters as a completion.
- Spawn dispatch response now includes subagent_ids + control hint.
- Control actions run synchronously (never backgrounded) and bypass
the spawn pause gate and depth limit; they also never consume the
per-turn subagent spawn cap, and remain usable once the cap is hit
(that is when stop matters most).
- Small-model robustness (found live with gpt-5.4-mini on Nous
Portal): tasks=[] alongside goal no longer trips the "Batch mode
requires at least 2 tasks" gate — treated as single-goal.
- CLI display: control calls render as "steer sa-…" / "list" instead
of an empty goal.
Live-tested E2E on Nous Portal (fable-5 + gpt-5.4-mini): full
spawn→list→steer→stop cycle, plus a steer-efficacy run where the
child acked the steer mid-essay and switched topics before finishing.
Adds delegation.worktree_isolation (default: false). When enabled, each
delegate_task child gets its own git worktree branched from the repo's
current HEAD under <repo>/.worktrees/subagent-<id>, its terminal session
starts there, and its goal message carries the isolation contract
(work + commit in the worktree; parent reviews/merges the branch).
- tools/subagent_worktree.py: clean-room implementation from Muse Code's
documented --subagent-worktree-isolation behavior (create per-child
worktree, finalize/inspect after run, auto-prune clean no-commit
worktrees, keep anything holding work).
- tools/delegate_tool.py: config gate + per-child setup in
_run_single_child; result entries gain a "worktree" field (path,
branch, commits, dirty, pruned) only when isolation engaged — the
default-off wire shape is byte-identical.
- Git-only + local-terminal-backend-only; non-git dirs, remote backends,
or any worktree failure degrade silently to shared-workspace behavior.
- Tests: tests/tools/test_subagent_worktree.py (15 tests, real git
repos) + E2E through _run_single_child with a real repo verified
parent-checkout isolation, branch reviewability, prune, and
default-off shape pinning.
- Docs: delegation feature page section + configuration.md key.
Two bugs reported on the docker terminal backend (desktop app, sandboxed
profiles with container_persistent: false):
1. A NEW chat's container inherited the PREVIOUS session's workspace,
bind-mounted rw at /workspace, because the mount source was the
process-global TERMINAL_CWD env var (written by the workspace picker,
outliving its session) and all sessions shared one 'default' container.
2. Every command failed with exit 126 because the desktop gateway recorded
the HOST launch directory as the session cwd, and each command was
prefixed with 'cd /Users/<user>/...' inside the container.
Fixes (class-wide, single owners):
- container_persistent: false + docker now keys containers PER SESSION:
fresh container per chat, removed at session close/idle. delegate_task
children share the parent's container via an explicit alias registry.
container_persistent: true keeps the documented ONE-long-lived-container
contract unchanged.
- _resolve_task_host_cwd() is the single owner of the cwd->/workspace mount
policy across all four env-creation sites; under isolation it refuses
process-global cwd sources and mounts only the session's own attached
workspace (tui_gateway now tags overrides with cwd_source).
- _resolve_command_cwd() gains the same host-path guard the env-creation
sites already had (#50636/#54447 sibling site): a recorded host cwd is
discarded on container backends instead of cd-ing every command into a
nonexistent path.
E2E-tested against real Docker: distinct containers per session, no stale
mount in a fresh session, no exit 126 from host cwd records, containers
removed at session teardown.
Four fix-forwards from the adversarial post-merge audit of the Aug 7
unreviewed merge batch:
- estop (#81148): is_engaged() now fails SAFE (engaged) on stat errors;
the gateway estop gate lets recognized slash commands and replies owned
by in-flight work (update prompts, clarify, slash-confirm, tool
approvals, running sessions) through instead of consuming them; new
gateway /pause [reason|off] command gives messaging-only operators an
in-band engage/resume path (busy_policy=dispatch so it works mid-run).
- cron monitor mode (#81138): execution-mode invariants (monitor x
no_agent, monitor_script x monitor_url, no_agent-requires-script) now
have ONE owner (_validate_job_mode_invariants) called from BOTH
create_job and update_job, so the create-time invariant can no longer
be silently violated through the update door.
- cron notepad (#81139): remove_job now clears the job's notepad rows
(clear_notepad was dead code -> orphaned KV state forever); clear is
best-effort and no-ops without creating notepad.db.
- delegation batch gate (#81141): template-marker regex narrowed to
multi-word placeholder shapes only (<feature name>, {file_path}) so
generics (Vec<T>), HTML tags, JSON snippets, glob braces and f-string
style no longer reject legitimate batches; duplicate-goal rejection
removed (best-of-N fan-outs are legitimate).
Per-task `output_schema` (JSON Schema object) on task items plus the
top-level single-goal form — a one-time static addition to the tool
schema (never varies per call).
- Child side: the schema is appended to the child's context as an
explicit OUTPUT CONTRACT block before spawn.
- Completion side: the parent validates the child's final answer with
jsonschema; on failure it sends exactly ONE bounded retry turn
carrying the validation errors verbatim (no schema re-paste).
- Result entries gain schema_valid (+ schema_retries, and schema_errors
on final failure) ONLY when a schema was requested; schema-less calls
keep a byte-identical result shape.
- Malformed schemas are rejected loudly at dispatch (coerce_output_schema
meta-validates via jsonschema's validator_for/check_schema).
- New helpers in tools/delegation_output_schema.py: coerce, contract
block, fence/prose-tolerant extraction+validation, retry message.
Pattern from: github/copilot-cli ctx.agent(prompt,{schema}) — PATTERN
ONLY, zero code/prompt text copied (proprietary); proven consumer:
delegate-task-output-patterns skill.
Tests: tests/tools/test_delegate_output_schema.py (24 tests — valid
first try, invalid->retry->valid, invalid twice -> schema_valid false +
errors surfaced, retry-exception degrade, no-schema legacy shape pin,
dispatch rejection, contract plumbing). Delegation suite: 221/221 green.
Each serialized result entry now carries cost_usd (rounded to 6 dp)
and cost_status (the child's session_cost_status — 'estimated',
'reported', 'included', or 'unknown') alongside tokens/api_calls/
duration, so the parent model can see what each delegation cost.
The internal _child_cost_usd field is still stripped before
serialization and the parent session cost rollup is untouched.
Tool schema is unchanged (byte-stable).
Inspired by: Perplexity Agent API result shape (idea-level)
Reject malformed tasks=[...] batches before any child agent is spawned:
- exact-duplicate goals (case/whitespace-normalized), error names both
task indices
- placeholder goals: bare 'TODO', bare 'task N', unexpanded <...> or
{...} template markers, or goals shorter than 10 chars after strip
- 1-task batches, with an error pointing the model at the single
`goal` form instead
All checks are batch-only — the single-goal form is exempt by design
(short goals like goal="test" are valid there). Error strings are
actionable: each tells the model exactly how to fix the call.
Tool schema is unchanged (byte-stable); validation is runtime-only in
the existing batch-validation region.
Existing tests using terse batch goals ("A"/"B"/"C") updated to
realistic distinct goals per the new contract.
Inspired by: MoonshotAI/kimi-code agent-swarm.md validation rules (MIT)
The turn finalizer already hands back steer text that queued after the
final tool batch — result["pending_steer"], with the comment "hand it
back to the caller so it can be delivered as the next user turn instead
of being silently lost." Every interactive surface honors that contract
(cli.py, gateway/run.py, tui_gateway/server.py all requeue it). The
delegation layer doesn't: _run_single_child never reads it, so a steer
queued into a delegated child that finishes first vanishes with no trace
in the completion entry. There is also no sanctioned sender: the registry
has interrupt_subagent() but no redirection-side mirror, and session.steer
cannot reach children (lazy watch sessions have agent=None, so it 4010s).
Complete the contract for delegated children — both halves:
- steer_subagent(subagent_id, text): redirection-side mirror of
interrupt_subagent(). Resolves the live child in _active_subagents and
queues text via AIAgent.steer(). True means queued, not delivered.
- missed_steer retention: when the child's result carries pending_steer,
_run_single_child names it on the completion entry (missed_steer field
plus a summary note) so the parent can re-issue the guidance instead of
trusting it landed. This is what makes adding a sender safe: without it
the finish-before-drain race silently loses the text — the exact loss
the finalizer contract exists to prevent.
- subagent.steer gateway RPC beside subagent.interrupt so programmatic
hosts (dashboard, voice layers, ACP bridges) get an in-tree caller;
catalogued in programmatic-integration.md.
- docs: "Steering a Running Subagent" section in delegation.md covering
the queued-vs-delivered semantics.
Tests: registry-level steer coverage (delivery, unknown id, empty text,
dead record, raising agent), the finish-before-drain race retaining
missed_steer, and the RPC contract (4000/4002 validation, queued and
rejected envelopes).
Top-level delegate_task runs in the background, and the 450s progress-stall
monitor only sees api_call_count / tool / last_activity_ts. Subagents use
non-streaming direct_api_call, which previously touched activity once and then
went silent — so a healthy local GGUF / long-prefill wait looked frozen and
was interrupted around ~450s as "Operation interrupted: waiting for model
response", even when child_timeout_seconds was raised. Refresh activity while
the inline request is open, and treat last_activity_ts advances as sync
heartbeat progress too.
Round-2 A/B (gpt-4o-mini, 6 reps) showed two passages could not survive
paraphrase: the DO-NOT-USE list needs the arrow-list shape with the
'no reasoning needed' qualifier (prose form regressed mechanical-work
routing 6/6->1/6), and the self-report rule needs the concrete
'claiming uploaded successfully may be wrong' framing (without it,
side-effect verification regressed 6/6->2/6). With both restored:
30/42 vs 30/42 on gpt-4o-mini and intent-parity on claude-haiku-4.5.
Final size: 1,900 chars (from 3,963).
The top-level delegate_task description repeated content the model already
receives through parameter descriptions: the concurrency limit (tasks param),
the full nesting clause (role param), context-passing guidance (goal/context
params), and background semantics (background param). Every API call paid for
the duplication (~4,000 chars).
The description now carries only what exists nowhere else in the schema:
use/don't-use routing (execute_code, cronjob), the no-poll rule, the
non-durability warning, the self-report verification contract with concrete
verbs, the language-passing example, the leaf blocked-tool list, and model
inheritance. 3,963 -> 1,704 chars (~570 tokens saved per API call), and the
top-level text is now static (dynamic limits flow only through the two param
descriptions, which are already rebuilt per get_definitions() call).
A/B benchmark across 4 models (gpt-4o, gpt-4o-mini, claude-haiku-4.5,
llama-3.3-70b) showed the naive compaction in PR #72813 regressed weaker
models on exactly the passages it cut (side-effect verification 8/8->0/8 on
gpt-4o-mini; language passing 3/3->0/3 on haiku-4.5). This version keeps
those benchmark-sensitive hooks verbatim.
Tests pin the contracts at keyword level (not prose-literal) plus a size
ceiling, and verify dynamic limits still reach the model via the tasks/role
param descriptions.
Refs #72737, supersedes the delegate_task half of PR #72813.
Makes interrupt-protected context compression cancellable by an explicit
user or lifecycle stop, without weakening protection against ordinary
incoming messages, voice interjections, or active-turn redirects.
Separates explicit hard cancellation from ordinary interrupt/redirect
state with a dedicated threading.Event; introduces
AuxiliaryExplicitCancellation as an attempt-local frozen-cause signal;
isolates the synchronous provider callback in a bounded daemon worker
during protected compression; atomically linearizes Codex timeout
cleanup against explicit cancellation; propagates hard cancellation
through child agents and explicit stop surfaces; serializes hard-cancel
admission against compression commit admission with
CompressionCommitFence; aborts before session rotation or late DB commit,
restores in-place transcript mutations and compressor state, and releases
the heartbeat and compression lease.
Based on #74449 by @suparious. Resolved merge conflicts in
agent/context_compressor.py (feasibility check + try/except) and
tui_gateway/methods_session.py.
AIAgent.__init__ calls set_current_session_id(self.session_id), which
mutated both the task-local ContextVar and the process-global os.environ.
Because _build_child_agent wraps construction in delegated_child_context(),
the ContextVar write is harmless (task-local), but the os.environ write
clobbered the parent's HERMES_SESSION_ID for the rest of the process —
leaking the child id into parent tools and subprocesses spawned after
the child was built.
Root cause of HermesPRDelegationSessionContext: parent
20260729_212118_5d797e dispatched child 20260730_160515_736ea1; later
parent terminal inherited HERMES_SESSION_ID=the child.
Fix: set_current_session_id() skips the process-global os.environ write
when called from within a delegated_child_context(). The child's own
tools and subprocesses still resolve their id through the ContextVar
(task-local), while the parent's process-wide env keeps the parent's
session identity. Root agents (CLI, gateway, cron) retain both paths.
Adds 7 regression tests covering single child, concurrent children (8
parallel), parent-tool observation after construction, and root-agent
session rotation backward compatibility. All pass; ruff clean.
Detached background delegation batches (_batch_runner) no longer honor
the foreground parent's interrupt flag — a busy-submit interrupt in the
TUI/desktop previously fabricated 'interrupted' results for background
children that should outlive the turn. Explicit cancellation still works
via _batch_interrupt.
Rebased onto current main from PR #65040; both interrupt-suppression
regression tests aligned with the current _session test helper.
Original work by @AtakanGs in #65040.
Root cause: delegate_task children run through three nested daemon-thread
layers (async-delegation executor -> per-child timeout executor -> the
interrupt worker interruptible_api_call spawns). After multi-day gateway
uptime the deepest layer wedges BEFORE the socket opens — the same
fingerprint as the gateway-cron hang (#62151): zero stale-detector output
(the worker never reaches dispatch), all providers, foreground/restart
works. The cron fix (should_use_direct_api_call) explicitly excluded
delegation 'for lack of evidence' — #60203 is that evidence.
- should_use_direct_api_call: extend the inline gate to delegated
children, detected via the delegation ContextVar set by
_run_single_child (platform='subagent' stamp as fallback). Scope
unchanged otherwise: chat_completions wire only; Codex/Anthropic/
Bedrock/MoA keep their established workers. Interrupts still work —
the inline path registers _active_request_abort, which interrupt()
invokes cross-thread (same mechanism the #72227 stall monitor uses).
- _dump_subagent_timeout_diagnostic: dump ALL thread stacks (bounded,
40), not just the conversation worker — a pre-HTTP wedge is
indistinguishable from a slow provider without seeing where the
nested helper threads sit.
Add timeout_seconds, timed_out_after_seconds, and timeout_phase to timeout
results so parent agents and users can distinguish timeouts before the
first LLM call from timeouts after one or more API calls.
Also attach diagnostic_path to the N>0 API-call timeout error message,
matching the existing zero-API-call timeout path.
Addresses part of #51690 and #17308.
Include last_activity_ts in the progress token sampled from each child.
_touch_activity ticks on every streamed chunk ('receiving stream
response'), every tool transition, and API-call start/completion — so a
child mid-stream on a long response is alive even though api_call_count
only advances when the call completes. Same liveness signal as the
compaction inactivity budget (PR #71508): if tokens are flowing it never
dies; staleness is measured from the last streamed token / tool activity
/ API call.
Replace the wall-clock timeout watchdog (from #60234) with progress-based
staleness detection, on by default with zero config:
- The async registry now accepts a progress_fn per dispatch; delegate_task
wires a sampler over the batch's child agents (api_call_count +
current_tool from get_activity_summary()).
- A single monitor thread sweeps running delegations: a child whose
progress token keeps advancing is never touched, no matter how long it
runs. A frozen token past the stale threshold (450s idle / 1200s
in-tool, mirroring the sync-path heartbeat monitor) marks the record
'stalling' and interrupts the child.
- A stalling child that unwinds within the grace window (120s) finalizes
through the NORMAL path, preserving its partial results. One that never
returns is force-finalized with a terminal 'stalled' completion event so
the owning session hears an outcome and the async slot frees.
- Late runner returns after force-finalization are deduped by the
begin/push/finish finalization split (kept from #60234).
Why not a timeout: delegation.child_timeout_seconds defaults to 0 by
deliberate design (DEFAULT_CHILD_TIMEOUT rationale) — a timeout-based
watchdog never arms for default configs, leaving the reported silent-
profile symptom (#60203) unfixed, and when armed it kills legitimately
slow heavy subagents mid-task. Progress detection distinguishes 'wedged
at first API call' from 'grinding through a 2h review'.
Builds on izumi0uu's finalization-atomicity work from #60234.
Async background delegation can leave gateway sessions holding only a dispatched handle when the detached runner wedges before it can return and enqueue a completion. Enforce the configured child timeout in the async registry so the parent observes a terminal timeout event and the async slot is released.
Constraint: Issue #60203 reports long-lived gateway processes with background child delegates that never produce completion events despite child_timeout_seconds being configured.
Rejected: Relying only on _run_single_child timeout handling | it cannot finalize the async registry when the outer runner thread itself never reaches normal completion.
Confidence: high
Scope-risk: narrow
Directive: Keep background delegation completion owned by the async registry whenever detached workers can outlive the caller's immediate control.
Tested: .venv/bin/python -m pytest tests/tools/test_async_delegation.py tests/tools/test_delegate_subagent_timeout_diagnostic.py tests/tools/test_delegate.py -q
Tested: .venv/bin/python -m ruff check tools/async_delegation.py tools/delegate_tool.py tests/tools/test_async_delegation.py
Tested: git diff --check
Not-tested: Multi-day real gateway degradation; covered with deterministic stuck-runner registry tests.
Wake-ups for kanban notifications and background delegation completions were
injected via handle_message() using a build_session_key()-derived key, which
can never match the raw X-Hermes-Session-Id key that api_server sessions run
under — so the wake landed in a session nobody was reading. On top of that,
ApiServerAdapter.send() reports failure without raising, and that was treated
as a successful delivery, so the notify cursor advanced past events that were
permanently lost; and background delegation was forced synchronous on
api_server since there was no way to wake the session afterward.
Fix: route wake-ups for non-push adapters through a self-post to
/v1/chat/completions with the original session id, treat non-raising send
failures as failures (rewind instead of advancing the cursor), and re-enable
background delegation whenever a session id is available to wake.
The origin session id is captured from the request-scoped api_server chat_id
binding rather than HERMES_SESSION_ID: constructing a child agent calls
set_current_session_id() with the subagent's internal id, clobbering that
variable right before dispatch would read it and misrouting the wake into
the subagent's own session.
Related: #56580, #64609, #53027, #63169, #56531, #50319, #64113