Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):
- Gate every scope-path branch on a new _IS_LINUX constant instead of
'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
never touches systemd code — no probe subprocess, no scope argv,
byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
legacy argv byte-identical, no unit recorded) and probe-returns-False
off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
wired into the on-demand windows-venv-e2e lane): real spawn_local on
windows-latest asserting jobs run exactly as before — spawned, output
captured, exit code correct, systemd path never reached even under
faked gateway identity.
Refs #70716, #71378.
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.
Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).
Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).
Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.
Closes#97654.
* fix(terminal): stop claiming a Linux environment — point at the env section; near-neutral tokens
* feat(terminal): default GIT_PAGER/PAGER=cat in session env; drop schema lines the runtime already enforces; fix pty backend claim
* refactor(terminal): unify notify_on_complete+watch_patterns into notify (bool|list); trim pipe-masking prose (runtime hint owns it)
* fix(terminal): background param referenced the unadvertised legacy arg name
* refactor(terminal): background-only modifiers (pty, notify) fail loud on foreground calls with corrected shape
* fix(execute_code): block the new notify arg in the sandbox terminal stub (foreground-only)
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.
- New config key delegation.surface_child_process_notifications
(default false = suppress). Flag true restores the previous behavior
exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
watch_disabled events whose task_id starts with 'sa-' when the flag
is false, logging at debug with session_id+task_id for diagnosis.
Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
crash the drain loop.
- Docs: delegation.md + configuration.md.
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
lifetime budget; the cap trips exactly at the Nth DELIVERED match and
promotes to notify_on_complete with the watch_disabled summary queued
right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
the global breaker drops the final match, so the user always learns why
watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
was already updated by #93532).
Refs #93513
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).
Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).
Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.
- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
— bounded, interrupt-safe wait over pending notify_on_complete
sessions; reconciles orphaned-pipe exits (#17327) each pass so a
wedged reader cannot burn the full bound; KeyboardInterrupt aborts
the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
session flush / cleanup (covers -q and -Q, i.e. the DM recipient
shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.
Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.
Fixes#90879
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.
Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.
Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
Factory Droid v0.175.0 made TaskOutput/TaskStop accept task-ID prefixes so
background tasks can be referenced without pasting the full ID. Hermes'
process tool had the same friction: every action required the exact
proc_<12-hex> session ID.
ProcessRegistry.get() now falls back to unique-prefix resolution when the
exact lookup misses: 'proc_4dae' or bare '4dae' resolves to
proc_4dae56ca81f6 when exactly one running/finished session matches.
Ambiguous or too-short (<4 suffix chars) prefixes still return None, so
callers keep their existing 'No process with ID ...' error and nothing is
ever picked arbitrarily. Exact IDs never pay the scan, and a full ID that
happens to prefix another always wins.
All process actions (poll/log/wait/kill/write/submit/close) route through
get(), so they all gain prefix support from the single change.
A delegated subagent that exhausts its per-child iteration budget
(delegation.max_iterations) still returns a summary, so the result carries
status='completed' even though the child's exit_reason is 'max_iterations' and
its work was cut off mid-task. The parent then reads 'completed', trusts the
partial summary, and only discovers the truncation by parsing the prose (where
the child happens to mention 'hit the iteration limit'). That wastes parent
turns and risks acting on incomplete work.
exit_reason is already computed authoritatively and threaded to every
parent-visible surface; it just wasn't reflected anywhere the parent reads at a
glance. This surfaces it:
- delegate_tool.py: add a parent-visible boolean 'truncated' (= exit_reason ==
'max_iterations') to each task entry, alongside the existing exit_reason.
- process_registry._format_async_delegation: for both the batch and single-task
paths, when truncated -> use a warning icon, append
'TRUNCATED: hit max_iterations — work may be incomplete' to the header/Status
line, and prefix the summary with an unmissable truncation notice. status
semantics are left unchanged (stays 'completed') so existing icon/summary
branch logic and ~10 tests asserting status=='completed' stay valid.
Tests: single-task truncated -> banner; single-task clean -> no banner; batch
marks only the truncated task, not its clean sibling. 23/23 in the async-
delegation suite.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
Plain type=completion events built in _run_process_watcher carried only
session_key (chat/thread routing) with no spawning-session stamp, so after
/new (or a session switch) a completion notification from the OLD session
was injected into the chat's NEW session. Main already solved this exact
class for async delegations via the _classify_completion_target pre-flight
(_USER_BOUNDARY_END_REASONS drop on user-closed sessions, deliver on
idle-ends, follow the compression-tip chain), but the gate only ran for
type=async_delegation events.
Kernel salvage of #16455:
- Stamp the spawning conversation's session-db id (HERMES_SESSION_ID via
session-scoped env) on the ProcessSession and the pending_watchers entry
at spawn time in tools/terminal_tool.py; persist it through the process
registry checkpoint/restore so recovered watchers keep the stamp.
- Thread the stamp into the completion_evt built by _run_process_watcher
(watcher entry first, ProcessSession fallback for recovered watchers).
- In _deliver_completion_notification, run the SAME pre-flight classifier
for stamped type=completion events: terminal -> drop with a log (output
stays available via process(action='log')), retry -> False so the
watcher re-polls, deliver -> proceed. The policy has exactly one owner
(_classify_completion_target); nothing is forked. Unstamped legacy
events keep today's deliver-always behavior, and the async-delegation
path is untouched.
Based on the session-boundary approach from #16455 by @Tosko4 (original PR
was over-scoped across adapters/slash-commands/cron; this lands the kernel
only).
Tests: completion from a /new-closed session is dropped; completion after
an idle-end still delivers; unstamped legacy event delivers; retry verdict
returns retryable False without adapter injection; async_delegation gate
unchanged; stamp survives checkpoint recovery.
Apply _redact_process_result() to completion and watch_match
notifications before enqueuing them in the completion_queue.
Previously, the explicit process tool path (poll/log/wait) applied
redact_terminal_output() via _redact_process_result(), but the
automatic notification delivery path (notify_on_complete, watch_patterns)
only applied strip_ansi(). This meant API keys, tokens, and other
secrets from background process output were injected into the LLM
conversation unmasked.
The fix ensures both code paths apply the same redaction, matching
the foreground terminal tool behavior.
format_process_notification had no case for watch_overflow_tripped /
watch_overflow_released, so a watch-pattern notification flood surfaced
as '[IMPORTANT: Background process exited (exit code ?)]' — a phantom
exit notification for a process that never existed — while the actual
'watch flood, N notifications suppressed' summary in the event's
message field was silently dropped. The gateway delivery path was
worse: _drain_gateway_watch_events retained only watch_match and
watch_disabled, discarding overflow events entirely before formatting.
Route both event types through the message field in the shared
formatter and the gateway formatter, and retain them in the gateway
drain.
* fix: warn agents off driving interactive console TUIs via pty on Windows
Driving 'gh auth login' (and other survey-style console TUIs) through a
pty background process on Windows silently hangs: these programs read
Win32 console key events via ReadConsoleInput, not the stdin byte
stream, so Enter keypresses submitted over process stdin never register.
The agent-visible symptom is a prompt frozen at 'Press Enter to open
browser...' while the user sees nothing, and a turn interrupt then kills
the process, invalidating any device code the user already entered on
github.com.
Two guidance fixes, both proven in a live session on Windows 10:
- agent/prompt_builder.py: extend _WINDOWS_BASH_SHELL_HINT to steer
agents toward non-interactive paths (flags, --with-token, config
files, curl-polled OAuth device flow) instead of answering console
prompts programmatically.
- skills/github/github-auth: document the pitfall and add the manual
OAuth device-flow procedure (curl against gh's public client_id,
poll for the token, finish with 'gh auth login --with-token'), which
succeeded first try after two interactive attempts hung.
* fix: send CRLF for Enter on Windows PTY submit; correct root cause in guidance
Review feedback (helix4u) was right on both counts:
1. Root cause correction. gh's 'Press Enter to open browser' prompt is
waitForEnter -> bufio.Scanner reading stdin, not a survey/console-API
prompt. The real bug is ours: submit_stdin appended a bare \n, and
through pywinpty/ConPTY a lone \n is not delivered as a line
terminator, so the child's blocking line read never returns. Verified
empirically against pywinpty 2.0.15 with a readline() child:
\n -> hang, \r -> line delivered, \r\n -> line delivered.
Fix: submit_stdin now appends \r\n for Windows PTY sessions (POSIX
PTYs and Popen pipes keep \n). Windows-only regression tests cover
the PTY and pipe branches.
2. Prompt hint rewritten: instead of claiming Windows console TUIs
cannot be driven, it now says to use process(submit) rather than raw
writes with bare \n, and to prefer non-interactive paths when a CLI
offers one.
3. Skill device flow rewritten as an executable script: parses the
device-code response, polls per the returned interval, handles
authorization_pending / slow_down (+5s per GitHub docs) /
expired_token / access_denied / unexpected responses, pipes the token
straight into gh without echoing it, and drops the undocumented
workflow scope (repo,read:org,gist is the documented minimum for
gh auth login --with-token). The pitfall note is narrowed to the
reproduced condition.
Two falsy-zero coercions in process_registry (salvaged from PR #60004,
credit @isheng-eqi; the EOF half of that PR landed separately in
893792c99):
- wait(timeout=0): schema says minimum=1 but the handler let 0 fall
through '0 or max_timeout' to the DEFAULT wait instead of rejecting.
- read_log(offset=0): conflated with the offset-unset default, silently
returning the TAIL of the log when the caller asked for the head.
Default is now offset=None; explicit 0 paginates from line one.
- Remove dead use_systemd_scope = False assignment (leftover from
the old try/except pattern, immediately overwritten).
- Update stale log label supervisor= -> in_supervised_gateway=
to match the renamed variable.
- Convert autouse _mark_gateway_process fixture to opt-in
_gateway_identity so negative tests start from a clean slate
instead of undoing the fixture's env/PID mocks.
- Parametrize 4 near-duplicate negative tests (2 scenarios x
pipe/PTY) into 2 parametrized tests, reducing ~130 lines to ~80.
76 tests pass, ruff clean, net -32 LOC.
systemd-run --scope does not give the invoked process a new session: the
worker keeps the parent's session and inherits its controlling terminal.
When the parent is an interactive TUI on a pts (INVOCATION_ID present ->
is_gateway_supervisor_process()=True), every background spawn drops the
worker into the same session as the foreground process group; the spawn
then stops the whole session (SIGTTIN/SIGTTOU family), observed as 5 dead
TUIs in state T ("Arrêté") on 2026-08-08.
Fix: popen_start_new_session = True in the systemd-scope branch of
spawn_local. The worker (and the systemd-run wrapper) get a private
session while the scope cgroup isolation is preserved - the scope is
attached to the invoked process, not to the spawning session.
Verified: simulated TUI (INVOCATION_ID) + ProcessRegistry.spawn_local ->
worker in hermes-worker-*.scope with sid != simulator sid, exits cleanly,
simulator stays alive (previously: same sid -> stopped).
When Hermes runs as a systemd gateway with MemoryHigh/MemoryMax limits,
local background terminal commands (terminal(background=true)) inherit the
gateway's cgroup. A memory-heavy executor (Codex, tests, Node) can push
the whole cgroup past MemoryMax and trigger systemd-oomd to kill the
ENTIRE gateway — taking down the messaging control plane and silently
losing the active turn.
Root cause: tools/process_registry.py::spawn_local() uses
start_new_session=True (creates a process session/group, NOT a resource
cgroup). The spawned process tree stays in the gateway's systemd cgroup.
Fix: when running under a service manager (detected via the existing
is_gateway_supervisor_process() helper), wrap the pipe-mode spawn command
in 'systemd-run --user --scope --unit=hermes-worker-<id>' so the worker
gets its own transient cgroup. An OOM in the worker then kills only the
worker, not the gateway.
The systemd-run availability is probed once (a no-op /bin/true in a
transient scope) and cached, because the binary can exist on PATH while
the user D-Bus session is unavailable (system services, containers). If
unavailable, fall back to the current start_new_session=True behavior
with a debug log.
Scope: this covers the common background pipe-mode path. PTY mode
(PtyProcess.spawn) is left as future work — it uses a different spawn
mechanism and is used for interactive CLI tools where cgroup isolation
has additional considerations.
_write_checkpoint persisted s.command verbatim to ~/.hermes/processes.json.
Recovery only uses command for display/logging (the process is already
running; adoption re-validates PID + start time, never re-runs the
command), so masking is lossless.
process(action='wait') hitting its window returned status='timeout'
with a terse note — models read it as an error and re-issued identical
waits (process is the #1 exact-duplicate tool call in production: 511
dupes in a 400k-msg window; wait is 57% of all process actions).
The timeout result now carries:
- process_running: true — machine-readable 'this is a status, not a
failure'
- an explicit note: 'Wait window of Ns elapsed — the process is still
running. This is not an error. Uptime: Ms.' plus the right next step:
when notify_on_complete is set, 'you will be notified on exit — do
more work instead of waiting again'; otherwise a pointer to
notify_on_complete for next time.
- the clamp note (requested > max) now composes with the status note
instead of replacing it.
Exited/interrupted results are unchanged.
kill_started_since duplicated kill_all's collect-under-lock/kill-outside-lock
loop line for line; it is now a thin delegate through new kill_all kwargs
(exclude_ids, source, consume_output). Public signatures unchanged — existing
callers and test monkeypatch seams keep working. kill_process's docstring now
names the deliberate consume_output=True exception for abandoned-turn reaping
so the deviation isn't 'fixed' later.
An agent turn can spawn a long-running background subprocess (e.g.
`next build`) and later be abandoned via inactivity timeout, /stop,
/new, or a client disconnect. Before this fix the gateway interrupted
the agent loop but never touched the subprocess: it kept running
inside the gateway's cgroup, unbounded, until memory pressure starved
the event loop and made every platform/cron look hung (#76115).
The process registry already knew how to kill a process tree — the
missing piece was per-turn ownership: nothing distinguished a process
that predates the turn (must survive), a process the turn started and
finished successfully (must survive), and a process an abandoned turn
left running (must be reaped).
- tools/process_registry.py: snapshot_running_ids() captures a turn's
starting baseline; kill_started_since() reaps only IDs created after
it, scoped to one task_id.
- gateway/turn_context.py: TurnContext carries process_task_id +
process_baseline so the timeout/interrupt paths can reach them.
- gateway/run.py: baseline is snapshotted right before the turn's
executor task starts; the inactivity-timeout path and the explicit
/stop|/new|disconnect interrupt path both reap via the same helper.
A daemon-thread watchdog backs up the asyncio-based timeout poll,
since a starved event loop is exactly the failure mode this bug
causes. The turn's own worker clears its ownership markers the
instant it finishes, closing a race where a /stop landing right
after normal completion could reap a background process the turn
deliberately left running.
Related but insufficient on their own: #37454 (cgroup ExecStopPost
reaper only fires on service restart) and #68915 (orphaned-pipe
grandchild detection, a registry bug not a turn-lifecycle gap).
Neither ties process cleanup to turn abandonment.
Port from openclaw/openclaw#112325: multibyte UTF-8 characters split
across a 4096-byte pipe or PTY read boundary were decoded statelessly
per chunk with errors='replace', corrupting both halves into U+FFFD
mojibake in background process output (poll/log/wait/completion
notifications). The foreground path already used an incremental decoder
(tools/environments/base.py::_wait_for_process); this applies the same
treatment to the background reader loops:
- _reader_loop (select and blocking paths): one
codecs.getincrementaldecoder('utf-8') per reader holds partial
sequences across chunks; the finally block flushes a truncated tail
as a single U+FFFD instead of dropping it.
- _pty_reader_loop: same treatment for ptyprocess byte chunks
(pywinpty str chunks pass through unchanged).
Genuinely invalid bytes keep errors='replace' behavior.
When a background terminal() command backgrounds its own long-lived
child (`node server.js &`, `sleep 300 &`), the grandchild inherits the
write end of the reader thread's stdout pipe. The direct bash child
exits promptly, but the pipe never reaches EOF while the grandchild
lives — so `_reader_loop`'s blocking `read1()` parked the thread
forever, `session.exited` never flipped on its own, and
`notify_on_complete` was silently lost. `_reconcile_local_exit`
(#17327) only runs lazily from poll()/wait(), so nothing autonomous
ever surfaced the exit; each occurrence also leaked a reader thread
and pipe fd for the grandchild's lifetime.
Fix: on POSIX, drain via select() with a short poll interval and stop
shortly after the direct child exits even if the pipe hasn't EOF'd —
the same pattern the foreground path uses in
tools/environments/base.py::_wait_for_process (#8340). Windows pipes
don't support select(), so the blocking path is kept there with the
existing lazy reconcile as the safety net; mocked/iterator stdout
streams (no usable fileno) also keep the historical path.
Fixes#68915
On Windows with Chinese locale (GBK), subprocess.run(text=True) without
explicit encoding causes UnicodeDecodeError crashes. This fix adds
encoding='utf-8', errors='replace' to all subprocess.run() and
subprocess.Popen() calls that use text=True across 76 non-test Python files.
Fixes#53428 (master tracker for Windows GBK locale crash).
Note: credential_pool.py and electron changes excluded per reviewer request —
those will be submitted as separate focused PRs.
Issue #68915: when the agent runs a compound command with trailing & (e.g.
`cd /app && node server.js &`), bash parses it as `(A && B) &` — a subshell
that holds the stdout pipe open forever when B is a long-running server.
The existing _rewrite_compound_background in terminal_tool.py correctly
rewrites this to `A && { B & }` to avoid the subshell fork, but it was only
applied in the foreground execute() path (tools/environments/base.py).
The background spawn_local() path bypasses base.py entirely and passed the
raw command directly to Popen/PTY, leaving the deadlock unmitigated.
Fix: apply _rewrite_compound_background in spawn_local() before the command
is passed to Popen or PTY spawn. Uses a lazy import to avoid circular
dependency (terminal_tool imports process_registry).
- PTY spawn path: now uses safe_command (rewritten)
- Popen spawn path: now uses safe_command (rewritten)
- Session.command still stores the original (unrewritten) command for display
- Simple `cmd &` is left unchanged (no subshell bug)
Tests: 4 regression tests verifying (1) compound is rewritten, (2) simple bg
is preserved, (3) multi-line compounds are rewritten, (4) session.command
stores original.
* feat(delegation): live-viewable subagent transcripts for delegate_task
Each child now streams an append-only, human-readable log to
<hermes_home>/cache/delegation/live/<delegation_id>/task-<n>.log while it
runs, and the dispatch return includes the paths so the caller can tail
them immediately instead of waiting blind for the consolidated summary.
- New tools/delegation_live_log.py: LiveTranscriptWriter (per-event append
+ flush, one-line rendering with truncation, never raises into the agent
loop), wrap_progress_callback (tees the child's existing
tool_progress_callback events into the log, preserves the _flush
contract), dispatch-time creation with pre-headered files so tail -f
attaches immediately, manifest.json (goals/task count/per-task status),
and 7-day retention pruning on new dispatches.
- delegate_task: wraps each child's progress callback with the writer;
sync results and background dispatch responses gain live_transcripts
(+ hint field on dispatch); per-task result entries carry
live_transcript; transcripts finalized with exit-reason markers.
- async_delegation: dispatch_async_delegation_batch accepts an optional
delegation_id so the live/ dir name matches the returned handle; the
completion event carries live_transcripts.
- process_registry: consolidated batch-completion block references each
task's live transcript path.
- Tool schema description documents the live_transcripts return surface;
docs gain a 'Live Transcripts' section with a tail -f example.
Placement under cache/delegation means the logs are mounted read-only
into remote terminal backends for free. Side-channel only: zero changes
to message content, so prompt caching is unaffected. Transcript-OUT only
— no overlap with the subagent control surfaces of PR #66046.
* fix(delegation): label the kickoff transcript line as user — it is the child's one user message
Apply positive-proof routing to every addressed notification in the registry and TUI poller while preserving ownerless legacy behavior and TUI delivery for poll-observed completions.
Remove the unused exact-key drain helper and cover ordinary success and failure, origin, compression-lineage, orphan, and poll-observed paths.
Complements NousResearch/hermes-agent#54785.
Three-layer companion to the salvaged CLI drain-ownership fix (#64240):
1. restore_undelivered_completions stamps restored=True (in-memory only)
on every durable completion re-enqueued at process start.
2. drain_notifications' legacy unfiltered branch re-queues restored
events instead of consuming them — a fresh process can no longer
adopt a dead session's delegation results (#64484). Same-process
keyless events keep the legacy behavior.
3. delegate_tool's async dispatch now falls back to the parent agent's
durable session_id when the approval-context key resolves empty (the
CLI case), so the CLI's new positive-ownership drain can actually
claim its own completions instead of failing closed on ''.