212 Commits

Author SHA1 Message Date
Teknium e01755f7d0 refactor(tools): process_registry — compact docstring layout (content unchanged) 2026-09-02 23:58:25 -07:00
Teknium 781287b535 refactor(tools): process_registry — _exit_fields, _stdin_op returns shared ok result 2026-09-02 23:55:21 -07:00
Teknium 7ef7c5d4e7 refactor(tools): process_registry — _config_seconds, _reap_untracked, inline drain append 2026-09-02 23:53:07 -07:00
Teknium 0000722290 refactor(tools): process_registry — drop stray blanks after decorators 2026-09-02 23:41:15 -07:00
Teknium a884d45321 refactor(tools): process_registry — hug call closers, squeeze body blanks, tighten header comments 2026-09-02 23:38:10 -07:00
Teknium beb1a8d81b refactor(tools): process_registry — contextlib.suppress for swallow-only try blocks, kill_all as sum, derive watcher checkpoint fields 2026-09-02 23:29:08 -07:00
Teknium 2c2935c4c8 refactor(tools): process_registry — _finish_reader/_owns_event/_status_head/_spawn_env helpers; notifications list building compacted 2026-09-02 23:23:25 -07:00
Teknium 870aea6e47 refactor(tools): process_registry — early-return watch limiter, _new_session, psutil gone-tuple, dedup shell-noise list 2026-09-02 23:09:56 -07:00
Teknium 6385b61370 refactor(tools): process_registry — pack ProcessSession constructor kwargs 2026-09-02 22:59:58 -07:00
Teknium 6ab5706f70 refactor(tools): process_registry — unified select/blocking reader loop, drop unused _watch_last_emit_at, tighter docstrings 2026-09-02 22:50:12 -07:00
Teknium eef71bb6e1 refactor(tools): process_registry — single reader read path, _WATCHER_ROUTE_KEYS, hugged signatures 2026-09-02 22:43:28 -07:00
Teknium 60e9009079 refactor(tools): process_registry — mark_exited/_output_tail/_stdin_op helpers, unified watch_disabled emitter, compact docs 2026-09-02 22:34:24 -07:00
Teknium 3bbec90f23 refactor(tools/environments): split local/base into output/wait/session-env siblings; extract docker_egress + remote_common; dedupe remote backends; compact process_registry 2026-09-02 14:45:15 -07:00
kshitijk4poor 75bcd86687 fix(process): keep the sandbox log poller from splitting a UTF-8 character across polls
Follow-up to #92164: the delta window now ends on a character boundary
(up to 3 trailing continuation bytes are held for the next poll), so
multibyte output no longer decodes to U+FFFD at the seam. Verified on
bash, dash and busybox sh; exhaustive-prefix regression test added.
2026-09-03 02:20:46 +05:30
Adolanium 8b681f70ea perf(process): poll sandbox job logs for new bytes only
The background-process poller for non-local backends ran `cat` on the
whole log file every two seconds, then threw away everything except the
part it had not seen yet. The offset it needed was already tracked one
line below, so the full read was pure waste.

Cost of one poll grew with the total output so far, which makes the cost
of a run grow with the square of its length. A job writing 10 MB over an
hour moved about 9 GB across the docker or SSH channel to deliver 10 MB
of output.

The poller now asks the shell for the file size and the bytes after the
offset in one command. Reading the size first and cutting the tail at
that same size keeps the two in step, so a file that grows mid-command
never sends a byte twice. A file that shrank was rotated or truncated,
so the offset drops back to 0 and the buffer is dropped.

The output buffer is now appended to rather than replaced, matching the
local reader loops, and the offset is counted in bytes because the shell
counts bytes.
2026-09-03 02:00:35 +05:30
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 5a8e8a6b87 fix(terminal): strict Linux-only gating for background-executor systemd scopes (#70716 follow-up)
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):

- Gate every scope-path branch on a new _IS_LINUX constant instead of
  'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
  never touches systemd code — no probe subprocess, no scope argv,
  byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
  legacy argv byte-identical, no unit recorded) and probe-returns-False
  off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
  wired into the on-demand windows-venv-e2e lane): real spawn_local on
  windows-latest asserting jobs run exactly as before — spawned, output
  captured, exit code correct, systemd path never reached even under
  faked gateway identity.

Refs #70716, #71378.
2026-09-01 02:32:53 -07:00
Teknium 5a4dbdec27 fix(delegation): subagent process notifications stay suppressed when the container key collapses
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.

Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).

Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
2026-08-31 07:28:18 -07:00
liuhao1024 8557e0a480 fix(tools): surface config-level model_not_found notices in delegation batch reports
When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).

Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
2026-08-30 22:19:10 -07:00
David Metcalfe c05d04fffb feat(delegation): surface config-level model_not_found notice in delegation batch reports
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.

Closes #97654.
2026-08-30 21:07:04 -07:00
Teknium 03e66c8cba polish(tool-search): stub-optimized openers for deferred tools — trigger+verb in the first ~60 chars (the catalog stub is the only ambient hint a deferred tool exists) 2026-08-29 18:13:20 -07:00
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium baa344dee7 refactor(process): schema diet — enum names the verbs, description keeps only non-obvious semantics; write-vs-submit trap teaching emphasized (306 -> 228 tok/call, -25%) (#97279) 2026-08-28 09:23:00 -07:00
Teknium 1dc552d5d1 refactor(terminal): honest schema, pager defaults in the env, unified notify arg (837 → 670 tok/call, −20%) (#95937)
* fix(terminal): stop claiming a Linux environment — point at the env section; near-neutral tokens

* feat(terminal): default GIT_PAGER/PAGER=cat in session env; drop schema lines the runtime already enforces; fix pty backend claim

* refactor(terminal): unify notify_on_complete+watch_patterns into notify (bool|list); trim pipe-masking prose (runtime hint owns it)

* fix(terminal): background param referenced the unadvertised legacy arg name

* refactor(terminal): background-only modifiers (pty, notify) fail loud on foreground calls with corrected shape

* fix(execute_code): block the new notify arg in the sandbox terminal stub (foreground-only)
2026-08-26 19:23:39 -07:00
Teknium afee35700e fix(delegation): suppress subagent-owned process notifications in parent chat by default
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.

- New config key delegation.surface_child_process_notifications
  (default false = suppress). Flag true restores the previous behavior
  exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
  watch_disabled events whose task_id starts with 'sa-' when the flag
  is false, logging at debug with session_id+task_id for diagnosis.
  Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
  events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
  crash the drain loop.
- Docs: delegation.md + configuration.md.
2026-08-25 21:59:04 -07:00
Teknium d8d1e18ab9 test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
2026-08-24 03:22:48 -07:00
chelsealong b3730153c3 fix(terminal): cap watch_patterns notifications over a process's lifetime
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
2026-08-24 03:22:48 -07:00
Teknium 9e18197745 fix(cli): one-shot runs linger for notify_on_complete background processes so Bot Mode replies survive parent exit
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).

Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.

- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
  — bounded, interrupt-safe wait over pending notify_on_complete
  sessions; reconciles orphaned-pipe exits (#17327) each pass so a
  wedged reader cannot burn the full bound; KeyboardInterrupt aborts
  the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
  session flush / cleanup (covers -q and -Q, i.e. the DM recipient
  shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
  kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.

Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.

Fixes #90879
2026-08-23 03:56:37 -07:00
Hermes Agent b95ec1cb5d fix(delegation): running subagents stay visible to list/steer across parent-agent rebuilds, and child-started process notifications carry delegation attribution
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.

Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.

Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
2026-08-18 00:01:56 -07:00
Teknium 2e4d771c69 Inspired by Factory Droid: accept unique ID prefixes in process tool lookups
Factory Droid v0.175.0 made TaskOutput/TaskStop accept task-ID prefixes so
background tasks can be referenced without pasting the full ID. Hermes'
process tool had the same friction: every action required the exact
proc_<12-hex> session ID.

ProcessRegistry.get() now falls back to unique-prefix resolution when the
exact lookup misses: 'proc_4dae' or bare '4dae' resolves to
proc_4dae56ca81f6 when exactly one running/finished session matches.
Ambiguous or too-short (<4 suffix chars) prefixes still return None, so
callers keep their existing 'No process with ID ...' error and nothing is
ever picked arbitrarily. Exact IDs never pay the scan, and a full ID that
happens to prefix another always wins.

All process actions (poll/log/wait/kill/write/submit/close) route through
get(), so they all gain prefix support from the single change.
2026-08-16 22:09:37 -07:00
Teknium dc2fe99ecf feat(delegation): mark max_iterations-truncated subagent results for the parent (#86641)
A delegated subagent that exhausts its per-child iteration budget
(delegation.max_iterations) still returns a summary, so the result carries
status='completed' even though the child's exit_reason is 'max_iterations' and
its work was cut off mid-task. The parent then reads 'completed', trusts the
partial summary, and only discovers the truncation by parsing the prose (where
the child happens to mention 'hit the iteration limit'). That wastes parent
turns and risks acting on incomplete work.

exit_reason is already computed authoritatively and threaded to every
parent-visible surface; it just wasn't reflected anywhere the parent reads at a
glance. This surfaces it:

- delegate_tool.py: add a parent-visible boolean 'truncated' (= exit_reason ==
  'max_iterations') to each task entry, alongside the existing exit_reason.
- process_registry._format_async_delegation: for both the batch and single-task
  paths, when truncated -> use a warning icon, append
  'TRUNCATED: hit max_iterations — work may be incomplete' to the header/Status
  line, and prefix the summary with an unmissable truncation notice. status
  semantics are left unchanged (stays 'completed') so existing icon/summary
  branch logic and ~10 tests asserting status=='completed' stay valid.

Tests: single-task truncated -> banner; single-task clean -> no banner; batch
marks only the truncated task, not its clean sibling. 23/23 in the async-
delegation suite.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
2026-08-14 21:19:15 -07:00
Teknium 7619564fbd fix(gateway): gate background-process completions on spawning-session boundary
Plain type=completion events built in _run_process_watcher carried only
session_key (chat/thread routing) with no spawning-session stamp, so after
/new (or a session switch) a completion notification from the OLD session
was injected into the chat's NEW session. Main already solved this exact
class for async delegations via the _classify_completion_target pre-flight
(_USER_BOUNDARY_END_REASONS drop on user-closed sessions, deliver on
idle-ends, follow the compression-tip chain), but the gate only ran for
type=async_delegation events.

Kernel salvage of #16455:

- Stamp the spawning conversation's session-db id (HERMES_SESSION_ID via
  session-scoped env) on the ProcessSession and the pending_watchers entry
  at spawn time in tools/terminal_tool.py; persist it through the process
  registry checkpoint/restore so recovered watchers keep the stamp.
- Thread the stamp into the completion_evt built by _run_process_watcher
  (watcher entry first, ProcessSession fallback for recovered watchers).
- In _deliver_completion_notification, run the SAME pre-flight classifier
  for stamped type=completion events: terminal -> drop with a log (output
  stays available via process(action='log')), retry -> False so the
  watcher re-polls, deliver -> proceed. The policy has exactly one owner
  (_classify_completion_target); nothing is forked. Unstamped legacy
  events keep today's deliver-always behavior, and the async-delegation
  path is untouched.

Based on the session-boundary approach from #16455 by @Tosko4 (original PR
was over-scoped across adapters/slash-commands/cron; this lands the kernel
only).

Tests: completion from a /new-closed session is dropped; completion after
an idle-end still delivers; unstamped legacy event delivers; retry verdict
returns retryable False without adapter injection; async_delegation gate
unchanged; stamp survives checkpoint recovery.
2026-08-14 01:09:42 -07:00
dsad 6c7cfd6621 fix(security): redact secrets in background process notifications
Apply _redact_process_result() to completion and watch_match
notifications before enqueuing them in the completion_queue.

Previously, the explicit process tool path (poll/log/wait) applied
redact_terminal_output() via _redact_process_result(), but the
automatic notification delivery path (notify_on_complete, watch_patterns)
only applied strip_ansi(). This meant API keys, tokens, and other
secrets from background process output were injected into the LLM
conversation unmasked.

The fix ensures both code paths apply the same redaction, matching
the foreground terminal tool behavior.
2026-08-14 01:08:13 -07:00
spfcraze d2d7766750 fix(tools,gateway): format watch_overflow events instead of dropping them
format_process_notification had no case for watch_overflow_tripped /
watch_overflow_released, so a watch-pattern notification flood surfaced
as '[IMPORTANT: Background process  exited (exit code ?)]' — a phantom
exit notification for a process that never existed — while the actual
'watch flood, N notifications suppressed' summary in the event's
message field was silently dropped. The gateway delivery path was
worse: _drain_gateway_watch_events retained only watch_match and
watch_disabled, discarding overflow events entirely before formatting.
Route both event types through the message field in the shared
formatter and the gateway formatter, and retain them in the gateway
drain.
2026-08-14 01:08:13 -07:00
Teknium 197a18314f fix: warn agents off driving interactive console TUIs via pty on Windows (#84364)
* fix: warn agents off driving interactive console TUIs via pty on Windows

Driving 'gh auth login' (and other survey-style console TUIs) through a
pty background process on Windows silently hangs: these programs read
Win32 console key events via ReadConsoleInput, not the stdin byte
stream, so Enter keypresses submitted over process stdin never register.
The agent-visible symptom is a prompt frozen at 'Press Enter to open
browser...' while the user sees nothing, and a turn interrupt then kills
the process, invalidating any device code the user already entered on
github.com.

Two guidance fixes, both proven in a live session on Windows 10:

- agent/prompt_builder.py: extend _WINDOWS_BASH_SHELL_HINT to steer
  agents toward non-interactive paths (flags, --with-token, config
  files, curl-polled OAuth device flow) instead of answering console
  prompts programmatically.
- skills/github/github-auth: document the pitfall and add the manual
  OAuth device-flow procedure (curl against gh's public client_id,
  poll for the token, finish with 'gh auth login --with-token'), which
  succeeded first try after two interactive attempts hung.

* fix: send CRLF for Enter on Windows PTY submit; correct root cause in guidance

Review feedback (helix4u) was right on both counts:

1. Root cause correction. gh's 'Press Enter to open browser' prompt is
   waitForEnter -> bufio.Scanner reading stdin, not a survey/console-API
   prompt. The real bug is ours: submit_stdin appended a bare \n, and
   through pywinpty/ConPTY a lone \n is not delivered as a line
   terminator, so the child's blocking line read never returns. Verified
   empirically against pywinpty 2.0.15 with a readline() child:
   \n -> hang, \r -> line delivered, \r\n -> line delivered.

   Fix: submit_stdin now appends \r\n for Windows PTY sessions (POSIX
   PTYs and Popen pipes keep \n). Windows-only regression tests cover
   the PTY and pipe branches.

2. Prompt hint rewritten: instead of claiming Windows console TUIs
   cannot be driven, it now says to use process(submit) rather than raw
   writes with bare \n, and to prefer non-interactive paths when a CLI
   offers one.

3. Skill device flow rewritten as an executable script: parses the
   device-code response, polls per the returned interval, handles
   authorization_pending / slow_down (+5s per GitHub docs) /
   expired_token / access_denied / unexpected responses, pipes the token
   straight into gh without echoing it, and drops the undocumented
   workflow scope (repo,read:org,gist is the documented minimum for
   gh auth login --with-token). The pitfall note is narrowed to the
   reproduced condition.
2026-08-12 01:15:17 -07:00
isheng-eqi fc09f1c695 fix(process): reject non-positive wait timeouts; distinguish log offset=0 from default
Two falsy-zero coercions in process_registry (salvaged from PR #60004,
credit @isheng-eqi; the EOF half of that PR landed separately in
893792c99):

- wait(timeout=0): schema says minimum=1 but the handler let 0 fall
  through '0 or max_timeout' to the DEFAULT wait instead of rejecting.
- read_log(offset=0): conflated with the offset-unset default, silently
  returning the TAIL of the log when the caller asked for the head.
  Default is now offset=None; explicit 0 paginates from line one.
2026-08-10 00:23:41 -07:00
kshitij 326bdfb7a2 refactor: clean up gateway scope identity predicate and tests
- Remove dead use_systemd_scope = False assignment (leftover from
  the old try/except pattern, immediately overwritten).
- Update stale log label supervisor= -> in_supervised_gateway=
  to match the renamed variable.
- Convert autouse _mark_gateway_process fixture to opt-in
  _gateway_identity so negative tests start from a clean slate
  instead of undoing the fixture's env/PID mocks.
- Parametrize 4 near-duplicate negative tests (2 scenarios x
  pipe/PTY) into 2 parametrized tests, reducing ~130 lines to ~80.

76 tests pass, ruff clean, net -32 LOC.
2026-08-09 21:43:32 +05:30
bgrablin aa32e81141 fix(process-registry): bind gateway scope identity to pid 2026-08-09 21:43:32 +05:30
bgrablin ff5dfdecef fix(process-registry): keep CLI workers off controlling tty 2026-08-09 21:43:32 +05:30
Driss NAAMANE 0e492a4840 fix(terminal): keep background workers in a private session under systemd scope (#70716)
systemd-run --scope does not give the invoked process a new session: the
worker keeps the parent's session and inherits its controlling terminal.
When the parent is an interactive TUI on a pts (INVOCATION_ID present ->
is_gateway_supervisor_process()=True), every background spawn drops the
worker into the same session as the foreground process group; the spawn
then stops the whole session (SIGTTIN/SIGTTOU family), observed as 5 dead
TUIs in state T ("Arrêté") on 2026-08-08.

Fix: popen_start_new_session = True in the systemd-scope branch of
spawn_local. The worker (and the systemd-run wrapper) get a private
session while the scope cgroup isolation is preserved - the scope is
attached to the invoked process, not to the spawning session.

Verified: simulated TUI (INVOCATION_ID) + ProcessRegistry.spawn_local ->
worker in hermes-worker-*.scope with sid != simulator sid, exits cleanly,
simulator stays alive (previously: same sid -> stopped).
2026-08-09 14:13:02 +05:30
Theophilus Chinomona 45aa902c18 fix(process_registry): surrogateescape-safe PTY stdin writes (#79178) 2026-08-08 12:31:19 -07:00
Dominic Bejar c5e032c804 fix(gateway): close ambiguous recovery cleanup gaps 2026-08-08 01:12:16 +05:30
Dominic Bejar 46b5314229 fix(terminal): harden scope fallback and memory override 2026-08-08 01:12:16 +05:30
Dominic Bejar b0346ba42a fix(terminal): align worker limit with local guard 2026-08-08 01:12:16 +05:30
Dominic Bejar 5f93083221 fix(terminal): bound isolated worker memory 2026-08-08 01:12:16 +05:30
Dominic Bejar 0690fd77c6 fix(terminal): make systemd cleanup gateway-safe 2026-08-08 01:12:16 +05:30
Dominic Bejar 69397937dd fix(terminal): serialize systemd scope capability probe 2026-08-08 01:12:16 +05:30
toprakeker 21de22a4ec fix(terminal): fully-qualified .scope unit name, exit-code check, already_exited cleanup (#70716) 2026-08-08 01:12:16 +05:30
toprakeker 7cfa90d90a fix(terminal): address review gaps — PTY isolation, unit-name kill, --quiet (#70716) 2026-08-08 01:12:16 +05:30
toprakeker 099eb73731 fix(terminal): isolate local background executors in their own systemd cgroup (#70716)
When Hermes runs as a systemd gateway with MemoryHigh/MemoryMax limits,
local background terminal commands (terminal(background=true)) inherit the
gateway's cgroup. A memory-heavy executor (Codex, tests, Node) can push
the whole cgroup past MemoryMax and trigger systemd-oomd to kill the
ENTIRE gateway — taking down the messaging control plane and silently
losing the active turn.

Root cause: tools/process_registry.py::spawn_local() uses
start_new_session=True (creates a process session/group, NOT a resource
cgroup). The spawned process tree stays in the gateway's systemd cgroup.

Fix: when running under a service manager (detected via the existing
is_gateway_supervisor_process() helper), wrap the pipe-mode spawn command
in 'systemd-run --user --scope --unit=hermes-worker-<id>' so the worker
gets its own transient cgroup. An OOM in the worker then kills only the
worker, not the gateway.

The systemd-run availability is probed once (a no-op /bin/true in a
transient scope) and cached, because the binary can exist on PATH while
the user D-Bus session is unavailable (system services, containers). If
unavailable, fall back to the current start_new_session=True behavior
with a debug log.

Scope: this covers the common background pipe-mode path. PTY mode
(PtyProcess.spawn) is left as future work — it uses a different spawn
mechanism and is used for interactive CLI tools where cgroup isolation
has additional considerations.
2026-08-08 01:12:16 +05:30