Commit Graph

189 Commits

Author SHA1 Message Date
kshitijk4poor 16fe50b1a2 refactor(gateway): inline the user-bus adoption gate at the run_gateway call site
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
2026-09-08 02:28:21 +05:30
HexLab98 802f0f97ad fix(gateway): adopt the user D-Bus session when systemd starts the gateway
A system-level unit (/etc/systemd/system, User=<someone>) is exec'd with neither
XDG_RUNTIME_DIR nor DBUS_SESSION_BUS_ADDRESS, and a process environment is fixed
at exec time. 'systemd-run --user --scope' therefore fails for the whole lifetime
of that gateway even after the user manager is up and /run/user/<uid>/bus is
reachable. That is the seam every restart-safe worker crosses
(restart_safe_gateway_child_argv), and it fails closed by design — so on headless
systemd installs every agent-driven cron job and every Kanban dispatch died at
launch, ~26ms in, with nothing but 'error' on the job row.

_ensure_user_systemd_env() already derives both values from our own uid and adopts
them only when the runtime dir is really ours and the socket really exists; it was
just wired exclusively to the systemctl management paths, never to the gateway's
own boot. Call it from run_gateway() — the single in-process boot every entry point
goes through — so the adoption precedes every worker-environment snapshot (cron
builds its env after the scope check, Kanban before it, so fixing this at the
dispatch seam would only fix one of them).

The fail-closed posture is unchanged: with no user manager at all the probe still
reports unavailable and dispatch still refuses. That refusal now names the remedy
in the message the operator actually reads (it is stored as the cron execution's
error), instead of only the symptom.

Fixes #104893
2026-09-08 02:28:21 +05:30
Teknium ef9239571d feat(delegation): report a child's exited-but-unread notify processes to the parent
A process that finishes while the child is alive needs no handoff, but if the child never
polls/waits/logs it, the result vanished: the completion notice is suppressed in the parent
and the child's summary never mentions it. Finalization now attaches exit code + output
tail as unread_completions, rendered in the parent's delegation notice.
2026-09-07 12:50:29 -07:00
Teknium 3c0d90e8ef feat(delegation): subagents hand background processes to the parent; leftovers are named, not trusted
A child's background processes are killed at its teardown and their
notify_on_complete notices are suppressed in the parent, yet the child's
terminal result still said `notify_on_complete: true` and the parent's
delegation notice said nothing about processes left behind. Orchestrators
believed "CI watcher running" and waited on a completion that could never
arrive (recurring in the Sep 7 campaign sessions).

- process_manage(action="handoff", session_id, data="<purpose>"), children
  only: process_registry.transfer_ownership flips owner_task_id/task_id/
  session_key to the parent under the registry lock, so the completion is
  stamped with the parent's owner at exit, passes the parent's sa- filter,
  and is reaped by the parent, not the child. Cap 3 per child; an exited,
  foreign, or non-child request is a tool error. The purpose rides the
  event as handoff_note and renders in the parent's notice.
- Child terminal(background=True, notify=True) now returns
  notify_on_complete=false plus a note: wait, kill, or hand off.
- _ChildRun.account_background_processes records handed_off_processes and
  orphaned_processes on the result before cleanup kills the leftovers; the
  parent's delegation block renders both.
2026-09-07 12:50:29 -07:00
Teknium 0522ae934e fix(tools): retain producer profile scope in process readers 2026-09-07 08:25:33 -07:00
Teknium fbed1d4584 fix(tools): keep retained terminal results scoped to their owner
Capture the durable parent session before output readers start, including CLI
and non-notifying spawns. Require that parent or its compression continuation
for retained reads; exact and prefix handles alone do not authorize access.

Live Linux terminal/one-shot linger/fresh-reader A/B: base loses results;
updated owner recovers both streams and exit 7. Unbound, foreign session,
delegated child, and other profile cannot recover the receipt. No notifications
are replayed. Full tools suite is queued behind the campaign test lock.

Follow-up to contributor salvage #104805 for #104511.
2026-09-07 08:25:33 -07:00
maximilliangrand 1770825481 fix(tools): register readers atomically with completion 2026-09-07 08:25:33 -07:00
maximilliangrand 633955408b fix(tools): keep retained result reads off live status scans 2026-09-07 08:25:33 -07:00
maximilliangrand b72e373e23 fix(tools): retain completed background process results across exit 2026-09-07 08:25:33 -07:00
maximilliangrand 5695ebc40f refactor(tools): extract process checkpoint persistence 2026-09-07 08:25:33 -07:00
Teknium afb4e080f6 Merge pull request #103496 from NousResearch/fix/goal-judge-own-processes-only
fix(goal): the judge sees only its own session's background processes; pid/session waits expire after 30 min (parked 3h22m on a grandchild's poller)
2026-09-06 12:03:05 -07:00
kshitijk4poor 78fc943ca4 docs(process_registry): say in the argv helper why OOMPolicy is absent 2026-09-06 14:57:33 +05:30
gkd2323c 611ee856c5 fix(process): drop OOMPolicy from systemd-run --scope argv
OOMPolicy is a service-unit property; systemd-run --user --scope
rejects it with 'Unknown assignment: OOMPolicy=kill'. The availability
probe therefore always failed in supervised Linux gateways, making
restart-safe cron worker dispatch (systemd scope) permanently
unavailable and falling back to in-cgroup workers.

MemoryMax + MemoryAccounting remain; OOMPolicy adds nothing for a
transient scope (no service manager to act on the OOM event).
2026-09-06 14:57:33 +05:30
Teknium 463292351f fix: mid-turn user message no longer waits behind a foreground terminal command
A message typed while the agent runs (CLI busy_input_mode=interrupt, gateway
priority redirect, ACP redirect) goes through AIAgent.redirect(). During tool
execution redirect() degrades to steer(), whose delivery rides the tool result
— so a long foreground command (a `sleep 285` CI poller, a build) parked the
user's message until it exited. The UI printed "Redirected current turn" while
nothing happened for minutes.

redirect() now also asks the tool workers to YIELD (tools/interrupt.request_yield).
The local terminal backend's wait loop honours it: the drain thread is stopped, the
still-running Popen is adopted by the process registry as a notify_on_complete
background session (ProcessRegistry.adopt_local — output so far seeds the buffer,
the registry reader continues from the pipe), and the tool returns immediately with
status "yielded_to_background" + session_id. The command is never killed; the
completion notification arrives as usual and process(poll/wait/log/kill) work on it.

Non-local backends and internal env.execute() consumers pass no yield_handler and
are unaffected; a stale yield bit is cleared with the interrupt bit per worker tid.
2026-09-05 13:51:06 -07:00
Teknium f8b87f5637 fix(goal): the judge only sees the goal's own background processes, and a pid/session wait barrier expires after 30 min
In a fan-out run the /goal loop parked for 3 h 22 min at the end on a
grandchild's poller (waiting_on_session=proc_21a6fe2369a1, 00:47 -> 04:09)
while the root itself had nothing running. Every one of the root's 7 logged
judge verdicts was WAIT on a child-owned process.

Two causes. gather_background_processes() was called with no task_id from
both goal-loop callers (CLI cli_loops_mixin, gateway run_goals), and the
registry's task_id is the CONTAINER key, which collapses to one value for
every agent in the process, so the judge's process list was every one of
~1,300 subagents' pollers. And a pid/session barrier had no ceiling: once
the judge said WAIT on a session that never exits, nothing resumed judging;
waiting_since was recorded and never read.

Now list_sessions() reports owner_task_id (the RAW spawning id the
registry already keeps for ownership checks), gather_background_processes
takes owner_task_id and both callers pass their own session id (CLI turns
register processes under self.session_id; gateway turns under
turn_ctx.session_id), and is_waiting() ages out a pid/session barrier
after _MAX_BARRIER_WAIT_S (30 min). Timed barriers keep their own deadline.

Tests: only the owning session's running processes are returned when
owner_task_id is given (unfiltered behaviour unchanged); a live barrier
older than the ceiling clears and judging resumes.
2026-09-05 00:30:19 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 707161e77b simplify(compat): file_tools/send_message_tool/cronjob_tools/process_registry — drop 34 re-exports, repoint 5 callers + 27 test files 2026-09-03 13:30:27 -07:00
Teknium 98c140bc4b simplify(compat): code_execution_tool/environments.local — drop 36 re-exports, repoint 5 callers + 14 test files 2026-09-03 13:24:04 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium d38bb1b3aa Merge simp/r3-33 (late tail) into hermes/simplify-codebase 2026-09-03 05:09:00 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium 23b9ffc4fa fix(integration): restore subprocess stdin=DEVNULL / utf-8 encoding guards and windows-footgun gates dropped by round-3 compaction
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
2026-09-03 02:46:19 -07:00
Teknium 6a6ee5bb67 refactor(tools): process_registry slice -25% LOC — shared exit/status/stdin/reader helpers, early-return watch limiter, suppress(), compact docs; schema byte-identical 2026-09-03 00:44:58 -07:00
Brooklyn Nicholson 83efdf5e5e [verified] fix(cron): close restart handoff races 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 3373e97693 [verified] fix(cron): preserve active runs across gateway restart 2026-09-03 12:40:26 +05:30
Teknium 5f90f7f8f6 refactor(tools): process_registry — walrus in drain/redact/prune, single-branch spawn_via_env tail 2026-09-03 00:08:38 -07:00
Teknium 7f59b62a1a refactor(tools): process_registry — pack result dicts, suppress() in drain 2026-09-03 00:06:22 -07:00
Teknium e01755f7d0 refactor(tools): process_registry — compact docstring layout (content unchanged) 2026-09-02 23:58:25 -07:00
Teknium 781287b535 refactor(tools): process_registry — _exit_fields, _stdin_op returns shared ok result 2026-09-02 23:55:21 -07:00
Teknium 7ef7c5d4e7 refactor(tools): process_registry — _config_seconds, _reap_untracked, inline drain append 2026-09-02 23:53:07 -07:00
Teknium 0000722290 refactor(tools): process_registry — drop stray blanks after decorators 2026-09-02 23:41:15 -07:00
Teknium a884d45321 refactor(tools): process_registry — hug call closers, squeeze body blanks, tighten header comments 2026-09-02 23:38:10 -07:00
Teknium beb1a8d81b refactor(tools): process_registry — contextlib.suppress for swallow-only try blocks, kill_all as sum, derive watcher checkpoint fields 2026-09-02 23:29:08 -07:00
Teknium 2c2935c4c8 refactor(tools): process_registry — _finish_reader/_owns_event/_status_head/_spawn_env helpers; notifications list building compacted 2026-09-02 23:23:25 -07:00
Teknium 870aea6e47 refactor(tools): process_registry — early-return watch limiter, _new_session, psutil gone-tuple, dedup shell-noise list 2026-09-02 23:09:56 -07:00
Teknium 6385b61370 refactor(tools): process_registry — pack ProcessSession constructor kwargs 2026-09-02 22:59:58 -07:00
Teknium 6ab5706f70 refactor(tools): process_registry — unified select/blocking reader loop, drop unused _watch_last_emit_at, tighter docstrings 2026-09-02 22:50:12 -07:00
Teknium eef71bb6e1 refactor(tools): process_registry — single reader read path, _WATCHER_ROUTE_KEYS, hugged signatures 2026-09-02 22:43:28 -07:00
Teknium 60e9009079 refactor(tools): process_registry — mark_exited/_output_tail/_stdin_op helpers, unified watch_disabled emitter, compact docs 2026-09-02 22:34:24 -07:00
Teknium 3bbec90f23 refactor(tools/environments): split local/base into output/wait/session-env siblings; extract docker_egress + remote_common; dedupe remote backends; compact process_registry 2026-09-02 14:45:15 -07:00
kshitijk4poor 75bcd86687 fix(process): keep the sandbox log poller from splitting a UTF-8 character across polls
Follow-up to #92164: the delta window now ends on a character boundary
(up to 3 trailing continuation bytes are held for the next poll), so
multibyte output no longer decodes to U+FFFD at the seam. Verified on
bash, dash and busybox sh; exhaustive-prefix regression test added.
2026-09-03 02:20:46 +05:30
Adolanium 8b681f70ea perf(process): poll sandbox job logs for new bytes only
The background-process poller for non-local backends ran `cat` on the
whole log file every two seconds, then threw away everything except the
part it had not seen yet. The offset it needed was already tracked one
line below, so the full read was pure waste.

Cost of one poll grew with the total output so far, which makes the cost
of a run grow with the square of its length. A job writing 10 MB over an
hour moved about 9 GB across the docker or SSH channel to deliver 10 MB
of output.

The poller now asks the shell for the file size and the bytes after the
offset in one command. Reading the size first and cutting the tail at
that same size keeps the two in step, so a file that grows mid-command
never sends a byte twice. A file that shrank was rotated or truncated,
so the offset drops back to 0 and the buffer is dropped.

The output buffer is now appended to rather than replaced, matching the
local reader loops, and the offset is counted in bytes because the shell
counts bytes.
2026-09-03 02:00:35 +05:30
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 5a8e8a6b87 fix(terminal): strict Linux-only gating for background-executor systemd scopes (#70716 follow-up)
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):

- Gate every scope-path branch on a new _IS_LINUX constant instead of
  'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
  never touches systemd code — no probe subprocess, no scope argv,
  byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
  legacy argv byte-identical, no unit recorded) and probe-returns-False
  off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
  wired into the on-demand windows-venv-e2e lane): real spawn_local on
  windows-latest asserting jobs run exactly as before — spawned, output
  captured, exit code correct, systemd path never reached even under
  faked gateway identity.

Refs #70716, #71378.
2026-09-01 02:32:53 -07:00
Teknium 5a4dbdec27 fix(delegation): subagent process notifications stay suppressed when the container key collapses
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.

Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).

Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
2026-08-31 07:28:18 -07:00
liuhao1024 8557e0a480 fix(tools): surface config-level model_not_found notices in delegation batch reports
When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).

Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
2026-08-30 22:19:10 -07:00
David Metcalfe c05d04fffb feat(delegation): surface config-level model_not_found notice in delegation batch reports
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.

Closes #97654.
2026-08-30 21:07:04 -07:00
Teknium 03e66c8cba polish(tool-search): stub-optimized openers for deferred tools — trigger+verb in the first ~60 chars (the catalog stub is the only ambient hint a deferred tool exists) 2026-08-29 18:13:20 -07:00
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium baa344dee7 refactor(process): schema diet — enum names the verbs, description keeps only non-obvious semantics; write-vs-submit trap teaching emphasized (306 -> 228 tok/call, -25%) (#97279) 2026-08-28 09:23:00 -07:00