Commit Graph

246 Commits

Author SHA1 Message Date
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium 561b053f79 perf(agents): run per-child timers on one shared scheduler thread
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog).  A
profiled session with ~130 children was carrying ~1000 threads.  All
of these timers now run on a single process-wide daemon thread.

- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
  one Condition-driven daemon thread.  schedule(fn, interval) -> handle;
  handle.cancel(wait=) blocks for an in-flight run like the old join.
  A callback returning False stops itself; a raising callback is logged
  at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
  scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
  idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
  stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
  _lease_refresh_interval; lease-lost / refresh-error interrupt paths
  and the stop-event fencing are unchanged; the join(timeout=1.0) is now
  cancel(wait=1.0) so the interrupt clear still runs after any in-flight
  tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
  schedule(); the poll body is _tick(), same sampling state machine.

Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132.  At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
2026-09-03 02:44:24 -07:00
Teknium 798470930f refactor(delegate): tool description text as joined constants (byte-identical) 2026-09-02 19:50:30 -07:00
Teknium defca3639a refactor(delegate): reflow comments/docstrings to 118 cols (word-identical, AST-identical) 2026-09-02 19:25:04 -07:00
Teknium 24ef289225 refactor(delegate): schema properties via _p() table (byte-identical schema) 2026-09-02 19:18:35 -07:00
Teknium 1283ccd855 refactor(delegate): docstring compaction by hand (every rule/why kept) 2026-09-02 19:16:07 -07:00
Teknium aa31f83995 refactor(delegate): origin — packed AIAgent kwargs, override table in _build_children, tighter preamble 2026-09-02 19:11:48 -07:00
Teknium eabec2ff50 refactor(delegate): AST-identical re-layout pass 2 2026-09-02 19:04:50 -07:00
Teknium 61f32ec330 refactor(delegate): drop unreferenced re-exports from origin (refs.py-verified) 2026-09-02 18:59:41 -07:00
Teknium fe9c26ee5d refactor(delegate): _ChildRun dataclass replaces _WorktreeReporter/_ChildWorkspace/_ChildFailure plumbing 2026-09-02 18:27:59 -07:00
Teknium 378692be4f refactor(delegate): fold _construct_child_agent/_announce_child_spawn into _build_child_agent; _quiet best-effort sites 2026-09-02 18:04:43 -07:00
Teknium b18228f377 refactor(delegate): AST-identical re-layout (pack signatures/call args, join split literals) 2026-09-02 17:59:28 -07:00
Teknium 901ba7db6f refactor(delegate): child_run best-effort blocks via _quiet; shared failure-entry tail 2026-09-02 17:54:08 -07:00
Teknium 64957b7e02 refactor(delegate): _Batch dataclass owns batch execution/dispatch; _quiet best-effort contextmanager 2026-09-02 17:50:16 -07:00
Teknium d77505c007 refactor(delegate): extract toolset resolution + task validation modules, move runtime resolution to config 2026-09-02 17:44:45 -07:00
Teknium 3da2ff88e3 refactor(delegate): pack re-import block; single blank line between top-level defs (AST-identical) 2026-09-02 17:03:55 -07:00
Teknium 987bc82574 refactor(delegate): compact docstrings/comments by hand (every WHY kept); collapse orchestrator toolset branches 2026-09-02 17:03:05 -07:00
Teknium 016f31375d refactor(delegate): join wrapped logical lines that fit in 120 cols (AST-identical) 2026-09-02 16:59:26 -07:00
Teknium e204008f10 refactor(delegate): _knob() unifies config>env>default int/float knobs; timeout diagnostic sections -> _diag_sizes/_diag_threads + attr table; drop dead _handle_child_wait_failure re-import 2026-09-02 16:57:22 -07:00
Teknium 338da30a45 refactor(delegate): _resolve_child_runtime returns AIAgent kwargs (drop _ChildRuntime; routing-filter table); dispatch reuses _detach_child/_signal_child_stop; control actions -> _CONTROL_OUTCOMES table; attribution lookup dedupe 2026-09-02 16:50:57 -07:00
Teknium 52419bda5d refactor(delegate): split delegate_task into _normalize_task_list/_coerce_task_schemas/_build_children/_execute_and_aggregate(_Batch); _build_child_agent -> _open_child_session_db/_construct_child_agent/_announce_child_spawn; _run_single_child -> _lease_child_credential/_await_child/_merge_late_steer 2026-09-02 16:46:39 -07:00
Teknium 05526b028a refactor(tools/delegate): split delegate_tool into child_run/config/dispatch/progress/registry/results; compact delegation helpers 2026-09-02 14:45:15 -07:00
Teknium 7840a0e2d9 feat: delegation batch tags read "set N" instead of a hex id slice
Interleaved subagent fan-outs were tagged with the first 4 hex chars of the
delegation id ([b2ac 3/9]), which is attributable but unreadable. Batches are
now numbered in order of appearance per process: [set 1 · 3/9], [set 2 · 1/7].
Desktop /agents already labels groups "Delegation N", so its duplicate hex
badge is dropped.
2026-09-02 10:12:54 -07:00
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 9387bf929c fix(delegate): drain abandoned-worker transports FD-safely on child timeout
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).

Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
  the shared client's pooled sockets (force_close_tcp_sockets — FD release
  stays with the owning worker), abort+poison of the cached per-request
  openai/anthropic wire clients, Codex app-server request_interrupt(), and
  the inline _active_request_abort hook. Never client.close(), never
  socket.close().
- delegate timeout path: after registering the deferred-close callback,
  run one immediate drain plus one 5s re-sweep (covers a connection opened
  between the interrupt and the first sweep). The settled read (EOF/EPIPE)
  lets the worker unwind, which triggers the deferred close on the worker's
  own thread — the only safe FD-release boundary. A worker that still never
  settles retains its resources rather than risking a cross-thread close.

Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.

Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.

Closes #94248
2026-09-01 12:07:52 -07:00
Leandro Piccione aa1d22670e fix(delegate): defer timed-out child teardown 2026-09-01 12:07:52 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
itskaism 5ce8f71553 fix(delegation): report schema-invalid child results as failed, not completed
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.

Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.

Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
2026-08-31 01:02:42 -07:00
itskaism 1b6ea1a2c2 fix(delegation): report failed children as failed, not completed
A subagent whose loop gave up on a structured failure (e.g. "API call
failed after 3 retries: HTTP 524") returns that error message as
final_response together with completed=False / failed=True /
failure_reason. _run_single_child derived the batch-entry status from
the summary alone (`elif summary and not _empty_sentinel: status =
"completed"`), so the non-empty error text made the batch report show
the task as "✓ status=completed" — the `failed` flag was never
consulted anywhere in delegate_tool.py. Only the "(empty)" sentinel was
mapped to failed.

Fix, at the single status-determination choke point both the single-task
and batch paths share:

- `failed=True` on the child result now wins over a non-empty summary:
  status = "failed".
- The child's classified failure_reason (rate_limit / billing /
  server_error / ...) is propagated onto the batch entry so the parent
  can tell a quota wall from a real task error without parsing prose.
- exit_reason for a structured failure is "error" instead of falling
  through to "max_iterations" (which also wrongly set truncated=True).

Successful children (completed=True, no failed flag) are untouched —
covered by an explicit control test alongside the regression test,
which is red on the old code and green with the fix.
2026-08-30 22:19:10 -07:00
David Metcalfe b4d5174385 fix(delegation): pin failure-status edge cases and document exit_reason enum
Per-finding verdicts from the cross-vendor review of fix/status-fix
(#97655/#97654):

[1] Flash NIT (real, cheap) — FIXED. Added test_error_without_failed_flag_
    marks_failed: an error string with the 'failed' key ABSENT (not False)
    must still be status=failed + exit_reason=error. The branch order
    (result.get('failed') or result.get('error')) already handles this; the
    test pins the error-alone path.

[2] GPT-OSS SHOULD-FIX — PINNED. Added test_empty_error_with_summary_is_
    completed: error='' is falsy so result.get('error') falls through to the
    summary-presence heuristic => status=completed. No code change; the
    existing branch is correct and the new test locks it in.

[3] GPT-OSS SHOULD-FIX — VERIFIED, NO CHANGE. Grepped every delegation
    exit_reason consumer:
      * tools/delegation_live_log.py finalize() prints exit_reason generically
        and only special-cases == 'max_iterations' for a readable suffix.
      * tools/process_registry.py derives truncated as
        (truncated or exit_reason == 'max_iterations') — gated, not exhaustive.
      * tools/async_delegation.py passes exit_reason through generically.
    The gateway/status.py, cron/scheduler.py and run_agent.py 'exit_reason'
    hits are a DIFFERENT field (turn_exit_reason / gateway exit reason), not
    the delegation result's exit_reason. No exhaustive if/elif over the enum
    missing an 'error' case, so nothing to add.

[4] GPT-OSS NIT — DONE. Enriched _run_single_child's docstring to enumerate
    status in {completed, interrupted, failed} and exit_reason in {completed,
    max_iterations, interrupted, error}, and added a compact enum comment at
    the result-entry construction. Verified the process_registry.py renderer
    comment (truncated <= exit_reason == 'max_iterations') still holds — the
    truncation flag is derived exactly that way, so no contradiction.

[5] GPT-OSS NIT — REJECTED. The proposed 'fallback for legacy dicts that
    explicitly set failed=False' is not adopted. No consumer produces a result
    dict with an explicit failed=False and no summary while relying on
    completed semantics: run_agent.py sets failed=True only on genuine failure
    and omits the key on success (no failed=False producer). Also, the
    proposed elif would reintroduce ambiguity (explicit failed=False + no
    summary => 'completed'?) and diverge from the conservative else => 'failed'.
    result.get('failed') is falsy for both explicit-False and absent, so no
    distinction exists to preserve; the else is the correct default.

Tests: 301 passed, 7 skipped (tests/tools -k 'delegate or process_registry').
TestDelegateFailedChildStatus: 6 passed.
2026-08-30 21:07:04 -07:00
David Metcalfe ec02d5179a fix(delegation): report provider-failed subagents as failed, not completed/max_iterations
A provider-rejected child (e.g. HTTP 400 "<model> is not a valid model ID")
returns completed=False with failed=True + an error string as its terminal
final_response. _run_single_child keyed status on summary presence alone and
assumed completed=False meant iteration-budget exhaustion, so such a child
was reported status=completed + exit_reason=max_iterations, rendering the
false '"TRUNCATED: hit max_iterations"' banner.

Consult the structured failure fields (failed / error) before falling back to
the summary-presence heuristic, and derive exit_reason honestly: failure ->
'error', interrupted -> 'interrupted', completed -> 'completed', and only
genuine budget exhaustion (completed=False, no failure) -> 'max_iterations'.
The 'truncated' flag stays keyed on exit_reason == 'max_iterations', so it is
now correct automatically. The batch renderer needed no change (the error
field is already plumbed into the result entry for the parent).

Closes #97655
2026-08-30 21:07:04 -07:00
Teknium 5a134383fe fix: failed subagents now surface a clean error to the user (CLI + gateway)
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.

- delegate_tool: result.failed now forces status 'failed' (with the
  error carried on the entry); new shared format_subagent_failure_line()
  renders one clean human-readable line (traceback -> exception message,
  length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
  terminal failure status deliver that line via _deliver_platform_notice
  BEFORE all progress-queue gates; tool_progress_callback is now always
  attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
2026-08-30 20:40:14 -07:00
Teknium 6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
Teknium 8a8aa850f1 fix(delegate): inherit endpoint-scoped capability map only on the parent's exact route
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
2026-08-30 05:16:10 -07:00
Teknium bacb90fe20 feat(delegation): honor delegation.request_overrides on all three resolution branches with explicit-over-runtime merge precedence
Completes the #90953 salvage on post-#98237 main:

- New _merge_request_overrides helper defines the precedence contract:
  explicit delegation.request_overrides merges OVER runtime/parent-derived
  overrides — explicit top-level keys win; extra_body is deep-merged one
  level so runtime extra_body keys survive unless redefined. Inputs are
  copy.deepcopy'd so transport-side mutation can't leak into config or the
  provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
  provider-alongside-base_url runtime overrides instead of being a separate
  return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
  delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
  (previously only when override_provider was set), enabling the inherit
  branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
  proofs, explicit-over-runtime precedence on the provider-alongside-base_url
  path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
  the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
2026-08-29 19:13:23 -07:00
fabiantax d3bfd2e9b1 feat(delegation): forward delegation.request_overrides on direct-endpoint branch
The direct base_url branch of _resolve_delegation_credentials returned no
request_overrides key, so a direct OpenRouter delegation (provider=custom,
base_url=openrouter.ai/api/v1) could not pass routing hints to its children.
The named-provider branch already forwards runtime request_overrides; this
gives the direct branch the same contract, honouring delegation.request_overrides
from config (dict → forwarded, anything else → None).

Primary use: extra_body.provider = {"sort": "throughput"} so delegation
children route to the fastest OpenRouter provider for their model, per the
fab-swarm throughput work (#901).
2026-08-29 19:13:23 -07:00
Teknium 7b3c9f86ed Merge origin/main — resolve todo-state seam onto the todo_list rename (accept legacy alias) 2026-08-29 18:55:29 -07:00
Ayush Nangia 5cace31708 fix(delegation): carry provider request_overrides through the base_url path (#65035) 2026-08-29 18:16:04 -07:00
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium 9dfbde19db refactor(delegate_task): tasks-only interface + depth-derived delegation (1,201 → 773 tok/call, −36%) (#96424)
* refactor(delegate_task): depth-derived delegation (role param retired), session-filtered restrictions, background unadvertised — 1,201->819 tok/call

* refactor(delegate_task): tasks[] is the only advertised shape — single task = one-entry array (legacy goal/context/output_schema stay handler-accepted)
2026-08-27 07:38:53 -07:00
Teknium 7526bd39a8 feat: every subagent's prompt embeds the workspace's project context files
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.

The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.

Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
Hermes Agent b95ec1cb5d fix(delegation): running subagents stay visible to list/steer across parent-agent rebuilds, and child-started process notifications carry delegation attribution
Control path: delegate_task(action=list/steer/stop) resolved ownership
purely through the _delegate_parent_ref weakref identity chain. The CLI
rebuilds its AIAgent mid-session (self.agent = None on route-signature
change, credential refresh, /model, MoA one-shots), so a running child's
chain pointed at a dead object and the child went invisible/unsteerable
while completion delivery (durable session-id routed) still worked.
Observed live 2026-08-17: deleg_88454b70 / sa-0-dc0100f4.

Fix: register each child with the owning conversation's durable session
id (owner_agent_session_id, the same spine delivery routes by) and add a
second ownership tier that matches it against the calling parent's
session_id with compression-lineage resolution on both sides. Foreign
sessions still fail closed.

Presentation path: background processes started BY a subagent (task_id ==
subagent_id) route their notify_on_complete notifications to the parent
conversation by design, but arrived as anonymous raw output walls. The
formatter now resolves the task_id against the live + recently-finished
subagent registry (bounded retention survives child completion) and adds
a provenance line (subagent id, delegation id, goal snippet), trimming
the output tail for subagent-owned processes. Parent-owned process
notifications are byte-identical to before.
2026-08-18 00:01:56 -07:00
kshitij ce93a398e8 refactor(delegate): extract the unproven-payload factory; drop the source-reading test
Phase 2c fold. The schema guard added in the previous commit read and
AST-parsed delegate_tool's source, which AGENTS.md:1514 bans outright ("Never
read source code in tests" -- it passes when the implementation is subtly
broken and fails on a correct refactor). Extracting the shared factory the rule
prescribes removes the duplication the AST test was invented to police, so one
change resolves both.

- subagent_worktree: new module-level `mark_worktree_payload_unproven()` +
  `unproven_worktree_payload()`. Both producers of this schema now call them,
  so the payload cannot drift and the note string exists once.
- delegate_tool: the finalize-raised fallback calls the factory instead of
  hand-building the dict (-16 lines). The re-import is guarded: the outer
  `except` can be entered because the `from tools import subagent_worktree`
  itself failed, in which case the name is unbound -- an inline fallback keeps
  the flag rather than raising NameError and losing it.
- Test replaced with a BEHAVIORAL equivalent: it calls the real factory and
  compares its key set against live `finalize_subagent_worktree()` output. Same
  contract, no source reading, refactor-proof, and it actually executes the
  code.

Also folded from the same review:

- Fail-closed on an unmeasurable commit count. With no `base_commit` the
  rev-list probe never ran, `commits` kept its unproven 0 default, and a clean
  tree still reached `git worktree remove --force` + `git branch -D` -- the
  exact bug class #88113 is about, on a public function that takes a
  caller-supplied dict. Now returns un-inspected instead, with a test driving a
  real child commit.
- Per-probe diagnostics: the note said only "rev-list/status non-zero". It now
  names WHICH probe failed, its exit code, and a bounded git stderr tail, so
  the parent (and the human) can act on first read.
- Dropped the redundant `inspection_ok` bool for a `failed: list` of reasons;
  removed the duplicated index-corruption block in favor of the existing
  `_break_git_index()` helper.

Validation: 21/21 tests/tools/test_subagent_worktree.py; ruff clean; ty clean
on subagent_worktree.py and 64-vs-64 unchanged on delegate_tool.py (all
pre-existing, verified against the base commit). All 6 guards mutation-checked
twice -- neutering the flag fails 6, reverting production to pre-fix main fails
the same 6. E2E on real git: clean still prunes; corrupt index keeps the work
and reports the real stderr; empty base_commit keeps a committed child.
2026-08-17 19:41:32 +05:30
kshitij 97c4f9eeec test(delegate): assert the unproven-state contract, not its prose
Review fold on the #88113 follow-up. The new guards asserted implementation
details that a strictly-better future change would break, and the second
producer of the payload schema had no coverage at all.

- The distinguishability test asserted the failure payload was byte-identical
  to the genuinely-clean one (`for key in commits/dirty/pruned: assertEqual`).
  That freezes the AMBIGUITY as a required property: emitting `commits: None`
  for "unknown" would improve exactly what #88113 is about and fail the test.
  Now asserts what the parent actually depends on -- both keep the worktree,
  and only the flag separates them.
- `assertNotIn("inspection_failed", ok_payload)` pinned key ABSENCE on the
  happy path, forbidding an always-present-but-False flag (a legitimately
  better JSON contract: stable key set for serializers). Now
  `assertFalse(...get("inspection_failed", False))` -- same coverage, tolerant
  of that refactor.
- `assertIn("UNKNOWN", note)` coupled tests to one word of English prose, and
  was not even a cross-producer contract: delegate_tool's note said "state
  unknown" (lowercase), so a copy-edit broke the implied convention. Tests now
  assert the note names the worktree AND branch -- the actionable part for a
  human -- and both producers' notes were aligned to read as one contract.
- The raises test never proved its patched seam ran (a future short-circuit
  before any git call would keep it green while proving nothing). Now checks
  `call_count` and mirrors the branch-survival + note-names-path legs its
  sibling had.
- NEW `WorktreePayloadSchemaTests`: commit 2's whole point is the schema the
  parent reads, but delegate_tool's fallback -- the second producer -- was
  verified only by reading. It now AST-parses the real fallback dict literal
  and compares against live `finalize_subagent_worktree()` output, so the two
  producers cannot drift and the pre-fix leak (repo_root/base_commit, missing
  commits/dirty/pruned) cannot come back.
- Docs/docstring drift: the flag has a second trigger (finalization itself
  raising, handled in delegate_tool), and the module docstring listed
  `inspection_failed` without `note`. Both corrected.
- Extracted the duplicated 5-line "corrupt the index" setup into
  `_break_git_index()` beside the file's other module-level helpers.

Validation: 19/19 tests/tools/test_subagent_worktree.py; ruff clean. New
schema guard mutation-checked -- reverting delegate_tool's fallback to the
pre-fix `dict(_worktree_info)` shape fails it. Restores checksum-verified.
2026-08-17 19:41:32 +05:30