Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.
Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
Carry the owning task's notification subscriptions independently of dependency
edges, within the creation transaction. Prefer its durable session over worker
and request-local sessions while preserving explicit overrides. Cover worker
CLI create and built-in decomposition, and retain conversation route anchors.
Auto-subscribe no longer upgrades an inherited passive subscription.
Slim adaptation of Christopher-Schulze's session-precedence fix in #85687,
expanded to durable subscription provenance and sibling creation paths.
Related: #85575, #85687
Validation: strict RED/GREEN (7 failing cases before; 7 passing after), then
58 Kanban test files: 383 passed, 2 skipped. Real dispatcher-spawn subprocess
probe covers direct, linked, unlinked, explicit-session, worker CLI, built-in
children and a plain CLI negative control, with recording transport only.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
Authorize route-only profiles using the ordered canonical route matcher and
served-profile set at both claim and delivery. Preserve secondary credential
boundaries, retry denied routes, and keep scope/parent anchors plus transport
provenance on synthetic wakes. Install the destination runtime scope rather
than inheriting the notifier's scope; a removed profile cannot wake as primary.
Slim forward-port of the direction in #101196/#101397 and #93863 (#93851).
Canonical scope_id takes precedence over the guild alias and live chat cache.
Two route invariants and two runtime-scope invariants reproduced red first.
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Bind settlement to the exact adapter task and event. Rejected or cancelled preparation refunds the existing claim; cancellation after the agent runner starts remains counted. Keep profile-scoped callback context and the manager replacement guard. A fire count is not outbound delivery proof. Drop departed routes rather than executing their stale schedule.
Slim accounting-invariant salvage of #93174; preserve current direct adapter dispatch instead of reviving its FIFO/inflight implementation. Prior art #92858.
Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Recover active watches from current persisted session origins and exact route
keys, reading heartbeat state off-loop in each source's profile. Failed scans
leave watches intact for the poller's next retry. Start the heartbeat poller
even when startup restores no watches.
Slim synthesis of #92660, #98310 and #98313, with earlier restart recovery
prior art from #92594. The integration poller calls restore_heartbeat_watches
on every poll, including empty registries.
Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.
Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
A process that finishes while the child is alive needs no handoff, but if the child never
polls/waits/logs it, the result vanished: the completion notice is suppressed in the parent
and the child's summary never mentions it. Finalization now attaches exit code + output
tail as unread_completions, rendered in the parent's delegation notice.
A child's background processes are killed at its teardown and their
notify_on_complete notices are suppressed in the parent, yet the child's
terminal result still said `notify_on_complete: true` and the parent's
delegation notice said nothing about processes left behind. Orchestrators
believed "CI watcher running" and waited on a completion that could never
arrive (recurring in the Sep 7 campaign sessions).
- process_manage(action="handoff", session_id, data="<purpose>"), children
only: process_registry.transfer_ownership flips owner_task_id/task_id/
session_key to the parent under the registry lock, so the completion is
stamped with the parent's owner at exit, passes the parent's sa- filter,
and is reaped by the parent, not the child. Cap 3 per child; an exited,
foreign, or non-child request is a tool error. The purpose rides the
event as handoff_note and renders in the parent's notice.
- Child terminal(background=True, notify=True) now returns
notify_on_complete=false plus a note: wait, kill, or hand off.
- _ChildRun.account_background_processes records handed_off_processes and
orphaned_processes on the result before cleanup kills the leftovers; the
parent's delegation block renders both.
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.
Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.
Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
Persist paused state, timestamp, reason and no first trigger in the original
locked creation write. Forward the same boolean contract across CLI, tool,
gateway API and dashboard API, validating at the store boundary. Preserve
explicit operator force-run behavior and normal enabled creation.
The live CLI probe also caught the command shim dropping failure return codes;
forward them so invalid creation reports exit 1 rather than success.
Credit earlier atomic-creation work in #78935 and #94952 and the focused
implementation in #104578. The broader manifest staging layer is not imported.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Chloé DuPont <321112755+misschloedupont@users.noreply.github.com>
Capture the durable parent session before output readers start, including CLI
and non-notifying spawns. Require that parent or its compression continuation
for retained reads; exact and prefix handles alone do not authorize access.
Live Linux terminal/one-shot linger/fresh-reader A/B: base loses results;
updated owner recovers both streams and exit 7. Unbound, foreign session,
delegated child, and other profile cannot recover the receipt. No notifications
are replayed. Full tools suite is queued behind the campaign test lock.
Follow-up to contributor salvage #104805 for #104511.
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.
Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Slim redo of #104546 and #104551: scan newest-first, match suppression only before payload separators, and keep error context. Covers wake gates and empty outputs without reading every historical file twice.
Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Use successful applied result records rather than requested operations, and keep staged writes silent. Include legacy delete/write messages.
Fixes#104506
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.
Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.
Refs #103974, #104058, #104904
Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions:
1. delegation.independent_completions (new, default false). #104299 made every
ungrouped task its own completion message, so a 15-task call woke the
orchestrator up to 15 times; one chain received 132 notices and answered
130 of them with "already incorporated". A multi-task call now returns as
ONE consolidated message unless the flag is on; `group` is inert until then.
2. Queued units were killed before they started. Units of one call share a
pool slot but the executor was still sized by slots, so with 15 units live
a new unit queued behind a full pool; the stale monitor's clock ran from
dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s`
when its thread finally came up (13 such lanes in one session). The
executor now grows to the number of live units and the stall clock arms
when the runner actually starts.
3. The tool text said "do not wait or poll — just continue" without saying
that completions are delivered only BETWEEN turns. A model that never ends
its turn (one 203-minute turn, 717 API calls) never received 40 finished
results. Tool description, dispatch note and completion header now say to
finish independent work, give a one-line status, and end the turn.
Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.
Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.
Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.
Co-authored-by: holny <holny@foxmail.com>
Salvage #99095 (e7ea53074ab2b64a1530641659399d9f1bb4b435), completing one-time migration for both boolean values and using the existing storage helpers. Hydration and local toggles never edit backend configuration. Fixes#99076.
Desktop-only backends now poll curator and personal/org skill sync without another long-lived loop. Respect active turns, the actual idle threshold, and messaging gateway ownership. Credit Jackal991 for the report and candidate #95453.
Wire the existing consent-aware, profile-keyed hook registrar at agent
construction, where the correct session home is already bound. This covers
serve and TUI agent construction without a startup-only registration or a
new helper that swallows registration failures.
Live isolated serve/WebSocket probes reproduce the missing registration on
base for both write_file and terminal, then verify each consented profile
blocks its configured tool while the unapproved profile still runs normally.
Repeated alpha construction does not duplicate hook callbacks.
Slim implementation of the agent-build placement proposed in #57020;
thanks also to the profile-scoped analysis in #102691.
Co-authored-by: grimmjoww578 <willies578@gmail.com>
Keep two invariants covering empty create/cwd/save/fork, genuine content and existing-row metadata. Preserve original authorship and avoid source-only legacy pruning: an empty ACP row does not prove its owner is dead. Native ACP wire plus a local streaming model fixture verifies the first turn and nonempty fork remain durable.