Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.
Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.
Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Slim redo of #104546 and #104551: scan newest-first, match suppression only before payload separators, and keep error context. Covers wake gates and empty outputs without reading every historical file twice.
Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Emit the one-shot reason marker outside the CLI facade; parse whole codes before falling back to legacy prose. Explicit coordination and unknown codes cannot be labeled target_busy.
Fixes#104784
Co-authored-by: William Echo <2054936695@qq.com>
Use successful applied result records rather than requested operations, and keep staged writes silent. Include legacy delete/write messages.
Fixes#104506
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Recover legacy stringified lists with a warning. Reject malformed shapes and nonstring members without admitting approvals or rewriting user config on read.
Fixes#104779
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.
Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.
Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698
Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
Salvage the bounded target-slot design from #104617, using mandatory
word separators to avoid ambiguous repeated matches. Preserve long
payload detection and execution-verb boundaries. Replace the three
candidate tests with two context-loader invariants and document the
heuristic's limits.
Fixes#104609
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
Reuse the effective gateway config and shared session platform policy before hashing model-facing metadata. Preserve original routing state and cover enabled/disabled redaction across all busy injection routes.
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.
Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.
Refs #103974, #104058, #104904
Document actual groups.promote/groups.demote parameters and required
old-writer fencing before confirmation. Demotion is a controlled rejoin
step, not an atomic promote-then-demote handover or log reconciliation.
Clarify replica coverage, confirmation meaning, and lineage readback.
Corrected redo of #104342; its nonexistent groups.peer methods and unsafe
handover ordering are not carried forward.
Fixes#104309
Refs #104904
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Salvage #101887 after native Actions 34097643131 reproduced WinError 32 using a ready Electron app with its cwd inside the live release. Reuse the existing install-scoped process cleanup before promotion and wait after forced termination, preserving rollback.
Co-authored-by: fangliquan <fangliquan@qq.com>
Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions:
1. delegation.independent_completions (new, default false). #104299 made every
ungrouped task its own completion message, so a 15-task call woke the
orchestrator up to 15 times; one chain received 132 notices and answered
130 of them with "already incorporated". A multi-task call now returns as
ONE consolidated message unless the flag is on; `group` is inert until then.
2. Queued units were killed before they started. Units of one call share a
pool slot but the executor was still sized by slots, so with 15 units live
a new unit queued behind a full pool; the stale monitor's clock ran from
dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s`
when its thread finally came up (13 such lanes in one session). The
executor now grows to the number of live units and the stall clock arms
when the runner actually starts.
3. The tool text said "do not wait or poll — just continue" without saying
that completions are delivered only BETWEEN turns. A model that never ends
its turn (one 203-minute turn, 717 API calls) never received 40 finished
results. Tool description, dispatch note and completion header now say to
finish independent work, give a one-line status, and end the turn.
Remount scope-owned credential state on Applies-to changes and bind onboarding requests to their initiating route. Cancel polling and invalidate late results when setup closes or reopens, without undoing writes already sent.
Co-authored-by: By JTT <29462570+jordan-thirkle@users.noreply.github.com>
Adapt the config-only portion of #104347; omit its environment flag and unrelated docs. Explicit update commands remain independent.
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.
Slim redo informed by #104274, #104283, and #104285.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>