When a typo'd delegation.model slug is rejected by the provider, every
subagent in the batch dies within a second carrying the provider's
rejection text as its summary while the per-task blocks keep labelling
it status=completed + TRUNCATED. The config-level root cause stays
buried in the batch dump (#97654).
Detect the rejection in the batch render path (summary/error text
matching a model_not_found pattern from agent.error_classifier AND
naming the configured delegation model id) and prepend a single
config-level notice with the model id, hit count, and the setting to
fix, before the per-task blocks.
A subagent whose loop gave up on a structured failure (e.g. "API call
failed after 3 retries: HTTP 524") returns that error message as
final_response together with completed=False / failed=True /
failure_reason. _run_single_child derived the batch-entry status from
the summary alone (`elif summary and not _empty_sentinel: status =
"completed"`), so the non-empty error text made the batch report show
the task as "✓ status=completed" — the `failed` flag was never
consulted anywhere in delegate_tool.py. Only the "(empty)" sentinel was
mapped to failed.
Fix, at the single status-determination choke point both the single-task
and batch paths share:
- `failed=True` on the child result now wins over a non-empty summary:
status = "failed".
- The child's classified failure_reason (rate_limit / billing /
server_error / ...) is propagated onto the batch entry so the parent
can tell a quota wall from a real task error without parsing prose.
- exit_reason for a structured failure is "error" instead of falling
through to "max_iterations" (which also wrongly set truncated=True).
Successful children (completed=True, no failed flag) are untouched —
covered by an explicit control test alongside the regression test,
which is red on the old code and green with the fix.
Follow-ups on the salvaged #97744 runner:
- tui_gateway/hosted_room_driver.py: HostedRoomRuntime.cancel() treated its
initial status read as truth, so a task transitioning queued->running (or
settling) between the read and the state call surfaced a transient
'running work requires acknowledged two-phase cancellation' /
StaleTaskError to the caller and failed groups.disband. Deterministic
repro on the PR head: test_client_event_id_cannot_squat_disband_receipt
failed 5/5 locally. cancel() now re-reads and re-routes on every
race-shaped failure (bounded retries), returns already-cancelled tasks
idempotently, and rejects truly terminal states honestly.
- methods_groups conflict resolution keeps both method sets: the replication
surface from #99047 (groups.replicate/replica_state/promote/demote) and
the runner surface from this layer (groups.stop/retry/approve).
- test_groups_replication_methods.py updated to the runner's stricter
create contract (2-6 profile-backed members, live worker service).
Per-finding verdicts from the cross-vendor review of fix/status-fix
(#97655/#97654):
[1] Flash NIT (real, cheap) — FIXED. Added test_error_without_failed_flag_
marks_failed: an error string with the 'failed' key ABSENT (not False)
must still be status=failed + exit_reason=error. The branch order
(result.get('failed') or result.get('error')) already handles this; the
test pins the error-alone path.
[2] GPT-OSS SHOULD-FIX — PINNED. Added test_empty_error_with_summary_is_
completed: error='' is falsy so result.get('error') falls through to the
summary-presence heuristic => status=completed. No code change; the
existing branch is correct and the new test locks it in.
[3] GPT-OSS SHOULD-FIX — VERIFIED, NO CHANGE. Grepped every delegation
exit_reason consumer:
* tools/delegation_live_log.py finalize() prints exit_reason generically
and only special-cases == 'max_iterations' for a readable suffix.
* tools/process_registry.py derives truncated as
(truncated or exit_reason == 'max_iterations') — gated, not exhaustive.
* tools/async_delegation.py passes exit_reason through generically.
The gateway/status.py, cron/scheduler.py and run_agent.py 'exit_reason'
hits are a DIFFERENT field (turn_exit_reason / gateway exit reason), not
the delegation result's exit_reason. No exhaustive if/elif over the enum
missing an 'error' case, so nothing to add.
[4] GPT-OSS NIT — DONE. Enriched _run_single_child's docstring to enumerate
status in {completed, interrupted, failed} and exit_reason in {completed,
max_iterations, interrupted, error}, and added a compact enum comment at
the result-entry construction. Verified the process_registry.py renderer
comment (truncated <= exit_reason == 'max_iterations') still holds — the
truncation flag is derived exactly that way, so no contradiction.
[5] GPT-OSS NIT — REJECTED. The proposed 'fallback for legacy dicts that
explicitly set failed=False' is not adopted. No consumer produces a result
dict with an explicit failed=False and no summary while relying on
completed semantics: run_agent.py sets failed=True only on genuine failure
and omits the key on success (no failed=False producer). Also, the
proposed elif would reintroduce ambiguity (explicit failed=False + no
summary => 'completed'?) and diverge from the conservative else => 'failed'.
result.get('failed') is falsy for both explicit-False and absent, so no
distinction exists to preserve; the else is the correct default.
Tests: 301 passed, 7 skipped (tests/tools -k 'delegate or process_registry').
TestDelegateFailedChildStatus: 6 passed.
When the configured Subagent Model is rejected by the provider (HTTP 400:
"<model> is not a valid model ID"), every subagent in a delegation batch dies
before doing any work, but the batch report only buried the cause inside each
per-task block. Detect the config-level case in the delegation batch renderer
(both the multi-task fan-out and single-task variants) and emit one actionable
notice at the top of the report naming the configured model + provider, and
pointing at Settings -> Advanced -> Subagent Model (hermes config get
delegation.model). The notice only fires when a result entry's error/summary
both matches a model_not_found phrase AND names the currently configured model,
so a stale task failing on a removed model isn't mis-attributed. Detection
loads the delegation config lazily and fails open (no notice) on any error.
When no fallback chain is configured, the notice calls out that no failover was
attempted. Renderer-only change: no changes to delegate_tool status derivation
or the result schema.
Closes#97654.
A provider-rejected child (e.g. HTTP 400 "<model> is not a valid model ID")
returns completed=False with failed=True + an error string as its terminal
final_response. _run_single_child keyed status on summary presence alone and
assumed completed=False meant iteration-budget exhaustion, so such a child
was reported status=completed + exit_reason=max_iterations, rendering the
false '"TRUNCATED: hit max_iterations"' banner.
Consult the structured failure fields (failed / error) before falling back to
the summary-presence heuristic, and derive exit_reason honestly: failure ->
'error', interrupted -> 'interrupted', completed -> 'completed', and only
genuine budget exhaustion (completed=False, no failure) -> 'max_iterations'.
The 'truncated' flag stays keyed on exit_reason == 'max_iterations', so it is
now correct automatically. The batch renderer needed no change (the error
field is already plumbed into the result entry for the parent).
Closes#97655
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.
- delegate_tool: result.failed now forces status 'failed' (with the
error carried on the entry); new shared format_subagent_failure_line()
renders one clean human-readable line (traceback -> exception message,
length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
terminal failure status deliver that line via _deliver_platform_notice
BEFORE all progress-queue gates; tool_progress_callback is now always
attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
Every participant gateway can now keep a durable copy of a hosted room's
ordered log and continue the room when its authority host is gone:
- gateway/hosted_room_replicas.py: replica store in root state.db.
ingest_page() persists authority-stamped groups.log pages idempotently,
refusing sequence gaps and authority-epoch regressions. promote_replica()
continues the room locally at epoch+1 with a lineage-proving
authority.claimed event; the stale owner is fenced everywhere the claim
replicates. demote_room() lets a returning stale authority fence itself
(authority.lost) upon observing a newer epoch, killing split-brain writes.
- tui_gateway/methods_groups.py: groups.replicate / groups.replica_state /
groups.promote / groups.demote RPC surface. Promotion requires
confirm=true — storage decides HOW takeover is atomic and provable, the
caller (user action now, lease/quorum driver later) decides WHEN it is
safe, matching the boundary blessed on #97681.
Validation: 20 new tests incl. a full failover round-trip (A hosts, B
replicates incrementally, A dies, B promotes with complete history, A
returns demoted and fenced); 69 total across the hosted-rooms area; E2E
with two real gateway stores and real install identities.
On the codex_app_server runtime the model's real working context is the
app-server's server-side thread: CodexAppServerSession is constructed with
no history and each turn submits only the new user message
(agent/codex_runtime.py), so Hermes' transcript is a mirror that is never
replayed into a thread. Every out-of-turn compression call site (gateway
session hygiene, gateway /compress) built a DETACHED agent whose
_codex_session was None, so the codex route bailed at its "no active codex
thread" guard and returned the transcript unchanged ("compressed 150 ->
150 msgs") — and hygiene's finally-clause then evicted the cached live
agent, destroying the only real context: the next turn spawned an empty
thread while Hermes still mirrored a full history.
Fix, per the documented compression.codex_app_server_auto contract:
* Session hygiene now routes codex_app_server sessions to
run_codex_hygiene_compaction(): in 'hermes' mode it compacts the LIVE
cached agent's thread via thread/compact/start (through the existing
codex route in _compress_context) and KEEPS that agent cached; 'native'
and 'off' skip cleanly with no eviction and no local fallback. A wedged
compaction records the persistent failure cooldown; success resets the
hygiene failure streak.
* Gateway /compress detects the codex_app_server runtime before building
a temporary compression agent and compacts the live thread with
force=True instead (a manual compress is an explicit user decision in
every mode). No live thread -> honest "nothing to compact" reply
instead of a mirror rewrite plus eviction.
* No mode ever runs the local transcript compressor on this runtime:
rewriting the mirror cannot shrink the thread, so the #73715-style
local fallback (including its force=True leak into native/off) is
deliberately not adopted.
Diagnosis of the mode-gate/no-thread deadlock builds on PR #73715.
Closes#73503
Co-authored-by: webtecnica <webtecnica@gmail.com>
PR #98628 removed _build_chunk_digests, so the two lean chunk-digest
cancellation tests reintroduced by the #97512 cherry-pick target a
deleted mechanism — removed. The #96775 stall-interrupt assertions now
match the stall_interrupted marker inside the strategy/kind-stamped
durable error instead of assuming it is the prefix.
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
Sabotage-verified regression tests: bounded-grace worker teardown on
ceiling (cooperative join + uninterruptible orphan with retained
lease), durable strategy/kind-stamped backoff that survives a simulated
gateway restart against a real temp SessionDB, success clearing the
backoff, superseded-attempt late results discarded, and the
transient-block signal (type-pinned against MagicMock agents).
Pin both AuxiliaryExplicitCancellation and commit-fence cancellation, keep early /stop cooldown-neutral, merge with a longer live deadline, and prove force=/compress still bypasses the automatic brake.
Follow-ups on top of the salvaged #97712 foundation:
- read_events() pages now include the room's authority stamp
(authority.gateway_id + authority.epoch), so a replicating participant
can persist lineage with every page and a future takeover layer can
fence stale authorities from replayed state alone.
- New regression test proves the replay page bound counts UTF-8 BYTES,
not characters: sabotaging LENGTH(CAST(.. AS BLOB)) back to
LENGTH(TEXT) and dropping the .encode('utf-8') guard fails the test;
the pre-existing multibyte test passed under that sabotage.
- _raise_room_not_found typed NoReturn so narrowing survives closures.
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).
Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):
1. skill_view(name, file_path='references') returned a raw
'[Errno 21] Is a directory' OS error. The local-skill branch gated
on target_file.exists(); a directory passes exists(), fell through
to read_text(), and raised. The plugin-skill sibling branch already
gated on is_file() — this aligns the local branch so a directory
request gets the same helpful not-found payload with
available_files listing instead of an OS error.
2. skill_manage rejected categorized names ('category/skill-name')
with 'not found in active profile'. _find_skill matched only the
bare directory name, while skill_view's own ambiguity hint tells
the caller to use exactly the categorized form — every call that
followed the hint failed. _find_skill now also matches the full
relative path of the skill dir, giving skill_manage resolution
parity with skill_view across edit/patch/delete/write_file/
remove_file.
Both fixes are covered by regression tests that fail on main.
test_explicit_blank_masks_leaked_cron_env_for_gateway_classification
used platform=api_server as an arbitrary gateway platform; api_server
is now intentionally excluded from gateway approval contexts
(unattended class). Switch to telegram — the test's subject is the
blank-cron-ContextVar masking, not platform policy.
Webhook sessions trigger the gateway approval branch because
HERMES_SESSION_PLATFORM is set, but the webhook adapter has no
send_exec_approval and no way to receive /approve replies. This
blocks the session for the full approval timeout (60-300 s) with
no human who can resolve it.
Fix: _is_gateway_approval_context() now returns False when the
session platform is 'webhook', falling through to the non-interactive
path (auto-approve with warning, or deny if cron).
Regression tests added for webhook, non-webhook gateway, cron, and
no-platform scenarios.
Follow-up to imsuperseller's #96740 (cherry-picked as the previous commit).
Widens _CACHE_BUSTING_CONFIG_KEYS with the other construction-baked
compaction-routing settings that had the same stale-cache shape:
compression.in_place, checkpoint_required, micro_compact,
micro_compact_every_n_turns, micro_compact_defrag_threshold_tokens.
Without these, a messaging-gateway session cached before a config edit
keeps the old compaction routing forever.
Not added (reported instead): abort_on_summary_failure, max_attempts,
protect_first_n, codex_gpt55_autoraise_notice, idle_compact_after_seconds
— behavior-tuning rather than routing, left for a deliberate pass.
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.
But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.
Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
Folded from PR #98345 (@ericmaddox): the one scenario its suite covered
that #91557's did not — an assistant message between the image-only user
message and the checkpoint, asserting post-prune ordering.
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.
Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.
Refs #90976 and #91477.
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).
Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.
Fixes#96995
* refactor(skills): shipped-set slim — 15 skills to optional, github six-way merge, pdf absorbs OCR+nano-pdf, channel-gated teams pipeline
Maintainer-directed shipped-skills curation (skills index 1,900 -> ~1,400
tok/call on desktop; every session pays the index, so this is a per-call
diet on all installs):
- optional-skills moves (installable via skills hub, history preserved):
creative comfyui/ascii-art/excalidraw/pretext/sketch/touchdesigner-mcp;
ALL of mlops (huggingface-hub, llama-cpp, serving-llms-vllm,
weights-and-biases, evaluating-llms-harness — subcategory structure
kept); research-paper-writing (55 supporting files, 17.3K-tok load);
openhue; blogwatcher (first taught the cronjob monitor-field watch
pattern + web_extract instead of pre-cron manual workflows)
- DELETED session-librarian (Aug-12 'inspired by Perplexity Computer'
port, never maintainer-intended; session_search covers discovery)
- github: six skills (auth, issues, pr-workflow, issue-to-pr,
code-review, repo-management) merged into ONE software-development/
github skill — routing body + complete per-workflow references;
benbarclay authorship credited; codebase-inspection rides along;
discipline pins from test_github_issue_to_pr_skill.py preserved
against the reference body in the new test_github_skill.py
- pdf absorbs ocr-and-documents + nano-pdf as references/ + scripts
(extract_pymupdf, extract_marker converted to the argparse house
standard its contract test enforces)
- NEW session_platforms frontmatter gate (metadata.hermes): hides a
skill from the index on gateway channels it is not for; fail-open on
unknown platform; teams-meeting-pipeline gated to [teams, cron]
- blocked-page-recovery: research -> new web category; trigger-first
description ('Use when a fetch fails: 403/429, paywall, WAF, bot
wall.') so the model actually reaches for it on blocked fetches
- docs regenerated via generate-skill-docs.py (195 pages); related_skills
swept repo-wide; tests: 1672 passed (2 openclaw failures pre-existing
on clean main, Windows-local)
* chore: ignore .skills_prompt_snapshot.json (local index cache, accidentally committed)
Bundled by #83063 into the always-shipped set, but it serves only
multi-agent kanban campaigns — too niche for every install's skills
index (~20 tok/call for all users). Moved to the optional catalog
(installable via /skills search + hub, official source) and renamed so
the trigger is legible at a glance: 'merge-reconciler' read like a
generic git helper; 'agent-merge-conflict-arbiter' says who it is for.
- skills/autonomous-ai-agents/merge-reconciler -> optional-skills/autonomous-ai-agents/agent-merge-conflict-arbiter
- frontmatter name + description updated (description within the 60-char hardline its own contract test enforces)
- contract test moved/renamed, 9/9 green
- kanban docs (en + zh-Hans) repointed; zero merge-reconciler refs remain