Fixes two related native-streaming bubble defects surfaced in production:
1. Duplicate bubble on long turns — when keep-alive already refreshed the
6-min reply window, the Layer-2 clock fallback still declined the finalize
frame and forced a proactive send(), duplicating the message. Skip the
clock fallback while keep-alive is active; intermediate-frame failures are
now fully fire-and-forget (only a failed FINAL frame falls back to send()).
2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
notice) called _reset_segment_state(), clearing the cumulative
_accumulated so the next delta + finalize frame carried only the few
chars accumulated after the reset. Native streaming now skips that
reset (commentary still posts as its own message via send()).
b) The adapter-side _BlockChunker.update() 'only grow' guard silently
dropped any cumulative snapshot shorter than its high-water mark, so
after a baseline reset the leading characters were stranded before
_emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
layer entirely; intermediate frames are pure identity-dedup, matching
the fire-and-forget model.
Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.
Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.
Includes the machinery intrinsic to native streaming on WeCom:
- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
frames update it, a finalize frame closes it; native-streaming adapters are
let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
long-connection mode has no documented edit-rate limit); an adapter-level
frame cap is retained. (Early builds gated frames behind a char throttle;
removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
(control vs normal) plus a per-chat token bucket to stay under WeCom's rate
limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
delivered the moment it is emitted; failures logged, not re-sent; delivery
marked once per turn), per-req_id reply queue with ack tracking, and the
timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
the prompt is the last thing on screen and never traps a lingering bubble;
eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
messages; image+text double-callback merged into one turn.
Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.