2 Commits

Author SHA1 Message Date
Teknium 341d5aebc6 Port from MoonshotAI/kimi-code#2647: read UTF-16 text files by transcoding to UTF-8
UTF-16 text files (Windows Notepad .txt, PowerShell > redirects) were
refused as binary: the terminal env decodes stdout as UTF-8 with
errors=replace, so their content arrived mangled with U+FFFD and
tripped the binary guard.

ShellFileOperations.read_file now probes the raw bytes via the
backend's Python when the binary guard fires: a BOM or the zero-byte
parity heuristic (derived from VS Code's encoding sniffer, tolerant of
mixed Latin/CJK content) identifies UTF-16 LE/BE, and the file is
transcoded to UTF-8 with CRLF normalized and the BOM stripped. Real
binaries (zeros at both parities), binary extensions, files over
10 MiB, and legacy 8-bit encodings (GBK, Big5) still refuse — a wrong
silent guess is worse than a clear refusal. Works on every shell
backend (local/docker/ssh) since the probe runs via python3 -c.

Tests run against a real LocalEnvironment (E2E, no mocks); sabotage
run confirmed 6/9 fail without the fix.
2026-08-16 22:08:28 -07:00
Teknium 93be7f0117 test(file-ops): end-to-end regression suite for the UTF-8-flagged-as-binary class
Real-backend coverage for the dupe-swarm cluster: truncated-CJK and
Cyrillic sample cuts, utf-8-sig BOM, genuine binaries (PNG/ELF magic,
NUL-in-text), empty files, UTF-16 both endians (read-only pin), plus the
sibling sites — read_file_raw (patch/V4A, #80221), patch_replace, and
content search (#80308).

Closes #76886 #77047 #77842 #80221 #80251 #80308 #80922
2026-08-08 12:34:06 -07:00