fix(tools): make json_parse tolerate UTF-8 BOM (salvage #57870)

json_parse used json.loads(strict=False), which relaxes control
characters but rejects a leading UTF-8 BOM (U+FEFF). Windows CLI
tools and some files prepend a BOM, causing JSONDecodeError on
otherwise valid JSON output.

Strip a leading BOM before calling json.loads when the input is
a string with a U+FEFF prefix.

Original PR by @woxinwuhen713-bit (#57870).

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
nankingjing
2026-07-05 01:55:34 +08:00
committed by Teknium
parent 2e227d74f4
commit f2feb6f37d
+5 -2
View File
@@ -473,9 +473,12 @@ _COMMON_HELPERS = '''\
# ---------------------------------------------------------------------------
def json_parse(text: str):
"""Parse JSON tolerant of control characters (strict=False).
"""Parse JSON tolerant of control characters and UTF-8 BOM (strict=False).
Use this instead of json.loads() when parsing output from terminal()
or web_extract() that may contain raw tabs/newlines in strings."""
or web_extract() that may contain raw tabs/newlines in strings,
or from tools/files that prepend a UTF-8 BOM (salvage #57870, credit @woxinwuhen713-bit)."""
if isinstance(text, str) and text.startswith(""):
text = text[1:]
return json.loads(text, strict=False)