fix(tools): make json_parse tolerate UTF-8 BOM (salvage #57870)
json_parse used json.loads(strict=False), which relaxes control characters but rejects a leading UTF-8 BOM (U+FEFF). Windows CLI tools and some files prepend a BOM, causing JSONDecodeError on otherwise valid JSON output. Strip a leading BOM before calling json.loads when the input is a string with a U+FEFF prefix. Original PR by @woxinwuhen713-bit (#57870). Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -473,9 +473,12 @@ _COMMON_HELPERS = '''\
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def json_parse(text: str):
|
||||
"""Parse JSON tolerant of control characters (strict=False).
|
||||
"""Parse JSON tolerant of control characters and UTF-8 BOM (strict=False).
|
||||
Use this instead of json.loads() when parsing output from terminal()
|
||||
or web_extract() that may contain raw tabs/newlines in strings."""
|
||||
or web_extract() that may contain raw tabs/newlines in strings,
|
||||
or from tools/files that prepend a UTF-8 BOM (salvage #57870, credit @woxinwuhen713-bit)."""
|
||||
if isinstance(text, str) and text.startswith(""):
|
||||
text = text[1:]
|
||||
return json.loads(text, strict=False)
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user