Description
XMLParser.parse() declares support for string | Uint8Array, but a plain Uint8Array containing valid UTF-8 XML produces an empty object. With validation enabled, it raises an error about the first digit of the byte array. The same XML succeeds as a string or a Node.js Buffer.
Input / code
import { XMLParser } from 'fast-xml-parser';
const parser = new XMLParser();
const xml = '<root>hello</root>';
const bytes = new TextEncoder().encode(xml);
console.log(parser.parse(xml)); // { root: 'hello' }
console.log(parser.parse(bytes)); // {}
console.log(parser.parse(bytes, true)); // Error: char '6' is not expected.:1:1
Expected output
Both byte-array calls should return { root: 'hello' }, matching the string input and the public TypeScript overloads.
Cause and proposed fix
The non-string input branch calls .toString(). A plain Uint8Array converts to a comma-separated list of byte values; a Node Buffer decodes UTF-8. Decode plain Uint8Array inputs as UTF-8 while retaining the existing Buffer conversion. Decoding the view also respects its byte offset and length.
The related type declaration change was discussed in #764. The earlier large-file/streaming request in #347 concerns a separate use case; this reproduction contains a small complete XML document.
Environment and validation
- fast-xml-parser master at
3617550adfb280989f482d662b7e9ece55a32a34, package version 5.11.1.
- Node.js v24.16.0 on Windows 11.
- Five regression cases fail against the original implementation, covering UTF-8 input, a Uint8Array subview, validation, and preserved node order.
- A local fix passes all seven input-type tests and the full suite (331 specs, zero failures, two pending).
Investigated with OpenAI Codex assistance; the reproductions and tests were run locally. A patch with regression coverage is prepared.
Description
XMLParser.parse()declares support forstring | Uint8Array, but a plainUint8Arraycontaining valid UTF-8 XML produces an empty object. With validation enabled, it raises an error about the first digit of the byte array. The same XML succeeds as a string or a Node.js Buffer.Input / code
Expected output
Both byte-array calls should return
{ root: 'hello' }, matching the string input and the public TypeScript overloads.Cause and proposed fix
The non-string input branch calls
.toString(). A plain Uint8Array converts to a comma-separated list of byte values; a Node Buffer decodes UTF-8. Decode plain Uint8Array inputs as UTF-8 while retaining the existing Buffer conversion. Decoding the view also respects its byte offset and length.The related type declaration change was discussed in #764. The earlier large-file/streaming request in #347 concerns a separate use case; this reproduction contains a small complete XML document.
Environment and validation
3617550adfb280989f482d662b7e9ece55a32a34, package version 5.11.1.Investigated with OpenAI Codex assistance; the reproductions and tests were run locally. A patch with regression coverage is prepared.