npm · @xmldom/xmldom
xmldom: Parser silently accepts a not-well-formed end tag whose name is followed by a line break and trailing content
xmldom's parser silently accepts a not-well-formed end tag whose valid name is followed by
trailing content — e.g. </a⏎junk>. The element is closed, the trailing content is discarded, and no
error is reported, even though the XML end-tag production allows only optional whitespace after the
name and both Chromium and Firefox reject such input as application/xml. An application that relies
on xmldom to reject not-well-formed input therefore receives a false "valid" result for a document the
specification and browsers consider malformed.
Across every affected version, an end tag whose valid Name is followed by trailing content before
> is silently accepted: the element is closed, the residue is dropped, and no error is reported. How
much leaks differs by line (see Affected Versions), but the observable weakness is the same.
On the current (0.9.x) line, the parser validates the end-tag name against the XML ETag production
with an anchored regular expression (^ QName S? $). That expression is compiled with the m
(multiline) flag by a shared builder, so $ matches at an interior line terminator: a valid name on
the first line satisfies the anchored production and any content after the line break escapes the
check. On 0.9.x the whitespace-separated variant (</a junk>) is already rejected; only the
line-terminator variant leaks. Older lines have no anchored end-tag validator at all, so they accept
both the line-terminator and the whitespace variant.
This is not content injection — the trailing content is dropped, and the resulting DOM is a normal
single-root document (<a/>). The security-relevant property is the silent acceptance of
not-well-formed input: xmldom's parse result disagrees with the specification and with browser XML
parsers, so any control that treats "xmldom parsed it without error" as "well-formed" is bypassed.
On the 0.9.x line, where the line-terminator variant specifically leaks:
m flag.^…$ under m are line anchors, not string anchors.^ QName S? $ is therefore satisfied by the first line alone, so
trailing content after a line terminator is neither matched nor rejected — the malformed end tag is
accepted and the residue silently discarded.The triggering line terminators are the ECMAScript LineTerminator set: U+000A, U+000D, U+2028, U+2029.
U+2028 and U+2029 are not XML whitespace, so they are non-conforming trailing content that nonetheless
leaks because the JavaScript $ anchor treats them as line boundaries under m.
Both maintained versions are affected and fixed:
0.9.x (>= 0.9.0, <= 0.9.11) → 0.9.12: the anchored end-tag validator's m flag leaks
the line-terminator variant. The whitespace variant (</a junk>) is already rejected on this line.0.8.x (>= 0.8.0, <= 0.8.14) → 0.8.15: no anchored end-tag validator at all — both the
line-terminator and the whitespace variant are silently accepted; the fix adds a residue check.release-0.7.x (<= 0.7.13) and the unscoped xmldom package (range *, last release 0.6.0) are
affected but will not be patched — they are end-of-life / unmaintained. They share the older-line
behavior (both variants silently accepted, residue dropped).
const { DOMParser, XMLSerializer } = require('@xmldom/xmldom');
const doc = new DOMParser().parseFromString('<a></a\njunk>', 'text/xml'); // no error
console.log(new XMLSerializer().serializeToString(doc));
// Observed: <a/> — the malformed end tag is accepted, "junk" silently discarded, no error reported.
// Expected (per XML spec / Chromium / Firefox): a parse error — the input is not well-formed.
The parser now reports a not-well-formed end tag whose valid name is followed by trailing content, on both 0.9.12 and 0.8.15 — previously it was accepted silently — and parsing recovers to the byte-identical DOM as before.
On 0.9.12 the anchored end-tag validator is corrected so a line break followed by trailing content no longer satisfies it, reported as a recoverable error when parsing as XML and a warning when parsing as HTML.
On 0.8.15, which previously performed no end-tag residue validation, an equivalent residue check is added, reported as a recoverable error in both XML and HTML.
Well-formed documents are unaffected.
Because the report is recoverable, the fix is non-breaking: no previously-parsed document begins to
throw and no serialized output changes. Consumers that want strict rejection can escalate the reported
error to a fatal one via the parser's error handler (onError in 0.9.12, errorHandler in
0.8.15). The two versions also differ in the whitespace-separated variant (</a junk>): 0.9.12
already rejected it with a fatal error in XML and continues to, while 0.8.15 — which validated
neither variant — now emits the same recoverable report for both the whitespace and line-break
variants.
By default the parser still recovers (it does not reject the document); the reported condition is a
recoverable error/warning, not a fatal error, to avoid changing the parsed output in a patch
release. Converting this and the other not-well-formed-acceptance cases to a consistent fatalError —
including the broader end-tag leniency on the older versions — is deferred to the next breaking release,
tracked at xmldom/xmldom#1074.
Is your project exposed to this? Stateward checks every dependency on every pull request and flags it only if your code actually reaches it.
Check my repoSources: CISA KEV (public domain), OSV.dev & GitHub Advisory Database (CC-BY-4.0), FIRST EPSS, NVD/CWE (public domain). Served live from the Stateward advisory database.