PyPI · nltk
NLTK: Pl196xCorpusReader has quadratic ReDoS on malformed TEI blocks
Pl196xCorpusReader still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.
nltk.corpus.reader.pl196x.TEICorpusView.read_block and Pl196xCorpusReader public methods3.9.4 and current source v3.10.0-rc2 both reproduced..*? whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.The parser uses regexes for paragraphs, sentences, and word tags across the whole <text> block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed <p> tags doubled, through normal public calls such as words() and tagged_words().
Preconditions
Steps
<text> block that contains many opening tags and no matching closing tags.Pl196xCorpusReader on that corpus.words() or tagged_words() and measure elapsed time as the malformed tag count doubles.Minimal reproducible excerpt
size=1000 0.014s
size=2000 0.057s
size=4000 0.231s
size=8000 0.927s
A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.
Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.
Is your project exposed to this? Stateward checks every dependency on every pull request and flags it only if your code actually reaches it.
Check my repoSources: CISA KEV (public domain), OSV.dev & GitHub Advisory Database (CC-BY-4.0), FIRST EPSS, NVD/CWE (public domain). Served live from the Stateward advisory database.