PyPI · vllm
vLLM: SSRF + arbitrary local file read in MiMoV2OmniMultiModalProcessor `_fetch_image` and audio loader bypass MediaConnector protections
vllm/transformers_utils/processors/mimo_v2_omni.py — the multimodal processor for MiMoV2OmniForCausalLM — issues requests.get(...) directly on user-supplied image and audio URL strings and Image.open(...) on user-supplied local paths, without the SSRF / allowed_local_media_path checks that vllm.multimodal.utils.MediaConnector was hardened with in GHSA-qh4c-xf7m-gxfc, GHSA-v359-jj2v-j536, and GHSA-pf3h-qjgv-vcpr.
This is the same bug class as those three published advisories, in a code path the patches missed. When a user passes a URL or local-file string through multi_modal_data (e.g. LLM.generate(multi_modal_data={"image": "http://..."})), the processor takes the unsanitized string and dispatches it without any URL-scheme allowlist, network-target allowlist, size cap, or local-path allowlist.
File: vllm/transformers_utils/processors/mimo_v2_omni.py (current main)
Sink 1 — image SSRF + local-file read (_fetch_image, lines 231–249):
def _fetch_image(src: Any) -> Image.Image:
if isinstance(src, Image.Image):
return _to_rgb(src)
if isinstance(src, bytes):
return _to_rgb(copy.deepcopy(Image.open(BytesIO(src))))
if isinstance(src, str):
if src.startswith(("http://", "https://")):
r = requests.get(src, timeout=30) # SSRF: no allowlist, follows redirects
r.raise_for_status()
return _to_rgb(copy.deepcopy(Image.open(BytesIO(r.content))))
if src.startswith("file://"):
return _to_rgb(Image.open(src[7:])) # arbitrary local file read
if src.startswith("data:image"):
...
return _to_rgb(Image.open(src)) # fallback also opens local files
raise ValueError(f"Unrecognized image source: {type(src)}")
Sink 2 — audio SSRF (around line 471):
elif audio.startswith(("http://", "https://")):
r = requests.get(audio, timeout=30) # SSRF: same pattern
r.raise_for_status()
file_obj = io.BytesIO(r.content)
Reachability. _fetch_image is invoked from MiMoVLProcessor.process_image:
def process_image(self, image: ImageInput) -> torch.Tensor:
kw = self._resolve_img_kw(image)
src = image.image
if isinstance(src, (str, bytes)):
src = _fetch_image(src)
...
MiMoVLProcessor is wrapped by MiMoV2OmniMultiModalProcessor and registered for the MiMoV2OmniForCausalLM model architecture (vllm/model_executor/models/mimo_v2_omni.py:1169). Whenever a user passes a string into multi_modal_data["image"] (or ["audio"]) for this model, the unsanitized URL/path reaches the sink.
Comparison to the recent fixes. The remediation pattern adopted in the three earlier advisories was to route every external resource fetch through MediaConnector, which checks allowed_local_media_path and applies SSRF protection before issuing the network request. chat_utils.py (lines 838, 902, 924, 963, 1053, 1081) already uses self._connector.fetch_image / fetch_audio / fetch_video. The model processor in mimo_v2_omni.py was added later and skipped the connector — it calls requests.get and Image.open directly. Result: the public OpenAI chat-completion path is protected, but library use (LLM.generate(multi_modal_data=...)), batch processing, and any other path that lets a string reach the processor receive no protection.
requests.get follows redirects and accepts any URL. An attacker who controls a multi_modal_data value can:http://169.254.169.254/latest/meta-data/iam/security-credentials/),http://127.0.0.1:<port>, http://10.x.y.z),Image.open.file://path (line 242) and the unguarded fallback Image.open(src) (line 248). Any file readable by the vLLM process is reachable through the model pipeline; with suitable formats this exposes /etc/passwd, ~/.aws/credentials, etc.Replace direct requests.get and bare Image.open paths with MediaConnector.fetch_image / fetch_audio_async (or pass the inputs through MediaConnector before they reach the processor):
# vllm/transformers_utils/processors/mimo_v2_omni.py
from vllm.multimodal.utils import MediaConnector
_connector = MediaConnector()
def _fetch_image(src):
if isinstance(src, Image.Image):
return _to_rgb(src)
if isinstance(src, bytes):
return _to_rgb(copy.deepcopy(Image.open(BytesIO(src))))
if isinstance(src, str):
return _to_rgb(_connector.fetch_image(src)) # delegates to the hardened path
raise ValueError(f"Unrecognized image source: {type(src)}")
Same change for the audio loader at line 471. This re-uses the SSRF allowlist, allowed_local_media_path policy, and size caps that the previous patches added.
Alternative: forbid str src from reaching the processor and require all multi-modal pre-processing to go through chat_utils.py / MediaConnector before hitting the model. Larger surface change, but completes the architectural fix.
Static review on vllm@main (HEAD as of 2026-04-30) — found by triaging the file list against the three recent SSRF advisories: the mimo_v2_omni.py processor, added after those fixes, reintroduced the same bypass class.
Ievgen Bondarenko — sactransport2000@gmail.com — GitHub @ibondarenko1
Is your project exposed to this? Stateward checks every dependency on every pull request and flags it only if your code actually reaches it.
Check my repoSources: CISA KEV (public domain), OSV.dev & GitHub Advisory Database (CC-BY-4.0), FIRST EPSS, NVD/CWE (public domain). Served live from the Stateward advisory database.