CVE-2026-100651: vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1
Summary
vLLM versions before 0.29.0 have a vulnerability where the /inference/v1/generate endpoint doesn't properly check if decoder prompts (the input text converted to tokens) are too long for the model when processing multimodal requests (requests with images, audio, or other non-text data). Certain multimodal processors skip this length validation entirely, allowing attackers to send excessively long token sequences that crash the system and cause a denial of service (making the service unavailable).
Solution / Mitigation
Fixed in vLLM version 0.29.0.
Vulnerability Details
6.5(medium)
EPSS: 0.0%
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
network
low
low
none
September 26, 2026
Classification
Affected Vendors
Related Issues
CVE-2026-47482: NVIDIA Triton Inference Server for Linux contains a vulnerability where an attacker can cause missing release of memory
CVE-2022-29200: TensorFlow is an open source platform for machine learning. Prior to versions 2.9.0, 2.8.1, 2.7.2, and 2.6.4, the implem
Original source: https://nvd.nist.gov/vuln/detail/CVE-2026-100651
First tracked: September 26, 2026 at 02:08 PM
Classified by LLM (prompt v3) · confidence: 95%