CVE-2026-71486: vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1
Summary
vLLM (a system for running and serving large language models) had a vulnerability in versions before 0.26.0 where certain API endpoints accepted user-supplied data that was processed before safety checks could limit resource usage. An authenticated attacker (someone with API access) could exploit this to consume excessive CPU and memory or generate oversized responses that bypass size restrictions.
Solution / Mitigation
Update vLLM to version 0.26.0 or later, where this issue is fixed.
Vulnerability Details
4.3(medium)
EPSS: 0.0%
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L
network
low
low
none
August 17, 2026
Classification
Affected Vendors
Related Issues
CVE-2026-47482: NVIDIA Triton Inference Server for Linux contains a vulnerability where an attacker can cause missing release of memory
CVE-2022-29200: TensorFlow is an open source platform for machine learning. Prior to versions 2.9.0, 2.8.1, 2.7.2, and 2.6.4, the implem
Original source: https://nvd.nist.gov/vuln/detail/CVE-2026-71486
First tracked: August 17, 2026 at 08:09 PM
Classified by LLM (prompt v3) · confidence: 95%