CVE-2026-93436: vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode di
Summary
vLLM (a software framework for running large language models) versions up to 0.29.0 has a memory cleanup bug in its decode workers (specialized processors that handle the generation phase of AI inference). Attackers can exploit this by sending requests with max_tokens=0 (asking for zero output tokens), which prevents the system from properly clearing temporary data, eventually consuming all available memory until the worker crashes and restarts.
Vulnerability Details
7.5(high)
EPSS: 0.0%
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
network
low
none
none
September 17, 2026
Classification
Taxonomy References
Affected Vendors
Related Issues
CVE-2026-47482: NVIDIA Triton Inference Server for Linux contains a vulnerability where an attacker can cause missing release of memory
CVE-2022-29200: TensorFlow is an open source platform for machine learning. Prior to versions 2.9.0, 2.8.1, 2.7.2, and 2.6.4, the implem
Original source: https://nvd.nist.gov/vuln/detail/CVE-2026-93436
First tracked: September 17, 2026 at 08:07 PM
Classified by LLM (prompt v3) · confidence: 95%