CVE-2026-105753: vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU
Summary
vLLM (a system for running large language models) has a bug in versions before 0.28.0 where its multimodal cache (a storage system that keeps frequently used media files) can store media in one part of the system but not another, causing crashes when that media is reused later. When a second request tries to use the same cached media, the system fails with an error message and becomes unavailable.
Solution / Mitigation
Update to vLLM version 0.28.0 or later, where this issue is fixed.
Vulnerability Details
6.5(medium)
EPSS: 0.0%
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
network
low
low
none
October 5, 2026
Classification
Taxonomy References
Affected Vendors
Related Issues
CVE-2026-47482: NVIDIA Triton Inference Server for Linux contains a vulnerability where an attacker can cause missing release of memory
CVE-2022-29200: TensorFlow is an open source platform for machine learning. Prior to versions 2.9.0, 2.8.1, 2.7.2, and 2.6.4, the implem
Original source: https://nvd.nist.gov/vuln/detail/CVE-2026-105753
First tracked: October 5, 2026 at 08:07 PM
Classified by LLM (prompt v3) · confidence: 92%