CVE-2026-94627: vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child reque
Summary
A vulnerability in vLLM Mooncake connector (a component that helps distribute AI model processing across systems) through version 0.29.0 fails to properly track ownership of GPU KV cache blocks (temporary storage on graphics processors for speeding up AI responses) when multiple child requests share the same transfer ID. Attackers can exploit this by sending completion requests with multiple prompts, causing orphaned cache blocks to build up until the system restarts, eventually blocking legitimate user requests from running.
Vulnerability Details
7.5(high)
EPSS: 0.0%
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
network
low
none
none
September 21, 2026
Classification
Taxonomy References
Affected Vendors
Related Issues
CVE-2026-47482: NVIDIA Triton Inference Server for Linux contains a vulnerability where an attacker can cause missing release of memory
CVE-2022-29200: TensorFlow is an open source platform for machine learning. Prior to versions 2.9.0, 2.8.1, 2.7.2, and 2.6.4, the implem
Original source: https://nvd.nist.gov/vuln/detail/CVE-2026-94627
First tracked: September 21, 2026 at 08:10 PM
Classified by LLM (prompt v3) · confidence: 92%