Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2025-47277: vLLM PyNcclPipe KV cache transfer exposed on all network interfaces
May 20, 2025CriticalVulnerabilitySecurityCVE-2025-47277vLLM versions 0.6.5 through 0.8.4 are affected by CVE-2025-47277, but only when the PyNcclPipe KV cache transfer integration is used with the V0 engine. PyTorch's TCPStore listens on all interfaces by default, so the address passed via --kv-ip did not restrict exposure as intended, leaving the PyNcclPipe communication interface reachable beyond the private network.
Fix: Fixed in 0.8.5, which limits the TCPStore socket to the private interface configured via --kv-ip.
NVD/CVE DatabaseCVE-2025-1975: Ollama denial of service through manifest array index validation in /api/pull
May 16, 2025MediumVulnerabilitySecurityCVE-2025-1975CVE-2025-1975 affects the Ollama server version 0.5.11. A malicious user can cause a Denial of Service by customizing manifest content and spoofing a service, which exploits improper validation of array index access (CWE-129) when a model is downloaded via the /api/pull endpoint. The flaw can crash the server.
NVD/CVE DatabaseCVE-2025-30165: vLLM unsafe pickle deserialization in multi-node ZeroMQ communication
May 6, 2025HighVulnerabilitySecurityCVE-2025-30165CVE-2025-30165 affects vLLM's V0 engine in multi-node deployments that use tensor parallelism. Secondary hosts subscribe to the primary host's XPUB socket over ZeroMQ and deserialize received data with pickle, which can be abused to execute code on a remote machine. The flaw is an escalation point if the primary host is compromised, and it can also be reached without primary-host access, for example through ARP cache poisoning. The V1 engine is not affected.
Fix: The maintainers decided not to fix this issue, because V0 has been off by default since v0.8.0 and the fix is fairly invasive. They recommend that users ensure their environment is on a secure network in case this pattern is in use.
NVD/CVE DatabaseCVE-2025-46560: vLLM multimodal tokenizer quadratic time complexity resource exhaustion
Apr 30, 2025MediumVulnerabilitySecurityCVE-2025-46560vLLM versions from 0.8.0 up to but not including 0.8.5 contain a performance vulnerability in the input preprocessing logic of the multimodal tokenizer. Placeholder tokens such as <|audio_|> and <|image_|> are replaced through inefficient list concatenation, giving the algorithm quadratic time complexity (O(n²)). A malicious actor can send specially crafted inputs to trigger resource exhaustion.
Fix: Fixed in 0.8.5.
NVD/CVE DatabaseCVE-2025-32444: vLLM remote code execution through pickle deserialization over ZeroMQ sockets
Apr 30, 2025CriticalVulnerabilitySecurityCVE-2025-32444CVE-2025-32444 affects vLLM versions starting from 0.6.5 and prior to 0.8.5 when the vLLM integration with mooncake is in use. Those versions use pickle-based serialization over unsecured ZeroMQ sockets, which are set to listen on all network interfaces, allowing remote code execution by an attacker who can reach them. vLLM instances without the mooncake integration are not vulnerable.
Fix: Fixed in 0.8.5.
NVD/CVE DatabaseCVE-2025-30202: vLLM denial of service and data exposure via ZeroMQ on multi-node deployment
Apr 30, 2025HighVulnerabilitySecurityCVE-2025-30202CVE-2025-30202 affects vLLM versions from 0.5.2 up to but not including 0.8.5 in multi-node deployments. The primary host binds an XPUB ZeroMQ socket to all interfaces, so any client with network access can connect unless a firewall blocks the port, and receives the internal vLLM state broadcast to secondary hosts. Connecting repeatedly without reading the published data can slow or block the publisher, causing denial of service.
Fix: Fixed in 0.8.5.
NVD/CVE DatabaseCVE-2025-0317: Ollama denial of service via crafted GGUF model file upload
Mar 20, 2025HighVulnerabilitySecurityCVE-2025-0317EPSS: 14.2%CVE-2025-0317 affects ollama/ollama versions <=0.3.14. A malicious user who can upload and create a customized GGUF model file on the Ollama server can trigger a division by zero error in the ggufPadding function. The server then crashes, resulting in a Denial of Service (DoS) attack. The weakness is classified as CWE-369 (Divide By Zero).
NVD/CVE DatabaseCVE-2025-0315: Ollama denial of service through crafted GGUF model file creation
Mar 20, 2025HighVulnerabilitySecurityCVE-2025-0315CVE-2025-0315 affects ollama/ollama versions up to and including 0.3.14. A malicious user can craft a customized GGUF model file, upload it to the Ollama server, and create it, causing the server to allocate unlimited memory. The result is a Denial of Service (DoS) condition, classified under CWE-770 (Allocation of Resources Without Limits or Throttling).
NVD/CVE DatabaseCVE-2025-0312: Ollama denial of service through crafted GGUF model file upload
Mar 20, 2025HighVulnerabilitySecurityCVE-2025-0312CVE-2025-0312 affects ollama/ollama versions <=0.3.14. A malicious user can craft a customized GGUF model file that, when uploaded and created on the Ollama server, triggers an unchecked null pointer dereference (CWE-476). The crash enables a Denial of Service attack over the remote network.
NVD/CVE DatabaseCVE-2024-9053: vllm-project vllm RPC server remote code execution via unsafe deserialization
Mar 20, 2025CriticalVulnerabilitySecurityCVE-2024-9053CVE-2024-9053 affects vllm-project vllm version 0.6.0. In the AsyncEngineRPCServer() RPC server entrypoints, run_server_loop() calls _make_handler_coro(), which passes received messages directly to cloudpickle.loads() without sanitization. The source states this can result in remote code execution by deserializing malicious pickle data.
NVD/CVE DatabaseCVE-2024-8063: Ollama divide by zero via crafted block_count in GGUF model import
Mar 20, 2025HighVulnerabilitySecurityCVE-2024-8063CVE-2024-8063 is a divide by zero flaw (CWE-369) in ollama/ollama version v0.3.3. It is triggered when GGUF models are imported with a crafted type for `block_count` in the Modelfile. The server crashes when it processes such a model, causing a denial of service (DoS).
NVD/CVE DatabaseCVE-2024-12055: Ollama out-of-bounds read via malicious gguf model file causes denial of service
Mar 20, 2025HighVulnerabilitySecurityCVE-2024-12055CVE-2024-12055 affects Ollama versions <=0.3.14. A malicious user can create a customized gguf model file and upload it to the public Ollama server. When the server processes the malicious model, it crashes, causing a Denial of Service (DoS). The root cause is an out-of-bounds read in the gguf.go file (CWE-125).
NVD/CVE DatabaseCVE-2024-11041: vllm-project vllm remote code execution in MessageQueue.dequeue() via pickle
Mar 20, 2025CriticalVulnerabilitySecurityCVE-2024-11041CVE-2024-11041 affects vllm-project vllm version v0.6.2. The MessageQueue.dequeue() API function passes received socket data directly to pickle.loads, which allows remote code execution. An attacker who sends a malicious payload to the MessageQueue can make the victim's machine run arbitrary code.
NVD/CVE DatabaseGHSA-v464-r2r9-www7: Ollama Vulnerable to Denial of Service (DoS) via Crafted GZIP
Mar 20, 2025HighVulnerabilitySecurityCVE-2024-12886CVE-2024-12886 affects the Ollama server in the Go module github.com/ollama/ollama, versions <= 0.3.14. A malicious API server can send a gzip bomb HTTP response, which causes an Out-Of-Memory condition and crashes the Ollama server. The flaw sits in the makeRequestWithRetry and getAuthorizationToken functions, which read the response body with io.ReadAll.
GitHub Advisory DatabaseCVE-2025-29783: vLLM unsafe deserialization over ZMQ/TCP when using Mooncake
Mar 19, 2025CriticalVulnerabilitySecurityCVE-2025-29783CVE-2025-29783 affects vLLM when it is configured to use Mooncake. Unsafe deserialization exposed directly over ZMQ/TCP on all network interfaces lets attackers execute remote code on distributed hosts. The flaw impacts any deployment that uses Mooncake to distribute KV across distributed hosts.
Fix: Fixed in 0.8.0.
NVD/CVE DatabaseCVE-2025-29770: vLLM unbounded filesystem cache growth via outlines guided decoding
Mar 19, 2025MediumVulnerabilitySecurityCVE-2025-29770CVE-2025-29770 affects vLLM's structured output path, which uses the outlines library's on-disk grammar cache, on by default. The unconditional use of that cache in vllm/model_executor/guided_decoding/outlines_logits_processors.py lets a user send many short requests with unique schemas, each adding a cache entry. Because outlines is also reachable per request through the OpenAI-compatible API server, this can exhaust filesystem space and cause a Denial of Service. The issue applies only to the V0 engine.
Fix: Fixed in 0.8.0.
NVD/CVE DatabaseCVE-2025-1953: vLLM AIBrix prefix caching uses insufficiently random values
Mar 4, 2025LowVulnerabilitySecurityCVE-2025-1953CVE-2025-1953 is a vulnerability in vLLM AIBrix 0.2.0, classified as problematic, in the Prefix Caching component. The flaw lies in pkg/plugins/gateway/prefixcacheindexer/hash.go, where manipulation leads to insufficiently random values (CWE-330). The source rates attack complexity as high and exploitation as difficult.
Fix: Upgrade to version 0.3.0, which the source says is able to address this issue.
NVD/CVE DatabaseCVE-2024-53880: NVIDIA Triton Inference Server integer overflow in model loading API
Feb 12, 2025MediumVulnerabilitySecurityCVE-2024-53880NVIDIA Triton Inference Server contains an integer overflow or wraparound vulnerability (CWE-190) in its model loading API. A user who loads a model with an extra-large file size can overflow an internal variable, which may lead to denial of service.
NVD/CVE DatabaseCVE-2025-25183: vLLM prefix caching hash collisions can reuse cache from other prompts
Feb 7, 2025LowVulnerabilitySecurityCVE-2025-25183CVE-2025-25183 affects vLLM, a high-throughput inference and serving engine for LLMs. Maliciously constructed statements can cause hash collisions in prefix caching, which relies on Python's built-in hash() function, so cached content generated from different prompts can be reused and interfere with later responses. Because hash(None) became a predictable constant in Python 3.12, an attacker who knows the prompts in use could populate the cache with a colliding prompt.
Fix: Fixed in version 0.7.2; all users are advised to upgrade. There are no known workarounds for this vulnerability.
NVD/CVE DatabaseCVE-2025-24357: vLLM deserialization of untrusted model weights via torch.load
Jan 27, 2025HighVulnerabilitySecurityCVE-2025-24357CVE-2025-24357 affects vLLM, a library for LLM inference and serving. In vllm/model_executor/weight_utils.py, the hf_model_weights_iterator function loads model checkpoints downloaded from Hugging Face using torch.load with weights_only defaulting to False. When torch.load loads malicious pickle data, it executes arbitrary code during unpickling. The weakness is classified as CWE-502, Deserialization of Untrusted Data.
Fix: This vulnerability is fixed in v0.7.0.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.