Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2025-48943: vLLM denial of service through invalid regex in structured output
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48943CVE-2025-48943 affects vLLM, an inference and serving engine for large language models, in versions 0.8.0 up to but excluding 0.9.0. An invalid regex supplied while using structured output triggers a ReDoS that crashes the vLLM server, a denial of service. The flaw is similar to GHSA-6qc9-v4r8-22xg/CVE-2025-48942, but applies to regex rather than a JSON schema.
Fix: Version 0.9.0 fixes the issue.
NVD/CVE DatabaseCVE-2025-48942: vLLM server crash via invalid json_schema in /v1/completions Guided Param
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48942CVE-2025-48942 affects vLLM, an inference and serving engine for large language models, in versions 0.8.0 up to but excluding 0.9.0. Sending an invalid json_schema as a Guided Param to the /v1/completions API crashes the vllm server. It is similar to GHSA-9hcf-v7m4-6m2j/CVE-2025-48943, but applies to regex instead of a JSON schema.
Fix: Version 0.9.0 fixes the issue.
NVD/CVE DatabaseCVE-2025-48887: vLLM regular expression denial of service in pythonic tool parser
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48887vLLM, an inference and serving engine for large language models (LLMs), contains a Regular Expression Denial of Service (ReDoS) flaw in `vllm/entrypoints/openai/tool_parsers/pythonic_tool_parser.py` in versions 0.6.4 up to but excluding 0.9.0. The tool call detection regex uses multiple nested quantifiers, optional groups and inner repetitions, so an attacker can trigger catastrophic backtracking, severely degrading performance or making the service unavailable.
Fix: Version 0.9.0 contains a patch for the issue.
NVD/CVE DatabaseCVE-2025-46722: vLLM hash collisions in multimodal image hashing via raw pixel bytes
May 29, 2025MediumVulnerabilitySecurityCVE-2025-46722vLLM versions from 0.7.0 before 0.9.0 contain a flaw in the MultiModalHasher class in vllm/multimodal/hasher.py. Its image hashing method serializes PIL.Image.Image objects using only obj.tobytes(), which omits metadata such as width, height and mode, so images of different dimensions with identical pixel bytes can produce the same hash. This can cause hash collisions, incorrect cache hits, and possible data leakage.
Fix: Fixed in 0.9.0.
NVD/CVE DatabaseCVE-2025-46570: vLLM timing side channel in prefix cache matching during prefill
May 29, 2025LowVulnerabilitySecurityCVE-2025-46570CVE-2025-46570 affects vLLM, an inference and serving engine for large language models, before version 0.9.0. When a new prompt is processed and the PageAttention mechanism finds a matching prefix chunk, prefill is faster, which shows up in TTFT (Time to First Token). These timing differences are large enough to be recognized and exploited.
Fix: Fixed in version 0.9.0.
NVD/CVE DatabaseCVE-2025-47277: vLLM PyNcclPipe KV cache transfer exposed on all network interfaces
May 20, 2025CriticalVulnerabilitySecurityCVE-2025-47277vLLM versions 0.6.5 through 0.8.4 are affected by CVE-2025-47277, but only when the PyNcclPipe KV cache transfer integration is used with the V0 engine. PyTorch's TCPStore listens on all interfaces by default, so the address passed via --kv-ip did not restrict exposure as intended, leaving the PyNcclPipe communication interface reachable beyond the private network.
Fix: Fixed in 0.8.5, which limits the TCPStore socket to the private interface configured via --kv-ip.
NVD/CVE DatabaseCVE-2025-1975: Ollama denial of service through manifest array index validation in /api/pull
May 16, 2025MediumVulnerabilitySecurityCVE-2025-1975CVE-2025-1975 affects the Ollama server version 0.5.11. A malicious user can cause a Denial of Service by customizing manifest content and spoofing a service, which exploits improper validation of array index access (CWE-129) when a model is downloaded via the /api/pull endpoint. The flaw can crash the server.
NVD/CVE DatabaseCVE-2025-30165: vLLM unsafe pickle deserialization in multi-node ZeroMQ communication
May 6, 2025HighVulnerabilitySecurityCVE-2025-30165CVE-2025-30165 affects vLLM's V0 engine in multi-node deployments that use tensor parallelism. Secondary hosts subscribe to the primary host's XPUB socket over ZeroMQ and deserialize received data with pickle, which can be abused to execute code on a remote machine. The flaw is an escalation point if the primary host is compromised, and it can also be reached without primary-host access, for example through ARP cache poisoning. The V1 engine is not affected.
Fix: The maintainers decided not to fix this issue, because V0 has been off by default since v0.8.0 and the fix is fairly invasive. They recommend that users ensure their environment is on a secure network in case this pattern is in use.
NVD/CVE DatabaseCVE-2025-46560: vLLM multimodal tokenizer quadratic time complexity resource exhaustion
Apr 30, 2025MediumVulnerabilitySecurityCVE-2025-46560vLLM versions from 0.8.0 up to but not including 0.8.5 contain a performance vulnerability in the input preprocessing logic of the multimodal tokenizer. Placeholder tokens such as <|audio_|> and <|image_|> are replaced through inefficient list concatenation, giving the algorithm quadratic time complexity (O(n²)). A malicious actor can send specially crafted inputs to trigger resource exhaustion.
Fix: Fixed in 0.8.5.
NVD/CVE DatabaseCVE-2025-32444: vLLM remote code execution through pickle deserialization over ZeroMQ sockets
Apr 30, 2025CriticalVulnerabilitySecurityCVE-2025-32444CVE-2025-32444 affects vLLM versions starting from 0.6.5 and prior to 0.8.5 when the vLLM integration with mooncake is in use. Those versions use pickle-based serialization over unsecured ZeroMQ sockets, which are set to listen on all network interfaces, allowing remote code execution by an attacker who can reach them. vLLM instances without the mooncake integration are not vulnerable.
Fix: Fixed in 0.8.5.
NVD/CVE DatabaseCVE-2025-30202: vLLM denial of service and data exposure via ZeroMQ on multi-node deployment
Apr 30, 2025HighVulnerabilitySecurityCVE-2025-30202CVE-2025-30202 affects vLLM versions from 0.5.2 up to but not including 0.8.5 in multi-node deployments. The primary host binds an XPUB ZeroMQ socket to all interfaces, so any client with network access can connect unless a firewall blocks the port, and receives the internal vLLM state broadcast to secondary hosts. Connecting repeatedly without reading the published data can slow or block the publisher, causing denial of service.
Fix: Fixed in 0.8.5.
NVD/CVE DatabaseCVE-2025-0317: Ollama denial of service via crafted GGUF model file upload
Mar 20, 2025HighVulnerabilitySecurityCVE-2025-0317EPSS: 14.2%CVE-2025-0317 affects ollama/ollama versions <=0.3.14. A malicious user who can upload and create a customized GGUF model file on the Ollama server can trigger a division by zero error in the ggufPadding function. The server then crashes, resulting in a Denial of Service (DoS) attack. The weakness is classified as CWE-369 (Divide By Zero).
NVD/CVE DatabaseCVE-2025-0315: Ollama denial of service through crafted GGUF model file creation
Mar 20, 2025HighVulnerabilitySecurityCVE-2025-0315CVE-2025-0315 affects ollama/ollama versions up to and including 0.3.14. A malicious user can craft a customized GGUF model file, upload it to the Ollama server, and create it, causing the server to allocate unlimited memory. The result is a Denial of Service (DoS) condition, classified under CWE-770 (Allocation of Resources Without Limits or Throttling).
NVD/CVE DatabaseCVE-2025-0312: Ollama denial of service through crafted GGUF model file upload
Mar 20, 2025HighVulnerabilitySecurityCVE-2025-0312CVE-2025-0312 affects ollama/ollama versions <=0.3.14. A malicious user can craft a customized GGUF model file that, when uploaded and created on the Ollama server, triggers an unchecked null pointer dereference (CWE-476). The crash enables a Denial of Service attack over the remote network.
NVD/CVE DatabaseCVE-2024-9053: vllm-project vllm RPC server remote code execution via unsafe deserialization
Mar 20, 2025CriticalVulnerabilitySecurityCVE-2024-9053CVE-2024-9053 affects vllm-project vllm version 0.6.0. In the AsyncEngineRPCServer() RPC server entrypoints, run_server_loop() calls _make_handler_coro(), which passes received messages directly to cloudpickle.loads() without sanitization. The source states this can result in remote code execution by deserializing malicious pickle data.
NVD/CVE DatabaseCVE-2024-8063: Ollama divide by zero via crafted block_count in GGUF model import
Mar 20, 2025HighVulnerabilitySecurityCVE-2024-8063CVE-2024-8063 is a divide by zero flaw (CWE-369) in ollama/ollama version v0.3.3. It is triggered when GGUF models are imported with a crafted type for `block_count` in the Modelfile. The server crashes when it processes such a model, causing a denial of service (DoS).
NVD/CVE DatabaseCVE-2024-12055: Ollama out-of-bounds read via malicious gguf model file causes denial of service
Mar 20, 2025HighVulnerabilitySecurityCVE-2024-12055CVE-2024-12055 affects Ollama versions <=0.3.14. A malicious user can create a customized gguf model file and upload it to the public Ollama server. When the server processes the malicious model, it crashes, causing a Denial of Service (DoS). The root cause is an out-of-bounds read in the gguf.go file (CWE-125).
NVD/CVE DatabaseCVE-2024-11041: vllm-project vllm remote code execution in MessageQueue.dequeue() via pickle
Mar 20, 2025CriticalVulnerabilitySecurityCVE-2024-11041CVE-2024-11041 affects vllm-project vllm version v0.6.2. The MessageQueue.dequeue() API function passes received socket data directly to pickle.loads, which allows remote code execution. An attacker who sends a malicious payload to the MessageQueue can make the victim's machine run arbitrary code.
NVD/CVE DatabaseGHSA-v464-r2r9-www7: Ollama Vulnerable to Denial of Service (DoS) via Crafted GZIP
Mar 20, 2025HighVulnerabilitySecurityCVE-2024-12886CVE-2024-12886 affects the Ollama server in the Go module github.com/ollama/ollama, versions <= 0.3.14. A malicious API server can send a gzip bomb HTTP response, which causes an Out-Of-Memory condition and crashes the Ollama server. The flaw sits in the makeRequestWithRetry and getAuthorizationToken functions, which read the response body with io.ReadAll.
GitHub Advisory DatabaseCVE-2025-29783: vLLM unsafe deserialization over ZMQ/TCP when using Mooncake
Mar 19, 2025CriticalVulnerabilitySecurityCVE-2025-29783CVE-2025-29783 affects vLLM when it is configured to use Mooncake. Unsafe deserialization exposed directly over ZMQ/TCP on all network interfaces lets attackers execute remote code on distributed hosts. The flaw impacts any deployment that uses Mooncake to distribute KV across distributed hosts.
Fix: Fixed in 0.8.0.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.