Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2026-24146: NVIDIA Triton Inference Server crash from insufficient input validation
Apr 7, 2026HighVulnerabilitySecurityCVE-2026-24146CVE-2026-24146 affects NVIDIA Triton Inference Server. Insufficient input validation combined with a large number of outputs can crash the server, which a successful exploit may use to cause denial of service. The source tags the weakness as CWE-789, Memory Allocation with Excessive Size Value.
NVD/CVE DatabaseCVE-2026-5530: Ollama server-side request forgery in model pull API
Apr 4, 2026MediumVulnerabilitySecurityCVE-2026-5530CVE-2026-5530 is a flaw in Ollama up to 18.1, located in the processing of server/download.go within the Model Pull API. Manipulating this processing can lead to server-side request forgery, and the attack can be launched remotely. The vendor was contacted early about the disclosure but did not respond.
NVD/CVE DatabaseGHSA-pq5c-rjhq-qp7p: vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
Apr 3, 2026MediumVulnerabilitySecurityCVE-2026-34755The `VideoMediaIO.load_base64()` method in vLLM (`vllm/multimodal/media/video.py`) splits `video/jpeg` data URLs on commas without any frame count limit, bypassing the `num_frames` default of 32 enforced on the `load_bytes()` path. A single request to `/v1/chat/completions` containing thousands of base64-encoded JPEG frames causes the server to decode them all into memory and crash with an out-of-memory condition.
GitHub Advisory DatabaseGHSA-pf3h-qjgv-vcpr: vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
Apr 3, 2026MediumVulnerabilitySecurityCVE-2026-34753A server-side request forgery flaw in the `download_bytes_from_url` function of vLLM's batch runner (`vllm/entrypoints/openai/run_batch.py`) lets anyone who controls batch input JSON make the server issue arbitrary HTTP or HTTPS requests. The function applies no hostname, IP, port or redirect validation, unlike the multimodal `MediaConnector` path, which uses a domain allowlist. The `file_url` field of the batch transcription and translation request models feeds the URL directly, so internal services such as cloud metadata endpoints reachable from the vLLM host can be targeted.
GitHub Advisory DatabaseGHSA-3mwp-wvh9-7528: vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
Apr 3, 2026MediumVulnerabilitySecurityCVE-2026-34756The vLLM OpenAI-compatible API server has a denial-of-service flaw in the `n` parameter of `ChatCompletionRequest` and `CompletionRequest`. These Pydantic models set no upper bound on `n`, and `_verify_args` in `vllm/sampling_params.py` checks only the lower bound, so an unauthenticated attacker can send one request with a very large `n`. The engine in `vllm/v1/engine/async_llm.py` then fans the request out into millions of copies, blocking the asyncio event loop and driving memory use up until the OS OOM-killer terminates the process.
GitHub Advisory DatabaseCVE-2026-34760: vLLM audio mono downmixing mismatch with ITU-R BS.775-4 standard
Apr 2, 2026MediumVulnerabilitySecurityResearchCVE-2026-34760vLLM, an inference and serving engine for large language models, versions 0.5.5 to before 0.18.0, uses Librosa's default numpy.mean for mono downmixing (to_mono). The ITU-R BS.775-4 standard specifies a weighted downmix, so audio heard by humans differs from audio processed by AI models such as those served by vLLM or Transformers.
Fix: This issue has been patched in version 0.18.0.
NVD/CVE DatabaseCVE-2026-27893: vLLM remote code execution through hardcoded trust_remote_code in model loading
Mar 26, 2026HighVulnerabilitySecurityCVE-2026-27893CVE-2026-27893 affects vLLM, an inference and serving engine for large language models, from version 0.10.1 before 0.18.0. Two model implementation files hardcode `trust_remote_code=True` when loading sub-components, overriding a user's explicit `--trust-remote-code=False` opt-out. This enables remote code execution through malicious model repositories even when remote code trust is disabled.
Fix: Version 0.18.0 patches the issue.
NVD/CVE DatabaseCVE-2026-24158: NVIDIA Triton Inference Server denial of service via HTTP endpoint
Mar 24, 2026HighVulnerabilitySecurityCVE-2026-24158CVE-2026-24158 affects the HTTP endpoint of NVIDIA Triton Inference Server. An attacker can cause a denial of service by sending a large compressed payload, and a successful exploit may lead to denial of service. The weakness is classified as CWE-789, Memory Allocation with Excessive Size Value.
NVD/CVE DatabaseCVE-2025-33254: NVIDIA Triton Inference Server state corruption leading to denial of service
Mar 24, 2026HighVulnerabilitySecurityCVE-2025-33254NVIDIA Triton Inference Server contains a vulnerability, CVE-2025-33254, where an attacker may cause internal state corruption. A successful exploit may lead to a denial of service. The weakness is classified as CWE-362, Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition').
NVD/CVE DatabaseCVE-2025-33238: NVIDIA Triton Inference Server Sagemaker HTTP server denial of service
Mar 24, 2026HighVulnerabilitySecurityCVE-2025-33238NVIDIA Triton Inference Server's Sagemaker HTTP server contains a vulnerability, tracked as CVE-2025-33238, that lets an attacker cause an exception. A successful exploit may lead to denial of service. The weakness is classified as CWE-362, a race condition from concurrent execution using a shared resource with improper synchronization. NVD has not yet provided an assessment.
NVD/CVE DatabaseGHSA-v359-jj2v-j536: vLLM has SSRF Protection Bypass
Mar 9, 2026MediumVulnerabilitySecurityCVE-2026-25960The SSRF protection fix in vLLM's `load_from_url_async` method, in `vllm/connections.py`, can be bypassed. The fix validates URLs with `urllib3.util.parse_url()`, but the actual requests use `aiohttp`, which relies on the `yarl` library. The two parsers handle backslash characters differently. A URL such as `https://httpbin.org\@evil.com/` passes validation with host `httpbin.org` but is requested from `evil.com`, which lets an attacker bypass the hostname allowlist and reach arbitrary internal or external services.
Fix: Fixed by https://github.com/vllm-project/vllm/pull/34743
GitHub Advisory DatabaseCVE-2026-22778: vLLM information leak of heap address via multimodal endpoint image errors
Feb 2, 2026CriticalVulnerabilitySecurityCVE-2026-22778EPSS: 11.2%CVE-2026-22778 affects vLLM from 0.8.3 before 0.14.1. An invalid image sent to the multimodal endpoint makes PIL raise an error, which vLLM returns to the client, leaking a heap address and cutting ASLR from about 4 billion guesses to about 8. The leak can be chained with a heap overflow in the JPEG2000 decoder of OpenCV or FFmpeg to achieve remote code execution.
Fix: Fixed in 0.14.1.
NVD/CVE DatabaseCVE-2026-0599: A vulnerability in huggingface/text-generation-inference version 3.3.6 allows unauthenticated remote attackers to…
Feb 2, 2026HighVulnerabilitySecurityCVE-2026-0599EPSS: 29.9%CVE-2026-0599 affects huggingface/text-generation-inference version 3.3.6. Unauthenticated remote attackers can trigger unbounded external image fetching during input validation in VLM mode, because the router performs a blocking HTTP GET on Markdown image links and reads the entire response into memory. This can exhaust network bandwidth, memory and CPU, and the default deployment, which lacks memory limits and authentication, could crash the host machine.
Fix: The issue is resolved in version 3.3.7.
NVD/CVE DatabaseCVE-2026-24779: vLLM server-side request forgery in MediaConnector via URL host bypass
Jan 27, 2026HighVulnerabilitySecurityCVE-2026-24779CVE-2026-24779 affects vLLM versions prior to 0.14.1. The `MediaConnector` class in its multimodal feature set has an SSRF flaw in the `load_from_url` and `load_from_url_async` methods, where two Python parsing libraries interpret backslashes differently, letting an attacker bypass the host name restriction. A successful attacker can make the vLLM server send arbitrary requests to internal network resources, which is especially serious in containerized deployments such as `llm-d`.
Fix: Fixed in 0.14.1 (the source states version 0.14.1 contains a patch for the issue).
NVD/CVE DatabaseCVE-2025-15063: Ollama MCP Server execAsync command injection leading to remote code execution
Jan 23, 2026HighVulnerabilitySecurityCVE-2025-15063CVE-2025-15063 is a command injection flaw in the execAsync method of the Ollama MCP Server. It stems from the lack of proper validation of a user-supplied string before the string is used in a system call. Remote attackers can exploit it without authentication to execute code in the context of the service account.
NVD/CVE DatabaseCVE-2026-22807: vLLM arbitrary code execution through Hugging Face auto_map model loading
Jan 21, 2026HighVulnerabilitySecurityCVE-2026-22807vLLM versions from 0.10.1 up to but not including 0.14.0 load Hugging Face `auto_map` dynamic modules during model resolution without checking `trust_remote_code`. An attacker who can influence the model repo or path, whether a local directory or a remote Hugging Face repo, can run arbitrary Python code on the vLLM host at server startup, before any request handling and without API access.
Fix: Fixed in 0.14.0.
NVD/CVE DatabaseCVE-2025-66960: Ollama denial of service through GGUF v1 string length parsing
Jan 21, 2026HighVulnerabilitySecurityCVE-2025-66960CVE-2025-66960 affects ollama v.0.12.10. A remote attacker can cause a denial of service through the readGGUFV1String function in fs/ggml/gguf.go, which reads a string length from untrusted GGUF metadata. The source classifies the weakness as CWE-20 Improper Input Validation and CWE-400 Uncontrolled Resource Consumption, and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-66959: ollama denial of service via GGUF decoder
Jan 21, 2026HighVulnerabilitySecurityCVE-2025-66959CVE-2025-66959 describes an issue in ollama v.0.12.10 that lets a remote attacker cause a denial of service through the GGUF decoder. The source links a GitHub issue (ollama/ollama issue 9820) and a third-party advisory that attributes the flaw to an unchecked length in the GGUF decoder copy, causing a panic. The CWE entries listed are CWE-20 (Improper Input Validation) and CWE-400 (Uncontrolled Resource Consumption), and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-15514: Ollama null pointer dereference in multi-modal image processing via /api/chat
Jan 12, 2026HighVulnerabilitySecurityCVE-2025-15514Ollama versions 0.11.5-rc0 through 0.13.5 contain a null pointer dereference in the multi-modal image processing code. Malformed base64 image data sent to the /api/chat endpoint makes mtmd_helper_bitmap_init_from_buf return NULL, which is dereferenced without a check, causing a segmentation fault that crashes the runner process. A remote attacker can cause a denial of service, leaving the model unavailable to all users until the service is restarted.
NVD/CVE DatabaseCVE-2026-22773: vLLM engine crash via crafted 1x1 pixel image in Idefics3 multimodal models
Jan 10, 2026MediumVulnerabilitySecurityCVE-2026-22773CVE-2026-22773 affects vLLM, an inference and serving engine for large language models, in versions from 0.6.4 before 0.12.0. A specially crafted 1x1 pixel image sent to a server running multimodal models that use the Idefics3 vision model implementation causes a tensor dimension mismatch. The resulting unhandled runtime error terminates the whole server.
Fix: This issue has been patched in version 0.12.0.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.