Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2026-24175: NVIDIA Triton Inference Server crash from malformed request header
Apr 7, 2026HighVulnerabilitySecurityCVE-2026-24175NVIDIA Triton Inference Server contains a vulnerability, CVE-2026-24175, classified as CWE-248 (Uncaught Exception). An attacker can crash the server by sending a malformed request header, which may lead to denial of service. NVD published the entry on 04/07/2026 and has not yet provided its own assessment.
NVD/CVE DatabaseCVE-2026-24174: NVIDIA Triton Inference Server crash from malformed request
Apr 7, 2026HighVulnerabilitySecurityCVE-2026-24174CVE-2026-24174 affects NVIDIA Triton Inference Server. An attacker can crash the server by sending a malformed request, which might lead to denial of service. The weakness is classified as CWE-681, Incorrect Conversion between Numeric Types, and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-24173: NVIDIA Triton Inference Server crash from malformed request
Apr 7, 2026HighVulnerabilitySecurityCVE-2026-24173NVIDIA Triton Inference Server contains a vulnerability, tracked as CVE-2026-24173, where an attacker can cause a server crash by sending a malformed request. The source states that a successful exploit might lead to denial of service. The weakness is classified as CWE-190, Integer Overflow or Wraparound, and NVD published the entry on 04/07/2026.
NVD/CVE DatabaseCVE-2026-24147: NVIDIA Triton Inference Server contains a vulnerability in triton server where an attacker may cause an information…
Apr 7, 2026MediumVulnerabilitySecurityCVE-2026-24147NVIDIA Triton Inference Server contains a vulnerability in its server component that lets an attacker cause information disclosure by uploading a model configuration. The source classifies it under CWE-22, Path Traversal. A successful exploit may lead to information disclosure or denial of service. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-24146: NVIDIA Triton Inference Server crash from insufficient input validation
Apr 7, 2026HighVulnerabilitySecurityCVE-2026-24146CVE-2026-24146 affects NVIDIA Triton Inference Server. Insufficient input validation combined with a large number of outputs can crash the server, which a successful exploit may use to cause denial of service. The source tags the weakness as CWE-789, Memory Allocation with Excessive Size Value.
NVD/CVE DatabaseCVE-2026-5530: Ollama server-side request forgery in model pull API
Apr 4, 2026MediumVulnerabilitySecurityCVE-2026-5530CVE-2026-5530 is a flaw in Ollama up to 18.1, located in the processing of server/download.go within the Model Pull API. Manipulating this processing can lead to server-side request forgery, and the attack can be launched remotely. The vendor was contacted early about the disclosure but did not respond.
NVD/CVE DatabaseGHSA-pq5c-rjhq-qp7p: vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
Apr 3, 2026MediumVulnerabilitySecurityCVE-2026-34755The `VideoMediaIO.load_base64()` method in vLLM (`vllm/multimodal/media/video.py`) splits `video/jpeg` data URLs on commas without any frame count limit, bypassing the `num_frames` default of 32 enforced on the `load_bytes()` path. A single request to `/v1/chat/completions` containing thousands of base64-encoded JPEG frames causes the server to decode them all into memory and crash with an out-of-memory condition.
GitHub Advisory DatabaseGHSA-pf3h-qjgv-vcpr: vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
Apr 3, 2026MediumVulnerabilitySecurityCVE-2026-34753A server-side request forgery flaw in the `download_bytes_from_url` function of vLLM's batch runner (`vllm/entrypoints/openai/run_batch.py`) lets anyone who controls batch input JSON make the server issue arbitrary HTTP or HTTPS requests. The function applies no hostname, IP, port or redirect validation, unlike the multimodal `MediaConnector` path, which uses a domain allowlist. The `file_url` field of the batch transcription and translation request models feeds the URL directly, so internal services such as cloud metadata endpoints reachable from the vLLM host can be targeted.
GitHub Advisory DatabaseGHSA-3mwp-wvh9-7528: vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
Apr 3, 2026MediumVulnerabilitySecurityCVE-2026-34756The vLLM OpenAI-compatible API server has a denial-of-service flaw in the `n` parameter of `ChatCompletionRequest` and `CompletionRequest`. These Pydantic models set no upper bound on `n`, and `_verify_args` in `vllm/sampling_params.py` checks only the lower bound, so an unauthenticated attacker can send one request with a very large `n`. The engine in `vllm/v1/engine/async_llm.py` then fans the request out into millions of copies, blocking the asyncio event loop and driving memory use up until the OS OOM-killer terminates the process.
GitHub Advisory DatabaseCVE-2026-34760: vLLM audio mono downmixing mismatch with ITU-R BS.775-4 standard
Apr 2, 2026MediumVulnerabilitySecurityResearchCVE-2026-34760vLLM, an inference and serving engine for large language models, versions 0.5.5 to before 0.18.0, uses Librosa's default numpy.mean for mono downmixing (to_mono). The ITU-R BS.775-4 standard specifies a weighted downmix, so audio heard by humans differs from audio processed by AI models such as those served by vLLM or Transformers.
Fix: This issue has been patched in version 0.18.0.
NVD/CVE DatabaseCVE-2026-27893: vLLM remote code execution through hardcoded trust_remote_code in model loading
Mar 26, 2026HighVulnerabilitySecurityCVE-2026-27893CVE-2026-27893 affects vLLM, an inference and serving engine for large language models, from version 0.10.1 before 0.18.0. Two model implementation files hardcode `trust_remote_code=True` when loading sub-components, overriding a user's explicit `--trust-remote-code=False` opt-out. This enables remote code execution through malicious model repositories even when remote code trust is disabled.
Fix: Version 0.18.0 patches the issue.
NVD/CVE DatabaseCVE-2026-24158: NVIDIA Triton Inference Server denial of service via HTTP endpoint
Mar 24, 2026HighVulnerabilitySecurityCVE-2026-24158CVE-2026-24158 affects the HTTP endpoint of NVIDIA Triton Inference Server. An attacker can cause a denial of service by sending a large compressed payload, and a successful exploit may lead to denial of service. The weakness is classified as CWE-789, Memory Allocation with Excessive Size Value.
NVD/CVE DatabaseCVE-2025-33254: NVIDIA Triton Inference Server state corruption leading to denial of service
Mar 24, 2026HighVulnerabilitySecurityCVE-2025-33254NVIDIA Triton Inference Server contains a vulnerability, CVE-2025-33254, where an attacker may cause internal state corruption. A successful exploit may lead to a denial of service. The weakness is classified as CWE-362, Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition').
NVD/CVE DatabaseCVE-2025-33238: NVIDIA Triton Inference Server Sagemaker HTTP server denial of service
Mar 24, 2026HighVulnerabilitySecurityCVE-2025-33238NVIDIA Triton Inference Server's Sagemaker HTTP server contains a vulnerability, tracked as CVE-2025-33238, that lets an attacker cause an exception. A successful exploit may lead to denial of service. The weakness is classified as CWE-362, a race condition from concurrent execution using a shared resource with improper synchronization. NVD has not yet provided an assessment.
NVD/CVE DatabaseGHSA-v359-jj2v-j536: vLLM has SSRF Protection Bypass
Mar 9, 2026MediumVulnerabilitySecurityCVE-2026-25960The SSRF protection fix in vLLM's `load_from_url_async` method, in `vllm/connections.py`, can be bypassed. The fix validates URLs with `urllib3.util.parse_url()`, but the actual requests use `aiohttp`, which relies on the `yarl` library. The two parsers handle backslash characters differently. A URL such as `https://httpbin.org\@evil.com/` passes validation with host `httpbin.org` but is requested from `evil.com`, which lets an attacker bypass the hostname allowlist and reach arbitrary internal or external services.
Fix: Fixed by https://github.com/vllm-project/vllm/pull/34743
GitHub Advisory Databaseggml.ai joins Hugging Face to ensure the long-term progress of Local AI
Feb 20, 2026InfoNewsIndustryggml.ai, the team behind llama.cpp, is joining Hugging Face to support the long-term development of local AI. The announcement says joint work will focus on tighter integration with the transformers library and better packaging and user experience for ggml-based software. The source is a commentary on the news that praises the move and the earlier llama.cpp release from March 2023.
Simon Willison's WeblogCVE-2026-22778: vLLM information leak of heap address via multimodal endpoint image errors
Feb 2, 2026CriticalVulnerabilitySecurityCVE-2026-22778EPSS: 11.2%CVE-2026-22778 affects vLLM from 0.8.3 before 0.14.1. An invalid image sent to the multimodal endpoint makes PIL raise an error, which vLLM returns to the client, leaking a heap address and cutting ASLR from about 4 billion guesses to about 8. The leak can be chained with a heap overflow in the JPEG2000 decoder of OpenCV or FFmpeg to achieve remote code execution.
Fix: Fixed in 0.14.1.
NVD/CVE DatabaseCVE-2026-0599: A vulnerability in huggingface/text-generation-inference version 3.3.6 allows unauthenticated remote attackers to…
Feb 2, 2026HighVulnerabilitySecurityCVE-2026-0599EPSS: 29.9%CVE-2026-0599 affects huggingface/text-generation-inference version 3.3.6. Unauthenticated remote attackers can trigger unbounded external image fetching during input validation in VLM mode, because the router performs a blocking HTTP GET on Markdown image links and reads the entire response into memory. This can exhaust network bandwidth, memory and CPU, and the default deployment, which lacks memory limits and authentication, could crash the host machine.
Fix: The issue is resolved in version 3.3.7.
NVD/CVE DatabaseCVE-2026-24779: vLLM server-side request forgery in MediaConnector via URL host bypass
Jan 27, 2026HighVulnerabilitySecurityCVE-2026-24779CVE-2026-24779 affects vLLM versions prior to 0.14.1. The `MediaConnector` class in its multimodal feature set has an SSRF flaw in the `load_from_url` and `load_from_url_async` methods, where two Python parsing libraries interpret backslashes differently, letting an attacker bypass the host name restriction. A successful attacker can make the vLLM server send arbitrary requests to internal network resources, which is especially serious in containerized deployments such as `llm-d`.
Fix: Fixed in 0.14.1 (the source states version 0.14.1 contains a patch for the issue).
NVD/CVE DatabaseCVE-2025-15063: Ollama MCP Server execAsync command injection leading to remote code execution
Jan 23, 2026HighVulnerabilitySecurityCVE-2025-15063CVE-2025-15063 is a command injection flaw in the execAsync method of the Ollama MCP Server. It stems from the lack of proper validation of a user-supplied string before the string is used in a system call. Remote attackers can exploit it without authentication to execute code in the context of the service account.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.