Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2025-63389: Ollama API endpoints missing authentication enable unauthorized model management
Dec 18, 2025CriticalVulnerabilitySecurityCVE-2025-63389CVE-2025-63389 is a critical authentication bypass in the API endpoints of the Ollama platform, affecting versions prior to and including v0.12.3. The platform exposes multiple API endpoints without requiring authentication, which lets remote attackers perform unauthorized model management operations. The weakness is classified as CWE-306, Missing Authentication for Critical Function.
NVD/CVE DatabaseCVE-2025-33201: NVIDIA Triton Inference Server denial of service from oversized payloads
Dec 3, 2025HighVulnerabilitySecurityCVE-2025-33201NVIDIA Triton Inference Server contains a vulnerability, CVE-2025-33201, classified as CWE-754 (Improper Check for Unusual or Exceptional Conditions). An attacker can trigger it by sending extra large payloads. A successful exploit may lead to denial of service.
NVD/CVE DatabaseCVE-2025-66448: vLLM remote code execution through Nemotron_Nano_VL_Config auto_map entry
Dec 1, 2025HighVulnerabilitySecurityCVE-2025-66448vLLM versions prior to 0.11.1 contain a critical remote code execution flaw in the Nemotron_Nano_VL_Config config class. When a model config includes an auto_map entry, the class calls get_class_from_dynamic_module(...) and instantiates the returned class, which fetches and runs Python from the referenced remote repository, even when trust_remote_code=False is set in vllm.transformers_utils.config.get_config. An attacker can publish a benign-looking frontend repo whose config.json points to a malicious backend repo, so loading the frontend silently executes the backend's code on the victim host.
Fix: Fixed in 0.11.1.
NVD/CVE DatabaseCVE-2025-62426: vLLM denial of service via chat_template_kwargs in chat endpoints
Nov 21, 2025MediumVulnerabilitySecurityCVE-2025-62426CVE-2025-62426 affects vLLM, an inference and serving engine for large language models, from version 0.5.5 to before 0.11.1. The /v1/chat/completions and /tokenize endpoints use the chat_template_kwargs request parameter in code before validating it against the chat template. Crafted values of that parameter can block the API server for long periods, delaying all other requests.
Fix: This issue has been patched in version 0.11.1.
NVD/CVE DatabaseCVE-2025-62372: vLLM engine crash via malformed multimodal embedding input shape
Nov 21, 2025MediumVulnerabilitySecurityCVE-2025-62372CVE-2025-62372 affects vLLM, an inference and serving engine for large language models, from version 0.5.5 before 0.11.1. Passing multimodal embedding inputs with the correct ndim but an incorrect shape, such as a wrong hidden dimension, can crash the vLLM engine serving multimodal models. The flaw is classified as CWE-129, Improper Validation of Array Index, and GitHub rates it CVSS 4.0 8.3 (High), with availability impact only.
Fix: This issue has been patched in version 0.11.1.
NVD/CVE DatabaseCVE-2025-62164: vLLM memory corruption through prompt embeddings in Completions API
Nov 21, 2025HighVulnerabilitySecurityCVE-2025-62164vLLM versions 0.10.2 up to but not including 0.11.1 contain a memory corruption flaw in the Completions API endpoint. The endpoint calls torch.load() on user-supplied prompt embeddings without sufficient validation, and PyTorch 2.8.0 disables sparse tensor integrity checks by default, so crafted tensors can trigger an out-of-bounds write during to_dense(). The result is a crash (denial-of-service) and potentially code execution on the server hosting vLLM.
Fix: This issue has been patched in version 0.11.1.
NVD/CVE DatabaseCVE-2025-33202: NVIDIA Triton Inference Server stack overflow from oversized payloads
Nov 11, 2025MediumVulnerabilitySecurityCVE-2025-33202NVIDIA Triton Inference Server for Linux and Windows contains a vulnerability, tracked as CVE-2025-33202, in which an attacker can cause a stack overflow by sending extra-large payloads. A successful exploit might lead to denial of service. The weakness is classified as CWE-121, Stack-based Buffer Overflow.
NVD/CVE DatabaseCVE-2025-6242: vLLM SSRF in MediaConnector via load_from_url methods
Oct 7, 2025HighVulnerabilitySecurityCVE-2025-6242CVE-2025-6242 is a Server-Side Request Forgery (SSRF) flaw in the MediaConnector class of the vLLM project's multimodal feature set. The load_from_url and load_from_url_async methods fetch and process media from user-provided URLs without adequate restrictions on target hosts, letting an attacker make the vLLM server send arbitrary requests to internal network resources. The weakness is classified as CWE-918, and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-59425: vLLM API key validation vulnerable to timing attack
Oct 7, 2025HighVulnerabilitySecurityCVE-2025-59425CVE-2025-59425 affects vLLM, an inference and serving engine for large language models, before version 0.11.0rc2. Its API key validation uses a string comparison that takes longer the more characters of the key are correct. Across many attempts, an attacker could use this timing difference to recover the key character by character and bypass authentication on deployments that rely on the built-in validation.
Fix: Version 0.11.0rc2 fixes the issue.
NVD/CVE DatabaseCVE-2025-23336: NVIDIA Triton Inference Server denial of service via misconfigured model
Sep 17, 2025MediumVulnerabilitySecurityCVE-2025-23336NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability, CVE-2025-23336, classified under CWE-20 (Improper Input Validation). An attacker can cause a denial of service by loading a misconfigured model.
NVD/CVE DatabaseCVE-2025-23329: NVIDIA Triton Inference Server memory corruption via shared memory region
Sep 17, 2025HighVulnerabilitySecurityCVE-2025-23329CVE-2025-23329 affects NVIDIA Triton Inference Server for Windows and Linux. An attacker who identifies and accesses the shared memory region used by the Python backend can cause memory corruption, which might lead to denial of service. NVD has not yet provided an assessment, and the weaknesses listed are CWE-787 (Out-of-bounds Write) and CWE-284 (Improper Access Control).
NVD/CVE DatabaseCVE-2025-23328: NVIDIA Triton Inference Server out-of-bounds write through crafted input
Sep 17, 2025HighVulnerabilitySecurityCVE-2025-23328NVIDIA Triton Inference Server for Windows and Linux contains CVE-2025-23328, a vulnerability in which an attacker can cause an out-of-bounds write (CWE-787) through a specially crafted input. A successful exploit might lead to denial of service. NVD had not yet provided an assessment at the time of the record.
NVD/CVE DatabaseCVE-2025-23316: NVIDIA Triton Inference Server Python backend remote code execution
Sep 17, 2025CriticalVulnerabilitySecurityCVE-2025-23316CVE-2025-23316 affects NVIDIA Triton Inference Server for Windows and Linux. A flaw in the Python backend lets an attacker achieve remote code execution by manipulating the model name parameter in the model control APIs, and successful exploitation may also cause denial of service, information disclosure, and data tampering. The weakness is classified as CWE-78, OS Command Injection, and NVD had not yet provided an assessment at the time of the source.
NVD/CVE DatabaseCVE-2025-23268: NVIDIA Triton Inference Server DALI backend improper input validation
Sep 17, 2025HighVulnerabilitySecurityCVE-2025-23268NVIDIA Triton Inference Server contains a vulnerability in its DALI backend, caused by improper input validation (CWE-20). According to the source, a successful exploit may lead to code execution. NVD has not yet provided an assessment, and the NVIDIA advisory is linked from the NVD page.
NVD/CVE DatabaseCVE-2025-48956: vLLM denial of service through oversized HTTP GET request header
Aug 21, 2025HighVulnerabilitySecurityCVE-2025-48956CVE-2025-48956 affects vLLM, an inference and serving engine for large language models, from 0.1.0 to before 0.10.1.1. A single HTTP GET request with an extremely large header, sent to an HTTP endpoint, triggers a Denial of Service that exhausts server memory and can crash the server or make it unresponsive. No authentication is required, so any remote user can exploit it.
Fix: Fixed in 0.10.1.1.
NVD/CVE DatabaseCVE-2025-44779: Ollama arbitrary file deletion via crafted packet to /api/pull
Aug 7, 2025MediumVulnerabilitySecurityCVE-2025-44779CVE-2025-44779 affects Ollama v0.1.33. According to the source, an attacker can delete arbitrary files by sending a crafted packet to the endpoint /api/pull. The NVD has not yet provided an assessment, and the CWE mappings are CWE-20 (Improper Input Validation) and CWE-552 (Files or Directories Accessible to External Parties).
NVD/CVE DatabaseCVE-2025-23335: NVIDIA Triton Inference Server integer underflow via model configuration
Aug 6, 2025MediumVulnerabilitySecurityCVE-2025-23335NVIDIA Triton Inference Server for Windows and Linux and the Tensor RT backend contain an integer underflow (CWE-191) that an attacker can trigger with a specific model configuration and a specific input. A successful exploit might lead to denial of service. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-23334: NVIDIA Triton Inference Server Python backend out-of-bounds read
Aug 6, 2025MediumVulnerabilitySecurityCVE-2025-23334NVIDIA Triton Inference Server for Windows and Linux contains an out-of-bounds read in its Python backend (CWE-125), tracked as CVE-2025-23334. An attacker can trigger it by sending a request, and a successful exploit might lead to information disclosure. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-23333: NVIDIA Triton Inference Server out-of-bounds read via shared memory
Aug 6, 2025MediumVulnerabilitySecurityCVE-2025-23333NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability in the Python backend, tracked as CVE-2025-23333 and classified as CWE-125 (Out-of-bounds Read). An attacker could manipulate shared memory data to cause an out-of-bounds read, and a successful exploit might lead to information disclosure. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-23331: NVIDIA Triton Inference Server memory allocation flaw via invalid request
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23331CVE-2025-23331 affects NVIDIA Triton Inference Server for Windows and Linux. A user who sends an invalid request can cause a memory allocation with an excessive size value, leading to a segmentation fault. A successful exploit might lead to denial of service.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.