Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2026-22807: vLLM arbitrary code execution through Hugging Face auto_map model loading
Jan 21, 2026HighVulnerabilitySecurityCVE-2026-22807vLLM versions from 0.10.1 up to but not including 0.14.0 load Hugging Face `auto_map` dynamic modules during model resolution without checking `trust_remote_code`. An attacker who can influence the model repo or path, whether a local directory or a remote Hugging Face repo, can run arbitrary Python code on the vLLM host at server startup, before any request handling and without API access.
Fix: Fixed in 0.14.0.
NVD/CVE DatabaseCVE-2025-66960: Ollama denial of service through GGUF v1 string length parsing
Jan 21, 2026HighVulnerabilitySecurityCVE-2025-66960CVE-2025-66960 affects ollama v.0.12.10. A remote attacker can cause a denial of service through the readGGUFV1String function in fs/ggml/gguf.go, which reads a string length from untrusted GGUF metadata. The source classifies the weakness as CWE-20 Improper Input Validation and CWE-400 Uncontrolled Resource Consumption, and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-66959: ollama denial of service via GGUF decoder
Jan 21, 2026HighVulnerabilitySecurityCVE-2025-66959CVE-2025-66959 describes an issue in ollama v.0.12.10 that lets a remote attacker cause a denial of service through the GGUF decoder. The source links a GitHub issue (ollama/ollama issue 9820) and a third-party advisory that attributes the flaw to an unchecked length in the GGUF decoder copy, causing a panic. The CWE entries listed are CWE-20 (Improper Input Validation) and CWE-400 (Uncontrolled Resource Consumption), and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-15514: Ollama null pointer dereference in multi-modal image processing via /api/chat
Jan 12, 2026HighVulnerabilitySecurityCVE-2025-15514Ollama versions 0.11.5-rc0 through 0.13.5 contain a null pointer dereference in the multi-modal image processing code. Malformed base64 image data sent to the /api/chat endpoint makes mtmd_helper_bitmap_init_from_buf return NULL, which is dereferenced without a check, causing a segmentation fault that crashes the runner process. A remote attacker can cause a denial of service, leaving the model unavailable to all users until the service is restarted.
NVD/CVE DatabaseCVE-2026-22773: vLLM engine crash via crafted 1x1 pixel image in Idefics3 multimodal models
Jan 10, 2026MediumVulnerabilitySecurityCVE-2026-22773CVE-2026-22773 affects vLLM, an inference and serving engine for large language models, in versions from 0.6.4 before 0.12.0. A specially crafted 1x1 pixel image sent to a server running multimodal models that use the Idefics3 vision model implementation causes a tensor dimension mismatch. The resulting unhandled runtime error terminates the whole server.
Fix: This issue has been patched in version 0.12.0.
NVD/CVE DatabaseCVE-2025-63389: Ollama API endpoints missing authentication enable unauthorized model management
Dec 18, 2025CriticalVulnerabilitySecurityCVE-2025-63389CVE-2025-63389 is a critical authentication bypass in the API endpoints of the Ollama platform, affecting versions prior to and including v0.12.3. The platform exposes multiple API endpoints without requiring authentication, which lets remote attackers perform unauthorized model management operations. The weakness is classified as CWE-306, Missing Authentication for Critical Function.
NVD/CVE DatabaseCVE-2025-33201: NVIDIA Triton Inference Server denial of service from oversized payloads
Dec 3, 2025HighVulnerabilitySecurityCVE-2025-33201NVIDIA Triton Inference Server contains a vulnerability, CVE-2025-33201, classified as CWE-754 (Improper Check for Unusual or Exceptional Conditions). An attacker can trigger it by sending extra large payloads. A successful exploit may lead to denial of service.
NVD/CVE DatabaseCVE-2025-66448: vLLM remote code execution through Nemotron_Nano_VL_Config auto_map entry
Dec 1, 2025HighVulnerabilitySecurityCVE-2025-66448vLLM versions prior to 0.11.1 contain a critical remote code execution flaw in the Nemotron_Nano_VL_Config config class. When a model config includes an auto_map entry, the class calls get_class_from_dynamic_module(...) and instantiates the returned class, which fetches and runs Python from the referenced remote repository, even when trust_remote_code=False is set in vllm.transformers_utils.config.get_config. An attacker can publish a benign-looking frontend repo whose config.json points to a malicious backend repo, so loading the frontend silently executes the backend's code on the victim host.
Fix: Fixed in 0.11.1.
NVD/CVE DatabaseCVE-2025-62426: vLLM denial of service via chat_template_kwargs in chat endpoints
Nov 21, 2025MediumVulnerabilitySecurityCVE-2025-62426CVE-2025-62426 affects vLLM, an inference and serving engine for large language models, from version 0.5.5 to before 0.11.1. The /v1/chat/completions and /tokenize endpoints use the chat_template_kwargs request parameter in code before validating it against the chat template. Crafted values of that parameter can block the API server for long periods, delaying all other requests.
Fix: This issue has been patched in version 0.11.1.
NVD/CVE DatabaseCVE-2025-62372: vLLM engine crash via malformed multimodal embedding input shape
Nov 21, 2025MediumVulnerabilitySecurityCVE-2025-62372CVE-2025-62372 affects vLLM, an inference and serving engine for large language models, from version 0.5.5 before 0.11.1. Passing multimodal embedding inputs with the correct ndim but an incorrect shape, such as a wrong hidden dimension, can crash the vLLM engine serving multimodal models. The flaw is classified as CWE-129, Improper Validation of Array Index, and GitHub rates it CVSS 4.0 8.3 (High), with availability impact only.
Fix: This issue has been patched in version 0.11.1.
NVD/CVE DatabaseCVE-2025-62164: vLLM memory corruption through prompt embeddings in Completions API
Nov 21, 2025HighVulnerabilitySecurityCVE-2025-62164vLLM versions 0.10.2 up to but not including 0.11.1 contain a memory corruption flaw in the Completions API endpoint. The endpoint calls torch.load() on user-supplied prompt embeddings without sufficient validation, and PyTorch 2.8.0 disables sparse tensor integrity checks by default, so crafted tensors can trigger an out-of-bounds write during to_dense(). The result is a crash (denial-of-service) and potentially code execution on the server hosting vLLM.
Fix: This issue has been patched in version 0.11.1.
NVD/CVE DatabaseCVE-2025-33202: NVIDIA Triton Inference Server stack overflow from oversized payloads
Nov 11, 2025MediumVulnerabilitySecurityCVE-2025-33202NVIDIA Triton Inference Server for Linux and Windows contains a vulnerability, tracked as CVE-2025-33202, in which an attacker can cause a stack overflow by sending extra-large payloads. A successful exploit might lead to denial of service. The weakness is classified as CWE-121, Stack-based Buffer Overflow.
NVD/CVE DatabaseCVE-2025-6242: vLLM SSRF in MediaConnector via load_from_url methods
Oct 7, 2025HighVulnerabilitySecurityCVE-2025-6242CVE-2025-6242 is a Server-Side Request Forgery (SSRF) flaw in the MediaConnector class of the vLLM project's multimodal feature set. The load_from_url and load_from_url_async methods fetch and process media from user-provided URLs without adequate restrictions on target hosts, letting an attacker make the vLLM server send arbitrary requests to internal network resources. The weakness is classified as CWE-918, and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-59425: vLLM API key validation vulnerable to timing attack
Oct 7, 2025HighVulnerabilitySecurityCVE-2025-59425CVE-2025-59425 affects vLLM, an inference and serving engine for large language models, before version 0.11.0rc2. Its API key validation uses a string comparison that takes longer the more characters of the key are correct. Across many attempts, an attacker could use this timing difference to recover the key character by character and bypass authentication on deployments that rely on the built-in validation.
Fix: Version 0.11.0rc2 fixes the issue.
NVD/CVE DatabaseCVE-2025-23336: NVIDIA Triton Inference Server denial of service via misconfigured model
Sep 17, 2025MediumVulnerabilitySecurityCVE-2025-23336NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability, CVE-2025-23336, classified under CWE-20 (Improper Input Validation). An attacker can cause a denial of service by loading a misconfigured model.
NVD/CVE DatabaseCVE-2025-23329: NVIDIA Triton Inference Server memory corruption via shared memory region
Sep 17, 2025HighVulnerabilitySecurityCVE-2025-23329CVE-2025-23329 affects NVIDIA Triton Inference Server for Windows and Linux. An attacker who identifies and accesses the shared memory region used by the Python backend can cause memory corruption, which might lead to denial of service. NVD has not yet provided an assessment, and the weaknesses listed are CWE-787 (Out-of-bounds Write) and CWE-284 (Improper Access Control).
NVD/CVE DatabaseCVE-2025-23328: NVIDIA Triton Inference Server out-of-bounds write through crafted input
Sep 17, 2025HighVulnerabilitySecurityCVE-2025-23328NVIDIA Triton Inference Server for Windows and Linux contains CVE-2025-23328, a vulnerability in which an attacker can cause an out-of-bounds write (CWE-787) through a specially crafted input. A successful exploit might lead to denial of service. NVD had not yet provided an assessment at the time of the record.
NVD/CVE DatabaseCVE-2025-23316: NVIDIA Triton Inference Server Python backend remote code execution
Sep 17, 2025CriticalVulnerabilitySecurityCVE-2025-23316CVE-2025-23316 affects NVIDIA Triton Inference Server for Windows and Linux. A flaw in the Python backend lets an attacker achieve remote code execution by manipulating the model name parameter in the model control APIs, and successful exploitation may also cause denial of service, information disclosure, and data tampering. The weakness is classified as CWE-78, OS Command Injection, and NVD had not yet provided an assessment at the time of the source.
NVD/CVE DatabaseCVE-2025-23268: NVIDIA Triton Inference Server DALI backend improper input validation
Sep 17, 2025HighVulnerabilitySecurityCVE-2025-23268NVIDIA Triton Inference Server contains a vulnerability in its DALI backend, caused by improper input validation (CWE-20). According to the source, a successful exploit may lead to code execution. NVD has not yet provided an assessment, and the NVIDIA advisory is linked from the NVD page.
NVD/CVE DatabaseCVE-2025-48956: vLLM denial of service through oversized HTTP GET request header
Aug 21, 2025HighVulnerabilitySecurityCVE-2025-48956CVE-2025-48956 affects vLLM, an inference and serving engine for large language models, from 0.1.0 to before 0.10.1.1. A single HTTP GET request with an extremely large header, sent to an HTTP endpoint, triggers a Denial of Service that exhausts server memory and can crash the server or make it unresponsive. No authentication is required, so any remote user can exploit it.
Fix: Fixed in 0.10.1.1.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.