Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2025-29770: vLLM unbounded filesystem cache growth via outlines guided decoding
Mar 19, 2025MediumVulnerabilitySecurityCVE-2025-29770CVE-2025-29770 affects vLLM's structured output path, which uses the outlines library's on-disk grammar cache, on by default. The unconditional use of that cache in vllm/model_executor/guided_decoding/outlines_logits_processors.py lets a user send many short requests with unique schemas, each adding a cache entry. Because outlines is also reachable per request through the OpenAI-compatible API server, this can exhaust filesystem space and cause a Denial of Service. The issue applies only to the V0 engine.
Fix: Fixed in 0.8.0.
NVD/CVE DatabaseCVE-2025-1953: vLLM AIBrix prefix caching uses insufficiently random values
Mar 4, 2025LowVulnerabilitySecurityCVE-2025-1953CVE-2025-1953 is a vulnerability in vLLM AIBrix 0.2.0, classified as problematic, in the Prefix Caching component. The flaw lies in pkg/plugins/gateway/prefixcacheindexer/hash.go, where manipulation leads to insufficiently random values (CWE-330). The source rates attack complexity as high and exploitation as difficult.
Fix: Upgrade to version 0.3.0, which the source says is able to address this issue.
NVD/CVE DatabaseCVE-2024-53880: NVIDIA Triton Inference Server integer overflow in model loading API
Feb 12, 2025MediumVulnerabilitySecurityCVE-2024-53880NVIDIA Triton Inference Server contains an integer overflow or wraparound vulnerability (CWE-190) in its model loading API. A user who loads a model with an extra-large file size can overflow an internal variable, which may lead to denial of service.
NVD/CVE DatabaseCVE-2025-25183: vLLM prefix caching hash collisions can reuse cache from other prompts
Feb 7, 2025LowVulnerabilitySecurityCVE-2025-25183CVE-2025-25183 affects vLLM, a high-throughput inference and serving engine for LLMs. Maliciously constructed statements can cause hash collisions in prefix caching, which relies on Python's built-in hash() function, so cached content generated from different prompts can be reused and interfere with later responses. Because hash(None) became a predictable constant in Python 3.12, an attacker who knows the prompts in use could populate the cache with a colliding prompt.
Fix: Fixed in version 0.7.2; all users are advised to upgrade. There are no known workarounds for this vulnerability.
NVD/CVE DatabaseCVE-2025-24357: vLLM deserialization of untrusted model weights via torch.load
Jan 27, 2025HighVulnerabilitySecurityCVE-2025-24357CVE-2025-24357 affects vLLM, a library for LLM inference and serving. In vllm/model_executor/weight_utils.py, the hf_model_weights_iterator function loads model checkpoints downloaded from Hugging Face using torch.load with weights_only defaulting to False. When torch.load loads malicious pickle data, it executes arbitrary code during unpickling. The weakness is classified as CWE-502, Deserialization of Untrusted Data.
Fix: This vulnerability is fixed in v0.7.0.
NVD/CVE DatabaseCVE-2024-39722: Ollama path traversal in api/push route exposes server file existence
Oct 31, 2024HighVulnerabilitySecurityCVE-2024-39722CVE-2024-39722 affects Ollama before 0.1.46. A path traversal flaw in the api/push route lets a remote party see which files exist on the server where Ollama is deployed. The weakness is classified as CWE-22, Improper Limitation of a Pathname to a Restricted Directory.
NVD/CVE DatabaseCVE-2024-39721: Ollama infinite goroutine via blocking file path in CreateModelHandler
Oct 31, 2024HighVulnerabilitySecurityCVE-2024-39721CVE-2024-39721 affects Ollama before 0.1.34. In the CreateModelHandler function, os.Open reads a file until completion, and the user-controlled req.Path parameter can be set to /dev/random, a blocking device. This causes the goroutine to run indefinitely, even after the client aborts the HTTP request. The weakness is classified as CWE-404, Improper Resource Shutdown or Release.
NVD/CVE DatabaseCVE-2024-39720: Ollama crash via malformed GGUF file in CreateModel route
Oct 31, 2024HighVulnerabilitySecurityCVE-2024-39720CVE-2024-39720 affects Ollama before 0.1.46. An attacker can upload a malformed GGUF file of just 4 bytes, starting with the GGUF custom magic header, using two HTTP requests. By pointing a custom Modelfile's FROM statement at that blob, the attacker crashes the application through the CreateModel route with a segmentation fault (SIGSEGV). The weakness is classified as CWE-125, Out-of-bounds Read.
Fix: The source links a fix commit in the ollama/ollama repository comparing v0.1.45 to v0.1.46, but it does not state the fix in words. Its stated mitigation is therefore: N/A -- no mitigation discussed in source.
NVD/CVE DatabaseCVE-2024-39719: Ollama file existence disclosure via api/create
Oct 31, 2024HighVulnerabilitySecurityCVE-2024-39719CVE-2024-39719 affects Ollama through 0.3.14. Calling the CreateModel route through api/create with a nonexistent path parameter returns a "File does not exist" error message to the caller, letting an attacker determine whether files exist on the server. The weakness is classified as CWE-209, Generation of Error Message Containing Sensitive Information.
NVD/CVE DatabaseCVE-2024-0116: NVIDIA Triton Inference Server out-of-bounds read by releasing shared memory
Oct 1, 2024MediumVulnerabilitySecurityCVE-2024-0116CVE-2024-0116 affects NVIDIA Triton Inference Server. A user can cause an out-of-bounds read (CWE-125) by releasing a shared memory region while it is still in use. A successful exploit may lead to denial of service.
NVD/CVE DatabaseCVE-2024-8768: vLLM denial of service through completions API with empty prompt
Sep 17, 2024HighVulnerabilitySecurityCVE-2024-8768CVE-2024-8768 is a flaw in the vLLM library. A completions API request with an empty prompt crashes the vLLM API server, causing a denial of service. The weakness is classified as CWE-617, Reachable Assertion, and NVD had not yet provided an assessment when published on 09/17/2024.
NVD/CVE DatabaseCVE-2024-45436: Ollama path traversal in ZIP extraction in extractFromZipFile
Aug 29, 2024HighVulnerabilitySecurityCVE-2024-45436CVE-2024-45436 affects extractFromZipFile in model.go in Ollama before 0.1.47. The flaw lets the function extract members of a ZIP archive outside the parent directory, which is classified as CWE-22 (Path Traversal). NVD has not yet provided an assessment.
Fix: Fix referenced in the source: upgrade to 0.1.47 or later. The source links the compare view v0.1.46...v0.1.47 and pull request #5314 as patch references.
NVD/CVE DatabaseCVE-2024-0103: NVIDIA Triton Inference Server for Linux incorrect resource initialization
Jun 13, 2024MediumVulnerabilitySecurityCVE-2024-0103CVE-2024-0103 affects NVIDIA Triton Inference Server for Linux. A user may cause an incorrect initialization of resource through a network issue, classified as CWE-NVD-CWE-Other. Successful exploitation may lead to information disclosure.
NVD/CVE DatabaseCVE-2024-0095: NVIDIA Triton Inference Server log injection leading to forged logs and commands
Jun 13, 2024CriticalVulnerabilitySecurityCVE-2024-0095NVIDIA Triton Inference Server for Linux and Windows contains CVE-2024-0095, a vulnerability where a user can inject forged logs and executable commands by inserting arbitrary data as a new log entry (CWE-117, Improper Output Neutralization for Logs). The source states that a successful exploit might lead to code execution, denial of service, escalation of privileges, information disclosure, and data tampering. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2024-37032: Ollama path traversal in digest validation when getting model path
May 31, 2024HighVulnerabilitySecurityCVE-2024-37032EPSS: 89.6%CVE-2024-37032 affects Ollama before 0.1.34. The software does not validate the format of the digest (sha256 with 64 hex digits) when getting the model path, so it mishandles inputs such as digests with fewer or more than 64 hex digits or an initial ../ substring. The weakness is classified as CWE-22, Path Traversal, and was published to NVD on 05/31/2024.
Fix: Fixed in 0.1.34. The source references the v0.1.33...v0.1.34 comparison and pull request 4175.
NVD/CVE DatabaseCVE-2024-3924: A code injection vulnerability exists in the huggingface/text-generation-inference repository, specifically within theā¦
May 30, 2024HighVulnerabilitySecurityCVE-2024-3924CVE-2024-3924 is a code injection flaw in the `autodocs.yml` workflow file of the huggingface/text-generation-inference repository. The workflow builds a package-install command from the untrusted `github.head_ref` value, so an attacker who forks the repository, names a branch with a malicious payload, and opens a pull request to the base repository can achieve arbitrary code execution within the GitHub Actions runner. The issue affects versions up to and including v2.0.0.
Fix: Fixed in version 2.0.0.
NVD/CVE DatabaseCVE-2024-0100: NVIDIA Triton Inference Server for Linux flaw corrupts system files
May 14, 2024MediumVulnerabilitySecurityCVE-2024-0100NVIDIA Triton Inference Server for Linux contains a vulnerability in its tracing API, where a user can corrupt system files. A successful exploit might lead to denial of service and data tampering. The weakness is classified as CWE-73, External Control of File Name or Path.
NVD/CVE DatabaseCVE-2024-0088: NVIDIA Triton Inference Server for Linux memory access flaw
May 14, 2024MediumVulnerabilitySecurityCVE-2024-0088EPSS: 18.9%NVIDIA Triton Inference Server for Linux contains an improper memory access flaw in its shared memory APIs, which a user can trigger through a network API. Successful exploitation might lead to denial of service and data tampering. The NVD entry, published 05/14/2024, lists CWE-787 (Out-of-bounds Write) and CWE-119 and has no NVD assessment yet.
NVD/CVE DatabaseCVE-2024-0087: NVIDIA Triton Inference Server for Linux arbitrary file logging location
May 14, 2024CriticalVulnerabilitySecurityCVE-2024-0087EPSS: 19.9%CVE-2024-0087 affects NVIDIA Triton Inference Server for Linux. A user can set the logging location to an arbitrary file, and if that file exists, logs are appended to it. The source says a successful exploit might lead to code execution, denial of service, escalation of privileges, information disclosure, and data tampering.
NVD/CVE DatabaseCVE-2024-34359: llama-cpp-python remote code execution via chat template in GGUF metadata
May 14, 2024CriticalVulnerabilitySecurityCVE-2024-34359EPSS: 26.0%llama-cpp-python, the Python bindings for llama.cpp, is affected by CVE-2024-34359. The Llama constructor in llama.py loads the chat template from a .gguf model's metadata and passes it to Jinja2ChatFormatter, which parses it with a non-sandboxed jinja2.Environment and renders it in __call__. A carefully constructed template payload therefore enables server-side template injection and remote code execution.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.