Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2025-23327: NVIDIA Triton Inference Server integer overflow via crafted inputs
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23327NVIDIA Triton Inference Server for Windows and Linux contains an integer overflow vulnerability (CWE-190) that an attacker can trigger through specially crafted inputs. A successful exploit might lead to denial of service and data tampering. The NVD assessment is not yet provided, and no affected versions are listed in the source.
NVD/CVE DatabaseCVE-2025-23326: NVIDIA Triton Inference Server integer overflow via crafted input
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23326NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability, CVE-2025-23326, where an attacker could cause an integer overflow through a specially crafted input. A successful exploit might lead to denial of service. The NVD assessment is not yet provided.
NVD/CVE DatabaseCVE-2025-23325: NVIDIA Triton Inference Server uncontrolled recursion via crafted input
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23325NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability (CWE-674, Uncontrolled Recursion) where an attacker can trigger uncontrolled recursion through a specially crafted input. A successful exploit might lead to denial of service. NVD has not yet provided an assessment, and the source does not list affected versions.
NVD/CVE DatabaseCVE-2025-23324: NVIDIA Triton Inference Server integer overflow via invalid request
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23324CVE-2025-23324 affects NVIDIA Triton Inference Server for Windows and Linux. A user who sends an invalid request can cause an integer overflow or wraparound (CWE-190), leading to a segmentation fault. A successful exploit might lead to denial of service.
NVD/CVE DatabaseCVE-2025-23323: NVIDIA Triton Inference Server integer overflow from invalid request
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23323NVIDIA Triton Inference Server for Windows and Linux contains CVE-2025-23323, an integer overflow or wraparound that a user can trigger by sending an invalid request. The overflow leads to a segmentation fault, and a successful exploit might cause denial of service. The entry is classified as CWE-190 and was published to NVD on 08/06/2025.
NVD/CVE DatabaseCVE-2025-23322: NVIDIA Triton Inference Server double free when stream is cancelled
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23322NVIDIA Triton Inference Server for Windows and Linux contains a double free vulnerability (CWE-415). Multiple requests could trigger it when a stream is cancelled before it is processed. A successful exploit might lead to denial of service.
NVD/CVE DatabaseCVE-2025-23321: NVIDIA Triton Inference Server divide by zero from invalid request
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23321NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability, tracked as CVE-2025-23321 and classified as CWE-369 (Divide By Zero). A user can trigger a divide by zero by sending an invalid request. A successful exploit might lead to denial of service.
NVD/CVE DatabaseCVE-2025-23320: NVIDIA Triton Inference Server Python backend shared memory limit exceeded
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23320NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability in the Python backend. An attacker could exceed the shared memory limit by sending a very large request, and a successful exploit might lead to information disclosure. The weakness is classified as CWE-209, Generation of Error Message Containing Sensitive Information.
NVD/CVE DatabaseCVE-2025-23319: NVIDIA Triton Inference Server Python backend out-of-bounds write via request
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23319NVIDIA Triton Inference Server for Windows and Linux contains an out-of-bounds write in its Python backend, which an attacker can trigger by sending a request. The source says a successful exploit might lead to remote code execution, denial of service, data tampering, or information disclosure. The NVD assessment is not yet provided, and the record is tagged CWE-787 and CWE-805.
NVD/CVE DatabaseCVE-2025-23318: NVIDIA Triton Inference Server Python backend out-of-bounds write
Aug 6, 2025HighVulnerabilitySecurityCVE-2025-23318NVIDIA Triton Inference Server for Windows and Linux contains a vulnerability in the Python backend, where an attacker could cause an out-of-bounds write (CWE-787). A successful exploit might lead to code execution, denial of service, data tampering, and information disclosure. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-23317: NVIDIA Triton Inference Server HTTP server reverse shell via crafted request
Aug 6, 2025CriticalVulnerabilitySecurityCVE-2025-23317NVIDIA Triton Inference Server contains a vulnerability in its HTTP server, where an attacker could start a reverse shell by sending a specially crafted HTTP request. The NVD entry states that a successful exploit might lead to remote code execution, denial of service, data tampering, or information disclosure, and it is classified as CWE-122 Heap-based Buffer Overflow.
NVD/CVE DatabaseCVE-2025-23311: NVIDIA Triton Inference Server stack overflow through crafted HTTP requests
Aug 6, 2025CriticalVulnerabilitySecurityCVE-2025-23311CVE-2025-23311 affects NVIDIA Triton Inference Server. An attacker can cause a stack overflow through specially crafted HTTP requests, which CWE-121 classifies as a stack-based buffer overflow. The source says a successful exploit might lead to remote code execution, denial of service, information disclosure, or data tampering.
NVD/CVE DatabaseCVE-2025-23310: NVIDIA Triton Inference Server stack buffer overflow from crafted inputs
Aug 6, 2025CriticalVulnerabilitySecurityCVE-2025-23310NVIDIA Triton Inference Server for Windows and Linux contains a stack buffer overflow, classified as CWE-121, that an attacker can trigger with specially crafted inputs. The source states that a successful exploit might lead to remote code execution, denial of service, information disclosure, and data tampering. NVD had not yet provided an assessment when the entry was published on 08/06/2025.
NVD/CVE DatabaseCVE-2025-51471: Ollama cross-domain token exposure in WWW-Authenticate realm handling
Jul 22, 2025MediumVulnerabilitySecurityCVE-2025-51471EPSS: 15.2%CVE-2025-51471 is a cross-domain token exposure flaw in server.auth.getAuthorizationToken in Ollama 0.6.7. A remote attacker who controls a malicious realm value in a WWW-Authenticate header returned by the /api/pull endpoint can steal authentication tokens and bypass access controls. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-48944: vLLM crash through malformed tool input on /v1/chat/completions
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48944CVE-2025-48944 affects vLLM versions 0.8.0 up to but excluding 0.9.0. The vLLM backend behind the /v1/chat/completions OpenAPI endpoint does not validate unexpected or malformed input in the "pattern" and "type" fields when the tools functionality is invoked. A single crafted request crashes the inference worker, which stays down until it is restarted.
Fix: Fixed in 0.9.0.
NVD/CVE DatabaseCVE-2025-48943: vLLM denial of service through invalid regex in structured output
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48943CVE-2025-48943 affects vLLM, an inference and serving engine for large language models, in versions 0.8.0 up to but excluding 0.9.0. An invalid regex supplied while using structured output triggers a ReDoS that crashes the vLLM server, a denial of service. The flaw is similar to GHSA-6qc9-v4r8-22xg/CVE-2025-48942, but applies to regex rather than a JSON schema.
Fix: Version 0.9.0 fixes the issue.
NVD/CVE DatabaseCVE-2025-48942: vLLM server crash via invalid json_schema in /v1/completions Guided Param
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48942CVE-2025-48942 affects vLLM, an inference and serving engine for large language models, in versions 0.8.0 up to but excluding 0.9.0. Sending an invalid json_schema as a Guided Param to the /v1/completions API crashes the vllm server. It is similar to GHSA-9hcf-v7m4-6m2j/CVE-2025-48943, but applies to regex instead of a JSON schema.
Fix: Version 0.9.0 fixes the issue.
NVD/CVE DatabaseCVE-2025-48887: vLLM regular expression denial of service in pythonic tool parser
May 30, 2025MediumVulnerabilitySecurityCVE-2025-48887vLLM, an inference and serving engine for large language models (LLMs), contains a Regular Expression Denial of Service (ReDoS) flaw in `vllm/entrypoints/openai/tool_parsers/pythonic_tool_parser.py` in versions 0.6.4 up to but excluding 0.9.0. The tool call detection regex uses multiple nested quantifiers, optional groups and inner repetitions, so an attacker can trigger catastrophic backtracking, severely degrading performance or making the service unavailable.
Fix: Version 0.9.0 contains a patch for the issue.
NVD/CVE DatabaseCVE-2025-46722: vLLM hash collisions in multimodal image hashing via raw pixel bytes
May 29, 2025MediumVulnerabilitySecurityCVE-2025-46722vLLM versions from 0.7.0 before 0.9.0 contain a flaw in the MultiModalHasher class in vllm/multimodal/hasher.py. Its image hashing method serializes PIL.Image.Image objects using only obj.tobytes(), which omits metadata such as width, height and mode, so images of different dimensions with identical pixel bytes can produce the same hash. This can cause hash collisions, incorrect cache hits, and possible data leakage.
Fix: Fixed in 0.9.0.
NVD/CVE DatabaseCVE-2025-46570: vLLM timing side channel in prefix cache matching during prefill
May 29, 2025LowVulnerabilitySecurityCVE-2025-46570CVE-2025-46570 affects vLLM, an inference and serving engine for large language models, before version 0.9.0. When a new prompt is processed and the PageAttention mechanism finds a matching prefix chunk, prefill is faster, which shows up in TTFT (Time to First Token). These timing differences are large enough to be recognized and exploited.
Fix: Fixed in version 0.9.0.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.