Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2026-24215: NVIDIA Triton Inference Server DALI backend uncontrolled resource consumption
May 20, 2026MediumVulnerabilitySecurityCVE-2026-24215CVE-2026-24215 affects the DALI backend of NVIDIA Triton Inference Server. An attacker could cause uncontrolled resource consumption, which might lead to denial of service. NVD has not yet provided an assessment, and the source names no affected versions.
NVD/CVE DatabaseCVE-2026-24214: NVIDIA Triton Inference Server DALI backend integer overflow
May 20, 2026HighVulnerabilitySecurityCVE-2026-24214CVE-2026-24214 is a vulnerability in the DALI backend of NVIDIA Triton Inference Server, classified as CWE-190 (Integer Overflow or Wraparound). An attacker could trigger an integer overflow in that backend. A successful exploit might lead to code execution, data tampering, or denial of service.
NVD/CVE DatabaseCVE-2026-24213: NVIDIA Triton Inference Server out-of-bounds read in DALI backend
May 20, 2026HighVulnerabilitySecurityCVE-2026-24213NVIDIA Triton Inference Server contains an out-of-bounds read in its DALI backend, tracked as CVE-2026-24213 and classified under CWE-125. According to the source, a successful exploit might lead to code execution, data tampering, denial of service, or information disclosure. NVD has not yet provided an assessment, and the affected software versions are not listed in the source text.
NVD/CVE DatabaseCVE-2026-24210: NVIDIA Triton Inference Server integer overflow leading to denial of service
May 20, 2026HighVulnerabilitySecurityCVE-2026-24210NVIDIA Triton Inference Server contains an integer overflow vulnerability, classified as CWE-190 (Integer Overflow or Wraparound). An attacker could trigger it, and a successful exploit might lead to denial of service. NVD has not yet provided an assessment, and the record was published 05/20/2026.
NVD/CVE DatabaseCVE-2026-24209: NVIDIA Triton Inference Server path traversal leading to denial of service
May 20, 2026HighVulnerabilitySecurityCVE-2026-24209NVIDIA Triton Inference Server contains a path traversal vulnerability, classified as CWE-22, that an attacker could trigger. The source states that a successful exploit might lead to denial of service. NVD has not yet provided an assessment, and the record was published on 05/20/2026.
NVD/CVE DatabaseCVE-2026-24208: NVIDIA Triton Inference Server path traversal leading to denial of service
May 20, 2026MediumVulnerabilitySecurityCVE-2026-24208NVIDIA Triton Inference Server contains a path traversal vulnerability, classified as CWE-22, that an attacker could trigger. A successful exploit might lead to denial of service. NIST has not yet provided an NVD assessment, and the source does not list affected versions.
NVD/CVE DatabaseCVE-2026-24207: NVIDIA Triton Inference Server authentication bypass
May 20, 2026CriticalVulnerabilitySecurityCVE-2026-24207NVIDIA Triton Inference Server contains an authentication bypass flaw, tracked as CVE-2026-24207 and classified as CWE-288 (Authentication Bypass Using an Alternate Path or Channel). A successful exploit might lead to code execution, escalation of privileges, data tampering, denial of service, or information disclosure. The NVD has not yet provided its own assessment.
NVD/CVE DatabaseCVE-2026-24206: NVIDIA Triton Inference Server authentication bypass via alternate path
May 20, 2026HighVulnerabilitySecurityCVE-2026-24206NVIDIA Triton Inference Server contains an authentication bypass vulnerability, tracked as CVE-2026-24206 and classified under CWE-288 (Authentication Bypass Using an Alternate Path or Channel). The source states that a successful exploit might lead to privilege escalation, denial of service, or information disclosure. NVD had not yet provided an assessment at the time of publication.
NVD/CVE DatabaseGHSA-v6qf-75pr-p96m: Open WebUI: Authenticated users can bypass model access control via exposed query parameter [AI-ASSISTED]
May 14, 2026MediumVulnerabilitySecurityCVE-2026-45365GHSA-v6qf-75pr-p96m affects Open WebUI. The bypass_filter parameter in the generate_chat_completion handlers in routers/openai.py and routers/ollama.py is bound from the query string by FastAPI, so an authenticated user can append ?bypass_filter=true to /openai/chat/completions or /ollama/api/chat. This skips the model access control check, letting any authenticated user invoke admin-restricted models through the server's configured connections.
GitHub Advisory DatabaseCVE-2026-8597: Amazon SageMaker Python SDK code execution via pickle model artifacts in S3
May 14, 2026HighVulnerabilitySecurityCVE-2026-8597Missing integrity verification in the Triton inference handler in Amazon SageMaker Python SDK v2 before v2.257.2 and v3 before v3.8.0 lets a remote authenticated actor with S3 write access to the model artifact path replace model artifacts with a crafted pickle payload. The payload is deserialized without verification, allowing code execution in inference containers.
Fix: Upgrade to Amazon SageMaker Python SDK v2.257.2 or v3.8.0, and rebuild any Triton models previously created with ModelBuilder using the updated SDK.
NVD/CVE DatabaseOllama Out-of-Bounds Read Vulnerability Allows Remote Process Memory Leak
May 10, 2026MediumNewsSecurityResearchers disclosed CVE-2026-7482 (CVSS 9.1), a heap out-of-bounds read in the GGUF model loader of Ollama before 0.17.1, codenamed Bleeding Llama by Cyera. A remote, unauthenticated attacker can upload a crafted GGUF file with an inflated tensor shape to the /api/create endpoint, leaking process memory such as environment variables, API keys, system prompts and other users' conversation data, which can then be pushed to an attacker-controlled registry via /api/push. The flaw likely affects over 300,000 servers.
Fix: Users are advised to apply the latest fixes, limit network access, audit running instances for internet exposure, and isolate and secure them behind a firewall. Deploying an authentication proxy or API gateway in front of all Ollama instances is also recommended, as the REST API does not provide authentication out of the box.
The Hacker NewsOllama vulnerability highlights danger of AI frameworks with unrestricted access
May 7, 2026MediumNewsSecurityIndustryCyera researchers disclosed CVE-2026-7482, dubbed Bleeding Llama, an out-of-bounds heap read in Ollama's model quantization pipeline triggered when GGUF files declare larger tensor sizes than their data. An unauthenticated attacker can upload a crafted file to the Ollama API endpoint and leak process memory, including system prompts, user messages and environment variables. The researchers estimate about 300,000 Ollama servers are exposed on the public internet.
Fix: Update to Ollama version 0.17.1, which includes a patch for this vulnerability. Deploy an authentication proxy or API gateway in front of all Ollama instances, and never expose them to the internet without IP access filters and firewalls. If an internet-accessible server was exposed, rotate API keys, tokens and credentials immediately. Isolate Ollama servers on local networks behind firewalls on secure network segments.
CSO OnlineGHSA-83vm-p52w-f9pw: vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters
May 6, 2026MediumVulnerabilitySecurityCVE-2026-44223A shape regression in vLLM's extract_hidden_states speculative decoding proposer, introduced when the KV connector interface was refactored in PR #37013, causes a RuntimeError that crashes the EngineCore process. Any request containing repetition_penalty, frequency_penalty, or presence_penalty triggers the crash, affecting deployments from v0.18.0 through v0.19.1 inclusive.
Fix: Fixed in PR #38610, first included in vLLM v0.20.0 (the fix slices the return value to sampled_token_ids[:, :1]). Workarounds: upgrade to vLLM v0.20.0 or later; if upgrading is not possible, avoid extract_hidden_states as the speculative decoding method on affected versions; or reject or strip the penalty parameters from incoming requests at an API gateway before they reach vLLM.
GitHub Advisory DatabaseGHSA-hpv8-x276-m59f: vLLM Vulnerable to Remote DoS via Special-Token Placeholders
May 5, 2026MediumVulnerabilitySecurityCVE-2026-44222Unauthenticated text-only prompts in vLLM that spell special tokens such as <|image_pad|> are interpreted as control tokens. When image or video placeholders arrive without matching data, get_input_positions_tensor in vllm/model_executor/layers/rotary_embedding.py indexes empty image_grid_thw/video_grid_thw grids without a bounds check, raising an unhandled IndexError that can terminate the worker. Reproduced on vLLM 0.10.0 with Qwen2.5-VL; a single request can trigger it.
Fix: Fixes: Changes associated with https://github.com/vllm-project/vllm/issues/32656
GitHub Advisory DatabaseCritical Bug Could Expose 300,000 Ollama Deployments to Information Theft
May 5, 2026MediumNewsSecuritySafetyCyera reports that a heap out-of-bounds read in Ollama's GGUF model loader, tracked as CVE-2026-7482 (CVSS 9.3) and dubbed Bleeding Llama, can be exploited without authentication to read sensitive heap data such as prompts, messages and environment variables. An attacker supplies a GGUF file with a tensor offset and size larger than the file, then uses Ollama's model push feature to exfiltrate the result, requiring three unauthenticated API calls. Cyera estimates about 300,000 Ollama servers are exposed on the public internet.
Fix: Fixed in Ollama version 0.17.1. Apply the fix as soon as possible, restrict network access to deployments, deploy an authentication proxy and network segmentation, and audit running instances for internet exposure.
SecurityWeekCVE-2026-7482: Ollama heap out-of-bounds read in GGUF model loader via /api/create
May 4, 2026CriticalVulnerabilitySecurityCVE-2026-7482CVE-2026-7482 affects Ollama before 0.17.1. A heap out-of-bounds read in the GGUF model loader lets an attacker send a crafted GGUF file to the /api/create endpoint, where declared tensor offset and size exceed the file's real length. During quantization in fs/ggml/gguf.go and server/quantization.go (WriteTo()), the server reads past the allocated heap buffer, and the leaked memory may include environment variables, API keys, system prompts and other users' conversation data. The attacker can exfiltrate it by pushing the resulting model artifact through /api/push to an attacker-controlled registry, and neither endpoint requires authentication in the upstream distribution.
NVD/CVE DatabaseCVE-2026-42249: Ollama for Windows remote code execution through update mechanism path traversal
Apr 29, 2026HighVulnerabilitySecurityCVE-2026-42249CVE-2026-42249 affects Ollama for Windows. Its update mechanism passes values from attacker-controlled HTTP response headers directly to filepath.Join without validation, so ../ sequences let files be written outside the update staging directory, including the Windows Startup directory. Chained with CVE-2026-42248 (Missing Signature Verification for Updates), this yields automatic, persistent code execution without user awareness. Versions 0.12.10 to 0.17.5 were tested and confirmed vulnerable; other versions were not tested.
NVD/CVE DatabaseCVE-2026-42248: Ollama for Windows update executables accepted without integrity verification
Apr 29, 2026HighVulnerabilitySecurityCVE-2026-42248Ollama for Windows does not verify the integrity or authenticity of downloaded update executables. Its Windows update verification routine unconditionally returns success, so no digital signature or trust check runs before update payloads are staged or executed, letting attacker-supplied executables be accepted and run. Because updates install silently, a malicious payload may be installed without user awareness. Versions 0.12.10 to 0.17.5 were tested and confirmed vulnerable; other versions were not tested but might also be affected.
NVD/CVE DatabaseCVE-2026-7141: vllm uninitialized resource in has_mamba_layers of KV Block Handler
Apr 27, 2026MediumVulnerabilitySecurityCVE-2026-7141A vulnerability in vllm up to 0.19.0 affects the function has_mamba_layers in vllm/v1/kv_cache_interface.py, part of the KV Block Handler component. Manipulating it leads to an uninitialized resource, and the attack can be launched remotely. The attack is rated high complexity and difficult to exploit, but a public exploit exists.
Fix: To fix this issue, it is recommended to deploy a patch. The patch is named 1ad67864c0c20f167929e64c875f5c28e1aad9fd.
NVD/CVE DatabaseCVE-2026-7020: Ollama path traversal in digest handling of Tensor Model Transfer
Apr 26, 2026MediumVulnerabilitySecurityCVE-2026-7020A path traversal flaw has been reported in Ollama up to 0.20.2, in the digestToPath function of x/imagegen/transfer/transfer.go within the Tensor Model Transfer Handler component. Manipulating the digest argument enables the traversal, and the attack can be launched remotely but has high complexity and is reported as difficult to exploit. A public exploit exists, and the vendor did not respond to early disclosure.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.