Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2026-73557: vLLM race condition bypasses sparse tensor validation in prompt_embeds
Aug 13, 2026HighVulnerabilitySecurityCVE-2026-73557vLLM, an inference and serving engine for large language models, contains a flaw from 0.20.2rc0 through 0.26.0 in safe_load_prompt_embeds in vllm/renderers/embed_utils.py. The function toggles the process-global torch.sparse.check_sparse_tensor_invariants setting, and concurrent prompt_embeds parts submitted to POST /v1/chat/completions can race that state, letting an invalid sparse tensor reach tensor.to_dense despite the CVE-2025-62164 guard when enable_prompt_embeds is enabled.
Fix: This issue is fixed in version 0.26.0.
NVD/CVE DatabaseCVE-2026-73556: vLLM regex denial of service via structured_outputs.regex on /v1/completions
Aug 13, 2026MediumVulnerabilitySecurityCVE-2026-73556CVE-2026-73556 affects vLLM, an inference and serving engine for large language models, before version 0.26.0. The structured_outputs.regex parameter in vllm/v1/structured_output/backend_lm_format_enforcer.py is passed to lmformatenforcer.RegexParser without compile_regex_with_timeout or validation. An unauthenticated /v1/completions request against the lm-format-enforcer backend can consume a CPU core and stall the structured-output engine path with a catastrophic regular expression.
Fix: Fixed in 0.26.0.
NVD/CVE DatabaseCVE-2026-73555: vLLM information disclosure via malformed JSON requests
Aug 13, 2026MediumVulnerabilitySecurityCVE-2026-73555vLLM versions prior to 0.26.0 leak internal details in validation error responses. The validation_exception_handler in vllm/entrypoints/openai/server_utils.py converts RequestValidationError objects with str(exc), and sanitize_message in vllm/entrypoints/utils.py fails to strip traceback-style file paths. An unauthenticated attacker sending malformed JSON to /v1/chat/completions, /v1/completions, /tokenize, or /detokenize can learn the OS username, home and virtual-environment paths, Python version, internal package structure, line numbers, and endpoint handler names.
Fix: Fixed in 0.26.0.
NVD/CVE DatabaseCVE-2026-27765: vLLM Hardware Plugin for Intel Gaudi denial of service
Aug 11, 2026MediumVulnerabilitySecurityCVE-2026-27765Improper input validation in some vLLM Hardware Plugin for Intel(R) Gaudi(R) software before version 0.16.0, within Ring 3: User Applications, may allow a denial of service. An authenticated adversary can exploit it with a low complexity attack via local access, requiring no user interaction. The impact is high availability on the vulnerable system, with no confidentiality or integrity impact.
Fix: Fixed in version 0.16.0.
NVD/CVE DatabaseCVE-2026-19334: NightTrek Ollama-mcp command injection through argument handling in src/index.ts
Aug 9, 2026MediumVulnerabilitySecurityCVE-2026-19334CVE-2026-19334 is a flaw in NightTrek Ollama-mcp up to commit 80cf2e17cfc144963a475b619093a2d13c13dbc9, located in an unspecified part of src/index.ts. Manipulating the name, modelfile, source or destination argument causes command injection. An attacker must have local access to exploit it. The product uses a rolling release, so no affected or fixed version numbers are available, and the project has not yet responded to an early issue report.
NVD/CVE DatabaseCVE-2026-47487: NVIDIA Triton Inference Server path traversal via model name in MLflow plugin
Aug 4, 2026MediumVulnerabilitySecurityCVE-2026-47487NVIDIA Triton Inference Server for Linux contains a path traversal flaw (CWE-22). A user can supply a path in the model name to the Triton MLflow plugin and reach files outside the model repository, allowing reads, writes or modifications. A successful exploit might lead to denial of service and information disclosure.
NVD/CVE DatabaseCVE-2026-65315: Ollama uncontrolled memory allocation in GGUF metadata parser via crafted file
Jul 21, 2026HighVulnerabilitySecurityCVE-2026-65315CVE-2026-65315 affects Ollama at HEAD f0078ae. The GGUF metadata parser allocates memory based on attacker-controlled length, count and dimension fields without checking them against the remaining file size. A remote attacker can upload a crafted GGUF file under 1KB through the blob upload, model create or pull API endpoints, causing an unrecoverable Go out-of-memory fatal error or makeslice panic. Because these bypass recovery middleware, the entire server process crashes.
NVD/CVE DatabaseCVE-2026-63086: text-generation-inference SSRF through image_url in chat completions endpoint
Jul 16, 2026HighVulnerabilitySecurityCVE-2026-63086CVE-2026-63086 affects text-generation-inference through 3.3.7. Its OpenAI-compatible multimodal chat completions endpoint lacks validation in fetch_image (router/src/validation.rs), so an unauthenticated attacker can supply a crafted image_url and make the server send arbitrary HTTP GET requests. Because the reqwest client follows redirects by default, redirect chains can bypass scheme checks and reach private, loopback, link-local and cloud metadata addresses for internal port scanning and credential theft.
NVD/CVE DatabaseCVE-2026-47475: NVIDIA TensorRT-LLM reachable assertion in OpenAI-compatible inference API
Jul 14, 2026MediumVulnerabilitySecurityCVE-2026-47475CVE-2026-47475 affects NVIDIA TensorRT-LLM's OpenAI-compatible inference API. An attacker can trigger a reachable assertion, classified as CWE-617, in the sampler thread. A successful exploit might lead to denial of service.
NVD/CVE DatabaseCVE-2026-24271: NVIDIA TensorRT-LLM resource allocation flaw in OpenAI-compatible inference API
Jul 14, 2026MediumVulnerabilitySecurityCVE-2026-24271NVIDIA TensorRT-LLM contains a vulnerability in its OpenAI-compatible inference API, tracked as CVE-2026-24271 and classified as CWE-770 (Allocation of Resources Without Limits or Throttling). An attacker could cause allocation of GPU resources without limits or throttling, which might lead to denial of service. The NVD assessment has not yet been provided, and the record was published 07/14/2026.
NVD/CVE DatabaseCVE-2026-24233: NVIDIA TensorRT-LLM for Linux unsafe deserialization in model weight loading
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-24233NVIDIA TensorRT-LLM for Linux contains a vulnerability in the restricted unpickler used for model weight deserialization. A local, unauthenticated attacker could cause deserialization of untrusted data (CWE-502), which might lead to code execution, escalation of privileges, data tampering, and information disclosure.
NVD/CVE DatabaseCVE-2026-47482: NVIDIA Triton Inference Server for Linux memory leak causing denial of service
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-47482CVE-2026-47482 affects NVIDIA Triton Inference Server for Linux. An attacker can cause missing release of memory after effective lifetime, classified as CWE-401. A successful exploit might lead to denial of service. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-47481: NVIDIA Triton Inference Server for Linux contains a vulnerability where an attacker can cause an authentication bypass…
Jul 14, 2026MediumVulnerabilitySecurityCVE-2026-47481CVE-2026-47481 affects NVIDIA Triton Inference Server for Linux and is classified as CWE-288, Authentication Bypass Using an Alternate Path or Channel. An attacker can bypass authentication through an alternative path or channel. A successful exploit might lead to code execution, escalation of privileges, information disclosure, and data tampering.
NVD/CVE DatabaseCVE-2026-47480: NVIDIA Triton Inference Server for Linux uncaught exception denial of service
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-47480NVIDIA Triton Inference Server for Linux contains a vulnerability, CVE-2026-47480, in which an attacker can cause an uncaught exception. The source states that a successful exploit might lead to denial of service. The weakness is classified as CWE-248 (Uncaught Exception), and NVD had not yet provided an assessment at publication on 07/14/2026.
NVD/CVE DatabaseCVE-2026-47479: NVIDIA Triton Inference Server for Linux uncontrolled resource consumption
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-47479NVIDIA Triton Inference Server for Linux contains a vulnerability, tracked as CVE-2026-47479 and classified as CWE-400 (Uncontrolled Resource Consumption). An attacker can cause uncontrolled resource consumption, which might lead to denial of service. NIST has not yet provided an NVD assessment, and the CVE was published on 07/14/2026.
NVD/CVE DatabaseCVE-2026-47478: NVIDIA Triton Inference Server for Linux use of expired file descriptor
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-47478NVIDIA Triton Inference Server for Linux contains a vulnerability, tracked as CVE-2026-47478 and classified as CWE-910 (Use of Expired File Descriptor). An attacker can cause the server to use an expired file descriptor, which might lead to denial of service.
NVD/CVE DatabaseCVE-2026-47477: NVIDIA Triton Inference Server for Linux stack-based buffer overflow
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-47477CVE-2026-47477 affects NVIDIA Triton Inference Server for Linux. An attacker can cause a stack-based buffer overflow (CWE-121) in it, which might lead to denial of service. NVD has not yet provided an assessment, and the source gives no affected versions.
NVD/CVE DatabaseCVE-2026-47476: NVIDIA Triton Inference Server for Linux uncontrolled resource consumption
Jul 14, 2026HighVulnerabilitySecurityCVE-2026-47476NVIDIA Triton Inference Server for Linux contains a vulnerability, CVE-2026-47476, classified as CWE-400 Uncontrolled Resource Consumption. An attacker can cause uncontrolled resource consumption, which might lead to denial of service. The NVD published the entry on 07/14/2026.
NVD/CVE DatabaseCVE-2026-15685: Ollama downloadBlob improper array index validation denial of service
Jul 13, 2026MediumVulnerabilitySecurityCVE-2026-15685CVE-2026-15685 is an improper validation of array index flaw in the downloadBlob function of Ollama. Remote attackers can exploit it without authentication to cause a denial-of-service condition, because user-supplied data is not properly validated, allowing a memory access past the end of an allocated array.
NVD/CVE DatabaseCVE-2026-15574: vllm-orchestrator-gateway logs authorization headers and chat payloads
Jul 13, 2026HighVulnerabilitySecurityPrivacyCVE-2026-15574CVE-2026-15574 affects the vllm-orchestrator-gateway component. Its production binary writes all incoming authorization headers and full chat payloads to persistent logs, including bearer tokens, personally identifiable information and secrets. Any user with logging privileges can read this data, which may let an attacker harvest credentials and conversation content.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.