Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2026-55514: vLLM crash via pure prompt embeds in /v1/completions requests
Jul 6, 2026HighVulnerabilitySecurityCVE-2026-55514CVE-2026-55514 affects vLLM, a library for LLM inference and serving, in versions from 0.12.0 to before 0.24.0. Sending a pure prompt embeds payload in a /v1/completions request to a model using M-RoPE makes EngineCore fail an assertion and crash, shutting down the entire server application. Any remote user authorized to make a /v1/completions request can trigger it.
Fix: This issue is fixed in version 0.24.0.
NVD/CVE DatabaseCVE-2026-54234: vLLM denial of service through speculative decoding over gRPC Generate
Jul 6, 2026HighVulnerabilitySecurityCVE-2026-54234CVE-2026-54234 affects vLLM before 0.24.0, an inference and serving engine for LLMs. A frontend-legal multi-request speculative decoding workload makes the rejection sampler emit a recovered token at the vocabulary size boundary, which becomes -1 and reaches the model's embedding and attention path, crashing the engine worker with a GPU device-side assertion. A remote client that can send requests through the public gRPC Generate and Abort endpoints can trigger this, aborting concurrent requests and causing a service-wide denial of service until the worker restarts.
Fix: Fixed in version 0.24.0.
NVD/CVE DatabaseCVE-2026-55646: vLLM audio transcription routes allow memory exhaustion via oversized uploads
Jul 6, 2026MediumVulnerabilitySecurityCVE-2026-55646From vLLM 0.22.0 to 0.23.0, the /v1/audio/transcriptions and /v1/audio/translations routes call request.file.read() and fully load an uploaded audio file into memory before the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB size limit (default 25 MB) is checked during speech-to-text preprocessing. An API caller who can reach these routes can submit an oversized multipart upload, causing memory allocation proportional to the file size before rejection, which can create memory pressure or terminate the process depending on deployment resource limits.
Fix: This issue is fixed in version 0.24.0.
NVD/CVE DatabaseCVE-2026-24266: NVIDIA Triton Inference Server for Linux use-after-free issue
Jul 1, 2026MediumVulnerabilitySecurityCVE-2026-24266NVIDIA Triton Inference Server for Linux contains a vulnerability, CVE-2026-24266, where an attacker can cause a use-after-free issue (CWE-416). A successful exploit might lead to denial of service. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-24264: NVIDIA Triton Inference Server for Linux improper handling of compressed data
Jul 1, 2026HighVulnerabilitySecurityCVE-2026-24264CVE-2026-24264 affects NVIDIA Triton Inference Server for Linux. An attacker can cause improper handling of highly compressed data, classified as CWE-409 (Improper Handling of Highly Compressed Data, Data Amplification). A successful exploit might lead to denial of service. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-54021: Open WebUI Ollama proxy routes let users reach unauthorized backends via url_idx
Jun 23, 2026MediumVulnerabilitySecurityCVE-2026-54021Open WebUI, a self-hosted AI platform, prior to 0.9.6 has several direct, index-addressed Ollama proxy routes that accept a caller-supplied url_idx path parameter and use it as a raw index into the admin-configured OLLAMA_BASE_URLS list. Access control checks only whether the user may use the requested model, not which backend receives the request, so any authenticated user can route requests to Ollama backends they were never authorized to reach, including internal, higher-privilege, or admin-disabled ones.
Fix: Fixed in 0.9.6.
NVD/CVE DatabaseCVE-2026-56340: vLLM multimodal embeddings processing missing sparse tensor validation
Jun 20, 2026HighVulnerabilitySecurityCVE-2026-56340vLLM versions 0.10.2 up to but not including 0.13.0 lack sparse tensor validation in multimodal embeddings processing. Because PyTorch disables sparse tensor invariant checks by default, crafted embedding requests with malformed (negative or out-of-bounds) tensor indices can crash the server or exhaust resources when the prompt-embeds feature is enabled. The source also notes potential out-of-bounds write (write-what-where) memory corruption. This continues CVE-2025-62164, whose earlier fix only disabled the feature by default.
NVD/CVE DatabaseCVE-2025-71379: vLLM regular expression denial of service in multiple components
Jun 20, 2026MediumVulnerabilitySecurityCVE-2025-71379CVE-2025-71379 affects vLLM versions >= 0.6.3 and < 0.9.0, which contain multiple regular expression denial of service (ReDoS) vulnerabilities. Regex patterns in vllm/lora/utils.py, the phi4mini tool parser, and the OpenAI-compatible serving chat endpoint are susceptible to catastrophic backtracking. An attacker submitting crafted input with nested or repeated structures can trigger severe CPU consumption and performance degradation, resulting in denial of service.
NVD/CVE DatabaseGHSA-6pr9-rp53-2pmc: vLLM: OOM Denial of Service via Audio Decompression Bomb
Jun 17, 2026MediumVulnerabilitySecurityCVE-2026-54233vLLM's `/v1/audio/transcriptions` endpoint limits the compressed upload size but not the decoded PCM output, so a 25MB OPUS file expands to about 14.9GB of float32 PCM at decode time, tested on v0.19.0. An unauthenticated attacker can exhaust server memory with a few concurrent requests, each a valid upload within the documented limit, because the audio decoder in `audio.py` accumulates all frames without a size check and `np.concatenate` allocates a second contiguous array.
Fix: A fix was merged in https://github.com/vllm-project/vllm/pull/44970
GitHub Advisory DatabaseGHSA-hgg8-fqqc-vfmw: vLLM: incomplete CVE-2026-22778 fix leaks PIL repr addresses via Anthropic router
Jun 17, 2026MediumVulnerabilitySecurityCVE-2026-54236The fix for CVE-2026-22778 / GHSA-4r2x-xpjr-7cvv (PRs #31987 and #32319) added `sanitize_message` at four OpenAI router exception sites, but later vLLM response paths still return `str(exc)` unsanitized. Malformed image bytes make PIL raise `UnidentifiedImageError` with a BytesIO object repr containing a heap address, which these paths echo to clients in `main` HEAD (`771e1e48b`, 2026-05-26). The paths listed are `vllm/entrypoints/anthropic/api_router.py` (lines 78 and 124), `vllm/entrypoints/anthropic/serving.py` (line 808), and `vllm/entrypoints/speech_to_text/realtime/connection.py` (lines 75 and 265).
GitHub Advisory DatabaseGHSA-5jv2-g5wq-cmr4: vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving
Jun 17, 2026MediumVulnerabilitySecurityPrivacyCVE-2026-53923vLLM's GGUF dequantize kernels in csrc/quantization/gguf/gguf_kernel.cu truncate tensor element counts because the to_cuda_ggml_t function pointer takes the count as a 32-bit int, while the caller passes m * n. The output tensors are allocated with torch::empty and left uninitialized, so when m * n exceeds INT_MAX the unfilled portion retains stale GPU memory. In multi-tenant serving this residual memory may contain other users' inference data, a silent confidentiality violation.
Fix: Change the int k parameter in the to_cuda_ggml_t typedef to int64_t k, which all dequantize functions inherit through the same typedef.
Hugging Face Security AdvisoriesGHSA-8jr5-v98p-w75m: vLLM: image EXIF Rotation & PNG tRNS Transparency Not Normalized, Causing Mismatch Between Model Input and Expectations
Jun 17, 2026MediumVulnerabilitySecurityGHSA-8jr5-v98p-w75m affects vLLM's image input handling in vllm/multimodal/image.py. The code never calls ImageOps.exif_transpose, so EXIF orientation is not normalized, and PNGs carrying tRNS transparency in P, L or RGB modes take the image.convert("RGB") path without being flattened against a background, so transparent or overlay content can become visible and change the pixels the model receives. Pillow also loads only the first frame of APNG and GIF files.
Fix: A fix for this vulnerability was merged here: https://github.com/vllm-project/vllm/pull/44974
GitHub Advisory DatabaseGHSA-7h4p-rffg-7823: vLLM: temperature=NaN and temperature=Infinity bypass validation and propagate to GPU kernels
Jun 17, 2026MediumVulnerabilitySecurityCVE-2026-54235vLLM's temperature validation in sampling_params.py uses comparison operators that evaluate to False for NaN and positive Infinity, so both values pass every guard. These values reach GPU sampling kernels, where they cause undefined behavior or CUDA errors that can crash the inference worker and degrade service for all concurrent users. The source notes that -Infinity is correctly rejected.
Fix: Add math.isfinite(self.temperature) check in _verify_args(). Reject non-finite float values with a 400 error. A fix was merged in https://github.com/vllm-project/vllm/pull/45116
GitHub Advisory DatabaseGHSA-94f4-hr76-p5j6: vLLM: OpenAI auth bypass
Jun 16, 2026CriticalVulnerabilitySecurityCVE-2026-48746GHSA-94f4-hr76-p5j6 is a vulnerability in vLLM's OpenAI-compatible API server that allows authentication bypass of the AuthenticationMiddleware. The flaw stems from starlette reconstructing the URL from an unfiltered Host header on ASGI servers such as uvicorn, letting an attacker control url.path and reach /v1 endpoints without the configured VLLM_API_KEY or --api-key. Instances behind an RFC-conforming web server such as nginx are not affected.
GitHub Advisory DatabaseGHSA-q8gq-377p-jq3r: vLLM: Security Check Bypass via assert Statement in Activation Function Loading Allows Arbitrary Code Execution
Jun 16, 2026HighVulnerabilitySecurityCVE-2026-41523An assert-based security check in vLLM's activation function loading, at vllm/model_executor/layers/pooler/activations.py:48, restricts which functions can be loaded from a HuggingFace model's config.json. When vLLM runs in Python optimized mode (python -O or PYTHONOPTIMIZE=1), Python strips the assert, so an attacker-published malicious model can pass an arbitrary function_name to resolve_obj_by_qualname() and execute code during model initialization. The attack requires the victim to load the malicious model and the model to use a cross-encoder architecture.
Fix: Suggested fix: replace the assert with an explicit conditional raise: if not function_name.startswith("torch.nn.modules."): raise ValueError("Loading of activation functions is restricted to torch.nn.modules for security reasons"). The source text ends mid-sentence at "A fix for this", so no released fixed version is stated.
Hugging Face Security AdvisoriesCVE-2026-5497: vLLM out-of-memory denial of service via unbounded video frames in data URLs
Jun 11, 2026HighVulnerabilitySecurityCVE-2026-5497vLLM versions 0.8.0 and later are vulnerable to an Out-of-Memory denial of service through the `VideoMediaIO.load_base64()` method. When processing `video/jpeg` data URLs, the method splits the base64 string on commas without enforcing a frame count limit. A single request to the OpenAI-compatible chat completions API containing thousands of frames can exhaust server memory and crash it, and no authentication is required.
NVD/CVE DatabaseGHSA-3ww4-5jv9-j5gm: vLLM's Artifact Pin Decay allows pinned deployments to load unpinned code, weights, and processors
Jun 10, 2026MediumVulnerabilitySecurityIndustryCVE-2026-47155GHSA-3ww4-5jv9-j5gm reports that vLLM's revision pinning does not consistently cover all artifacts loaded for a model. Deployments that set `--revision` or `--code-revision` can still load dynamic code, GGUF files, image processors, retrieval side weights, or same-repository subfolder weights and config from an unpinned or default revision, so the pin does not describe the full artifact set served. The source states this is not a `trust_remote_code=False` bypass, unauthenticated RCE, or a demonstrated artifact compromise. The strongest example is Kimi-Audio, whose `whisper-large-v3` subfolder loads outside the configured revision.
Hugging Face Security AdvisoriesCVE-2026-4944: vllm-project/vllm remote code execution via trust_remote_code
May 28, 2026HighVulnerabilitySecurityCVE-2026-4944vllm-project/vllm version 0.14.1 hardcodes `trust_remote_code=True` in two model implementation files, `vllm/model_executor/models/nemotron_vl.py` and `vllm/model_executor/models/kimi_k25.py`. This overrides a user's explicit `--trust-remote-code=False` setting, allowing remote code execution through malicious HuggingFace model repositories. The issue is an incomplete fix for CVE-2025-66448 and CVE-2026-22807 and is particularly impactful for deployments loading NemotronVL or KimiK25 models.
NVD/CVE DatabaseCVE-2026-9540: vllm-project vllm denial of service in OpenAI-compatible serving path
May 26, 2026MediumVulnerabilitySecurityCVE-2026-9540CVE-2026-9540 affects vllm-project vllm 0.19.0 and lies in unspecified processing of the OpenAI-compatible Serving Path. Remote manipulation of this component leads to denial of service, and the exploit is publicly available. The CVSS 4.0 base score from CNA VulDB is 5.5 (MEDIUM), and the weakness is CWE-404, Improper Resource Shutdown or Release.
Fix: The pull request to fix this issue awaits acceptance. No fixed version is stated.
NVD/CVE DatabaseCVE-2026-24215: NVIDIA Triton Inference Server DALI backend uncontrolled resource consumption
May 20, 2026MediumVulnerabilitySecurityCVE-2026-24215CVE-2026-24215 affects the DALI backend of NVIDIA Triton Inference Server. An attacker could cause uncontrolled resource consumption, which might lead to denial of service. NVD has not yet provided an assessment, and the source names no affected versions.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.