Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
222 items
CVE-2026-103663: Ollama path traversal in layer digest validation of /api/pull
Oct 8, 2026CriticalVulnerabilitySecurityCVE-2026-103663Ollama's /api/pull endpoint is vulnerable to path traversal because the digestToPath function does not sufficiently validate layer digests. An unauthenticated remote attacker can supply a traversal sequence as a layer digest to write a malicious binary outside the model store. Where the server process can write to /usr/lib/ollama, the file is loaded and executed on the next restart, giving remote code execution as root.
Fix: Fixed in version 0.35.0.
NVD/CVE DatabaseCVE-2026-105922: vLLM denial of service in Penalty Handler token bin counting
Oct 6, 2026MediumVulnerabilitySecurityCVE-2026-105922A security flaw in vllm-project vLLM up to 0.31.0 affects the function get_token_bin_counts_and_mask in vllm/model_executor/layers/utils.py, part of the Penalty Handler component. Remote exploitation leads to denial of service, and a public exploit exists. The project was notified through an issue report but has not yet responded.
NVD/CVE DatabaseCVE-2026-105775: vLLM out-of-bounds read in conv_ssm_forward via Completions Request Handler
Oct 6, 2026MediumVulnerabilitySecurityCVE-2026-105775A security vulnerability in vllm-project vLLM up to 0.31.0 affects the function conv_ssm_forward in vllm/model_executor/layers/mamba/mamba_mixer2.py, within the Completions Request Handler component. Manipulation of this component leads to an out-of-bounds read, and the attack can be carried out remotely. The exploit has been publicly disclosed and may be used, and the project was informed through an issue report but has not yet responded.
NVD/CVE DatabaseGHSA-x6mc-67gf-chw4: vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105758An unauthenticated attacker can exhaust memory in the vLLM API-server process on Qwen2-VL and Qwen3-VL deployments by raising the request-level media_io_kwargs.video.max_frames and fps values. The Qwen video samplers never read num_frames, so the num_frames ceiling in PR #51969 does not stop this path, while the same knobs were already capped for GLMGAVideoBackend in commit 8b6de0eb9. In the reporter's test, 74 extra bytes of JSON sent over POST /tokenize raised peak RSS from 2 271 MiB to 13 629 MiB.
GitHub Advisory DatabaseGHSA-58v5-2m8f-94pr: vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105760vLLM's OpenAI-compatible chat endpoint accepts request-level `media_io_kwargs`, so a caller can select the GLMGA video sampler and set large `fps` and `max_frames` values. GLMGA builds and deduplicates an attacker-sized index list before frame reads, so a tiny valid video with compact JSON options can consume disproportionate CPU and memory in the shared media-loading executor and delay unrelated media requests. The demonstrated impact is partial denial of service, with no memory corruption, data disclosure or code execution claimed.
Fix: Validate request-level video sampling options before dispatching work to the media executor. Enforce conservative absolute limits for `fps`, `max_frames`, and especially the computed candidate count.
GitHub Advisory DatabaseGHSA-ph72-cqr5-qpp7: vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105754vLLM versions up to and including 0.25.1 are affected through the disaggregated scale-out transport. The POST /inference/v1/generate route decodes a caller-supplied features object (kwargs_data, mm_hashes, mm_placeholders) and passes it to the engine as if it came from the trusted render step, without validating it against the active model's renderer output. Any authenticated caller can forge one field to crash the shared EngineCore process, poison or disclose cross-request encoder-cache state, or drop placeholder masks during replay.
GitHub Advisory DatabaseGHSA-85xf-c7hm-whqw: vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105757vLLM versions up to and including 0.25.1 contain three structured-output request paths that raise an uncaught, engine-fatal exception instead of a per-request validation error. The exception escapes into EngineCore's busy loop and triggers a fatal _send_engine_dead(), so one malformed structured-output request denies service to all concurrent and subsequent users of that engine. Site 1 is a valid duplicate-root EBNF grammar that the latched xgrammar backend fails to compile, and the completion check catches only TimeoutError. The advisory states that these sites are distinct from the fixes in GHSA-6qc9-v4r8-22xg and GHSA-8wr5-jm2h-8r4f.
GitHub Advisory DatabaseGHSA-2phq-3phc-84px: vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integrity break and induced errors on `/score` and `/rerank`
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105755vLLM versions up to and including 0.25.1 cache late-interaction query embeddings on `/score` and `/rerank` under a key derived from the caller-supplied `X-Request-Id` header when flash late interaction is enabled, which is the default for supported models. A concurrent request reusing the victim's header value replaces the victim's cached query embedding, so the victim's documents are scored against the attacker's query, and one request can also force the other into a late-interaction cache-miss error. The source states the lower bound of affected versions is not confirmed and maintainers can confirm it.
GitHub Advisory DatabaseGHSA-2823-qmq8-rwvj: vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP deployments — uncaught downstream `ValueError` denial of service
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105756vLLM's OpenAI-compatible request models (Completions, Chat Completions, Responses) accept a client-supplied `cache_salt` validated only as a non-empty string. With the built-in LMCache-MP KV connector enabled, the value reaches the scheduler's cache lookup unguarded, and the downstream LMCache `IPCCacheServerKey.__post_init__` raises an uncaught `ValueError` for `@`, `/`, `\`, NUL or lengths over 128 characters. Because the exception is never caught on the scheduling path, a single request such as `cache_salt="/"` kills EngineCore for all users.
GitHub Advisory DatabaseCVE-2026-105759: vLLM Rust frontend memory exhaustion via unique HTTP method labels in metrics
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105759vLLM, an inference and serving engine for large language models, is affected before version 0.30.0 by a flaw in the Rust frontend's track_http_metrics middleware. The middleware records the raw HTTP method token as a Prometheus label on requests reaching registered routes, so an unauthenticated attacker sending unique arbitrary method tokens to unguarded routes such as /tokenize can permanently create label sets. This increases process memory usage and enlarges the /metrics response until the service or monitoring path is exhausted.
Fix: Fixed in version 0.30.0.
NVD/CVE DatabaseCVE-2026-105753: vLLM denial of service via multimodal cache hash reuse
Oct 5, 2026MediumVulnerabilitySecurityCVE-2026-105753CVE-2026-105753 affects vLLM, an inference and serving engine for large language models, prior to 0.28.0. The default mirrored multimodal LRU cache can record a media hash in the frontend sender cache before engine admission, while the engine receiver cache never gets the payload if the request is rejected. A later request reusing that media hash makes the receiver cache hit an assertion, causing an availability failure in the shared service.
Fix: Fixed in version 0.28.0.
NVD/CVE DatabaseCVE-2026-105752: vLLM prefix cache isolation bypass in Harmony tool continuations
Oct 5, 2026LowVulnerabilitySecurityCVE-2026-105752vLLM versions prior to 0.30.0 drop the cache_salt value when rebuilding the next-turn engine input for Harmony tool continuations submitted through POST /v1/responses. The continuation prefix lands in the global unsalted cache namespace even when the caller enabled salting, and on deployments with prefix caching enabled (the default), an authenticated tenant can use the cached_tokens_per_turn count to determine whether a victim's low-entropy post-tool prefix was previously processed, defeating salted prefix cache isolation.
Fix: Fixed in version 0.30.0.
NVD/CVE DatabaseCVE-2026-103241: vLLM denial of service through Gemma4UnifiedParser
Sep 30, 2026MediumVulnerabilitySecurityCVE-2026-103241A flaw in the Gemma4UnifiedParser component of vllm-project vLLM up to 0.26.0, located in rust/src/parser/src/unified/gemma4.rs, can be triggered remotely through a manipulated input to cause a denial of service. Public exploit code has been published, so the flaw can be used by attackers.
Fix: Upgrade to version 0.29.1rc0. The patch is commit 3439bad37e68ba9755a46f4f6b44a4aeaf1f60a9. Upgrading the affected component is advised.
NVD/CVE DatabaseCVE-2026-102697: Ollama agent mode Bash tool approval bypass via shell operators
Sep 29, 2026HighVulnerabilitySecurityIndustryCVE-2026-102697Ollama versions 0.14.0 before 0.31.2 contain an incorrect authorization flaw in the experimental agent mode Bash tool approval mechanism, which fails to properly parse shell syntax. An attacker who can influence model output through prompt injection can append control operators such as semicolons or logical operators to an approved command, executing additional shell commands and bypassing the session approval requirement.
NVD/CVE DatabaseGHSA-456v-xq2p-r4cj: code-ollama: `grep_search` Command Injection via Unescaped `$()` Shell Substitution (CWE-78)
Sep 28, 2026HighVulnerabilitySecurityThe `grep_search` tool in `code-ollama` builds a shell command by interpolating attacker-controlled `pattern` and `path` arguments and runs it through `child_process.exec()`. Its sanitization escapes only backslashes and double quotes, so `$()` and backtick substitution pass through, allowing arbitrary OS command execution with the privileges of the local user. Because `grep_search` is treated as read-only, it runs automatically in Plan mode without an approval prompt.
GitHub Advisory DatabaseCVE-2026-100654: vLLM denial of service via stop_token_ids in completion endpoints
Sep 26, 2026MediumVulnerabilitySecurityCVE-2026-100654vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints, validating only that the values are integers rather than that each id falls within the model vocabulary or logits range. When min_tokens is greater than 0, an out-of-range id reaches a CUDA index_put_ operation and triggers a device-side assertion. An authenticated API user can send one malformed request that returns 500 Internal Server Error and leaves EngineCore in a fatal state, so later requests fail until the service restarts (denial of service).
NVD/CVE DatabaseCVE-2026-100653: vLLM unpinned Hugging Face artifact loads for FunAudioChat and Tarsier2
Sep 26, 2026MediumVulnerabilitySecurityIndustryCVE-2026-100653CVE-2026-100653 affects vLLM versions 0.22.1 through 0.28.0. The operator-supplied model revision pin (--revision / --code-revision) is not applied to several Hugging Face artifact loads for the FunAudioChat and Tarsier2 architectures, including the WhisperFeatureExtractor and speech_tokenizer loads in vllm/model_executor/models/funaudiochat.py and the Qwen2VLConfig.from_pretrained call in vllm/model_executor/models/qwen2_vl.py. Pinned deployments therefore resolve these artifacts from the repository's default revision, so upstream changes can alter audio preprocessing, speech tokenizer behavior, or Tarsier2 configuration without any change to the pin.
Fix: Fixed in 0.28.0.
NVD/CVE DatabaseCVE-2026-100652: vLLM stop_token_ids validation flaw in Rust HTTP and gRPC frontends
Sep 26, 2026MediumVulnerabilitySecurityCVE-2026-100652vLLM versions 0.22.0 through 0.23.0 do not validate stop_token_ids against vocabulary bounds in the Rust HTTP and gRPC frontends, so out-of-vocabulary token IDs reach MinTokensLogitsProcessor. An attacker can send requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures, which leave EngineCore in a fatal state that requires a service restart.
NVD/CVE DatabaseCVE-2026-100651: vLLM denial of service via overlong prompt on disaggregated generate endpoint
Sep 26, 2026MediumVulnerabilitySecurityCVE-2026-100651vLLM versions before 0.29.0 do not enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate when a request includes a 'features' multimodal payload. The GenerateRequest.token_ids list is not checked against model_config.max_model_len, and processors reporting skip_prompt_length_check=True (such as Nemotron Parse, Whisper, and FireRedLID) bypass validation entirely. A client that can reach the endpoint on an affected configuration can submit an overlong token_ids list and crash the worker, causing denial of service.
Fix: Fixed in 0.29.0.
NVD/CVE DatabaseCVE-2026-100650: vLLM media fetching exhausts memory and bandwidth before size limits apply
Sep 26, 2026MediumVulnerabilitySecurityCVE-2026-100650vLLM through 0.29.0 fetches and fully reads remote or inline media before enforcing its documented media controls, the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB audio size cap and the --limit-mm-per-prompt item limits. Across four ingress paths, including the chat completions audio_url and base64 path and the unauthenticated Rust frontend POST /tokenize route, a remote attacker can force memory and outbound bandwidth use proportional to an attacker-chosen body size or media item count before rejection. The result is pre-inference memory and bandwidth exhaustion, a denial of service. The source states no code execution or data disclosure impact.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.