Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2026-100649: vLLM resource-limit bypass in PyNvVideoCodec decoder allocation
Sep 26, 2026LowVulnerabilitySecurityCVE-2026-100649vLLM before 0.29.0 contains a resource-limit bypass in PyNvVideoCodec decoder allocation, tracked as CVE-2026-100649. Sampler subclass shadowing lets requests increment counters independently, so an unauthenticated attacker can select different sampler subclasses in video requests to exceed configured decoder limits and exhaust unaccounted GPU memory.
NVD/CVE DatabaseCVE-2026-100648: vllm audio decoding bypasses file size limit in multimodal chat
Sep 26, 2026MediumVulnerabilitySecurityCVE-2026-100648vLLM versions before 0.29.0 do not enforce the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB limit during multimodal chat audio decoding. Unauthenticated clients can submit oversized audio files through chat endpoints, consuming excessive memory and CPU during decoding.
NVD/CVE DatabaseCVE-2026-100647: vLLM denial of service through oversized cache_salt parameter
Sep 26, 2026MediumVulnerabilitySecurityCVE-2026-100647vLLM versions before 0.29.0 have a denial-of-service flaw in the cache_salt parameter accepted on OpenAI-compatible and Anthropic API endpoints. The parameter has no maximum length validation and is processed on the single EngineCore scheduler thread. Unauthenticated attackers can send multi-hundred-megabyte salt values that trigger expensive pickle serialization and SHA-256 hashing, stalling the scheduler and denying service to all concurrent requests.
NVD/CVE DatabaseCVE-2026-94627: vLLM Mooncake connector flaw enables memory exhaustion
Sep 21, 2026HighVulnerabilitySecurityCVE-2026-94627The vLLM Mooncake connector through 0.29.0 does not properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. An attacker can submit completion requests with multiple prompts to trigger GPU memory exhaustion, as orphaned KV cache blocks accumulate until the process restarts and legitimate requests can no longer execute.
NVD/CVE DatabaseCVE-2026-94626: vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatible completion…
Sep 21, 2026HighVulnerabilitySecurityCVE-2026-94626vLLM through 0.29.0 does not validate the tp_size parameter in kv_transfer_params on its OpenAI-compatible completion endpoints. Attackers can supply arbitrary tp_size values in prefill/decode disaggregated deployments, allocating unbounded memory and exhausting it, which triggers a kernel OOM-kill of the decode worker process.
NVD/CVE DatabaseCVE-2026-94625: vLLM resource exhaustion in MooncakeConnector through rejected prefill requests
Sep 21, 2026MediumVulnerabilitySecurityCVE-2026-94625CVE-2026-94625 affects vLLM through 0.29.0 in MooncakeConnector. Rejected prefill requests leave ownerless transfer placeholders that are never reclaimed, so an attacker can send such requests to exhaust sender task pools. Valid requests can then be delayed by up to 480 seconds while health checks keep returning success.
NVD/CVE DatabaseCVE-2026-94624: vLLM denial of service in P2P KV offloading via kv_transfer_params
Sep 21, 2026HighVulnerabilitySecurityCVE-2026-94624vLLM through 0.29.0 has a denial of service flaw in peer-to-peer KV offloading when OffloadingConnector uses TieringOffloadingSpec with a peer-to-peer secondary tier. Attackers can supply arbitrary remote host and port values in kv_transfer_params to create unreachable peer sessions, which retain ZeroMQ sockets until the context quota is exhausted. The resulting uncaught ZMQError crashes EngineCore and halts all inference.
NVD/CVE DatabaseCVE-2026-94623: vLLM denial of service through NIXL connector prefix caching
Sep 21, 2026HighVulnerabilitySecurityCVE-2026-94623CVE-2026-94623 affects vLLM through 0.29.0, specifically the NIXL connector's prefix caching implementation in prefill/decode disaggregated deployments. Completion requests containing multiple prompts of varying lengths trigger an assertion failure in NixlBaseConnectorWorker._apply_prefix_caching, because block counts are not properly validated. The resulting crash makes the decode worker unavailable until it is restarted.
NVD/CVE DatabaseCVE-2026-94622: vLLM denial of service through NIXL connector metadata handling
Sep 21, 2026HighVulnerabilitySecurityCVE-2026-94622vLLM versions through 0.29.0 contain a denial of service flaw in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete kv_transfer_params dictionary entries to trigger an uncaught KeyError in EngineCore scheduling, which terminates the decode engine and makes all routed requests fail until a manual restart.
NVD/CVE DatabaseCVE-2026-93989: vLLM bad_words token index validation flaw corrupts concurrent request logits
Sep 19, 2026LowVulnerabilitySecurityCVE-2026-93989vLLM through 0.29.0 does not properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). An attacker can supply out-of-bounds token indices that corrupt the logits memory of concurrent requests, so different in-flight HTTP requests return incorrect tokens.
NVD/CVE DatabaseCVE-2026-93841: vLLM memory corruption in Triton _bincount_kernel via prompt token IDs
Sep 18, 2026LowVulnerabilitySecurityCVE-2026-93841vLLM through 0.29.0 has a memory corruption flaw in the Triton _bincount_kernel. Prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. An attacker can submit multimodal audio requests with tokens equal to vocabulary size, causing out-of-bounds writes that corrupt concurrent requests' sampler state and alter repetition penalty behavior.
NVD/CVE DatabaseCVE-2026-93840: vLLM allowed_token_ids validation flaw lets requests sample outside allowlists
Sep 18, 2026LowVulnerabilitySecurityCVE-2026-93840vLLM versions before 0.29.0 validate allowed_token_ids against the tokenizer length instead of the model output logits width, in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass this check, which causes LogitBiasState to corrupt GPU logits state so that concurrent requests can sample tokens outside their allowlists.
Fix: Fixed in 0.29.0.
NVD/CVE DatabaseCVE-2026-93592: vLLM crash via negative token IDs in /v1/embeddings and /pooling endpoints
Sep 18, 2026HighVulnerabilitySecurityCVE-2026-93592vLLM versions before 0.28.0 do not validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints. An unauthenticated attacker can crash the engine by submitting a negative token ID, which triggers a CUDA device-side assertion that poisons the GPU context and makes all later requests fail until the process restarts.
Fix: Fixed in vLLM 0.28.0. The source does not state an upgrade path or workaround beyond this version.
NVD/CVE DatabaseCVE-2026-93436: vLLM memory exhaustion via max_tokens=0 requests in prefill/decode deployments
Sep 17, 2026HighVulnerabilitySecurityCVE-2026-93436vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
NVD/CVE DatabaseCVE-2026-69147: vLLM GPU memory exhaustion through request-selected video decoder backend
Sep 16, 2026MediumVulnerabilitySecurityCVE-2026-69147vLLM versions prior to 0.28.0 let Chat Completions and Responses request bodies select the pynvvideocodec video backend through media_io_kwargs.video.video_backend, even when startup configuration chose a software decoder. The engine budgets decoder GPU memory only from static configuration, so the request-selected backend's CUDA context, decoder surfaces and decoded-frame allocations escape the KV-cache budget. An attacker who can submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes or denial of service.
Fix: Fixed in 0.28.0.
NVD/CVE DatabaseCVE-2026-57173: vLLM input_audio path allows memory exhaustion
Sep 16, 2026MediumVulnerabilitySecurityCVE-2026-57173vLLM versions prior to 0.24.0 have an input_audio handling flaw on the /v1/chat/completions endpoint. The path calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder, so the duration guard used by /v1/audio/transcriptions is bypassed. An unauthenticated client can submit a small compressed audio input that expands into a very large float32 PCM allocation, causing an out-of-memory worker crash. Inline data URLs are also not bounded by VLLM_AUDIO_FETCH_TIMEOUT. The issue affects deployments serving an audio-capable model, and authentication changes only deployment-specific reachability.
Fix: Fixed in version 0.24.0.
NVD/CVE DatabaseCVE-2026-92365: vLLM inefficient algorithmic complexity in thinking budget state handling
Sep 16, 2026MediumVulnerabilitySecurityCVE-2026-92365A vulnerability was found in vllm-project vllm up to 0.29.0. The flaw sits in unspecified functionality of vllm/v1/sample/thinking_budget_state.py and is triggered by manipulation that results in inefficient algorithmic complexity. It can be exploited remotely.
Fix: The pull request to fix this issue awaits acceptance.
NVD/CVE DatabaseCVE-2026-92220: vLLM resource consumption through MoRIIO acknowledgement handler request_id
Sep 15, 2026MediumVulnerabilitySecurityCVE-2026-92220A vulnerability in vllm-project vLLM 0.26.0 and 0.27.0 affects the MoRIIO Acknowledgement Handler in vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py, specifically MoRIIOConnectorScheduler.request_finished, MoRIIOConnectorWorker.get_finished and MoRIIOWrapper._handle_release_message. Manipulating the request_id or kv_transfer_params argument causes resource consumption, and the attack can be launched remotely. The project was notified early through a pull request but had not responded at the time of disclosure.
NVD/CVE DatabaseCVE-2026-90878: vLLM resource consumption through chat_template in /v1/chat/completions
Sep 15, 2026MediumVulnerabilitySecurityCVE-2026-90878A vulnerability in vllm-project vLLM up to 0.27.1 affects an unspecified part of the /v1/chat/completions endpoint, in the Jinja Template Rendering component. Manipulating the chat_template argument causes resource consumption, and the attack can be launched remotely. The exploit is publicly disclosed and may be used.
Fix: A pull request to fix this issue is awaiting acceptance. No fixed version is stated in the source.
NVD/CVE DatabaseCVE-2026-90713: vLLM denial of service in tiktoken vocab file handling
Sep 14, 2026LowVulnerabilitySecurityCVE-2026-90713A flaw in vllm-project vLLM up to 0.29.0 affects the function TiktokenTokenizer::new in rust/src/text/src/backend/hf/mod.rs, within the tiktoken vocab File Handler component. Manipulating this function causes a denial of service, and the attack requires local access. A public exploit exists, and a pull request to fix the issue is awaiting acceptance.
Fix: The source states that a pull request to fix this issue awaits acceptance. It does not name a fixed version or a workaround.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.