Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
217 items
CVE-2026-90555: vLLM audio sample rate validation flaw in transcription endpoint
Sep 12, 2026MediumVulnerabilitySecurityCVE-2026-90555vLLM versions before 0.28.0 fail to validate audio sample rate headers in the transcription endpoint, letting authenticated clients bypass duration checks. An attacker can submit forged FLAC headers with inflated sample rates to trigger excessive memory allocation and crash the API server process, affecting all tenants.
Fix: Fixed in 0.28.0.
NVD/CVE DatabaseCVE-2026-90554: vLLM denial of service through unbounded audio decoding from video input
Sep 12, 2026MediumVulnerabilitySecurityCVE-2026-90554vLLM versions 0.10.2 up to but not including 0.28.0 apply no audio decode-size or duration limit when extracting audio from video input for NanoNemotronVL models. The _extract_audio_from_videos function in nano_nemotron_vl.py calls load_audio_pyav without the max_duration_s or max_decode_bytes parameters, so VLLM_MAX_AUDIO_DECODE_DURATION_S and VLLM_MAX_AUDIO_DECODE_BYTES are not enforced. When such a model is served with use_audio_in_video=True, an attacker supplying a small, highly compressed video can make the server allocate gigabytes of memory during audio decoding, causing a denial of service.
Fix: Fixed in vLLM 0.28.0.
NVD/CVE DatabaseCVE-2026-90553: vLLM remote code execution in LlavaOnevision2 processor loader
Sep 12, 2026HighVulnerabilitySecurityCVE-2026-90553vLLM before 0.28.0 contains a remote code execution flaw in the LlavaOnevision2 processor loader. The loader ignores the trust_remote_code parameter when it loads remote processor classes, so an attacker who crafts a malicious model with arbitrary code in processing_llava_onevision2.py can run that code with vLLM process authority, even when trust_remote_code is set to False.
NVD/CVE DatabaseGHSA-fxg7-897c-57mp: Nuxt Ollama: Public Runtime Config Exposes Ollama API Key to Browser Clients
Sep 9, 2026HighVulnerabilitySecurityCVE-2026-59158nuxt-ollama@1.2.26 merges all module options, including api_key, into Nuxt's public runtime config (runtimeConfig.public.ollama) in src/module.ts. Nuxt serializes that namespace into the SSR HTML window.__NUXT__ payload, so any unauthenticated HTTP client that fetches a page can read the Ollama cloud API key in plaintext and use it against the Ollama API at the operator's expense.
Fix: Recommended remediation: Move api_key to the private runtime config and remove it from the browser composable. Split _options so api_key is written to runtimeConfig.ollama, while publicOptions without api_key are merged into runtimeConfig.public.ollama. The api_key should then only be consumed in the server-side utility (src/runtime/server/utils/useOllama.ts) via useRuntimeConfig().ollama.api_key.
GitHub Advisory DatabaseCVE-2026-47625: NVIDIA Triton Inference Server for Linux missing authorization flaw
Sep 8, 2026HighVulnerabilitySecurityCVE-2026-47625NVIDIA Triton Inference Server for Linux contains a vulnerability in which an attacker could abuse missing authorization. A successful exploit might lead to information disclosure, data tampering, and denial of service.
NVD/CVE DatabaseCVE-2026-16497: NVIDIA Triton Inference Server for Linux excessive iteration denial of service
Sep 8, 2026HighVulnerabilitySecurityCVE-2026-16497NVIDIA Triton Inference Server for Linux contains a vulnerability (CVE-2026-16497) that lets an attacker cause excessive iteration. A successful exploit might lead to denial of service.
NVD/CVE DatabaseCVE-2026-86289: Ollama integer overflow in GGUF decoder readGGUFV1String
Sep 7, 2026MediumVulnerabilitySecurityCVE-2026-86289CVE-2026-86289 affects Ollama up to 0.31.1. An integer overflow in the readGGUFV1String function of fs/ggml/gguf.go, in the GGUF Decoder component, can be triggered remotely. A public exploit exists.
Fix: Upgrading to version 0.31.2-rc1 addresses this issue. The patch is named 67b6a1c2d45321e0cb3c04a18073f9818de7724b. Upgrading the affected component is recommended.
NVD/CVE DatabaseCVE-2026-85180: Ollama redirect validation flaw when pulling tensor-layer models
Sep 3, 2026HighVulnerabilitySecurityCVE-2026-85180Ollama fails to validate redirect destinations when pulling tensor-layer models, letting unauthenticated attackers redirect blob downloads to arbitrary hosts. An attacker controlling a registry can serve a malicious tensor-layer manifest and cause the server to issue GET requests to internal hosts, including cloud metadata endpoints.
NVD/CVE DatabaseCVE-2026-37237: vLLM denial of service via unbounded media fetch from remote URLs
Aug 28, 2026MediumVulnerabilitySecurityCVE-2026-37237vLLM versions up to and including 0.17.0 are affected by CVE-2026-37237. The AsyncMediaIO.fetch_audio and AsyncMediaIO.fetch_image functions in multimodal/inputs.py fetch user-supplied media URLs with aiohttp and call r.read() without enforcing a maximum response size. A remote attacker can exhaust server memory by supplying a URL to an arbitrarily large file, causing a Denial of Service.
NVD/CVE DatabaseCVE-2026-78684: vLLM DeepStream backend classification flaw enables denial of service
Aug 25, 2026MediumVulnerabilitySecurityCVE-2026-78684CVE-2026-78684 affects vLLM before 0.27.0, which fails to classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. An unauthenticated attacker can activate DeepStream at request time, initialize the process-wide GPU decode pool, and submit video that bypasses resource controls, causing partial denial of service for concurrent requests. VulnCheck rates it CVSS 4.0 6.9 (MEDIUM); NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-47630: NVIDIA Triton Inference Server for Linux absolute path traversal
Aug 18, 2026MediumVulnerabilitySecurityCVE-2026-47630NVIDIA Triton Inference Server for Linux contains an absolute path traversal vulnerability, tracked as CWE-36. A successful exploit might lead to code execution. The NVD assessment has not yet been provided, and NVIDIA published the advisory under its product-security repository.
NVD/CVE DatabaseCVE-2026-47629: NVIDIA Triton Inference Server for Linux improper input validation
Aug 18, 2026HighVulnerabilitySecurityCVE-2026-47629NVIDIA Triton Inference Server for Linux contains an improper input validation flaw, tracked as CVE-2026-47629 and classified as CWE-20. A successful exploit might lead to denial of service. NVD has not yet provided an assessment, and the source names no affected versions.
NVD/CVE DatabaseCVE-2026-47628: NVIDIA Triton Inference Server for Linux resource allocation without limits
Aug 18, 2026HighVulnerabilitySecurityCVE-2026-47628NVIDIA Triton Inference Server for Linux contains a vulnerability, tracked as CVE-2026-47628 and classified as CWE-770 (Allocation of Resources Without Limits or Throttling). An attacker could cause an allocation of resources without limits, and a successful exploit might lead to denial of service. NVD had not yet provided an assessment at the time of the record.
NVD/CVE DatabaseCVE-2026-47627: NVIDIA Triton Inference Server for Linux path traversal
Aug 18, 2026CriticalVulnerabilitySecurityCVE-2026-47627CVE-2026-47627 affects NVIDIA Triton Inference Server for Linux and is classified as CWE-22, Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal'). The source states that an attacker could cause path traversal, and a successful exploit might lead to denial of service. NVD had not yet provided an assessment when the entry was published on 08/18/2026.
NVD/CVE DatabaseCVE-2026-47606: NVIDIA Triton Inference Server for Linux absolute path traversal
Aug 18, 2026MediumVulnerabilitySecurityCVE-2026-47606NVIDIA Triton Inference Server for Linux contains an absolute path traversal vulnerability, tracked as CWE-36. A successful exploit might lead to code execution and information disclosure. NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2026-73560: vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the MiMoV2OmniMultiModalProcessor…
Aug 17, 2026MediumVulnerabilitySecurityCVE-2026-73560vLLM, an inference and serving engine for large language models, is affected prior to version 0.26.0. The MiMoV2OmniMultiModalProcessor in vllm/transformers_utils/processors/mimo_v2_omni.py passes attacker-controlled image and audio strings through _fetch_image, requests.get and Image.open instead of MediaConnector, which bypasses the allowed_media_domains and allowed_local_media_path protections. This allows server-side requests and reads of arbitrary files accessible to the vLLM process.
Fix: Fixed in version 0.26.0.
NVD/CVE DatabaseCVE-2026-71486: vLLM resource exhaustion through derender endpoints
Aug 17, 2026MediumVulnerabilitySecurityCVE-2026-71486vLLM versions prior to 0.26.0 expose the /v1/completions/derender and /v1/chat/completions/derender endpoints, which accept caller-supplied GenerateResponse objects. The OnlineDerenderer and tokenizer.decode process these structures before max_model_len, max_tokens, max_num_seqs, or response-size limits are enforced, so an authenticated API client can consume excessive CPU and memory and produce oversized responses.
Fix: Fixed in 0.26.0.
NVD/CVE DatabaseCVE-2026-73559: vLLM completions endpoint resource exhaustion through unbounded prompt list
Aug 13, 2026MediumVulnerabilitySecurityCVE-2026-73559vLLM versions 0.19.0 through 0.26.0 accept an unbounded list of prompts in the /v1/completions CompletionRequest.prompt field. The preprocessing path and serving code create one engine generator and response slot per prompt, so a single request from an authenticated API client can exhaust CPU, memory, async scheduling capacity, engine request slots and response buffering.
Fix: Fixed in 0.26.0.
NVD/CVE DatabaseCVE-2026-73558: vLLM integer overflow in activation kernel leaks batched inference results
Aug 13, 2026MediumVulnerabilitySecurityCVE-2026-73558CVE-2026-73558 affects vLLM, an inference and serving engine for large language models, prior to 0.27.0. An integer overflow in blockIdx.x * 2 * d in activation_kernels.cu can cause act_and_mul_kernel to consume another batched user's input. A request in the same inference batch may then receive a partial or complete copy of another user's inference result.
Fix: Fixed in version 0.27.0.
NVD/CVE DatabaseCVE-2026-73557: vLLM race condition bypasses sparse tensor validation in prompt_embeds
Aug 13, 2026HighVulnerabilitySecurityCVE-2026-73557vLLM, an inference and serving engine for large language models, contains a flaw from 0.20.2rc0 through 0.26.0 in safe_load_prompt_embeds in vllm/renderers/embed_utils.py. The function toggles the process-global torch.sparse.check_sparse_tensor_invariants setting, and concurrent prompt_embeds parts submitted to POST /v1/chat/completions can race that state, letting an invalid sparse tensor reach tensor.to_dense despite the CVE-2025-62164 guard when enable_prompt_embeds is enabled.
Fix: This issue is fixed in version 0.26.0.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.