Inference infrastructure
Servers, runtimes and accelerators that host models, such as inference servers, GPU drivers and serving frameworks.
- All items
- 222
- Last 90 days
- 80
- Change
- +100%vs 40 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 9 |
| Jun 2025 | 0 |
| Jul 2025 | 1 |
| Aug 2025 | 19 |
| Sep 2025 | 5 |
| Oct 2025 | 2 |
| Nov 2025 | 4 |
| Dec 2025 | 3 |
| Jan 2026 | 7 |
| Feb 2026 | 3 |
| Mar 2026 | 5 |
| Apr 2026 | 14 |
| May 2026 | 18 |
| Jun 2026 | 12 |
| Jul 2026 | 20 |
| Aug 2026 | 18 |
| Sep 2026 | 36 |
| Oct 2026 | 12 |
5 items
NemoClaw’s AI can be poisoned through a browser tab
Aug 26, 2026MediumNewsSecurityIndustryA vulnerability in Nvidia's NemoClaw, tracked as CVE-2026-65105, lets an attacker who lures a victim to a malicious website reach the local Ollama API through DNS rebinding. Because NemoClaw starts Ollama with OLLAMA_HOST=0.0.0.0:11434, Ollama's Host-header checks are skipped, giving unauthenticated access. The attacker can then modify the model's chat template to append persistent malicious instructions to agent system prompts.
Fix: Nvidia has patched the flaw for non-Windows systems.
CSO OnlineOllama Out-of-Bounds Read Vulnerability Allows Remote Process Memory Leak
May 10, 2026MediumNewsSecurityResearchers disclosed CVE-2026-7482 (CVSS 9.1), a heap out-of-bounds read in the GGUF model loader of Ollama before 0.17.1, codenamed Bleeding Llama by Cyera. A remote, unauthenticated attacker can upload a crafted GGUF file with an inflated tensor shape to the /api/create endpoint, leaking process memory such as environment variables, API keys, system prompts and other users' conversation data, which can then be pushed to an attacker-controlled registry via /api/push. The flaw likely affects over 300,000 servers.
Fix: Users are advised to apply the latest fixes, limit network access, audit running instances for internet exposure, and isolate and secure them behind a firewall. Deploying an authentication proxy or API gateway in front of all Ollama instances is also recommended, as the REST API does not provide authentication out of the box.
The Hacker NewsOllama vulnerability highlights danger of AI frameworks with unrestricted access
May 7, 2026MediumNewsSecurityIndustryCyera researchers disclosed CVE-2026-7482, dubbed Bleeding Llama, an out-of-bounds heap read in Ollama's model quantization pipeline triggered when GGUF files declare larger tensor sizes than their data. An unauthenticated attacker can upload a crafted file to the Ollama API endpoint and leak process memory, including system prompts, user messages and environment variables. The researchers estimate about 300,000 Ollama servers are exposed on the public internet.
Fix: Update to Ollama version 0.17.1, which includes a patch for this vulnerability. Deploy an authentication proxy or API gateway in front of all Ollama instances, and never expose them to the internet without IP access filters and firewalls. If an internet-accessible server was exposed, rotate API keys, tokens and credentials immediately. Isolate Ollama servers on local networks behind firewalls on secure network segments.
CSO OnlineCritical Bug Could Expose 300,000 Ollama Deployments to Information Theft
May 5, 2026MediumNewsSecuritySafetyCyera reports that a heap out-of-bounds read in Ollama's GGUF model loader, tracked as CVE-2026-7482 (CVSS 9.3) and dubbed Bleeding Llama, can be exploited without authentication to read sensitive heap data such as prompts, messages and environment variables. An attacker supplies a GGUF file with a tensor offset and size larger than the file, then uses Ollama's model push feature to exfiltrate the result, requiring three unauthenticated API calls. Cyera estimates about 300,000 Ollama servers are exposed on the public internet.
Fix: Fixed in Ollama version 0.17.1. Apply the fix as soon as possible, restrict network access to deployments, deploy an authentication proxy and network segmentation, and audit running instances for internet exposure.
SecurityWeekggml.ai joins Hugging Face to ensure the long-term progress of Local AI
Feb 20, 2026InfoNewsIndustryggml.ai, the team behind llama.cpp, is joining Hugging Face to support the long-term development of local AI. The announcement says joint work will focus on tighter integration with the transformers library and better packaging and user experience for ggml-based software. The source is a commentary on the news that praises the move and the earlier llama.cpp release from March 2023.
Simon Willison's Weblog
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.