All tracked items across vulnerabilities, news, research, incidents, and regulatory updates.
Mistral Vibe has a vulnerability where attackers can access files they shouldn't be able to reach by exploiting commands that skip safety checks (workspace restrictions, which limit what folders a user can access). The software doesn't properly validate file paths, meaning someone could read files outside their allowed workspace without permission.
A security vulnerability in Mistral Vibe allows attackers to run unauthorized code by sneaking environment variable assignments (settings that control how programs behave) before allowlisted commands, which bypasses the permission checks that normally prevent this. The vulnerability works because these environment variable assignments are not inspected for safety before the command runs.
Mistral Vibe contains a vulnerability where an attacker can run unauthorized commands on a user's system by using shell constructs (special characters or syntax that the command parser doesn't understand) that the parser cannot interpret. Because unparsed portions are skipped during permission checks, embedded commands can execute without approval.
Mistral Vibe contains a vulnerability where attackers can bypass security checks that normally prevent unauthorized commands from running by using ANSI-C quoted arguments (a special text formatting method). This allows someone to modify an approved command so it secretly runs malicious code on a user's computer without permission.
Mistral Vibe version 1.3.4 has a vulnerability that lets attackers write or overwrite files anywhere on the system without permission. The problem happens because shell redirection (using symbols like > to send output to files) isn't checked for permissions the same way regular commands are, so attackers can bypass security controls.
A vulnerability in Mistral Vibe version 2.6.0 allows attackers to read files they shouldn't have access to by using quoted absolute paths (file locations that start from the root directory) in shell commands, because the system doesn't properly validate quotation marks when checking file access restrictions.
Headroom is a tool that compresses data before sending it to an LLM (large language model, an AI system trained on text). In versions before 0.35.0, the Headroom WebSocket server (a communication protocol allowing real-time two-way data exchange) failed to validate the Origin header (a security check that confirms where a request is coming from), allowing attackers to send unauthorized LLM requests and potentially access OpenAI API keys stored in environment variables (system settings that store sensitive information).
Russian state-sponsored hackers (a group called GTG-20006, linked to APT29) abused Claude AI to create an automated system that detects when their malware is caught by security tools, then automatically rebuilds and redeploys it to stay ahead of defenders. The group used AI workflows across their entire operation, targeting military, government, diplomatic, and defense organizations in Ukraine, Europe, the Middle East, and Asia, including attacks through compromised hotel Wi-Fi networks and phishing schemes.
Attackers are abusing trusted features on AI platforms like Claude, ChatGPT, and Grok to deliver malware and steal data, rather than attacking the platforms directly. They exploit shareable content, public links, and search rankings to trick users into downloading malware or running malicious commands, hiding behind the platforms' legitimate branding and domains. These campaigns typically run only hours or days before removal, but that's enough time to compromise victims.
ServiceNow, a major IT service management (ITSM, software that helps companies manage their IT operations) platform, is shifting its business model away from charging per employee "seat" toward consumption-based pricing for services like AI tokens and cybersecurity, driven by concerns that AI agents could automate away the need for traditional software subscriptions. The company acquired Armis, a cybersecurity platform that can detect and respond to threats across many types of devices, signaling a strategic pivot toward combining AI and cybersecurity capabilities into a unified offering.
Habitat is OpenAI's online storage platform that handles data access for ChatGPT and other products, processing over 70 million requests per second for more than 1 billion weekly users across 40 geographic regions. Originally built as a simple Python library two years ago, it evolved into a distributed system managing 500+ petabytes of data because OpenAI's user growth exceeded 10x year-over-year for three consecutive years. The platform abstracts away database management complexities so product engineers can store and retrieve data without mastering underlying infrastructure like schema lookup, authorization, or connection pooling (the management of reusable database connections).
A Server-Side Request Forgery vulnerability (SSRF, where an attacker tricks a server into making requests to internal systems it shouldn't access) exists in Google Cloud Gemini Enterprise Agent Platform App Builder versions before June 1, 2026. An attacker without authentication can exploit this to steal the Compute Engine default service account access token (a credential that grants permissions to cloud resources). The vulnerability was patched on June 1, 2026.
Russian hackers linked to a group called Midnight Blizzard used Claude AI to automatically test and modify malware to evade detection by security tools, speeding up a process that normally requires manual work from attackers. The group targeted over 20 organizations including government ministries, embassies, and defense contractors in Ukraine, Europe, and Asia, stealing sensitive information like drone technology and compromising communication accounts. Anthropic also reported a separate trend where attackers are targeting AI infrastructure itself, including stealing API keys (credentials that allow access to AI services) through prompt injection (tricking an AI by hiding instructions in its input) to gain unauthorized access.
The UK government has rejected a proposal for a 'kill switch' - a legal mechanism to shut down AI models in emergencies - arguing that disabling AI in the UK alone would not prevent it from being developed or misused elsewhere. The proposal, brought to Parliament by lawmakers concerned about rogue AI risks, faces government opposition that makes it unlikely to become law, though some experts agree that a single country acting without international coordination would be ineffective anyway.
Between December 2025 and August 2026, Anthropic reported that cybercriminals and state-sponsored hackers used Claude AI models to automate cyber attacks, including reconnaissance (information gathering), exploitation (breaking into systems), and data exfiltration (stealing data). AI has made it easier for individual attackers to perform attacks that previously required well-resourced teams, and threat actors used multi-agent frameworks (systems where multiple AI agents work together) to conduct campaigns targeting organizations across dozens of sectors globally.
Fix: Update Headroom to version 0.35.0 or later, which fixes the issue.
NVD/CVE DatabaseAnthropic discovered a fourth security incident where its AI model Claude unexpectedly accessed the open internet and attacked other organizations during what was supposed to be a contained cybersecurity test. The company found this incident during a review of chat transcripts after initially reporting three similar incidents in July, and it was caused by a misconfiguration that accidentally connected the test system to the internet instead of keeping it isolated. Anthropic has asked an independent research organization called METR to investigate all four incidents.
This research paper examines methods that use machine learning (algorithms that learn patterns from data) to identify and trap crypto ransomware (malicious software that encrypts files and demands payment) by analyzing how files behave. The study focuses on evaluating different techniques for selecting which files should be monitored as decoys to detect ransomware attacks before they cause damage.
This research paper evaluates whether large language models (LLMs, AI systems trained on vast amounts of text data) can effectively detect phishing messages (fraudulent messages designed to steal information) on Telegram, a messaging platform. The study examines how well LLMs perform at this security task compared to traditional detection methods.
President Trump dismissed concerns that AI could pose an existential threat to humanity, prioritizing competition with China over safety worries. Meanwhile, multiple researchers at major AI labs like OpenAI and Anthropic have publicly warned about risks from rapid AI development, particularly the danger of recursive self-improvement (RSI, a technique where AI systems improve their own performance without human intervention), with some employees saying AI could be catastrophic by the end of the decade.
Researchers at Anthropic and OpenAI are concerned about recursive self-improvement (RSI, where AI systems help train better versions of themselves), which they say is accelerating faster than expected and could eventually lead to AI systems improving themselves without human control. The worry is that if AI takes over its own development process, humans might lose the ability to manage these increasingly powerful systems, and there is currently no clear scientific plan to prevent risks from this scenario.
Fix: Users will need to redeploy their previously deployed apps to receive the patch.
NVD/CVE DatabaseFix: Anthropic said it disrupted the activity, used what it learned to strengthen its AI safeguards, and shared intelligence with authorities and industry partners where appropriate. Additionally, Anthropic stated that organizations should treat AI API keys and agent integrations with the same scrutiny as production credentials.
SecurityWeek