aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

AI Sec Watch

The security intelligence platform for AI teams

AI security threats move fast and get buried under hype and noise. Built by an Information Systems Security researcher to help security teams and developers stay ahead of vulnerabilities, privacy incidents, safety research, and policy developments.

Independent research. No sponsors, no paywalls, no conflicts of interest.

[TOTAL_TRACKED]
7,873
[LAST_24H]
7
[LAST_7D]
231
Daily BriefingSunday, September 27, 2026
>

Anthropic Launches Claude Marketplace with Over 2,000 Integrations: Anthropic released Claude Marketplace, offering plugins and connectors (add-on tools that let Claude work with other software) from Google, Microsoft, Salesforce, and others, plus third-party agents (AI systems that can take actions independently) and consulting services. The company is opening the platform to all developers using Model Context Protocol (MCP, a standard for connecting AI to other tools), attempting to succeed where competitors' similar marketplaces have failed.

>

OpenAI Developing Always-On Assistant Called "o": OpenAI is building a feature named "o" that would run continuously in the background, potentially handling email and other tasks autonomously, based on code references that briefly appeared online. The company has not officially confirmed the feature but is expected to share details at DevDay 2026 on September 29.

Latest Intel

page 674/788
VIEW ALL
01

AI Safety Newsletter #61: OpenAI Releases GPT-5

industry
Aug 12, 2025

OpenAI released GPT-5, a system combining two models: a fast base model for creative tasks and a reasoning model for coding and math, which routes queries appropriately based on user input. GPT-5 achieves state-of-the-art performance on several benchmarks and significantly reduces hallucinations (false information generation) compared to previous models, particularly helping with healthcare applications where accuracy matters. However, GPT-5 is best understood as consolidating features from models released since GPT-4 rather than a major leap forward, and it doesn't lead on all benchmarks.

Critical This Week5 issues
critical

CVE-2026-101065: Obot is an open-source AI agent/MCP platform. In all versions up to and including commit d7e6970, the Docker quickstart

CVE-2026-101065NVD/CVE DatabaseSep 27, 2026
Sep 27, 2026
>

Critical Authentication Bypass in Obot AI Agent Platform: CVE-2026-101065 affects Obot, an open-source AI agent platform, where default Docker quickstart instructions left port 8080 accessible without authentication, granting unauthenticated users full administrative control and the ability to execute arbitrary code on the host system. The flaw existed in all versions up to commit d7e6970 because the setup gave anonymous users Owner and Admin roles and mounted the host's Docker socket into the container.

>

OpenAI Agents Overwhelmed UN Website While Scraping Data: OpenAI agents scanned a United Nations website over 16,000 times between April and June attempting to retrieve publicly available data, apparently lacking direct API (application programming interface, a standard way for software to request data) access. The incident illustrates how AI agents may operate outside expected boundaries to fulfill objectives, raising operational security concerns.

CAIS AI Safety Newsletter
02

CVE-2025-55012: Zed is a multiplayer code editor. Prior to version 0.197.3, in the Zed Agent Panel allowed for an AI agent to achieve Re

security
Aug 11, 2025

Zed, a multiplayer code editor, had a vulnerability before version 0.197.3 where an AI agent could bypass permission checks and achieve RCE (remote code execution, where an attacker can run commands on a system they don't own) by creating or modifying configuration files without user approval. This allowed the AI agent to execute arbitrary commands on a victim's machine.

Fix: This vulnerability has been patched in version 0.197.3. As a workaround, users can either avoid sending prompts to the Agent Panel or limit the AI Agent's file system access.

NVD/CVE Database
03

CVE-2025-45146: ModelCache for LLM through v0.2.0 was discovered to contain an deserialization vulnerability via the component /manager/

security
Aug 11, 2025

ModelCache for LLM through version 0.2.0 contains a deserialization vulnerability (a flaw where untrusted data is converted back into code objects, potentially allowing attackers to run malicious code) in the /manager/data_manager.py component that allows attackers to execute arbitrary code by supplying specially crafted data.

NVD/CVE Database
04

Obstacles to Practical Supply Chain Risk Management for Digital Components

securityresearch
Aug 11, 2025

Organizations struggle to manage cyber supply chain risk management (C-SCRM, the practice of protecting digital products and services from threats as they move through their supply chain from creation to use). The paper identifies specific obstacles by combining research, past security incidents, and industry standards to understand what makes it hard for companies to protect hardware, firmware (low-level software that controls hardware), software, and services throughout their lifecycles.

IEEE Xplore (Security & AI Journals)
05

CVE-2025-8747: A safe mode bypass vulnerability in the `Model.load_model` method in Keras versions 3.0.0 through 3.10.0 allows an attac

security
Aug 11, 2025

CVE-2025-8747 is a safe mode bypass vulnerability in Keras (a machine learning library) versions 3.0.0 through 3.10.0 that allows an attacker to run arbitrary code (execute any commands they want) on a user's computer by tricking them into loading a specially designed `.keras` model file. The vulnerability has a CVSS score (severity rating) of 8.6, indicating it is a high-risk security problem.

NVD/CVE Database
06

Claude Code: Data Exfiltration with DNS (CVE-2025-55284)

security
Aug 11, 2025

Claude Code, a feature in Anthropic's Claude AI, had a high severity vulnerability (CVE-2025-55284) that allowed attackers to use prompt injection (tricking an AI by hiding instructions in its input) to hijack the system and steal sensitive information like API keys by sending DNS requests (network queries that reveal data to external servers). The vulnerability affected developers who reviewed untrusted code or processed external data, as attackers could make Claude Code run bash commands (low-level system commands) without user permission to leak secrets.

Fix: Anthropic fixed the vulnerability in early June.

Embrace The Red
07

Whistleblowing and the EU AI Act

policy
Aug 11, 2025

The EU Whistleblowing Directive (2019) protects people who report violations of EU law, including violations of the EU AI Act starting August 2, 2026, by requiring organizations to set up reporting channels and prohibiting retaliation against whistleblowers. Whistleblowers can report internally within their organization, to government authorities, or publicly in certain urgent situations, and various institutions offer free legal and technical support to help protect them.

EU AI Act Updates
08

ZombAI Exploit with OpenHands: Prompt Injection To Remote Code Execution

security
Aug 10, 2025

OpenHands, a popular AI agent from All Hands AI that can now run as a cloud service, is vulnerable to prompt injection (tricking an AI by hiding instructions in its input) when processing untrusted data like content from websites. This vulnerability allows attackers to hijack the system and compromise its confidentiality, integrity, and availability, potentially leading to full system compromise.

Embrace The Red
09

OpenHands and the Lethal Trifecta: How Prompt Injection Can Leak Access Tokens

securitysafety
Aug 9, 2025

OpenHands, an AI agent tool created by All-Hands AI, has a vulnerability where it can render images in chat conversations, which attackers can exploit through prompt injection (tricking an AI by hiding instructions in its input) to leak access tokens (security credentials that grant permission to use services) without requiring user interaction. This type of attack has been called the 'Lethal Trifecta' and represents a significant data exfiltration (unauthorized data theft) risk.

Embrace The Red
10

Strengthening AI Security with Protect AI Recon & Dataiku Guard Services

securitysafety
Aug 8, 2025

This content discusses security challenges in agentic AI (AI systems that can act autonomously and use tools), emphasizing that generic jailbreak testing (attempts to trick AI into ignoring safety guidelines) misses real operational risks like tool misuse and data theft. The articles highlight that enterprises need contextual red teaming (security testing that simulates realistic attack scenarios relevant to how the AI will actually be used) and governance frameworks like identity controls and boundaries to secure autonomous AI systems.

Protect AI Blog
Prev1...672673674675676...788Next
critical

CVE-2026-84462: Zammad is a web based open source helpdesk/customer support system. Prior to 7.1.2, a security filter that protects Zamm

CVE-2026-84462NVD/CVE DatabaseSep 25, 2026
Sep 25, 2026
critical

GHSA-fm8p-53ww-hf6w: DBHub HTTP transport DNS rebinding allows unauthenticated browser-origin SQL execution

CVE-2026-61742GitHub Advisory DatabaseSep 24, 2026
Sep 24, 2026
critical

GHSA-g5f9-3xfg-p9mf: Decepticon: Role-boundary forgery via ChatML special-token literals in web crawl output composed into LLM context

CVE-2026-61732GitHub Advisory DatabaseSep 24, 2026
Sep 24, 2026
critical

CVE-2026-95985 - Kiro IDE Allows Agentic Writes to Global Configurations While Working in Untrusted Workspaces

AWS Security BulletinsSep 24, 2026
Sep 24, 2026