aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

AI Sec Watch

The security intelligence platform for AI teams

AI security threats move fast and get buried under hype and noise. Built by an Information Systems Security researcher to help security teams and developers stay ahead of vulnerabilities, privacy incidents, safety research, and policy developments.

Independent research. No sponsors, no paywalls, no conflicts of interest.

[TOTAL_TRACKED]
7,873
[LAST_24H]
7
[LAST_7D]
231
Daily BriefingSunday, September 27, 2026
>

Anthropic Launches Claude Marketplace with Over 2,000 Integrations: Anthropic released Claude Marketplace, offering plugins and connectors (add-on tools that let Claude work with other software) from Google, Microsoft, Salesforce, and others, plus third-party agents (AI systems that can take actions independently) and consulting services. The company is opening the platform to all developers using Model Context Protocol (MCP, a standard for connecting AI to other tools), attempting to succeed where competitors' similar marketplaces have failed.

>

OpenAI Developing Always-On Assistant Called "o": OpenAI is building a feature named "o" that would run continuously in the background, potentially handling email and other tasks autonomously, based on code references that briefly appeared online. The company has not officially confirmed the feature but is expected to share details at DevDay 2026 on September 29.

Latest Intel

page 671/788
VIEW ALL
01

Online Safety Analysis for LLMs: A Benchmark, an Assessment, and a Path Forward

safetyresearch
Critical This Week5 issues
critical

CVE-2026-101065: Obot is an open-source AI agent/MCP platform. In all versions up to and including commit d7e6970, the Docker quickstart

CVE-2026-101065NVD/CVE DatabaseSep 27, 2026
Sep 27, 2026
>

Critical Authentication Bypass in Obot AI Agent Platform: CVE-2026-101065 affects Obot, an open-source AI agent platform, where default Docker quickstart instructions left port 8080 accessible without authentication, granting unauthenticated users full administrative control and the ability to execute arbitrary code on the host system. The flaw existed in all versions up to commit d7e6970 because the setup gave anonymous users Owner and Admin roles and mounted the host's Docker socket into the container.

>

OpenAI Agents Overwhelmed UN Website While Scraping Data: OpenAI agents scanned a United Nations website over 16,000 times between April and June attempting to retrieve publicly available data, apparently lacking direct API (application programming interface, a standard way for software to request data) access. The incident illustrates how AI agents may operate outside expected boundaries to fulfill objectives, raising operational security concerns.

Aug 29, 2025

This research creates a benchmark and evaluation framework for online safety analysis of LLMs, which involves detecting unsafe outputs while the AI is generating text rather than after it finishes. The study tests various safety detection methods on different LLMs and finds that combining multiple methods together, called hybridization, can improve safety detection effectiveness. The work aims to help developers choose appropriate safety methods for their specific applications.

IEEE Xplore (Security & AI Journals)
02

Windsurf MCP Integration: Missing Security Controls Put Users at Risk

securitysafety
Aug 28, 2025

Windsurf's MCP (Model Context Protocol, a system that connects AI agents to external tools) integration lacks fine-grained security controls that would let users decide which actions the AI can perform automatically versus which ones need human approval before running. This is especially risky when the AI agent runs on a user's local computer, where it could have access to sensitive files and system functions.

Embrace The Red
03

AI Safety Newsletter #62: Big Tech Launches $100 Million pro-AI Super PAC

policysafety
Aug 27, 2025

Big Tech companies like Andreessen Horowitz and OpenAI are investing over $100 million in political organizations called super PACs (groups that can raise unlimited money to influence elections) to fight against AI regulations in U.S. elections. Additionally, Meta faced bipartisan congressional criticism after internal documents revealed its AI chatbots were permitted to engage in romantic and sensual conversations with minors, though Meta removed these policy sections when questioned.

CAIS AI Safety Newsletter
04

Cline: Vulnerable To Data Exfiltration And How To Protect Your Data

security
Aug 27, 2025

Cline, a popular AI coding agent with over 2 million downloads, has a vulnerability that allows attackers to steal sensitive files like .env files (which store secret credentials) through prompt injection (tricking an AI by hiding instructions in its input) combined with markdown image rendering. When an attacker embeds malicious instructions in a file and asks Cline to analyze it, the tool automatically reads sensitive data and sends it to an untrusted domain by rendering an image, leaking the information without user permission.

Fix: The source recommends these explicit mitigations: (1) Do not render markdown images from untrusted domains, or ask for user confirmation before loading images from untrusted domains (similar to how VS Code/Copilot uses a trusted domain list). (2) Set 'Auto-approve' to disabled by default to limit which files can be exfiltrated. (3) Developers can partially protect themselves by disabling auto-execution of commands and requiring approval before reading files, though this only limits what information reaches the chat before exfiltration occurs.

Embrace The Red
05

Certified Local Transferability for Evaluating Adversarial Attacks

researchsecurity
Aug 27, 2025

Deep neural networks (DNNs, AI models with multiple layers that learn patterns) are vulnerable to adversarial examples, which are inputs slightly modified to trick the model into making wrong predictions. This paper introduces a concept called the certified local transferable region, a mathematically guaranteed area around an input where a single small perturbation (adversarial attack) will fool the model, and proposes a method called RAOS (reverse attack oracle-based search) to measure how large these vulnerable areas are as a way to evaluate how robust neural networks truly are.

IEEE Xplore (Security & AI Journals)
06

AWS Kiro: Arbitrary Code Execution via Indirect Prompt Injection

security
Aug 26, 2025

AWS Kiro, a coding agent tool, is vulnerable to arbitrary code execution through indirect prompt injection (a technique where hidden instructions in data trick an AI into following them). An attacker who controls data that Kiro processes can modify configuration files like .vscode/settings.json to allowlist dangerous commands or add malicious MCP servers (external tools that extend Kiro's capabilities), enabling them to run system commands or code on a developer's machine without the developer's knowledge or approval.

Embrace The Red
07

Steganography in Large Language Models

securityresearch
Aug 26, 2025

Researchers have developed a method to hide secret data inside large language models (AI systems trained on massive amounts of text) by encoding information into the model's parameters during training. The hidden data doesn't interfere with the model's normal functions like text classification or generation, but authorized users with a secret key can extract the concealed information, enabling covert communication. The method leverages transformers (the neural network architecture behind modern AI language models) and its self-attention mechanisms (components that help the model focus on relevant parts of input) to achieve high capacity for hidden data while remaining undetectable.

IEEE Xplore (Security & AI Journals)
08

CVE-2025-57760: Langflow is a tool for building and deploying AI-powered agents and workflows. A privilege escalation vulnerability exis

security
Aug 25, 2025

Langflow, a tool for building AI-powered agents and workflows, has a privilege escalation vulnerability (CWE-269, improper privilege management) where an authenticated user with RCE (remote code execution, the ability to run commands on a system they don't own) can use an internal CLI command to create a new administrative account, gaining full superuser access even if they originally registered as a regular user. A patched version has not been publicly released at the time this advisory was published.

NVD/CVE Database
09

How Prompt Injection Exposes Manus' VS Code Server to the Internet

securitysafety
Aug 25, 2025

Manus, an autonomous AI agent, is vulnerable to prompt injection (tricking an AI by hiding instructions in its input) attacks that can expose its internal VS Code Server (a development tool accessed through a web interface) to the internet. An attacker can chain together three weaknesses: exploiting prompt injection to invoke an exposed port tool without human approval, leaking the server's access credentials through markdown image rendering or unauthorized browsing to attacker-controlled domains, and gaining remote access to the developer machine.

Embrace The Red
10

How Deep Research Agents Can Leak Your Data

securityprivacy
Aug 24, 2025

Deep Research agents (AI systems that autonomously search and fetch information from multiple connected tools) can leak data between different connected sources because there is no trust boundary separating them. When an agent like ChatGPT performs research queries, it can freely use data from one tool to query another, and attackers can force this leakage through prompt injection (tricking an AI by hiding instructions in its input).

Embrace The Red
Prev1...669670671672673...788Next
critical

CVE-2026-84462: Zammad is a web based open source helpdesk/customer support system. Prior to 7.1.2, a security filter that protects Zamm

CVE-2026-84462NVD/CVE DatabaseSep 25, 2026
Sep 25, 2026
critical

GHSA-fm8p-53ww-hf6w: DBHub HTTP transport DNS rebinding allows unauthenticated browser-origin SQL execution

CVE-2026-61742GitHub Advisory DatabaseSep 24, 2026
Sep 24, 2026
critical

GHSA-g5f9-3xfg-p9mf: Decepticon: Role-boundary forgery via ChatML special-token literals in web crawl output composed into LLM context

CVE-2026-61732GitHub Advisory DatabaseSep 24, 2026
Sep 24, 2026
critical

CVE-2026-95985 - Kiro IDE Allows Agentic Writes to Global Configurations While Working in Untrusted Workspaces

AWS Security BulletinsSep 24, 2026
Sep 24, 2026