aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

Browse All

All tracked items across vulnerabilities, news, research, incidents, and regulatory updates.

to
Export CSV
9703 items

The future of AI regulation is courting the strangest, most anxious bedfellows

infonews
policy
Jun 10, 2026

This article discusses AI regulation efforts in Washington, D.C., noting that various political figures and stakeholders with differing interests are coming together to shape AI policy. The piece frames these unexpected political alliances as complex and contentious, comparing the current regulatory landscape to chaos.

The Verge (AI)

GHSA-cxh2-4639-vmc5: OpenTelemetry Operator for Kubernetes's ServiceMonitor bearerTokenFile reads arbitrary local file and sends contents as bearer auth

highvulnerability
security
Jun 10, 2026
CVE-2026-47701

OpenTelemetry Operator's TargetAllocator has a vulnerability where a tenant who can create or update a ServiceMonitor (a Kubernetes resource that tells Prometheus what to monitor) can trick the Collector into reading arbitrary files from its pod and sending them as authentication credentials to an attacker-controlled endpoint. This allows attackers to steal the Collector's service account token (a credential that proves the pod's identity to Kubernetes) and potentially access sensitive cluster information or files.

GHSA-j9rx-rppg-6hh4: Anyquery has Path Traversal through `clear_plugin_cache`, Allowing Arbitrary Directory Deletion

highvulnerability
security
Jun 10, 2026
CVE-2026-47253

Anyquery versions up to 0.4.4 contain a path traversal vulnerability in the `clear_plugin_cache` function, which accepts user input and passes it directly to file deletion commands without proper validation. An attacker with API access can use sequences like `../../../../tmp/target` to escape the intended cache directory and delete arbitrary directories on the server.

GHSA-3ww4-5jv9-j5gm: vLLM's Artifact Pin Decay allows pinned deployments to load unpinned code, weights, and processors

mediumvulnerability
security
Jun 10, 2026
CVE-2026-47155

vLLM has a vulnerability called Artifact Pin Decay where revision pinning (locking a model to a specific version) doesn't consistently apply to all files and code that a model needs. When operators use `--revision` to lock their deployment to a reviewed version, vLLM can still load related files like weights, image processors, and configuration from the unpinned default version, breaking the safety guarantee that a pinned deployment serves only reviewed code.

Microsoft restricts Claude Fable for employees over data retention concerns

infonews
securityprivacy

DiffusionGemma: 4x faster text generation

infonews
research
Jun 10, 2026

DiffusionGemma is an experimental open AI model that uses text diffusion (a method that generates multiple words at once instead of one at a time) to achieve up to 4x faster text generation on GPUs compared to traditional language models. Unlike standard LLMs that predict words sequentially, DiffusionGemma generates entire blocks of 256 tokens in parallel, making it useful for speed-critical tasks like real-time editing and code completion, though with lower output quality than standard models.

Turn specs into evals for any agent with ASSERT

infonews
researchsafety

Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable

infonews
safety
Jun 10, 2026

Anthropic released Fable, a limited version of its cybersecurity AI model Mythos, with guardrails (safety restrictions) that block requests related to cybersecurity and biology topics to prevent misuse for creating malware or biological weapons. However, cybersecurity researchers complain the restrictions are overly broad and keyword-based, rejecting even legitimate tasks like code reviews and secure coding practices, though experts acknowledge this is an early-stage approach that may improve over time.

AI Agents Are Becoming Enterprise Workers. Who Secures Them?

infonews
securitysafety

TAPGuard: A Semantic-Aware Graph Framework for TAP Rule Cascading Threat Detection

inforesearchPeer-Reviewed
research

Dual Attention Guided Defense Against Malicious Edits

inforesearchPeer-Reviewed
safety

FIT-Print: Toward False-Claim-Resistant Model Ownership Verification via Targeted Fingerprint

inforesearchPeer-Reviewed
security

SOOM: A Schedule-Search-Based Operator Obfuscation Method Against Model Extraction Attacks

inforesearchPeer-Reviewed
security

Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector

inforesearchPeer-Reviewed
security

Mad or Impossible to Be Mad? Rethinking Load Manipulation Threats in Renewable-Integrated Power Grids and Defenses

inforesearchPeer-Reviewed
security

PAAS: A Policy-Adaptive Anonymity Scheme for Cross-Domain Virtual Asset Service Providers

inforesearchPeer-Reviewed
security

MEC-Dedup: Secure data deduplication for mobile users in edge-assisted cloud storage systems

inforesearchPeer-Reviewed
security

A provably secure identity-based aggregate signcryption scheme for Vehicle-to-Infrastructure communication in VANETs

inforesearchPeer-Reviewed
security

CISO Forum Webinar Today: 2026 Mid-Year Review

infonews
securitypolicy

PRC-linked influence operations are targeting AI debates in the US

infoincident
securitypolicy
Previous188 / 486Next

Fix: PR #5104 adds a `DenyFSAccessThroughSMs` feature that causes the Target Allocator to drop ServiceMonitor and PodMonitor endpoints that reference arbitrary files on the filesystem. When enabled, endpoints with `bearerTokenFile`, `tlsConfig.caFile`, `tlsConfig.certFile`, or `tlsConfig.keyFile` are dropped from the produced scrape configuration while remaining endpoints are kept.

GitHub Advisory Database
GitHub Advisory Database
Hugging Face Security Advisories
Jun 10, 2026

Microsoft is restricting employee access to Claude Fable 5, Anthropic's new AI model, because of concerns about its data retention requirements. While the model is available to external GitHub Copilot and Foundry customers, Microsoft employees cannot access it through their internal tools because Claude Fable 5 does not operate under Zero Data Retention (ZDR, a policy where user data is not stored after interactions) like other Claude models do.

The Verge (AI)

Fix: For applications requiring maximum quality, the source recommends deploying standard Gemma 4 instead. Additionally, the source states that you can improve DiffusionGemma's performance on specific tasks through fine-tuning.

DeepMind Safety Research
Jun 10, 2026

ASSERT is an open-source framework that automatically converts written behavior requirements into evaluation tests for AI systems (like chatbots or agents). Instead of manually creating tests, ASSERT takes plain-language specifications and generates test scenarios, metrics, and scorecards to check whether an AI system behaves as intended, addressing the problem that generic evaluation metrics often miss application-specific requirements.

Microsoft Security Blog

Fix: Anthropic offers a Cyber Verification Program that approved cybersecurity professionals can join to gain fewer limitations on using Claude for cybersecurity work. Additionally, the source notes that Fable is programmed to fall back to Claude Opus 4.8 when it hits a guardrail, allowing users to continue their work with a less restricted model version.

TechCrunch (Security)
Jun 10, 2026

AI agents are now being deployed in companies to automate business workflows, such as managing customer renewal requests by reading emails, accessing CRM (customer relationship management, a database of customer information) data, and taking actions like drafting responses and updating records. Unlike simple text generators, these agents actively read sensitive business data, use system credentials (login information that grants access), and call external tools, which creates new security challenges that organizations need to address.

Check Point Research
safety
Jun 10, 2026

This research proposes TAPGuard, a framework for detecting cascading threats in Trigger-Action Programming (TAP, a system where one event automatically triggers another action, commonly used in smart home devices). The framework uses large language models (AI systems trained on text) to understand the semantic meaning (the actual intent and meaning, not just the structure) of automation rules and identifies two types of threats: explicit ones from direct device interactions and implicit ones from rules sharing environmental variables that shouldn't interact. TAPGuard performs better than existing methods at catching these dangerous rule combinations.

IEEE Xplore (Security & AI Journals)
research
Jun 10, 2026

Text-to-image diffusion models (AI systems that generate images from text descriptions) can be misused to create fake or harmful images, and current defenses using hidden noise patterns don't work well enough. Researchers propose DANP (Dual Attention-Guided Noise Perturbation), a defense method that adds imperceptible perturbations (tiny, invisible changes) to confuse the model's understanding by manipulating how it pays attention to different parts of the image and how it predicts noise during generation.

Fix: The proposed solution is the Dual Attention-Guided Noise Perturbation (DANP) immunization method, which works by: adding imperceptible perturbations across multiple timesteps, using dynamic thresholds to identify text-relevant and irrelevant regions, reducing attention in relevant areas while increasing it in irrelevant ones to misdirect malicious edits, and maximizing the discrepancy between injected noise and the model's predicted noise to further interfere with generation. The authors state this method achieves state-of-the-art performance against malicious edits.

IEEE Xplore (Security & AI Journals)
research
Jun 10, 2026

Existing model fingerprinting techniques (methods that create unique digital signatures to prove ownership of AI models) are vulnerable to false claim attacks, where attackers can fraudulently claim they own models they didn't create. This paper introduces FIT-Print, a targeted fingerprinting approach that uses optimization to create verifiable signatures resistant to these false claims, offering two specific methods (bit-wise FIT-ModelDiff and list-wise FIT-LIME) that achieved 100% success in preventing false ownership claims while maintaining accurate ownership verification.

Fix: The paper proposes FIT-Print, a targeted fingerprinting paradigm that 'actively counters false claim attacks' by leveraging 'optimization to transform the fingerprint into a verifiable, targeted signature.' Two specific black-box fingerprinting methods are introduced: 'bit-wise FIT-ModelDiff' which 'utilizes output distances' and 'list-wise FIT-LIME' which utilizes 'feature attributions as robust model signatures.' The framework demonstrated '100% defense success rate' against false claim attacks and '100% ownership verification rate.'

IEEE Xplore (Security & AI Journals)
research
Jun 10, 2026

Researchers created SOOM, a defense method that obfuscates (hides or disguises) deep learning operators to protect against model extraction attacks, where attackers reverse-engineer compiled neural network code to recreate trainable models. Built on TVM (a deep learning compiler), SOOM uses a machine learning cost model to scramble how operators work while keeping inference fast, achieving a 89% failure rate against extraction attacks with minimal performance slowdown.

Fix: The source proposes SOOM itself as the mitigation: a schedule-search-based operator obfuscation method built on TVM that constructs an obfuscation space for deep learning operators and uses a security-aware learned cost model based on XGBoost gradient boosted trees to generate obfuscated executable code for various deep learning operators, balancing security objectives with performance requirements.

IEEE Xplore (Security & AI Journals)
research
Jun 10, 2026

Researchers developed new methods to perform targeted adversarial attacks (carefully crafted inputs designed to trick AI systems into producing specific harmful outputs) on Large Vision-Language Models, which are AI systems that process both images and text. The attack methods exploit a component called the projector (a part of the model that helps align visual and text information) to make attacks more precise and effective, allowing attackers to modify specific parts of an image while leaving other parts unchanged, and these attacks were shown to work against commercial AI systems like Google Gemini and OpenAI GPT.

IEEE Xplore (Security & AI Journals)
Jun 10, 2026

Researchers discovered a new threat called ToLaR (threat of loads and renewables) that can attack modern power grids with renewable energy sources like solar panels and batteries more efficiently than previous attacks. By manipulating both the electricity demand side and the renewable energy generation side simultaneously, attackers can cause grid instability with only 10% of the resources needed for older attack methods, potentially dropping grid frequency to dangerous levels.

IEEE Xplore (Security & AI Journals)
Jun 10, 2026

PAAS is a system designed to help virtual asset service providers (VASPs, companies that handle cryptocurrencies or digital assets) follow financial regulations that require sharing customer information while still protecting user privacy across multiple companies. The system uses cryptographic techniques to prevent linking a user's anonymous account to their real identity, allow authorities to trace transactions when needed based on specific policies, and safely handle situations where users move assets between different providers.

IEEE Xplore (Security & AI Journals)
Jun 10, 2026

MEC-Dedup is a security approach for mobile users storing data in cloud systems that use edge computing (processing done on devices near the user rather than in distant data centers). The system addresses risks that arise when multiple users' identical files are deduplicated (combined into one copy to save space), which could let attackers identify sensitive information. The research proposes methods to keep user data secure while still allowing the efficiency gains of deduplication in edge-assisted cloud storage.

Elsevier Security Journals
Jun 10, 2026

This academic paper presents a new cryptographic method for secure communication between vehicles and infrastructure in VANETs (vehicular ad hoc networks, which are temporary networks formed by moving vehicles). The scheme uses identity-based aggregate signcryption (a technique that combines digital signatures for authentication with encryption for confidentiality, while processing multiple messages together), and the authors claim to have mathematically proven it cannot be broken by attackers.

Elsevier Security Journals
Jun 10, 2026

This webinar announcement discusses how attackers are using AI to exploit vulnerabilities more quickly, and how security teams can defend using AI-driven tools. Key topics include protecting against Shadow AI (unmonitored use of generative AI in business units) and building AI governance frameworks to manage AI risks in organizations.

SecurityWeek
Jun 10, 2026

OpenAI discovered and banned two clusters of ChatGPT accounts likely from China that were running covert influence operations (hidden campaigns to manipulate public opinion) to shape American debates about AI policy. One cluster spread false claims that data centers were raising electricity prices, while the other criticized US tariffs while excluding China's leader from discussions, and OpenAI is publishing these findings to help the industry, governments, and the public identify and stop similar foreign manipulation attempts.

OpenAI Blog