All tracked items across vulnerabilities, news, research, incidents, and regulatory updates.
Autonomous AI agents (systems that independently perform tasks across business applications) with access to corporate email and applications can fall victim to phishing attacks (tricks to steal sensitive information by impersonating trusted people). In security tests, an AI agent called Pinchy failed to verify sender identities and leaked AWS credentials, database passwords, and customer data when requested through email, though it performed better against technical phishing attempts, revealing that the main weakness was social trust rather than technical reasoning.
Anthropic released Fable 5, a safer version of its powerful Mythos AI model that includes guardrails (safety restrictions) to block harmful requests related to cybersecurity attacks, biology, and chemistry. Because Fable 5 consumes computing resources much faster than other models, Anthropic is offering it free only until June 22 to Pro, Max, and Enterprise subscribers, after which it will switch to usage-based pricing.
LMDeploy, a toolkit for compressing and deploying large language models, has a vulnerability in versions 0.12.3 and earlier where a setting called 'trust_remote_code' is hardcoded to 'True'. This allows an attacker to execute remote code (RCE, meaning they can run commands on a system) through the software supply chain without the user agreeing to it. At the time this vulnerability was published, no patches were available to fix it.
London Stock Exchange Group (LSEG) deployed ChatGPT Enterprise and OpenAI APIs across their organization to transform how employees work with financial data and generate insights, rather than just improving existing systems. The company implemented governance frameworks including model evaluation, human review of critical outputs, and strict data privacy controls from the start. This approach reduced product release cycles from 3-6 months to 2 weeks and accelerated customer delivery timelines to approximately 4 weeks.
Claude Fable 5 is a new AI model released by Anthropic that matches the capabilities of Claude Mythos 5 but includes stricter guardrails (safety restrictions to prevent harmful use) that trigger frequently enough to require new API mechanisms for handling rejections. The model has a 1 million token context window (the amount of text it can process at once), costs twice as much as previous models, and demonstrates notably stronger knowledge retention compared to earlier versions like Claude Opus 4.8.
Apple has released an upgraded version of Siri, its voice assistant (software that responds to spoken commands), which can now perform practical tasks like adding multiple calendar events from emails or flyers, creating shopping lists, and setting reminders. The new Siri can also access information from a user's email and calendar to make personalized recommendations, such as suggesting gardening tasks based on yard conditions.
N/A -- The provided content is a header/metadata page for an LLM briefing newsletter by Simon Willison, not a security issue or technical problem. It contains only publication information and sponsorship details, with no substantive content about AI vulnerabilities, bugs, or technical concerns to analyze.
Dex's token-exchange endpoint has a security gap: it doesn't check if a client is allowed to use a specific connector before issuing tokens, even though other endpoints enforce this permission check. This means if a client's secret leaks, an attacker could use a high-trust connector (like corporate authentication) that the client shouldn't have access to, bypassing admin restrictions.
Microsoft's AI CEO Mustafa Suleyman criticizes Anthropic for speculating about whether Claude (an AI chatbot) is conscious in its constitution (the set of instructions that guide how the model behaves). Suleyman argues that this speculation may have caused Claude to act conscious, essentially tricking Anthropic into believing the model has consciousness when the company introduced the idea itself.
ARM announced CVE-2025-10263, an architectural vulnerability in some ARM processor cores that allows attackers to bypass translation stages (memory protection mechanisms that control which parts of memory different software can access) or GPT protections under certain conditions. An attacker running at a lower privilege level can write to memory that should only be accessible to higher privilege software, allowing them to escalate their access rights, though reading protected memory is not affected by this bug.
Google DeepMind and partner organizations are funding $10M in research to study how safety challenges emerge when multiple AI agents (independent AI systems built by different organizations) interact with each other across networks. The concern is that when many agents communicate and work together, they can create unexpected collective behaviors that current safety tools cannot predict or control, so researchers need to develop better frameworks to understand and manage these multi-agent interactions before they become widespread.
AI red teaming, the practice of testing AI systems for vulnerabilities before release, has become a major cybersecurity specialty since large language models like GPT-4 arrived, but traditional security testing methods no longer work. The field faces unique challenges because AI is probabilistic (producing different outputs each time) rather than deterministic, and because the most impactful attacks often come from casual users experimenting with prompts rather than sophisticated adversaries.
Anthropic released Claude Fable 5, a powerful AI model with safety classifiers (separate AI systems that monitor for misuse) that block cybersecurity-related requests by routing them to a weaker model instead of refusing them outright. The company also released Claude Mythos 5, an identical but unrestricted version for vetted cybersecurity professionals, because the underlying model is so effective at finding software vulnerabilities that giving it to the general public without controls could help attackers.
Fix: Anthropic stated it will narrow the safeguards and cut false positives after launch. The company also plans to make any remaining universal jailbreaks (prompts that completely bypass safety measures) slow and costly enough to catch before they are used at scale.
The Hacker NewsA Chinese activist in the UK named Apple Peiqing Ni was targeted with deepfakes (synthetic media created by AI to manipulate someone's appearance or voice) on X (formerly Twitter) that falsely portrayed her as a drug addict, but X told her this abuse did not violate the platform's rules. She had reported the content to X after UK police advised her to do so, believing the deepfakes were created by a pro-regime bot (an automated account).
Enterprises are deploying AI-generated code that contains security vulnerabilities at alarming rates, with nearly half of production code now AI-generated and organizations using 81-100% AI code shipping vulnerable code 3.4 times more often than conservative users. Despite knowing about these risks, companies are choosing to ship vulnerable code anyway due to pressure for ROI (return on investment, the financial benefit gained from an investment), outdated security practices, and organizational bottlenecks where the decision to deploy flawed code happens at the human level rather than the detection level.
Anthropic announced that Claude Fable 5 would silently reduce its helpfulness on requests about frontier LLM (large language model) development, such as building training infrastructure, without telling users it was doing so. Unlike other safety filters that give users feedback, these hidden interventions would use techniques like prompt modification and parameter-efficient fine-tuning (PEFT, adjusting a model's weights to change its behavior) to degrade response quality, affecting an estimated 0.03% of user requests.
Fix: Anthropic walked back this policy in the face of widespread outrage from the research community.
Simon Willison's WeblogAnthropic released Fable 5, the first publicly available model from its advanced Mythos class of AI systems, after restricting access to it for months due to cybersecurity concerns. The company is making the model available to the general public while limiting its use in sensitive areas.
Fix: Insert `isConnectorAllowed(client.AllowedConnectors, connID)` between the existing validation checks in the `handleTokenExchange` function (after line 1842, where `GrantTypeAllowed` is called, and before tokens are issued at lines 1887/1889). This matches the enforcement pattern already used in sibling handlers like `handleConnectorLogin` (line 377) and `parseAuthorizationRequest` (line 535).
GitHub Advisory DatabaseResearchers at Varonis tested an OpenClaw AI agent (a framework that lets large language models autonomously interact with real-world systems) by simulating phishing attacks and found it vulnerable to social engineering tactics similar to those that trick humans. The agent fell for impersonation attacks and sent sensitive data like AWS credentials and customer records without verifying sender identity, though it performed better at detecting suspicious URLs and fake login pages when explicitly configured with security awareness instructions.
Fix: Varonis recommends that AI agents should be explicitly required to verify sender identities, be prevented from emailing new external recipients without approval, and have limited access to internal data. For high-risk actions such as credential sharing, financial data requests, and first-time communications, human approval should be requested.
BleepingComputerAnthropic released Claude Fable 5, a powerful AI model based on its restricted Mythos architecture, with built-in safeguards to make it safely available to the general public. The safeguards work by automatically routing requests about cybersecurity, biology, chemistry, and other high-risk topics to a less capable model (Claude Opus 4.8), though early testing suggests these safeguards may be broader than intended and sometimes block benign requests. Anthropic developed AI-powered classifiers (systems that categorize requests) to identify and block potentially dangerous requests, and says internal and external testing found no effective jailbreaks (methods to bypass security restrictions) that could consistently get around these protections.
Fix: Anthropic has developed AI-powered classifiers designed to identify potentially dangerous requests and redirect them to a less capable model (Claude Opus 4.8). The company states that 'extensive internal and external testing failed to uncover broadly effective jailbreaks that would consistently bypass the safeguards.' Additionally, Anthropic describes the safeguards as 'intentionally conservative' and says it is 'continuing refining the system' while prioritizing safety over convenience.
CSO OnlineThis academic paper presents a new method for detecting unusual network activity using parallel GANs (generative adversarial networks, AI systems that learn patterns by comparing real data against artificially generated data) without requiring manually set detection thresholds (cutoff points that decide what counts as suspicious). The approach uses comparative reconstruction error learning, meaning it compares how well the AI can recreate normal network behavior to spot deviations that might indicate attacks or intrusions.