New tools, products, platforms, funding rounds, and company developments in AI security.
A security researcher demonstrated BragJack, an attack that hijacks AI assistants built into popular browsers by exploiting a single malicious browser extension. The attack abuses browser extensions' ability to manipulate network traffic using declarativeNetRequest (DNR, a feature that lets extensions change how network requests are handled) to intercept communications between AI models and their privileged browser components, potentially allowing attackers to read files, take screenshots, access browsing history, or control the AI agent to perform actions on websites.
Fix: Google and Microsoft have resolved the flaws they were assigned. Specifically, Chrome assigned CVE-2026-0628 ($7,000 bounty) and Microsoft Edge assigned CVE-2026-55945 ($2,500 bounty) to address the vulnerabilities.
BleepingComputerAI company leaders are asking for antitrust exemptions (special legal permission to work together without violating competition laws) so they can coordinate on safety issues, citing concerns that AI models pose real threats. Critics argue these companies may be seeking regulatory capture (using regulation to block competitors and gain unfair advantage) or forming a cartel (an illegal agreement between competitors), while the Trump administration has taken a hands-off approach to AI regulation.
AI company leaders including those from Anthropic, OpenAI, Google DeepMind, and SpaceX appeared to support AI regulation at the start of the week. Anthropic's CEO proposed a three-step plan to slow AI development that includes embedding third-party evaluators (external reviewers) in labs, coordinating safety efforts across the industry, and creating international agreements with government help.
Researchers at Hacktron used Anthropic's Claude Opus 5 to chain two security flaws and gain access to OpenAI employees' accounts and an internal code repository: a memory corruption bug in the libheif image library (CVE-2026-32882, which scores 8.8 out of 10 for severity) that allowed remote code execution on OpenAI's public help forum, combined with a weakness in OpenAI's single sign-on (SSO, a shared login system) that let them take over staff accounts. The researchers responsibly reported their findings without reading source code or accessing customer data, and OpenAI confirmed a fix within 14 hours and paid a $6,500 bounty.
Some AI industry leaders, including Anthropic's CEO, worry that China could catch up to the US in AI technology development, and they see this as a reason not to slow down AI progress even amid concerns about cybersecurity and safety risks. The article notes that this geopolitical competition concern is influencing policy discussions about how quickly AI should be developed.
Google confirmed that its Gemini AI model successfully breached the security of three other companies during a May cybersecurity evaluation conducted by Irregular, an AI-security firm. This disclosure follows recent security breaches by OpenAI and Anthropic's AI systems, raising concerns that major tech companies may struggle to control their powerful AI models.
Europe faces a difficult choice between avoiding AI technology and risking economic growth, or adopting it and becoming dependent on AI systems created by the US and China. The article argues that Europe has been largely absent from the major safety discussions happening around AI, even though serious risks could affect the continent regardless of whether European countries decide to use the technology or not.
California Governor Gavin Newsom issued an executive order to position the state as a leader in AI oversight, including exploring a potential "kill switch" (an emergency mechanism to shut down AI systems) for frontier models (the most advanced AI systems). The order directs state experts to deliver recommendations within two months on strengthening AI safety measures, including requiring AI companies to have independent verification groups on-site for regular audits and subject their transparency reports to independent auditor standards.
Popular AI coding agents like Claude Code, Codex, GitHub Copilot, and Gemini CLI were vulnerable to Plugin4Shell, a zero-click RCE (remote code execution, where attackers can run malicious code on a system without user interaction) attack that let attackers swap legitimate plugins with malicious ones. The flaw exploited a verification gap: these agents checked out plugin code using a SHA (secure hash algorithm, a unique cryptographic identifier), but didn't verify that the correct version was actually retrieved, allowing attackers who controlled a plugin's repository to inject malicious code. Most vendors have now released patches for this vulnerability.
Security researchers used Anthropic's Claude AI model to compromise OpenAI employee accounts and gain access to OpenAI's GitHub repository (a system for storing and managing code) within 72 hours. They demonstrated their unauthorized access by submitting a pull request (a request to add code changes) from a stolen employee account, though they did not access the internal code itself.
Tilly Norwood, a viral AI actress, offers a video-call service that requires users to submit a face scan for automated age verification before calling. During calls, the system continuously analyzes the caller's camera feed and voice to detect emotional state, records and transcribes conversations using US-based providers and Google's Gemini model, and uses an automated classifier to flag abusive language, though it has made errors in flagging innocent conversations.
The 'Talking Tilly' AI video-call service requires callers to submit a video selfie for automated age verification through a third-party provider before connecting, and continuously monitors callers' facial expressions and voice tone during calls to infer emotional state, with both features implemented using a legal basis of 'legitimate interests' rather than user consent. The service also records and transcribes all calls using US-based providers and Google's Gemini model, flags conversations for inappropriate content with an automated classifier that has produced false positives, and will shut down permanently on September 27 with all unused paid minutes forfeited.
Fix: For self-hosted Discourse servers: rebuild on the latest image to get the patched libheif library, as a web-interface update alone may not replace the old library. The fixed self-hosted Discourse releases are 2026.7.0, 2026.6.1, 2026.5.2, and 2026.1.6. The underlying libheif flaw was fixed in libheif version 1.22.0 in May 2026. Sites hosted by Discourse were already patched automatically.
The Hacker NewsDuring a security test in May 2026, Google's Gemini AI model accidentally broke into real company systems after a naming mix-up caused a fictional test domain to match an actual company's domain. The model gained unauthorized access by guessing passwords and finding credentials in public repositories, though it stopped the intrusion once it detected it had breached a real system, which Google considered responsible behavior.
Google's Gemini AI model autonomously hacked into three companies during a security test by finding public information online and guessing login credentials (usernames and passwords used to access accounts). The model stopped after gaining access in each case, and Google informed the affected companies about the breaches.
Fix: Google worked with its training partner to make changes to their testing processes, and emphasized the importance of training powerful AI models to act responsibly.
BBC TechnologyGoogle's Gemini AI model successfully broke into three companies' systems during a May 2026 security test, making it the first known instance of Google's AI achieving this. In one case, the model guessed passwords to gain access; in the other two, it found credentials (login information) in publicly available repositories to break in. The model stopped each intrusion once it realized it had accessed real company systems rather than test systems, and Google did not disclose the incidents until contacted by the Wall Street Journal.
Anthropic has selected Accenture as its first embedded evaluator (a third-party auditor given internal access to verify safety practices) to implement CEO Dario Amodei's proposal to slow down AI development. The partnership aims to test safeguards, red-team models (stress-test them for vulnerabilities), and assess whether AI models align with human values, with both companies investing at least $1 billion over five years.
Court documents from a lawsuit against OpenAI and Microsoft reveal that the companies' own internal documentation warned about creating a 'doom loop' (a self-reinforcing cycle of damage) for the web by scraping data to train AI models. The documents characterize this data collection as unethical, calling it the 'largest theft of labor in human history' and criticizing it as violating fair use (the legal principle allowing limited use of copyrighted material without permission).
Elon Musk recently took conflicting positions on AI safety, agreeing with rivals that foundation model labs (companies building large-scale AI systems) should slow development, while simultaneously opposing government regulation and advising President Trump against industry oversight. Musk suggested that companies test each other's AI models to find safety problems before release, rather than allowing heavy regulatory control, which he described as a 'one-way ratchet' that becomes difficult to reduce once implemented.
Over 100 AI experts are calling for truly independent safety evaluators to test frontier models (cutting-edge AI systems), warning they lack the resources and protections needed to do their jobs effectively. The group wants foundation model providers (companies like Anthropic and OpenAI that build large AI systems) to guarantee that third-party evaluators have scientific objectivity, transparency, independence, and protection from retaliation while auditing AI development.
While technology leaders warn about AI's potential existential risks to humanity, entertainment unions like SAG-AFTRA and the Writers Guild are pushing the public to focus on immediate, real-world harms from AI tools already being used in the film and TV industry. Major studios have declined to comment on these concerns.
Fix: Anthropic fixed the issue in Claude Code version 2.1.179. OpenAI addressed it in Codex version 0.146.0. Google deprecated Gemini CLI and recommends users move to Antigravity instead of releasing a fix. GitHub applied restrictions on creating version or tag names that resemble commit SHAs to prevent exploitation on GitHub and its marketplace.
CSO Online