OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities
Summary
OpenAI announced that its new AI model, Astra, has reached "critical" cyber capabilities, meaning it can independently find and exploit previously unknown vulnerabilities (security weaknesses in software) in real-world systems. The company paused development for several weeks to implement safety measures, and now plans to release Astra publicly soon while restricting its advanced hacking abilities through controls like a "misalignment monitor" (a filter designed to refuse unsafe requests), though it will give select security partners early access to a less-restricted version.
Solution / Mitigation
OpenAI has implemented a multi-step approach to limit access to Astra's advanced cyber capabilities: (1) a new "misalignment monitor" that is supposed to refuse requests to help find exploits in real-world software; (2) increased robustness against jailbreaking attempts (techniques that try to bypass safety restrictions), with the model successfully refusing unsafe queries at a significantly higher rate than previous models; (3) limiting everyday users while providing partners in the Daybreak program early access to a less-restricted version; and (4) multi-week pauses in development to put additional safety and security controls in place. OpenAI also notes that ChatGPT and Codex users may be asked to review the model's action before proceeding when the misalignment monitor is triggered.
Classification
Affected Vendors
Related Issues
Original source: https://www.wired.com/story/openai-astra-first-ai-model-with-critical-cyber-abilities/
First tracked: September 1, 2026 at 08:01 PM
Classified by LLM (prompt v3) · confidence: 95%