GPT-6 Astra Scores 100% on ExploitBench as OpenAI Blocks PoC Exploit Requests
Summary
OpenAI released GPT-6 Astra, a new AI model that scored 100% on ExploitBench (a test measuring how well an AI can turn known software vulnerabilities into working exploits), compared to 78.5% for the previous model. To prevent misuse, the released version refuses requests to create proof-of-concept exploits (working examples of attacks), though OpenAI plans to expand access with fewer safeguards for defensive security work in coming weeks. The model also includes stronger safety measures like jailbreak resistance and detection systems to catch misalignment.
Solution / Mitigation
OpenAI limited the released version of Astra to 'secure code review and patching, while refusing to comply with prompts related to creating proof-of-concept (PoC) exploits for vulnerabilities.' The company also added 'stronger model robustness to better tackle jailbreaks, more context to its monitoring systems, and extra safeguards to help detect and contain misalignment.' Additionally, safety checks are in place that 'proceed with care commensurate with its risk' in sensitive environments.
Classification
Affected Vendors
Related Issues
Original source: https://thehackernews.com/2026/09/gpt-6-astra-scores-100-on-exploitbench.html
First tracked: September 4, 2026 at 08:01 AM
Classified by LLM (prompt v3) · confidence: 92%