{"data":{"id":"df28a8d6-e77f-4b94-98db-307843cabb71","title":"GPT-6 Astra Scores 100% on ExploitBench as OpenAI Blocks PoC Exploit Requests","summary":"OpenAI released GPT-6 Astra, a new AI model that scored 100% on ExploitBench (a test measuring how well an AI can turn known software vulnerabilities into working exploits), compared to 78.5% for the previous model. To prevent misuse, the released version refuses requests to create proof-of-concept exploits (working examples of attacks), though OpenAI plans to expand access with fewer safeguards for defensive security work in coming weeks. The model also includes stronger safety measures like jailbreak resistance and detection systems to catch misalignment.","solution":"OpenAI limited the released version of Astra to 'secure code review and patching, while refusing to comply with prompts related to creating proof-of-concept (PoC) exploits for vulnerabilities.' The company also added 'stronger model robustness to better tackle jailbreaks, more context to its monitoring systems, and extra safeguards to help detect and contain misalignment.' Additionally, safety checks are in place that 'proceed with care commensurate with its risk' in sensitive environments.","labels":["security","safety"],"sourceUrl":"https://thehackernews.com/2026/09/gpt-6-astra-scores-100-on-exploitbench.html","publishedAt":"2026-09-04T06:47:52.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"info","attackType":[],"issueType":"news","affectedPackages":null,"affectedVendors":["OpenAI"],"affectedVendorsRaw":["OpenAI","GPT-6 Astra","ChatGPT Plus","ChatGPT Pro","ChatGPT Business","ChatGPT Enterprise","Microsoft Azure","Amazon Web Services (AWS) Bedrock"],"classifierModel":"claude-haiku-4-5-20251001","classifierPromptVersion":"v3","cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"patchAvailable":null,"disclosureDate":"2026-09-04T06:47:52.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.92,"researchCategory":null,"atlasIds":null}}