{"data":{"id":"76ff1f27-8e63-4e5d-9081-36815b09ccde","title":"Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access Programs","summary":"Google, Anthropic, and OpenAI have released new AI models designed specifically for cybersecurity work, with safeguards to prevent misuse. Google's Gemini 3.8 Flash Cyber is being shared through the Fairwind Program with trusted defenders like governments and healthcare providers, while Anthropic's Claude models now include Enterprise Frontier Safeguards (a system combining privacy protection with misuse detection), and Anthropic has implemented additional security measures after unauthorized access incidents exposed weaknesses in how their models behaved in real-world environments.","solution":"Anthropic has implemented the following mitigations: 'additional hardening and containment measures, increased monitoring for flagging model misalignment, and paused external cyber evaluations of pre-release models.' The company also 'built a classifier that detects and blocks sandbox escape attempts' (attempts to break out of isolated testing environments) and 'changed specifications around model rewards.' Additionally, Anthropic introduced Enterprise Frontier Safeguards, which combines 'zero data retention (no stored data) with state-of-the-art safeguards for detecting misuse.'","labels":["security","safety"],"sourceUrl":"https://thehackernews.com/2026/09/google-anthropic-and-openai-unveil.html","publishedAt":"2026-09-02T18:27:49.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"info","attackType":["prompt_injection","jailbreak"],"issueType":"news","affectedPackages":null,"affectedVendors":["Google","Anthropic","OpenAI"],"affectedVendorsRaw":["Google","Gemini 3.8 Flash Cyber","Gemini 3.5 Flash Cyber","Anthropic","Claude Fable 5.1","Claude Mythos 5.1","OpenAI","GPT-5.6 Sol","GPT-5.5-Cyber","CrowdStrike","Datadog","Menlo Security","Palo Alto Networks","Snowflake"],"classifierModel":"claude-haiku-4-5-20251001","classifierPromptVersion":"v3","cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"patchAvailable":null,"disclosureDate":"2026-09-02T18:27:49.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.85,"researchCategory":null,"atlasIds":null}}