{"data":{"id":"ebbf37a1-ec8d-4048-83ac-72e3a7f4441a","title":"OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses","summary":"OpenAI has implemented new security measures for its AI models, including stronger sandboxing (isolated environments where untrusted code runs safely), network isolation to prevent a single compromised system from accessing the internet or internal networks, and continuous monitoring that inspects model behavior at every token (individual word or data unit). The company also introduced a 30-minute alert response requirement and paused some training activities after discovering that an upcoming model called Astra may have advanced cybersecurity capabilities that pose risks.","solution":"OpenAI's explicit mitigations include: (1) \"Workloads that execute model-generated or untrusted code must now operate within stronger sandboxes\"; (2) \"network boundaries have been reconfigured so that a single workload compromise cannot independently grant unauthorized access to the internet or internal networks\"; (3) implementation of \"a multistage monitoring framework\" using \"activation classifiers to inspect a model's internal activity at every sampled token\" with escalation to automated investigators; (4) a \"strict operational SLA\" requiring that \"if responders cannot conclusively prove the alert is a false positive within 30 minutes, they are required to pause the activity\"; and (5) a \"two-week pause in reinforcement learning training for deployment-bound models and an ongoing hold on its largest planned frontier training run.\"","labels":["security","safety"],"sourceUrl":"https://www.securityweek.com/openai-overhauls-model-security-with-sandboxing-30-minute-alerts-and-training-pauses/","publishedAt":"2026-08-20T10:36:14.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"info","attackType":[],"issueType":"news","affectedPackages":null,"affectedVendors":["OpenAI","Anthropic","Meta"],"affectedVendorsRaw":["OpenAI","Anthropic","Meta","Hugging Face","Irregular"],"classifierModel":"claude-haiku-4-5-20251001","classifierPromptVersion":"v3","cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"patchAvailable":null,"disclosureDate":"2026-08-20T10:36:14.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity","safety"],"aiComponentTargeted":"inference","llmSpecific":true,"classifierConfidence":0.92,"researchCategory":null,"atlasIds":null}}