Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models
Summary
SentinelOne created a benchmark test using the Fast16 malware (a 2005 Windows program designed to sabotage Iran's nuclear weapons development) to evaluate how well frontier AI models can conduct long-horizon reverse-engineering, which is the process of analyzing software to understand how it works. GPT-5.6 Sol was the only model tested that completed all eight stages of the investigation, while other models like GPT-5.5, GLM-5.2, and Anthropic's Opus struggled with what researchers call "project-scale recovery," or the ability to fix errors and trace their consequences throughout an investigation. The researchers concluded that human oversight remains essential because even the best-performing AI made technical mistakes and needed human analysts to validate conclusions.
Solution / Mitigation
According to SentinelLabs researchers, "the best current use [of these AI models] is supervised investigative agency, with human analysts defining objectives, exposing blind spots, and retaining final publication authority." The source emphasizes that "Senior reverse engineers remain essential" to oversee AI-assisted investigations.
Classification
Affected Vendors
Related Issues
Original source: https://www.securityweek.com/nuclear-sabotage-malware-benchmark-trips-up-most-frontier-ai-models/
First tracked: July 23, 2026 at 02:01 PM
Classified by LLM (prompt v3) · confidence: 85%