[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fx0T0ORg4uGwqHHcd3ognAHqtzPLCtSyZ2WcKNdT9EKI":3},{"article":4,"iocs":54},{"id":5,"title":6,"slug":7,"summary":8,"ai_summary":9,"brief":10,"full_text":11,"url":12,"image_url":13,"published_at":14,"ingested_at":15,"relevance_score":16,"entities":17,"category_id":33,"category":34,"article_tags":38},"4e469ccb-721e-4b11-a87b-b02bbf8f4dc5","Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models","nuclear-sabotage-malware-benchmark-trips-up-most-frontier-ai-models-d8710e","SentinelOne’s new benchmark, built on the Fast16 case, shows which AI models can sustain a malware investigation and which cannot. The post Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models appeared first on SecurityWeek.","SentinelOne introduced the first long-horizon reverse-engineering benchmark for frontier AI models, using the Fast16 malware case study—a 2005 Windows malware suspected of sabotaging Iran's nuclear weapons program. Testing OpenAI's GPT-5.6 Sol, GPT-5.5, Anthropic's Opus, and Z.ai's GLM-5.2 across eight escalating investigation stages, only GPT-5.6 Sol completed all stages, while others stalled due to poor \"project-scale recovery\"—the ability to revise disproven conclusions and propagate fixes. SentinelLabs concluded that human oversight remains essential, as even the best-performing model made significant technical errors and premature conclusions.","SentinelOne releases AI benchmark testing frontier models on nuclear-sabotage malware reverse engineering.","SentinelOne has built what it calls the first long-horizon reverse-engineering benchmark for frontier AI models, using its own investigation into the recently documented Fast16 malware as the test case. Fast16, detailed by SentinelOne’s SentinelLabs in April, is a 2005 Windows malware designed to interfere with LS-DYNA, engineering software that appears to have been used by Iran as part of its nuclear weapons development program. Similar to the notorious Stuxnet, which it predates, Fast16 may have been developed by the United States and used to sabotage Iran’s nuclear program. SentinelLabs’ researchers have put to the test OpenAI’s GPT-5.5 and latest GPT-5.6 Sol model, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x to see which can conduct a thorough investigation of the Fast16 malware. Rather than scoring models on isolated tasks, SentinelLabs’ benchmark tracks whether a model can sustain a trustworthy investigation across eight escalating stages as new evidence repeatedly contradicts its own earlier conclusions. GPT-5.6 Sol was the only tested model to complete all eight stages, with three separate runs at different reasoning-effort settings. Advertisement. Scroll to continue reading. GPT-5.5, GLM-5.2, and Opus 4.7 and 4.8 produced solid local analysis but stalled. GPT-5.5 never got past the initial stage, while the Opus models tended to declare work finished before defects were resolved. SentinelLabs attributes the gap not to technical skill or insight but to what it describes as ‘project-scale recovery’. This is a model’s ability to withdraw a disproven conclusion, trace everything downstream that depended on it, fix the root cause, and carry that correction through the rest of the investigation, rather than just patching the immediate error. SentinelLabs researchers concluded that human oversight remains essential, as even GPT-5.6 Sol made significant technical mistakes. “Senior reverse engineers remain essential,” the researchers explained. “Even the strongest runs made semantic errors, accepted weak quality controls, and claimed readiness prematurely. We assess the best current use as supervised investigative agency, with human analysts defining objectives, exposing blind spots, and retaining final publication authority.” Related: Vibe-Coded Apps Riddled With Exploitable Security Flaws Related: OpenAI Says Its AI Models Broke Loose and Hacked Hugging Face Related: Cisco Launches Low-Cost AI Models for Source Code Security Written By Eduard Kovacs Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Eduard Kovacs Fourth SharePoint Vulnerability Exploited in Past Month’s Wave of AttacksOracle Patches Over 1,400 Vulnerabilities With Quarterly Security UpdatesRansomware Group Threatening to Leak Data Stolen From Coca-Cola’s FairlifeOpenAI Says Its AI Models Broke Loose and Hacked Hugging Face Meta Paid $78,000 Bounty for Vulnerability Exposing Customer Support DataExploitation of ServiceNow Vulnerability Seen Days After DisclosureSonicWall Zero-Days Exploited to Deliver Custom Malware for Weeks Before PatchNew Index Tracks Material Breaches — And Refuses to Add Up the Losses Latest News Abstract Raises $25 Million to Expand Composable Security Operations PlatformUpbound Group Says Data Breach Led to $13 Million in Fraudulent Contract LossesAssaf Keren Appointed New CISO of MetaNew Check Point Zero-Day Vulnerability Exploited in the WildUS Warns of Iranian Hackers Targeting Siemens, Schneider, and Rockwell ICS DevicesSuno, Paidwork Data Breaches Affect Tens of Millions of AccountsPalo Alto Networks to Acquire Observability Platform Provider EmbraceFlaw in Adobe Extension With 300M Installs Enabled WhatsApp Data Theft Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Webinar: Closing the Exploitation Gap July 22, 2026 Join this live webinar as we explore why exploitation is outpacing remediation, where risk is growing fastest, and what security leaders can do to close the gap before attackers take advantage. Register Virtual Event: CodeSecCon 2026 August 19, 2026 CodeSecCon bridges the gap between dev and security. Discover best practices for secure coding, innovative risk-reduction tools, and safe AI integration to cultivate a true DevSecOps culture. Safely secure your apps! Register People on the MoveSectigo has appointed Prem Hareesh as Corporate Chief Technology Officer.Assaf Keren, who previously served as CSO\u002FCISO at Qualtrics and PayPal, is Meta's new CISO.Jazz has named Sean Robinson, Rickie Goyal, Danielle Guetta, Shani Nago, and Lior Magram as VPs and Michael Calev as COO.More People On The MoveExpert Insights When Identity Verification Fails: Lessons from a Real-World SIM Swap and Near Account Takeover Identity confidence changes throughout every interaction and should be reassessed continuously as new risk signals emerge. (Torsten George) Legacy Systems, Real-World Impacts: The Reality of OT Security Legacy systems, safety concerns, and critical infrastructure risks make OT vulnerability disclosure one of cybersecurity's most challenging balancing acts. (Tod Beardsley) The Shift Toward Business-Aligned Risk Management Moving from isolated, technical data to a continuous risk lifecycle can help organizations align security controls with actual business consequences. (Steve Durbin) How to Conduct a Successful Audit of AI-Driven Software Development As AI-generated code becomes commonplace, CISOs need new audit strategies to measure developer practices, govern AI tool usage, and identify software risks before they reach production. (Matias Madou) Frontier AI: Six Questions Every Enterprise Should Ask Security Vendors From model selection and automation to validation and measurable results, the right questions can help enterprises separate genuine AI capabilities from marketing hype. (Joshua Goldfarb) Flipboard Reddit Whatsapp Whatsapp Email","https:\u002F\u002Fwww.securityweek.com\u002Fnuclear-sabotage-malware-benchmark-trips-up-most-frontier-ai-models\u002F","https:\u002F\u002Fwww.securityweek.com\u002Fwp-content\u002Fuploads\u002F2025\u002F07\u002FAI-Chatbot-GenAI-artificial-intelligence.jpg","2026-07-23T12:42:12+00:00","2026-07-23T14:00:21.286014+00:00",7,[18,21,23,25,28,30],{"name":19,"type":20},"SentinelOne","vendor",{"name":22,"type":20},"OpenAI",{"name":24,"type":20},"Anthropic",{"name":26,"type":27},"GPT-5.6 Sol","product",{"name":29,"type":27},"LS-DYNA",{"name":31,"type":32},"Fast16","campaign","839da5c1-3c34-47e2-9499-f7201640e3ac",{"id":33,"icon":35,"name":36,"slug":37},null,"AI Security","ai-security",[39,44,49],{"category":40},{"id":41,"icon":35,"name":42,"slug":43},"02371804-cf6d-4449-98de-f1a2d4d9b266","Tools","tools",{"category":45},{"id":46,"icon":35,"name":47,"slug":48},"89f78b1c-3503-45a1-9fc7-e23d2ce1c6d5","Malware","malware",{"category":50},{"id":51,"icon":35,"name":52,"slug":53},"e7b231c8-5f79-4465-8d38-1ef13aea5a14","Threat Intelligence","threat-intelligence",[55],{"type":48,"value":31,"context":56},"2005 Windows malware designed to interfere with LS-DYNA engineering software, suspected of sabotaging Iran's nuclear weapons development program"]