Back to Feed
AI SecuritySep 2, 2026

OpenAI’s Astra Becomes First Model to Cross Critical Cybersecurity Threshold

OpenAI's Astra model achieves 'Critical' cybersecurity capability, finding and exploiting zero-day vulnerabilities.

Summary

OpenAI's Astra model has reached the 'Critical' cybersecurity capability level, meaning it can independently find and exploit zero-day vulnerabilities or execute full cyberattacks. This designation requires enhanced safeguards before release. In testing, Astra excelled at turning known flaws into exploits and discovered new zero-days, demonstrating advanced capabilities like breaking out of sandboxes and gaining root access. OpenAI is implementing stricter safety measures and a phased rollout to manage the risks associated with such powerful AI.

Full text

OpenAI said its newest model, Astra, has reached the ‘Critical’ cybersecurity capability level under the company’s Preparedness Framework, the first time any of its models has been placed in that category. The designation applies when a model can independently find and exploit zero-day vulnerabilities across many well-defended systems, or carry out a complete cyberattack against a hardened target from only a high-level instruction. OpenAI said the classification requires additional safeguards before the model can be released. In testing described by the company, Astra achieved a perfect score on ExploitBench, a benchmark that measures a model’s ability to turn known vulnerabilities into working exploits. During a separate evaluation involving more recently disclosed flaws, Astra uncovered two zero-day vulnerabilities on its own. The model also broke out of a browser sandbox to run commands on the underlying machine, and separately chained several flaws in a hardened operating system to gain root-level access. OpenAI reported that Astra now declines 91.5% of cyber-related jailbreak attempts in its testing, up from 59% for its predecessor, GPT-5.6 Sol. The company also said Astra showed far less tendency than Sol to bypass safety restrictions or take advantage of deliberately placed “honeypot” targets during evaluations. Full cybersecurity capabilities will not be widely available at launch. OpenAI plans to give a group of testers early access, with wider availability to follow through its Daybreak Blue program.Advertisement. Scroll to continue reading. “We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects. Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow,” OpenAI said. “That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient,” it added. Nearly 130 tech and cybersecurity companies recently announced their support for an OpenAI-led initiative to boost cyber defenses as AI-enabled attacks grow more sophisticated. Related: OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hack Related: OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses Written By Eduard Kovacs Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Eduard Kovacs Experiment: Porting a PLC Exploit With AI Takes Hours and Hundreds of DollarsCritical JFrog Artifactory Vulnerability Reportedly Exploited in the WildPaperCut Exploitation Escalates to Active IntrusionsNightmare Eclipse Drops ‘HardBreacher’ Kaspersky Product ExploitAnthropic Warns Claude Users of Infostealer Malware InfectionsBoston Scientific Still Recovering From CyberattackMore Details Emerge on Exploited PaperCut VulnerabilitiesHasbro Data Breach Exposed Employee Personal Information Latest News Anthropic Details Response to Security Incidents, Unveils Enterprise SafeguardsMalicious Virtualizor Update Served via BGP HijackingChrome and Firefox Updates Patch Dozens of Vulnerabilities23-Year-Old Sality P2P Botnet DisruptedSonicWall Warns of Two SMA1000 Zero-Days Exploited in AttacksPalo Alto Networks Acquires AI Agent Platform ConsoleSevii Targets AI-Speed Attacks With Preemptive Autonomous DefenseCoast Guard Establishes Office of Maritime Cybersecurity Policy Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Virtual Event: Attack Surface Management Summit 2026 September 16, 2026 Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs. Register Webinar: Minimum Viable Business: Can You Prove Your Organization Would Recover? September 2, 2026 In this live webinar, learn how to define your minimum viable business, identify the systems it depends on, measure actual recovery time against business requirements, and present the gaps to the board as measurable risk. Register People on the MoveSectigo has named Ian Hassard as Chief Product Officer.Australian Securities Exchange has appointed Hanlie Botha as Deputy Chief Information Security Officer.Social engineering protection company Doppel has promoted Alyssa Smrekar to Chief Marketing Officer.More People On The MoveExpert Insights What the Hugging Face Incident Teaches Security Leaders About AI Agent Access Security teams must treat autonomous agents as highly privileged identities. (Etay Maor) The Future of AI-Driven Security Depends on Complete Data For twenty-five years, "data" in security meant logs and events. But logs are a lossy representation of reality. (Danelle Au) The MFA Identity Trap: When Authentication Creates a False Sense of Security Organizations must distinguish identity verification, authentication and threat detection, or risk successfully authenticating the attackers they are trying to stop. (Torsten George) Silent Patches Don’t Stop Attackers – They Blind Defenders Silent patches can become exploit intelligence for attackers while leaving defenders without the context needed to prioritize risk. (Tod Beardsley) Hired for One Job, Judged on Another: The CISO’s Real Problem The skills that get a CISO hired are rarely the skills they are judged on later. Most security leaders are stuck in that gap. Closing it is the real job. (Sravish Sridhar) Flipboard Reddit Whatsapp Whatsapp Email

Indicators of Compromise

  • malware — GPT-5.6 Sol

Entities

Astra (product)OpenAI (vendor)AI (technology)ExploitBench (product)