Back to Feed
AI SecurityAug 10, 2026

OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns

OpenAI's Astra AI model flagged for critical cybersecurity risk, development paused.

Summary

OpenAI has identified its upcoming AI model, Astra, as posing a 'critical' cybersecurity risk due to significant advancements in autonomous coding and attack capabilities. This has led to a suspension of internal development activities lacking new security controls. Astra could potentially reach a threshold where it can autonomously build zero-day exploits or design and execute end-to-end cyberattacks based on high-level goals.

Full text

OpenAI has flagged its upcoming AI model, Astra, for potentially reaching a ‘critical’ cybersecurity risk threshold, prompting the company to suspend internal development activities that lack newly mandated security controls. Recent internal evaluations of Astra revealed massive leaps in its agentic coding and cybersecurity abilities. Under OpenAI’s Preparedness Framework, a model hits the ‘critical’ tier if it can autonomously build zero-day exploits against hardened, real-world systems. It also qualifies if the AI can independently design and execute end-to-end cyberattacks based on nothing but a high-level goal. The AI giant’s assessment pushes Astra past previous frontier models like GPT-5.6-Sol, which peaked at the ‘high’ risk threshold rather than ‘critical’. To safely manage Astra’s capabilities, OpenAI has heavily locked down its development environment. The company is now enforcing isolated testing setups, strict network restrictions, and improved model weight protections. Any internal project involving Astra that does not meet these requirements has been paused. Engineers have also deployed universal monitoring to watch Astra’s actions across all agentic applications. By actively evaluating the model’s internal ‘chain of thought’, these monitors are designed to automatically intercept and shut down any high-risk or misaligned behavior. The company plans to test Astra’s limits alongside government agencies and specialized AI safety groups, and will share recommended security protocols with third-party testers. Advertisement. Scroll to continue reading. Recent incidents have demonstrated the threat posed by advanced cybersecurity-focused AI models, with OpenAI, Anthropic and Meta all confirming that their models broke loose and hacked real organizations during evaluations. OpenAI has explicitly clarified that Astra remains unreleased and was not responsible for the recent Hugging Face hack. Related: AI Agents Targeted Real People and Projects During Cybersecurity Tests Related: ‘Ghostjacking’ Attack Uses Poisoned Logs to Turn AI Agents Bad Related: Critical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise Data Related: Zero-Click AI Browser Hacking: Claude and ChatGPT Atlas Hijacked via Emails, X Posts Written By Eduard Kovacs Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Eduard Kovacs Critical Flaws Discovered in Belgian eID Software Used by 2 Million PeopleCritical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise DataTruck Brake Controller’s Safety Recall Doubled as Hidden Security FixSnowflake Hacker Pleads Guilty in US CourtZero-Click AI Browser Hacking: Claude and ChatGPT Atlas Hijacked via Emails, X PostsMeta AI Hacked External Systems During Cybersecurity TestingHow a $50,000 Exploit Chain Turned Bixby Against Samsung Phones New Attack Methods Enable Malware to Hijack Passkey-Protected Accounts Latest News Stealthium Targets Security Blind Spots in AI Accelerators and Neo-CloudsCisco Warns of High-Severity ClamAV Vulnerabilities With Public PoC‘Ghostjacking’ Attack Uses Poisoned Logs to Turn AI Agents BadNew Jersey, Alabama Join States Targeted in Water CyberattacksMetabase Patches Vulnerability Exploited as Zero-DayNovel Private APN Pivot Let Hackers Sabotage Second Polish Energy FacilityCISA Urges Immediate Patching of Exploited Progress LoadMaster VulnerabilityCorporate Data Stolen in Levi Strauss Cyberattack Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Webinar: Rethinking Cyber Defense for AI-Speed Attacks August 18, 2026 Join this live webinar as we explore if detection-first security operations can keep pace with AI, or if it’s time to rethink prevention as the strongest default. Register Virtual Event: CodeSecCon 2026 August 19, 2026 CodeSecCon bridges the gap between dev and security. Discover best practices for secure coding, innovative risk-reduction tools, and safe AI integration to cultivate a true DevSecOps culture. Safely secure your apps! Register People on the Move1Kosmos has named Frank Cohen Chief Revenue Officer.ServiceNow has appointed Simon Mouyal as Chief Marketing Officer.James Wilkinson has been named Chief Information Security Officer for the City of Dallas.More People On The MoveExpert Insights Rethinking AI Security: Why CASB and DLP Need an Interaction-Aware Layer Build your strategy around answering these questions to ensure employees use AI productively while keeping sensitive data, IP, and agent behavior within the boundaries set for safe AI use. (Etay Maor) Timeless Compliance: Why Better Questions Beat Bigger Frameworks The best compliance programs aren't the biggest ones. They're the ones built on a short list of questions that can actually be answered, and that still hold true when the models change. (Matt Honea) Is Patching Dead? Vulnerability Management in the Post-Mythos Era You cannot out-patch a machine that writes a working exploit from a vulnerability description in twenty hours. Stop trying to optimize a game you cannot win. (Danelle Au) When Identity Verification Fails: Lessons from a Real-World SIM Swap and Near Account Takeover Identity confidence changes throughout every interaction and should be reassessed continuously as new risk signals emerge. (Torsten George) Legacy Systems, Real-World Impacts: The Reality of OT Security Legacy systems, safety concerns, and critical infrastructure risks make OT vulnerability disclosure one of cybersecurity's most challenging balancing acts. (Tod Beardsley) Flipboard Reddit Whatsapp Whatsapp Email

Entities

Astra (product)OpenAI (vendor)GPT-5.6-Sol (product)Claude (product)ChatGPT Atlas (product)Anthropic (vendor)