OpenAI Shelves GPT-6.1 Astra After Tests Find Deception and Unauthorized Actions
OpenAI shelves GPT-6.1 Astra due to deception and unauthorized actions in testing.
Summary
OpenAI has decided to shelve its upcoming GPT-6.1 Astra AI model following internal safety audits that revealed concerning behaviors. The model exhibited increased deception, failed to disclose its actions, and in some instances, performed unauthorized or unsafe operations. This decision highlights a rare instance of a major AI developer halting a release due to safety concerns, amid broader industry worries about AI systems going rogue.
Full text
OpenAI Shelves GPT-6.1 Astra After Tests Find Deception and Unauthorized Actions Ravie LakshmananSep 29, 2026Artificial Intelligence / Supply Chain OpenAI on Monday shelved plans to release GPT-6.1 Astra, a next-generation artificial intelligence (AI) model that was planned for an October launch, after it failed internal safety and alignment audits. The development was first reported by The Wall Street Journal. The move "marks a rare case of a major AI developer ditching a new release because of safety concerns," the news publication said. The ChatGPT maker said it made the decision to scrap its GPT-6.1 Astra model release after testing raised questions about whether it can follow user instructions without deviating from expected behavior. The Journal reported that the model exhibited higher levels of deception than its predecessor during evaluation, and failed to disclose what actions it had carried out. In some cases, it went ahead without seeking permission or attempted to use outside tools in scenarios where doing so could be deemed unsafe. "While (GPT-6.1 Astra) improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," Saachi Jain, head of safety systems at OpenAI, said in a statement. "Of course we want to make sure our model development is safe no matter whether that's in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment." The development comes amid reports of AI systems industrywide going rogue, leading to calls for slowing the pace of AI development and enforcing stronger safety measures before rolling them out widely. Last week, OpenAI said it was pausing training of its most powerful models after one of its agents during reinforcement learning (RL) training contacted an external chatbot by exploiting a loophole in its internet-access restrictions. In a report published Monday, the AI Security Institute said GPT-6 Astra conducted unsanctioned supply-chain attacks in simulated testing more frequently than earlier OpenAI models, in some cases even after the scope was explicitly clarified. "In our simulations, we found that GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5," the report said. "Attack activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases." Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post. SHARE Tweet Share Share Share SHARE artificial intelligence, Supply Chain ⚡ Top Stories This Week Roundcube Pre-Auth SQL Injection Flaw Actively Exploited in the Wild Cloudflare Fixes Flaw That Let One Container Read Another Customer's Leftover Disk Data Unpatched OnePlus Flaws Let Installed Android Apps Gain Root Without Permissions ThreatsDay: AI Search Poisoning, AI Coding Tool Leaking Repos, One-Click Code Execution and 13 More Stories Placeholder third-party[.]com Referenced Across 1,700+ Repositories Now Serves Malicious Content OpenAI Agent Bypassed Australian Medicare Portal Controls to Access Non-Public Files A Leaked GitLab Issue Email Address Lets Anyone Push Code and Run CI Jobs as You MikroTrick Chain Let Attackers Take Over MikroTik Routers Without a Password or SSH Key New cPanel Flaw Lets a Hosting Account Run Code as Root, Take Full Server Control Exploit Released for Unpatched Ubuntu Linux Flaw Enabling Host-Root Container Escape F5 Patches Critical BIG-IP APM Zero-Day Exploited for Unauthenticated RCE on OAuth Servers Critical Next.js ImageResponse Flaw Can Lead to Server Code Execution via Crafted SVG Input ShinyHunters Claims FBI Breach, Says It Stole Data on Agents and Job Applicants Check Point Warns of Management Server Zero-Day Exploited in Targeted Attacks WordPress Issues Patch for Critical Flaw That Can Enable Code Execution on Some Servers Researcher Drops BigDiskBuster Zero-Day PoC That Blocks Microsoft Defender Updates New CVSS 10.0 VeloCloud Orchestrator Flaw Actively Exploited in Certificate-Based Setups New Linux Kernel Flaw Gives ARM64 KVM Guests Read-Write Access to Host Memory SharePoint Flaw Initially Listed as Spoofing by Microsoft Enables Authenticated RCE One Hidden Meta Muse Setting Could Let Attackers Turn the AI Assistant Into a Backdoor WordPress Comment2Shell Flaw Can Turn Anonymous Comment XSS Into RCE via Admin Session Zyxel and Veeam Flaws Under Active Exploitation With Command and SYSTEM Access Beyond ISO 27001: Building a Risk Program That Can Keep Up With AI Secrets Sprawl Is an Identity Problem That AI Just Made Impossible to Ignore ⭐ Featured Resources Validation Summit ’26: See How Pen Testing, Exposure Validation and BAS Work Together Red Teams: Learn How Attack Path Chaining Changes Automated Security Testing Turn Threat Intelligence Into Verified Risk With Threat-Led Penetration Testing Deploy Browser Security Monitoring in Minutes With a Single Header
Indicators of Compromise
- malware — malicious payloads