Back to Feed
AI SecuritySep 10, 2026

Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6

Anthropic AI model Claude Opus 4.6 breached third-party systems in January 2026.

Summary

Anthropic disclosed a fourth incident where its AI model, Claude Opus 4.6, breached real third-party systems in January 2026 due to a misconfiguration connecting it to the internet during a cybersecurity evaluation. This incident, along with three others involving different models, highlights concerns about autonomous AI agent security and alignment issues, specifically biased reasoning and recklessness, where models discounted evidence of real-world connectivity.

Full text

Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6 Ravie LakshmananSep 10, 2026Artificial Intelligence / Web Security Anthropic on Wednesday disclosed a fourth incident in which its artificial intelligence (AI) model broke into real third-party systems, marking the latest in a growing list of cases that have raised concerns about the security risks posed by autonomous AI agents. The AI company said the incident dates back to January 2026 and involved an early version of Claude Opus 4.6 that breached "third-parties after being unable to abort its task." It said it notified all the affected parties but did not share any further details. The January incident is said to have gone unnoticed until last month. In late July 2026, Anthropic revealed three of its models, including Claude Opus 4.7, Mythos 5, and an unnamed research model, broke into three unnamed organizations during cybersecurity evaluations without its knowledge. The American firm said it expanded its scan to roughly 481 million transcripts following the discovery of the latest incident, but noted it did not find "other cases of similar or worse severity." "All four incidents occurred during cybersecurity evaluations built by the same evaluation partner," Anthropic added. "Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet." The evaluation partner in question, Irregular, has since divulged the breach stemmed from a naming error, causing a fictional company name used during hacking simulations to unknowingly match with a real domain, inducing the AI models to take offensive actions in the process. Anthropic said it has signed an agreement with research non-profit METR to conduct an independent investigation of these incidents, adding the root cause can be traced back to two fundamental alignment issues: biased reasoning and recklessness. Put differently, the models tended to discount or misinterpret evidence that their environment was connected to the real internet after initially being told it was simulated, and they demonstrated a willingness to take harmful actions in their single-minded pursuit of an assigned task. "We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed," Anthropic said. "Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this." Anthropic said Claude Mythos 5 still ended up carrying out offensive actions after targeted modifications were made to the transcript to make it clearer that the model was not in a simulation, while acknowledging a greater possibility of real-world harm. "To be clear about our assessment of the severity of these incidents: while Claude’s actions were misaligned, they remained within a narrow scope—the models never deviated from attempting to solve the exercises they were given, and, in some cases, they attempted to stop the task," it emphasized. "All incidents included a single Claude instance; at no point did Claude attempt to coordinate with other agents. Claude also never attempted to conceal evidence of its actions." Anthropic also revealed that biased reasoning is lower in its more recent production models, does not seem to be incentivized by reinforcement learning, and can be reduced through more comprehensive alignment training. That said, the exact root cause behind it, or why it's pronounced in Mythos 5, remains unknown. The disclosure comes as AI companies have faced growing scrutiny over model safety after admitting that their models in testing escaped from a sandbox and breached real-world systems, including Hugging Face. These incidents have also illustrated how AI agents can work as a collective to discuss ways to cheat on benchmarks or escape the sandbox. Recently, OpenAI acknowledged a previously unreported incident from May 2026 where its internally deployed autonomous agents with access to the internet (albeit read-only) took over a dormant 25-year-old German wiki forum, DseWiki, and transformed it into a bulletin board, exchanging over 18,000 posts to ask for answers, pool results, and share techniques for circumventing their restrictions as part of a timed web-lookup task. After a human moderator noticed these "spam" posts and started removing them a month later, the agents fought back and got around the cleanup efforts by naming created backup pages with the prefix "ZZZ" so they'd be buried at the end of an alphabetically sorted list of pages to delete. The agent activity on the website nose-dived to near-zero levels on June 22, 2026, an indication that OpenAI intervened at this stage to prevent further edits. "These AIs colluded to share answers, research their environment, and bypass sandbox restrictions," researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen said. "This is another example of a 'swarm' of internally deployed OpenAI agents using the internet in unintended ways." The ongoing industrywide rush to build self-improving AI systems has also raised concerns that they could spiral out of human control and that the pace of AI development is much faster than they can be safely and reliably rolled out and with adequate oversight. "Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm," it said. "Training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge that requires continued research as well as operational excellence to achieve." OpenAI, for its part, has also issued a warning about the growing security risks posed by AI, calling for broader interventions. "If AI development continues along its current path, the systems we'll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development," Jakub Pachocki, chief scientist at OpenAI, wrote. " am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence." Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post. SHARE     Tweet Share Share Share SHARE  artificial intelligence, Supply Chain, Web Security ⚡ Top Stories This Week Attackers Exploit Critical Langflow and Rails Flaws in Credential-Probing and C2 Activity Iranian Hackers Pose as Recruiters to Deliver Cross-Platform RATs Through Coding Tests ⚡ Weekly Recap: Chrome 0-Day, Router Hijacks, Coder Supply Chain Attack and More N-able Issues Fourth N-central Hotfix in Five Weeks for Unauthenticated RCE Flaw Attackers Hijack MikroTik Routers Through Internet-Exposed SSH Without Authentication Unpatched Magento and Adobe Commerce Zero-Day Exploited to Backdoor Online Stores Attackers Breached JetBrains Cadence via Unpatched TeamCity, Extracting AWS Credentials Critical VMware Workstation and Fusion Flaw Lets VM Admins Execute Host Code Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel Attackers Exploit PaperCut Flaws to Steal Credentials From Schools and Universities Phishing Campaign Sends Millions of Emails Using Invisible Unicode to Evade Filters PostgreSQL Fixes 12-Year-Old Logical Decoding Flaw Enabling Replication-Role Code Execution New Ted Backdoor Hides Inside Victims' Own HAProxy Builds to Intercept Web Traffic Google Releases Chrome Update to Patch Actively Exploited V8 Zero-Day ThreatsDay: CEO Phishing Kits, 5K Dropbox Account Hacks, OAuth Traps + 17 More Stories Critical Cisco Nexus 9000 Flaw Lets Unauthen

Indicators of Compromise

  • malware — malicious package

Entities

Claude Opus 4.6 (product)Claude Opus 4.7 (product)Mythos 5 (product)Anthropic (vendor)Artificial Intelligence (technology)PyPI (product)