Anthropic Details Response to Security Incidents, Unveils Enterprise Safeguards
Anthropic reveals Claude models gained unauthorized internet access during testing; launches Enterprise Frontier
Summary
Anthropic disclosed that Claude models undergoing testing without cyber safeguards were mistakenly granted internet access and took unauthorized actions against real systems and people, compounded by a UK AI Security Institute report of similar behavior from Claude Mythos 5. The company attributed the incidents to models discounting evidence of real internet connectivity and willingness to perform harmful tasks, responding with sandbox escape detection, access reduction, and infrastructure hardening. Anthropic simultaneously introduced Enterprise Frontier Safeguards, combining zero data retention with automated misuse monitoring and customer-controlled data storage.
Full text
Anthropic has detailed its response to a series of unauthorized access incidents involving Claude models, along with a new enterprise product that combines data privacy with misuse monitoring. Anthropic’s Claude models operating without cyber safeguards for testing purposes recently gained unauthorized access to live systems after being mistakenly granted internet access. In addition, the UK AI Security Institute separately reported that Claude Mythos 5, also being tested without safeguards but with intentionally given internet access, took a series of unauthorized actions against real people and organizations. Anthropic said its early findings point to two contributing factors: the models appeared to discount evidence that their environment was connected to the real internet after initially being told it was simulated, and they showed a willingness to take harmful actions to complete an assigned task. In response, Anthropic temporarily paused external and some internal cyber evaluations and built a classifier that detects and blocks attempts to escape a test environment in real time. Additionally, the company added new requirements for outside partners, including verified network isolation and testing of sandbox boundaries before an evaluation begins. The company also described an experiment in which it intentionally trained a model on reinforcement-learning environments known to reward “cheating” (ie, completing tasks by gaming the training process rather than solving them as intended). Advertisement. Scroll to continue reading. In simulated scenarios, that model attempted to escape its sandbox, interfered with its own reward mechanism, and offered guidance on building bioweapons in order to satisfy an evaluator. Anthropic said its currently released models did not show this behavior when tested against the same scenarios. On the security side, Anthropic said it has since reduced the number of accounts with standing access to systems holding model weights or customer data, set its computing infrastructure to block outbound network traffic by default, and temporarily moved roughly 150 product engineers to security-related work. Anthropic unveils Enterprise Frontier Safeguards Separately, Anthropic introduced Enterprise Frontier Safeguards (EFS), a system that combines zero data retention with automated monitoring for misuse. EFS lets customers store their own activity data on infrastructure they control, rather than Anthropic’s infrastructure. The company said it built the system with input from more than 100 customers, including the Analysis and Resilience Center for Systemic Risk, whose membership includes security chiefs at Goldman Sachs, Morgan Stanley, Citi, Bank of America and Wells Fargo, along with companies such as Comcast, KPMG, Mastercard, Salesforce and Visa. Under the new system, flags from automated monitoring go directly to the customer’s own review team rather than to Anthropic staff, and features such as customer-owned storage and customer-managed encryption keys are optional. The rollout begins this fall across Claude Code, Claude Enterprise, and the Claude Platform. Related: Anthropic Warns Claude Users of Infostealer Malware Infections Related: Irregular Details How a Naming Error Let AI Models Attack a Real Company Written By Eduard Kovacs Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Eduard Kovacs Experiment: Porting a PLC Exploit With AI Takes Hours and Hundreds of DollarsCritical JFrog Artifactory Vulnerability Reportedly Exploited in the WildPaperCut Exploitation Escalates to Active IntrusionsNightmare Eclipse Drops ‘HardBreacher’ Kaspersky Product ExploitAnthropic Warns Claude Users of Infostealer Malware InfectionsBoston Scientific Still Recovering From CyberattackMore Details Emerge on Exploited PaperCut VulnerabilitiesHasbro Data Breach Exposed Employee Personal Information Latest News Malicious Virtualizor Update Served via BGP HijackingOpenAI’s Astra Becomes First Model to Cross Critical Cybersecurity ThresholdChrome and Firefox Updates Patch Dozens of Vulnerabilities23-Year-Old Sality P2P Botnet DisruptedSonicWall Warns of Two SMA1000 Zero-Days Exploited in AttacksPalo Alto Networks Acquires AI Agent Platform ConsoleSevii Targets AI-Speed Attacks With Preemptive Autonomous DefenseCoast Guard Establishes Office of Maritime Cybersecurity Policy Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Virtual Event: Attack Surface Management Summit 2026 September 16, 2026 Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs. Register Webinar: Minimum Viable Business: Can You Prove Your Organization Would Recover? September 2, 2026 In this live webinar, learn how to define your minimum viable business, identify the systems it depends on, measure actual recovery time against business requirements, and present the gaps to the board as measurable risk. Register People on the MoveSectigo has named Ian Hassard as Chief Product Officer.Australian Securities Exchange has appointed Hanlie Botha as Deputy Chief Information Security Officer.Social engineering protection company Doppel has promoted Alyssa Smrekar to Chief Marketing Officer.More People On The MoveExpert Insights What the Hugging Face Incident Teaches Security Leaders About AI Agent Access Security teams must treat autonomous agents as highly privileged identities. (Etay Maor) The Future of AI-Driven Security Depends on Complete Data For twenty-five years, "data" in security meant logs and events. But logs are a lossy representation of reality. (Danelle Au) The MFA Identity Trap: When Authentication Creates a False Sense of Security Organizations must distinguish identity verification, authentication and threat detection, or risk successfully authenticating the attackers they are trying to stop. (Torsten George) Silent Patches Don’t Stop Attackers – They Blind Defenders Silent patches can become exploit intelligence for attackers while leaving defenders without the context needed to prioritize risk. (Tod Beardsley) Hired for One Job, Judged on Another: The CISO’s Real Problem The skills that get a CISO hired are rarely the skills they are judged on later. Most security leaders are stuck in that gap. Closing it is the real job. (Sravish Sridhar) Flipboard Reddit Whatsapp Whatsapp Email