Back to Feed
AI SecuritySep 3, 2026

Capsule Security Launches ‘AI Circuit Breaker’ to Stop Rogue Agents

Capsule Security launches AI Circuit Breaker to stop rogue AI agents in real-time.

Summary

Capsule Security has released an 'AI Circuit Breaker' designed to prevent autonomous AI agents from executing actions outside their intended scope. The solution uses specialized AI models, trained with NVIDIA Nemotron 3 Ultra, to detect rogue behavior in milliseconds, offering a faster and more efficient alternative to reviewing actions with large general-purpose models. This aims to provide a crucial runtime security layer for increasingly capable AI agents.

Full text

Providing security against anomalous autonomous agents requires instantaneous detection, evaluation and action – a circuit breaker for AI rather than a circuit breaker for electricity. Such a ‘device’ has now been developed and released by Capsule Security. The firm was founded in 2025 by Naor Paz (CEO) and Lidan Hazout (CTO). They had seen how autonomous AI agents create a dangerous security gap and decided to build a missing runtime security layer. Their latest solution was announced on September 2, 2026, and is described as an ‘AI circuit breaker’. It is designed to prevent damage from agents operating outside their intended scope, in real-time. “The defining AI security risk is no longer only what people can do with agents. It is what autonomous agents can decide to do by themselves,” explains Paz in announcing the product. “When software can reason, use tools and take action, a wrong decision can become a real-world incident in seconds. Human trust in AI depends on our ability to stop that action before it happens.” The problem is not simply the reach of autonomous agents; it is also the speed at which they operate. While it may be possible to review an agent’s behavior before it happens, doing so commonly introduces latency that may be too slow and too expensive for the agent’s intended action. Capsule’s solution uses its own specialized AI for real-time intervention. The firm used NVIDIA Nemotron 3 Ultra to support the training process, combining real agent traces, human review and adversarial examples designed to teach its AI models the boundary between authorized and rogue behavior.Advertisement. Scroll to continue reading. They developed two models capable of providing strong detection without the cost and latency of sending every agent action to a large general-purpose model for review. The more accurate model attained 96.9% detection accuracy, compared with 86% for the strongest third-party model evaluated. The models could also make a decision in as little as 71 milliseconds, meaning they could operate within the agent’s workflow without creating any meaningful delay. Capsule then reduced the infrastructure for its larger model, reducing its memory requirement by almost 50%. The result is Capsule’s own evaluator running within the agent’s execution path able to evaluate the agent’s intention before the action takes place and stop it before execution if necessary. This is the AI circuit breaker. “The model evaluates an agent’s intended action immediately prior to execution, giving organizations the ability to allow, flag or block it in real time. This creates an independent control layer for agents that can access sensitive data, write code, operate infrastructure and interact with other systems,” says Capsule. “Post-incident monitoring only identifies the problem after the damage has occurred.” Capsule claims 98% efficiency for the circuit breaker’s decision maker when tested against StepShield – an independent academic benchmark that can be used to measure whether security systems can identify and stop rogue agent behavior before damage occurs. The key lesson from Capsule’s work is that specialized Small Language Models (SLMs) are the key to safely scaling trusted agentic workflows across the enterprise. “Moving beyond general-purpose models to specialized, efficient detectors allows organizations to secure their agentic workflows without sacrificing speed, cost, or performance.” It claims. Related: AI Agent Firewall Startup AIR Security Emerges From Stealth With $50 Million Related: OpenLeash Adds a Human Check to Risky AI Agent Actions Related: UK Government Rolls Out Agentic AI Defense Plan Alongside Industry Pledge Related: Critical Vulnerability Exposes GitHub Agentic Workflows to Prompt Injection Written By Kevin Townsend Kevin Townsend is a Senior Contributor at SecurityWeek. He has been writing about high tech issues since before the birth of Microsoft. For the last 15 years he has specialized in information security; and has had many thousands of articles published in dozens of different magazines – from The Times and the Financial Times to current and long-gone computer magazines. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Kevin Townsend Sevii Targets AI-Speed Attacks With Preemptive Autonomous DefenseThink You’ve Eliminated Chinese AI? Check the Model’s Lineage, Cisco SaysCISO Conversations: Chris Wheeler – Trust Is the Job, From the Navy to the C-SuiteIran-Linked Hackers Shut Down UK Power Plant for Four DaysEncrypted Prompts Bypass AI Safety Guardrails in Grok and GeminiNew Phishing Toolkit Uses Passkeys to Maintain Access After Password ResetsSurveillance – Everything You Wanted to Know, But Were Afraid to AskCISO Conversations: Nico Waisman – From Self-Taught Hacker to AI-Driven Offensive Security at XBOW Latest News Manchester Airports Group Data on 8.8 Million People Leaked After Ransom RefusalHiddenLayer Raises $100 Million for AI Runtime SecurityAI Agent Firewall Startup AIR Security Emerges From Stealth With $50 Million153 Million Driver License Images Offered on Dark WebOver 3 Million WordPress Sites Affected by Migration Plugin VulnerabilityCisco Warns of Unpatched Secure Email Flaws, Patches Critical Switch VulnerabilitiesOpenLeash Adds a Human Check to Risky AI Agent ActionsUK Moves to Block High-Risk Tech Suppliers From Critical Infrastructure Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Virtual Event: Attack Surface Management Summit 2026 September 16, 2026 Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs. Register Webinar: Minimum Viable Business: Can You Prove Your Organization Would Recover? September 2, 2026 In this live webinar, learn how to define your minimum viable business, identify the systems it depends on, measure actual recovery time against business requirements, and present the gaps to the board as measurable risk. Register People on the MoveTom Bonos has been named Chief Revenue Officer at Sumo Logic.Axonius has appointed Chris Jones as CTSO and Dan Schoenbaum as SVP of Business Development.Optiv has appointed Sean Forkan as Chief Revenue Officer (CRO).More People On The MoveExpert Insights What the Hugging Face Incident Teaches Security Leaders About AI Agent Access Security teams must treat autonomous agents as highly privileged identities. (Etay Maor) The Future of AI-Driven Security Depends on Complete Data For twenty-five years, "data" in security meant logs and events. But logs are a lossy representation of reality. (Danelle Au) The MFA Identity Trap: When Authentication Creates a False Sense of Security Organizations must distinguish identity verification, authentication and threat detection, or risk successfully authenticating the attackers they are trying to stop. (Torsten George) Silent Patches Don’t Stop Attackers – They Blind Defenders Silent patches can become exploit intelligence for attackers while leaving defenders without the context needed to prioritize risk. (Tod Beardsley) Hired for One Job, Judged on Another: The CISO’s Real Problem The skills that get a CISO hired are rarely the skills they are judged on later. Most security leaders are stuck in that gap. Closing it is the real job. (Sravish Sridhar) Flipboard Reddit Whatsapp Whatsapp Email

Entities

AI Circuit Breaker (product)Capsule Security (vendor)Nemotron 3 Ultra (product)NVIDIA (vendor)Small Language Models (technology)