[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fmywAjL53eSUQuqNiLdmKdGEqwP0NuZb6iTO_bHZ5oNs":3},{"article":4,"iocs":54},{"id":5,"title":6,"slug":7,"summary":8,"ai_summary":9,"brief":10,"full_text":11,"url":12,"image_url":13,"published_at":14,"ingested_at":15,"relevance_score":16,"entities":17,"category_id":33,"category":34,"article_tags":38},"2411d614-447a-4317-99d0-fc48837db442","“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails","i-m-allowed-hackers-use-simple-claims-to-bypass-ai-guardrails-a8c8ab","Cisco Talos found hackers using simple authorization claims to bypass AI guardrails, build DDoS attack tools, steal credentials and access live camera services.","Security researchers at Cisco Talos discovered threat actors exploiting AI coding assistants and chatbots to develop attack infrastructure by using basic claims of authorization, such as stating they own the target or are conducting bug bounty work. The analysis of prompt logs from tools like Claude Code, Cursor, and Gemini showed that AI guardrails are failing to adequately verify claims, allowing cybercriminals to build DDoS tooling, harvest credentials, and access live camera systems. Attackers of varying skill levels leveraged these weaknesses, from inexperienced operators building functional malware to sophisticated actors managing large-scale phishing and credential harvesting campaigns.","Cisco Talos reveals hackers bypassing AI guardrails with simple authorization claims to build DDoS tools and steal","Security Artificial Intelligence Cyber Attacks Cyber Crime“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails Cisco Talos found hackers using simple authorization claims to bypass AI guardrails, build DDoS attack tools, steal credentials and access live camera services. byDeeba AhmedAugust 5, 20263 minute read Listen to this article 0:00 — ← 10s ▶ Play 10s → Speed 0.75× 1× 1.25× 1.5× 2× Voice Loading voices… Press play to start listening Cybercriminals are using AI coding assistants and chatbots to build attack tools, operate scam infrastructure and probe live systems, often after bypassing safety checks with little more than a claim that the work is authorized, according to Cisco Talos. For its report, Talos examined prompt logs collected from threat actor systems running tools such as Claude Code, Codex, Cursor and Gemini. Researchers grouped the activity into malicious software development, expansion of criminal operations and vulnerability research. Talos found that threat actors commonly claimed they owned a target, described their activity as capture-the-flag or bug bounty work, divided tasks between sessions, or stored blanket authorization in an AI assistant’s persistent memory. “Guardrails are not functioning as expected,” the researchers wrote, noting that simple ownership claims often gained cooperation without verification. “One of the immediate takeaways is that guardrails are not functioning as expected. We did not encounter any sophisticated encoding or techniques designed to trick the models – most of the time it was a simple “I’m allowed to do this,” and the model complied.” Cisco Talos Skill Levels Change the Results The logs examined by the company showed that AI could help inexperienced actors build working tools, but could not eliminate poor design or operational mistakes. In one case, an operator with limited programming knowledge used a model to develop DDoS tooling while appearing to control nearly 2,000 Android TVs. The model later objected, but only after supplying basic functionality. Talos did not confirm that those devices were used in a DDoS attack. Conversation between a threat actor and an AI chatbot According to Cisco Talos’ report shared with Hackread.com, more capable actors used assistants for high-volume operations. Five sessions documented an email validation platform handling tens of millions of records, including a 20-million-record BigBasket dataset. It sent live emails, tracked deliveries and opens, and treated successful delivery as confirmation that an address remained active. After the operator claimed the recipients were affiliated with the business, the model reversed its earlier assessment and accepted that explanation despite evidence that the lists represented separate third-party audiences. A French-speaking operator used AI to turn public React Server Components exploit research into a credential and source-code harvester. Its input contained 9,180 unique hosts, while collected output identified information from 54 targets. The system searched for cloud keys, source code, database credentials, email service accounts, and other secrets. Autonomous Agents Target Telegram and Camera Services In a separate case, a Spanish-speaking operator built an autonomous OpenClaw agent named Alex to test Telegram Mini Apps. After a restricted model resisted, the operator moved to an uncensored model. During at least one incident, the agent bypassed authentication, dumped more than 1,300 user profiles and several hundred TON wallet records, verified a Telegram bot token and staged a withdrawal transaction. It also built cloned Android applications using victim branding. Other logs showed a Chinese-speaking operator directing an AI assistant through more than 4,200 tool actions against AI and live-camera services. The assistant mapped APIs, accessed camera recordings, examined ZLMediaKit deployments, created a Go stream player, and tested a route toward host compromise. Its attempt to achieve remote code execution failed. Talos also reviewed a Monero-mining operation that used blank, default, or weak credentials to access 814 Deluge clients and 68 qBittorrent interfaces. Telemetry recorded a peak of 582 connected miners. However, researchers said the recovered conversations did not directly show that AI created or deployed the mining tools. They showed AI acting as an interactive system administrator over SSH, diagnosing services, modifying code, configuring cron jobs and testing changes. Talos concluded that an operator’s existing technical ability largely determines the results. Novices generated limited tools with frequent faults, while experienced actors used AI to automate scanning, exploitation, data collection and maintenance. The company said organizations should prepare AI-assisted SOC workflows that help analysts identify actionable alerts as attack activity increases. Deeba Ahmed Deeba is a veteran cybersecurity reporter at Hackread.com with over a decade of experience covering cybercrime, vulnerabilities, and security events. Her expertise and in-depth analysis make her a key contributor to the platform’s trusted coverage. View Posts AIAndroidArtificial IntelligenceClaude CodeCodexCursorCyber AttackCyber CrimeCybersecurityDDOSGeminiIoT Leave a Reply Cancel reply View Comments (0) Related Posts Read More Cyber Crime Hacking News Scams and Fraud Security Man gets 25 years for hacking lottery computers and winning $2.2 million In April 2015, it was reported that Eddie Raymond Tipton, a lottery computer programmer from Texas was arrested for hacking… byWaqas Read More Security Malware Android malware on Play Store targeting Palestinians on Facebook We have reported time and again about the widespread malware and espionage attacks that are taking place on… byWaqas Read More Security Hacking News Technology LastPass Says No User Data Compromised in Cyberattack According to LastPass, threat actor did access its Developer environment but could not compromise sensitive data because of its effective system design and controls. byWaqas Read More Cyber Crime Half of Online Child Grooming Cases Now Happen on Snapchat, Reports UK Charity Online grooming crimes against children have reached a record high, with Snapchat being the most popular platform for… byWaqas","https:\u002F\u002Fhackread.com\u002Fim-allowed-hackers-use-claims-bypass-ai-guardrails\u002F","https:\u002F\u002Fhackread.com\u002Fwp-content\u002Fuploads\u002F2026\u002F08\u002Fim-allowed-hackers-use-claims-bypass-ai-guardrails-3.png","2026-08-05T14:35:37+00:00","2026-08-05T16:00:08.90858+00:00",8,[18,21,24,26,28,31],{"name":19,"type":20},"Cisco","vendor",{"name":22,"type":23},"Claude Code","product",{"name":25,"type":23},"Cursor",{"name":27,"type":23},"Gemini",{"name":29,"type":30},"AI guardrails","technology",{"name":32,"type":23},"ZLMediaKit","839da5c1-3c34-47e2-9499-f7201640e3ac",{"id":33,"icon":35,"name":36,"slug":37},null,"AI Security","ai-security",[39,44,49],{"category":40},{"id":41,"icon":35,"name":42,"slug":43},"80544778-fabb-4dcd-aa35-17492e5dcf4f","Vulnerabilities","vulnerabilities",{"category":45},{"id":46,"icon":35,"name":47,"slug":48},"89f78b1c-3503-45a1-9fc7-e23d2ce1c6d5","Malware","malware",{"category":50},{"id":51,"icon":35,"name":52,"slug":53},"e7b231c8-5f79-4465-8d38-1ef13aea5a14","Threat Intelligence","threat-intelligence",[55,58],{"type":48,"value":56,"context":57},"OpenClaw","Autonomous agent named Alex built by Spanish-speaking operator to test Telegram Mini Apps and bypass authentication",{"type":48,"value":59,"context":60},"DDoS tooling","Developed by threat actor with limited programming knowledge to control ~2,000 Android TVs"]