[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fh3EkcS1AdWb1X5C5Al96jvut_Bs8UI3qaNaO3AfJ9_Y":3},{"article":4,"iocs":55},{"id":5,"title":6,"slug":7,"summary":8,"ai_summary":9,"brief":10,"full_text":11,"url":12,"image_url":13,"published_at":14,"ingested_at":15,"relevance_score":16,"entities":17,"category_id":32,"category":33,"article_tags":37},"4a9c072f-c74c-45d9-a5f5-8dfea93197b2","OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models' Reasoning","openai-anthropic-google-api-flaw-let-weaker-ai-models-decode-stronger-models-rea-685f5a","A newly disclosed flaw in the way OpenAI, Anthropic, and Google carried hidden AI reasoning between API calls let researchers recover internal reasoning and secrets from session logs, including API keys and passwords. The weakness affected encrypted reasoning objects used by the providers' reasoning APIs, where a block created in one session could be replayed into another and, during testing,","A vulnerability in the reasoning APIs of OpenAI, Anthropic, and Google allowed researchers to extract sensitive information, including API keys and passwords, from session logs. The flaw involved encrypted reasoning objects that could be replayed into different sessions or even handed to weaker models, effectively acting as a 'fuzzy' decoder. While the issue has been mitigated and is no longer reproducible, the researchers demonstrated potential for model distillation, data extraction, and hidden prompt injection.","Flaw in OpenAI, Anthropic, Google APIs allowed weaker models to decode stronger models' reasoning.","OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models' Reasoning Swati KhandelwalAug 12, 2026Vulnerability \u002F Artificial Intelligence A newly disclosed flaw in the way OpenAI, Anthropic, and Google carried hidden AI reasoning between API calls let researchers recover internal reasoning and secrets from session logs, including API keys and passwords. The weakness affected encrypted reasoning objects used by the providers' reasoning APIs, where a block created in one session could be replayed into another and, during testing, even handed to a weaker model in the same provider family to make it reveal the hidden content. The team behind the paper Stealing Reasoning Traces from Proprietary LLM APIs demonstrated four abuse paths: stealing proprietary reasoning for model distillation, extracting private data from other users' published traces, recovering harmful content concealed behind a safe visible answer, and hiding prompt injections inside opaque reasoning blocks. Across 6,708 public agent trajectories, the team decoded 315,320 thinking blocks. After excluding benchmark sources, it counted 704 distinct privacy artifacts from genuine user sessions, including 62 API keys, 33 passwords, 24 access tokens, and seven private keys. The cross-user attack did not provide arbitrary access to private chats. It required obtaining an encrypted reasoning block, such as one published in an agent log, and API access to a compatible model from the same provider. The researchers disclosed the findings to the affected model providers, Microsoft and Hugging Face, and say the demonstrated attacks stopped working after mitigations. Their reproducibility statement says the main extraction attack is no longer reproducible as of August 2026. The report does not document malicious exploitation in the wild. Developers are advised to strip reasoning blocks and opaque reasoning fields from shared traces and avoid committing raw API transcripts even when the visible text has been sanitized. The problem starts with a design meant to preserve reasoning across API calls when conversation state is managed manually or statelessly. OpenAI can return encrypted reasoning items that applications replay with manually managed history, Anthropic carries full reasoning in an encrypted signature, and Google uses encrypted thought signatures. These objects preserve reasoning state without exposing the underlying plaintext directly to the client. The encryption itself was not cracked, and the attack did not require obtaining an encryption key. It relied on intact opaque blocks being accepted and processed by the provider. During testing, the paper found those objects portable across sessions, users, and models, allowing a weaker compatible model to act as what the authors call a \"fuzzy\" decoder: Claude Haiku 4.5 for Claude traces, GPT-5.6 Luna for GPT traces, and Gemini Robotics ER-1.6 for Gemini traces. The decoder was prompted to transcribe reasoning produced by a stronger model. That cross-user behavior turns published agent logs into the sharper security problem. Of the 704 non-benchmark artifacts the team recovered, 64 appeared only in hidden reasoning and nowhere in the visible trace. Sanitizing the readable conversation could therefore leave secrets inside an opaque block that another account was able to replay. The exposure the study demonstrates is bounded: it lands on developers who published raw agent logs with the reasoning objects intact, one identifiable group rather than every API user, and not necessarily the only one at risk. The same portability also enabled an invisible prompt-injection proof of concept. The team crafted an opaque reasoning block that carried a malicious instruction and later replayed it into an unrelated task, causing the receiving model to add an attacker-directed upload action without putting the injected instruction in visible text. The authors caution that they do not have ground-truth plaintext for the proprietary reasoning, so they cannot guarantee every reconstructed trace is an exact copy. Their fidelity checks relied on reasoning-token counts and qualitative comparisons, with extracted lengths generally tracking the providers' reported thinking-token counts. Current vendor documentation shows that encrypted reasoning remains part of these APIs, but handling has changed. OpenAI still tells developers to replay encrypted reasoning items when manually managing stateless history, while Google says its backend manages thought compatibility when a session switches models. Anthropic now says thinking blocks are tied to the model that produced them and should be stripped when switching models because other models ignore them. Several questions the disclosure raises are left open by the public record. No public acknowledgment of the flaw from any of the three providers has surfaced so far, and none has tied its current documentation to this research, so the account that the demonstrated attacks no longer work rests on the researchers' own reproducibility statement rather than on vendor confirmation. The same record shows the team decoded hundreds of thousands of reasoning blocks already sitting in public repositories, yet it does not address whether those already-published blocks remain decodable, a separate question from whether fresh attacks still succeed. The work builds on May research by Johns Hopkins cryptographer Matthew Green, who showed that encrypted reasoning blocks could be replayed across sessions and accounts but stopped short of a reliable secret-extraction technique. Green says he reported the replay behavior to OpenAI and Anthropic through their bug-bounty programs; in his account, OpenAI called the report unreproducible and Anthropic said it did not see security implications in the replay or side-channel behavior. The new paper turns that replay behavior into a broader extraction method and documents the privacy consequences at scale. Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post. SHARE     Tweet Share Share Share SHARE  AI Security, API Security, Application Security, artificial intelligence, Data Exposure, Privacy, Prompt Injection, Vulnerability ⚡ Top Stories This Week Azure Cosmos DB Flaw Exposed Platform-Wide Key That Could Access Any Database Anthropic Says Claude Mistook the Open Internet for a CTF and Breached Three Organizations Researchers Report 84 Flaws in 4G and 5G Cores, Including a Session Hijacking Flaw Cheap Android TV Boxes Pose as Phones and Turn Owners’ Broadband Into Proxies N-able Says Attackers Take Over N-central Servers After Initial Fix Proves Incomplete Google Password Manager Attacks Could Let Malware Hijack Passkey-Protected Accounts New cPanel Critical Flaw Could Let Hosting Customers Run SQL as Database Root Keyv-Linked npm Worm Poisons Hundreds of Packages, Plants Claude Code and VS Code Hooks Claude Mythos 5 Tried to Backdoor a Real Open-Source Project in Testing, Then Vouched for Itself Critical Gitea Flaw Let Unauthenticated Attackers Read Server Files via Org-Mode Markup Poison Claude Sells Discounted Claude Access While Its Operator Sees Every Customer Prompt Over 250 ClickFix Domains Use Browser Fingerprinting to Hide macOS Malware Lures Chinese-Made Zbtlink Routers Ship With Backdoor That Opens Unauthenticated Root Shells Apple iCloud Private Relay Can Expose Real IPs Through WebKit Proxy Bypasses ThreatsDay: Odysseus RCE, Samsung One-Click Takeover, iCloud Backdoor Fight + 27 More Stories New Interrupt Injection Attack Can Bypass Spectre v2 Defenses on Intel and AMD CPUs New Zapscape KVM Flaw Could Let Privileged L1 Guest Code Escape to Linux Hosts New NatJack Attacks Hijack TCP Sessions and Spoof DNS by Manipulating NAT Tables 18-Year-Old Linux SCTP Flaw Could Let Local Users Gain Root and Escape Containers New WordPress Pre-Auth XSS Could Lead to PHP Code Execution - Patch ASAP Metabase Zero-Da","https:\u002F\u002Fthehackernews.com\u002F2026\u002F08\u002Fopenai-anthropic-google-api-flaw-let.html","https:\u002F\u002Fblogger.googleusercontent.com\u002Fimg\u002Fb\u002FR29vZ2xl\u002FAVvXsEiZP21lJn1EY_0JXWBul8gpBgRD_ryI4ACYqFiu6icbKCgIMyta9UqjQrtnzjTkO1Yi9tdzTaQw6X949lzMwtQPKUPPIxA5_EwFjJqjMDwBtFSc56vaCyybwJNcsjVZMfMW4KHQ_VVCjz3AVeovFEtsI8ENZTQmQuc9rDIF5yd0n9uciJmwqlWS_EKf5mo\u002Fs1600\u002Fai-models.jpg","2026-08-12T11:47:38+00:00","2026-08-12T14:00:16.090881+00:00",8,[18,21,23,25,27,29],{"name":19,"type":20},"OpenAI","vendor",{"name":22,"type":20},"Anthropic",{"name":24,"type":20},"Google",{"name":26,"type":20},"Microsoft",{"name":28,"type":20},"Hugging Face",{"name":30,"type":31},"Claude Haiku 4.5","product","80544778-fabb-4dcd-aa35-17492e5dcf4f",{"id":32,"icon":34,"name":35,"slug":36},null,"Vulnerabilities","vulnerabilities",[38,43,45,50],{"category":39},{"id":40,"icon":34,"name":41,"slug":42},"2e06f76c-d5b9-4f54-9eef-4d3447b10730","Breaches","breaches",{"category":44},{"id":32,"icon":34,"name":35,"slug":36},{"category":46},{"id":47,"icon":34,"name":48,"slug":49},"839da5c1-3c34-47e2-9499-f7201640e3ac","AI Security","ai-security",{"category":51},{"id":52,"icon":34,"name":53,"slug":54},"e7b231c8-5f79-4465-8d38-1ef13aea5a14","Threat Intelligence","threat-intelligence",[56,59,62],{"type":57,"value":30,"context":58},"malware","Used as a decoder model for Claude traces.",{"type":57,"value":60,"context":61},"GPT-5.6 Luna","Used as a decoder model for GPT traces.",{"type":57,"value":63,"context":64},"Gemini Robotics ER-1.6","Used as a decoder model for Gemini traces."]