Back to Feed
AI SecurityOct 1, 2026

OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI Associates

OpenAI disrupted a campaign to extract AI model reasoning, linked to China-based Moonshot AI.

Summary

OpenAI has disrupted a coordinated campaign aimed at illicitly extracting protected reasoning from its AI models. The activity, attributed to individuals associated with Chinese AI company Moonshot AI, involved manipulating model interactions to reproduce protected reasoning. OpenAI has since implemented additional mitigations and banned fraudulent accounts.

Full text

OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI Associates Ravie LakshmananOct 01, 2026Artificial Intelligence / Vulnerability OpenAI on Wednesday said it identified and disrupted a coordinated distillation campaign that was designed to illicitly extract protected reasoning from its artificial intelligence (AI) models. A "core cluster of the activity," going back to the first week of July, has been attributed to individuals associated with Moonshot AI, a Chinese AI company based in Beijing. It did not cite any technical evidence to back this assessment, likely owing to security reasons. "The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations," OpenAI said. "Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service." The activity is said to have begun on July 1, 2026, initially at a low volume before it spiked on July 24 and 25, 2026, to 16,000 attempted requests using a relevant extraction pattern from over 4,000 users. Upon further investigation, the company said it identified related "prompt-pattern activity" across more than 15,000 users. The campaign was fully disrupted on July 28, 2026. The AI upstart characterized the activity as adversarial distillation, one that involves the systematic and unauthorized use of one model's outputs to help train, reproduce, or improve another model. OpenAI said it has since deployed additional mitigations to combat this attack and banned the fraudulent accounts engaged in the activity. In addition, OpenAI said it closed a "pathway" that made it possible for some who already possessed another user's encrypted reasoning to replay it and recover its contents, alongside adding checks to detect and hold streamed output that might expose reasoning. In a study published in August 2026, a group of researchers found an architectural vulnerability that made the encrypted reasoning traces "fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem," which an attacker could exploit to develop a scalable decryption jailbreak and circumvent anti-distillation mechanisms. "By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly," researchers from MATS Research, ELLIS Institute Tübingen, and Synk said. Furthermore, it allows for large-scale private data extraction, opens the door for invisible prompt injections by embedding malicious payloads entirely within encrypted blocks, and inadvertently reveals hazardous information hidden within the reasoning process, even if the model's final, visible output rejects a harmful request. Given that protected reasoning offers insights into how a model works its way through a task, extracting this information can reveal sensitive data and help others reproduce the model's capabilities, OpenAI added. "Adversarial distillation poses safety and national security risks," the company said. "Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model's user-facing outputs." "At scale, distillation can also accelerate the transfer of advanced capabilities without requiring the same investment in safety. These concerns become heightened as models gain capabilities in dual-use domains." This is not the first time Moonshot AI has faced distillation accusations. Last month, rival Anthropic accused Moonshot AI of stealthily relaying customer requests to Claude as opposed to processing them using Kimi, and then displaying responses from Claude back to users. The company is also alleged to have retained a subset of these exchanges to train its chain-of-thought (CoT) model. The activity has been tracked under the moniker GTG-16002. Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post. SHARE     Tweet Share Share Share SHARE  artificial intelligence, data security, Vulnerability ⚡ Top Stories This Week Roundcube Pre-Auth SQL Injection Flaw Actively Exploited in the Wild Cloudflare Fixes Flaw That Let One Container Read Another Customer's Leftover Disk Data Unpatched OnePlus Flaws Let Installed Android Apps Gain Root Without Permissions ThreatsDay: AI Search Poisoning, AI Coding Tool Leaking Repos, One-Click Code Execution and 13 More Stories Placeholder third-party[.]com Referenced Across 1,700+ Repositories Now Serves Malicious Content OpenAI Agent Bypassed Australian Medicare Portal Controls to Access Non-Public Files A Leaked GitLab Issue Email Address Lets Anyone Push Code and Run CI Jobs as You MikroTrick Chain Let Attackers Take Over MikroTik Routers Without a Password or SSH Key New cPanel Flaw Lets a Hosting Account Run Code as Root, Take Full Server Control Exploit Released for Unpatched Ubuntu Linux Flaw Enabling Host-Root Container Escape F5 Patches Critical BIG-IP APM Zero-Day Exploited for Unauthenticated RCE on OAuth Servers Critical Next.js ImageResponse Flaw Can Lead to Server Code Execution via Crafted SVG Input ShinyHunters Claims FBI Breach, Says It Stole Data on Agents and Job Applicants Check Point Warns of Management Server Zero-Day Exploited in Targeted Attacks WordPress Issues Patch for Critical Flaw That Can Enable Code Execution on Some Servers Researcher Drops BigDiskBuster Zero-Day PoC That Blocks Microsoft Defender Updates New CVSS 10.0 VeloCloud Orchestrator Flaw Actively Exploited in Certificate-Based Setups New Linux Kernel Flaw Gives ARM64 KVM Guests Read-Write Access to Host Memory SharePoint Flaw Initially Listed as Spoofing by Microsoft Enables Authenticated RCE One Hidden Meta Muse Setting Could Let Attackers Turn the AI Assistant Into a Backdoor WordPress Comment2Shell Flaw Can Turn Anonymous Comment XSS Into RCE via Admin Session Zyxel and Veeam Flaws Under Active Exploitation With Command and SYSTEM Access Beyond ISO 27001: Building a Risk Program That Can Keep Up With AI Secrets Sprawl Is an Identity Problem That AI Just Made Impossible to Ignore ⭐ Featured Resources Validation Summit ’26: See How Pen Testing, Exposure Validation and BAS Work Together Red Teams: Learn How Attack Path Chaining Changes Automated Security Testing Turn Threat Intelligence Into Verified Risk With Threat-Led Penetration Testing Deploy Browser Security Monitoring in Minutes With a Single Header

Indicators of Compromise

  • malware — adversarial distillation

Entities

OpenAI (vendor)Moonshot AI (vendor)Anthropic (vendor)Claude (product)Kimi (product)