China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies
China-based AI firms conduct industrial-scale knowledge distillation campaigns against US AI models.
Summary
US agencies NSA, CISA, and FBI have issued a joint advisory detailing how China-based AI companies, including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, are systematically extracting proprietary functionalities from US AI models through large-scale knowledge distillation. These campaigns, likely with Chinese government awareness, aim to accelerate their own AI development and bridge technological gaps by violating terms of service and using proxy networks to bypass restrictions. The advisory recommends enhanced detection, response alterations, and cross-organization intelligence sharing.
Full text
Cybersecurity Advisory China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies Release DateSeptember 08, 2026 Alert CodeAA26-251A China-Based Artificial Intelligence Companies Conducting Industrial-Scale Disti… Related topics: Cybersecurity Best Practices , Nation-State Threats , Cyber Threats and Response Executive summary China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy. While “distillation” is recognized as a legitimate and useful technique in AI research, China-based AI companies are engaging in aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of U.S. frontier AI models. The National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) (hereafter referred to as the authoring agencies) are releasing this joint Cybersecurity Advisory to alert organizations about these malicious activities and techniques and recommend mitigations to reduce their potential impact. Likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024. DeepSeek has conducted organized campaigns since at least 2024 targeting reasoning capabilities, specialized optimizations, and domain-specific functions to train its R1 and V3 models. Alibaba leveraged industrial-scale distillation to improve the company’s Qwen family of AI models. Moonshot AI, MiniMax, Stepfun, and Z.AI also engaged in malicious knowledge distillation of U.S. AI companies’ models. China-based AI companies route distillation requests through multiple pathways to gain unauthorized access, consequently violating U.S. AI companies’ terms of use. These pathways include native application programming interfaces (APIs), remote cloud providers, and third-party aggregators that automatically obfuscate user metadata to avoid detection. Further, China-based AI companies use a gray market of proxies known as “transfer stations” to bypass U.S. AI companies’ geographic restrictions, breach terms of use, evade safeguards, and undermine traceability. China-based AI companies achieve cost savings for their industrial-scale distillation campaigns through bulk procurement of the U.S. AI companies’ premium subscriptions shared across teams of developers. Advanced industrial-scale distillation tactics include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures. China-based AI companies that conduct industrial-scale distillation against U.S. AI models see significantly shorter AI development timelines and reduced financial expenditures in training a frontier model. China-based AI companies deliberately distribute operations across multiple providers, platforms, and pathways to avoid single-point detection. They also attempt to distill the best capabilities and proprietary features of each U.S. frontier model to train their China-based AI models. This represents systematic extraction of proprietary functionalities and capabilities threatening U.S. technological leadership. Addressing industrial-scale distillation merits a coordinated response across the AI ecosystem, including effective information-sharing, spanning the U.S. Government, private industry, and allied nations. The authoring agencies recommend U.S. AI companies take three immediate actions: Implement comprehensive detection and mitigation: Detect anomalous and malicious prompts, accounts, networks, and behaviors. Additionally, monitor subscription-to-usage ratios, immediate maximum usage from new accounts, and enterprise-scale throughput patterns. Deploy targeted response changes: Subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs to companies conducting industrial-scale distillation campaigns. Establish cross-organization intelligence sharing: Correlate activity across model providers, cloud platforms, and API aggregators to reveal distributed campaigns. Attribution Since at least late 2024, China-based AI companies, including DeepSeek (DeepSeek Artificial Intelligence Technology Research Co., Ltd.), Moonshot AI (Beijing Moonshot Technology Co., Ltd.), Alibaba Group, MiniMax (Shanghai MiniMax Co., Ltd.), StepFun (Shanghai Jieyue Xingchen Intelligence Technology Co., Ltd.), and Z.AI, have conducted high-volume knowledge distillation campaigns against several U.S. AI companies. The sheer scale of these campaigns and their sophistication indicate that distillation is not a supplement to these companies’ AI model development, but the critical core of it. Likely with the knowledge of the Chinese government, the China-based AI sector has turned to a comprehensive distillation strategy in an attempt to bridge the technological and performance gaps between their AI models and U.S. frontier AI models. To access U.S. AI companies’ application programming interfaces (APIs), China-based AI companies use a gray market of API proxies known as “transfer stations” to bypass U.S. AI companies’ regional restrictions, breach terms of use, evade safeguards, and undermine traceability. DeepSeek DeepSeek has been conducting an organized distillation campaign against U.S. AI companies’ frontier AI models since at least late 2024 to generate synthetic training data for its models, including R1, released in early 2025. The company targeted specific knowledge domains to extract proprietary functionality and reasoning capabilities to reduce their compute and research costs. DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.1 Between late 2024 and mid-2025, DeepSeek distilled specialized training data and capabilities from the following U.S. frontier AI company models to train their R1 and V3 models: Claude 3.7 Claude Sonnet 4 Claude Sonnet 4.5 Claude Opus 4.1 Gemini 2.5 Pro Preview Gemini 2.5 Flash Preview GPT-4 GPT-4o GPT-4 Mini GPT-4 Nano GPT-5 Grok 4 The specific knowledge and capabilities distilled included: Legal specialization optimization API rule-driven tasks Writing using CoT drafts Agentic functions Question and answer optimization Coach/assistant capabilities Functional creation optimization Supervised fine-tuning (SFT) optimization Creative and occupational writing optimization Moonshot AI Moonshot AI has conducted a widespread distillation campaign against U.S. frontier AI companies since at least mid-2025. Notably, Moonshot AI extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model. The company has used the following models to distill SFT optimization, reinforcement learning (RL), software engineering, and math capabilities: Claude Opus 4.1 Claude Sonnet 3.7 Claude Sonnet 4 Claude Sonnet 4.5 Claude Sonnet 4.5 Thinking Claude Fable 5 GPT-oss-20b GPT-3 GPT-4o GPT-4o mini GPT-5 GPT-5 Codex GPT-5 Pro Gemini 2.5 Flash Gemini 2.5 Flash-Image Gemini 2.5 Pro Nano Banana Grok Code Fast-1 Other companies Several other China-based AI companies, including Alibaba, MiniMax, StepFun, and Z.AI have also leveraged distillation techniques to build their AI models. In late 2025, Alibaba distilled Claude-4, Claude Opus, Claude Sonnet, and GPT-5 to improve their AI models’ software engineering skills, customer service dialogue functionality, image/character creati