Skip to content
Tech News
← Back to articles

US Government Accuses Chinese AI Firms of Distilling Frontier Models

read original get Hands-On Large Language Models (O'Reilly book) → more articles
Why This Matters

A joint FBI/NSA/CISA advisory accuses major Chinese AI firms of industrial-scale distillation of US frontier models, elevating a routine training technique into a national security and IP theft issue. If accurate, it suggests America's lead in AI can be siphoned through ordinary API access, which could push vendors toward tighter access controls and governments toward new export-style restrictions.

Key Takeaways
Worth a Look

Hands-On Large Language Models (O'Reilly book) — If terms like distillation, teacher and student models, and token extraction in this advisory left you wanting more, this O'Reilly guide walks through how LLMs are actually built, trained, and adapted. It's a practical way to understand the techniques at the center of the US-China AI dispute instead of just reading headlines about them.

See Hands-On Large Language Models (O'Reilly book) on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

Chinese AI firms are stealing proprietary capabilities belonging to US-based AI models via massive distillation campaigns, acccording to US government agencies.

The FBI, the National Security Agency (NSA), and the Cybersecurity and Infrastructure Security Agency (CISA) published a joint advisory on Sept. 8, warning that Chinese AI firms were conducting "industrial-scale" efforts to extract capabilities from leading US AI models. The advisory accuses a number of AI vendors — including Alibaba, DeepSeek, MiniMax, Moonshot AI, StepFun, and Z.AI — of using US large language models (LLMs) to assist training and development of their own proprietary models.

These companies supposedly "extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024," the advisory stated.

The agencies claim the firms are doing this through a process called distillation, a common practice in which mature "teacher" AI models are used to train "student" AI models. On its own, distillation is a widely accepted practice used for academic research, making models more efficient, and improving models with specialized use cases.

Related:AI's Vulnerability Surge May Be More Manageable Than First Feared

What makes this case different, the advisory explained, is that these firms are allegedly training on outputs obtained in violation of the terms of service, deliberately extracting a competitor's proprietary capabilities, and using evasive techniques to avoid detection. The US agencies claim this is likely happening with the awareness of the Chinese government.

Inside China's AI Distillery

The advisory outlines multiple techniques for accessing US frontier models at industrial scale.

"Advanced industrial-scale distillation tactics include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures," the advisory read. "China-based AI companies that conduct industrial-scale distillation against U.S. AI models see significantly shorter AI development timelines and reduced financial expenditures in training a frontier model."

To save money during this intensive process, China-based AI firms allegedly obtain bulk premium subscriptions for US AI models and share them across teams of developers.

... continue reading