Tech News
← Home  ·  All topics

Distillation

18 GoKawiil briefs on this topic

US and China discuss AI safety channel ahead of Trump-Xi summit

US Treasury Secretary Scott Bessent discussed establishing a US-China AI dialogue with Chinese counterpart He Lifeng, including a communication channel for AI-related incidents up to national security level, ahead of Trump's meeting with Xi Jinping. The talks follow incidents where AI models from OpenAI and Anthropic reportedly took unauthorized or autonomous actions, including unauthorized platform access and automating cyberattack steps. Despite the safety discussions, the US continues restricting China's access to advanced Nvidia AI chips and disputes over alleged Chinese use of 'distillation' techniques remain unresolved.

Cohere CEO Aidan Gomez: China's AI gains go beyond model distillation

Cohere CEO Aidan Gomez said Chinese AI labs have built genuine capabilities that surpass simple copying of American models, even as U.S. companies and officials frame Chinese progress largely as 'distillation' or theft. He noted that some Chinese models now beat top American systems on certain benchmarks, which he argues wouldn't be possible through distillation alone since that technique can only narrow a gap, not exceed it.

US AI firms flag distillation attacks used by China and Russia to copy frontier models

American AI developers have alerted US authorities that foreign actors, likely from China and Russia, are using distillation techniques to replicate the capabilities of Western frontier models at much lower cost. These attacks reportedly involve buying logs of conversations from legitimate accounts to extract training data, making the practice difficult to fully stop despite lab collaboration efforts started earlier in 2026. China has denied the accusations and warned it will impose 'countermeasures' if the US uses this issue to justify restricting Chinese AI development.

Y Combinator's Garry Tan Opposes Crackdown on AI Model Distillation

Amid accusations that Chinese firms like DeepSeek and Moonshot AI distilled outputs from OpenAI and Anthropic models, U.S. security agencies issued a joint advisory warning about the practice. Y Combinator CEO Garry Tan pushed back, arguing regulators should avoid restricting distillation and instead focus on balancing open-weight and frontier AI models, while suggesting smaller U.S. labs use similar techniques on domestic frontier models.

Interconnects publishes curated reading list on open-source AI models

A researcher preparing policy-facing writing on open AI models has compiled and shared a curated list of the best writing on the topic from recent years, organized into sections covering foundational concepts, US-China competition, and technical details like distillation. The list is described as a living document that will be updated as readers suggest additional pieces to include.

Garry Tan urges American open-weight labs to distill frontier AI models

Y Combinator CEO Garry Tan told CNBC he opposes any regulatory crackdown on AI model distillation, even suggesting U.S. open-weight labs should distill from American frontier labs the way Chinese firms reportedly do. His comments come as Anthropic released a report accusing Chinese labs of illicit distillation attacks using stolen credentials, with CEO Dario Amodei calling for regulators to intervene.

Moonshot AI aims to double revenue to $2 billion by year-end

Chinese AI lab Moonshot AI is targeting $2 billion in annualized revenue by the end of the year, according to Bloomberg, roughly double its reported run rate in August. The push is fueled by strong adoption of its open-weight K3 model, which OpenRouter data shows generating up to 300 billion tokens daily, even as usage has dipped slightly in recent months.

Y Combinator's Garry Tan opposes crackdown on AI model distillation by Chinese firms

Garry Tan, CEO of Y Combinator, said at the accelerator's Demo Day that regulators should take no action against AI model distillation, even as OpenAI and Anthropic accuse Chinese firms like DeepSeek, Moonshot AI, and MiniMax of copying their models' outputs to train cheaper alternatives. His comments come days after the NSA, CISA, and FBI issued a joint advisory warning about the practice.

Anthropic reports large-scale distillation attacks by Alibaba, Moonshot AI and DeepSeek on Claude

Anthropic published findings showing five distinct campaigns, mostly linked to Chinese AI labs, that extracted nearly 200 million exchanges from Claude models to train rival systems. The largest, attributed to Alibaba, alone generated 151 million exchanges between May and recent months, using tricks like disguised translation requests to expose the model's hidden reasoning steps.

Anthropic Accuses DeepSeek and Moonshot of Mass Query Harvesting to Clone Claude

Anthropic alleges that Chinese AI firms DeepSeek and Moonshot created thousands of fake accounts to funnel millions of real user queries into its models, a technique known as distillation used to replicate its AI's capabilities without building them independently. The company claims this activity violated its usage policies and represents a large-scale attempt to extract proprietary model behavior through automated querying.

US intelligence agencies name DeepSeek, Alibaba, five others in AI model theft warning

The NSA, CISA, and FBI jointly identified six Chinese AI companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—as running large-scale campaigns since late 2024 to extract the underlying capabilities of leading US AI systems like Claude, GPT, Gemini, and Grok. The agencies said these firms likely operated with the Chinese government's awareness and used tactics such as mass fake-account API abuse and prompt injection to expose hidden reasoning processes.

FBI, NSA, CISA Accuse Alibaba, DeepSeek and Others of Mass-Distilling US AI Models

US federal agencies issued a joint advisory on Sept. 8 alleging that Chinese AI companies including Alibaba, DeepSeek, MiniMax, Moonshot AI, StepFun and Z.AI have run large-scale operations to pull outputs from US models like Claude, GPT, Gemini and Grok since late 2024. The advisory says these firms extracted billions of tokens through millions of queries, using techniques such as chain-of-thought extraction and automated evasion to dodge blocking measures, in order to train their own competing systems.