179: 蒸馏风暴:一场无人公开谈论的技术竞赛
- 01Distillation Is the AI Industry's Open Secret
- 02The Two Eras of Distillation: Compression vs. Capability Climbing
- 03OpenAI's o1 Was the Inflection Point That Made Distillation Explosive
- 04ByteDance's Zhang Yiming Made a Rare, Principled Stand Against Distillation
- 05The Organizational Cost of Distillation Is As Dangerous As the Technical Cost
- 06Anthropic Is Actively Fighting Back With Detection Infrastructure
晚点聊 LateTalk, Episode 179 | Hosts: Manchi & Honghao
1. Key Themes
Distillation Is the AI Industry's Open Secret
Distillation — using a powerful "teacher" model's outputs to train a weaker "student" model — has quietly become one of the most consequential and contested practices in AI development. Everyone is doing it, but no one publicly admits it.
"This is indeed a topic that the industry is paying close attention to right now. Everyone is talking about distillation. But all of them go watch this closely — yet no one will talk about it openly." [00:01:35.510]
The Two Eras of Distillation: Compression vs. Capability Climbing
Distillation originally served a narrow purpose — compressing large cloud models into smaller, edge-deployable ones (e.g., for autonomous vehicles). The current era has flipped the goal entirely: distillation is now used aggressively to climb toward frontier capability.
"The early primary purpose of distillation was not to become stronger to catch up — it was to compress the capability of a larger model into a smaller model... The goal was not to become stronger but to be more practical." [00:12:19.170]
"What we're seeing now is distillation for the purpose of becoming stronger. So who started this trend of distilling to become stronger?" [00:12:49.930]
OpenAI's o1 Was the Inflection Point That Made Distillation Explosive
Two specific technical developments from OpenAI's o1 dramatically increased distillation's ROI: (1) it demonstrated that large-scale RL in post-training could teach reasoning strategies, and (2) it created richer "chain-of-thought" inference-time data that became prime raw material for distillation.
"o1 revealed that doing large-scale reinforcement learning during post-training can teach the model reasoning strategies... And after o1, there was a new discovery — that giving the model more compute during inference, letting it generate longer chains of thought step by step, could also continuously improve performance. This chain-of-thought process became better raw material for distillation." [00:08:51.250]
ByteDance's Zhang Yiming Made a Rare, Principled Stand Against Distillation
In an internal meeting at Seed (ByteDance's AI lab), Zhang Yiming explicitly prohibited distillation — an unusual stance given ByteDance's enormous compute resources and the competitive pressure it faces. He argued distillation can only bring you close to, never beyond, the teacher model.
"Zhang Yiming, in an internal meeting at Seed, mentioned that he does not allow ByteDance to do this distillation. And we previously also learned that ByteDance for a long time did not engage in distillation behaviors." [00:14:20.190]
"Zhang Yiming himself mentioned: if you go down the distillation path, at most you can continuously approach the other party — it is very difficult to achieve true surpassing." [00:17:44.910]
The Organizational Cost of Distillation Is As Dangerous As the Technical Cost
Beyond the technical ceiling argument, the more subtle and damning critique is that distillation warps organizational incentives — it crowds out the risky, long-horizon research that actually creates breakthroughs, and makes top researchers feel their innovative work is undervalued.
"If you place a lot of energy and focus on distillation... then within the organization, projects and people doing longer-term exploration, taking more uncertain paths — they may not get enough resources or recognition. I think recognition is very important. Top researchers still need an environment where more innovative research can be encouraged and seen." [00:18:43.470]
"Zhang Yiming may not be the most technically knowledgeable, but he is most likely the best at understanding human nature." [00:19:41.750]
Anthropic Is Actively Fighting Back With Detection Infrastructure
Rather than relying on legal tools alone, Anthropic has built behavioral detection systems to catch distillation in real time — identifying coordinated cross-account querying and chain-of-thought extraction patterns.
"Anthropic said they have established classifiers that identify distillation traffic and a behavioral fingerprinting system. They can detect cross-account coordinated repeated querying and chain-of-thought extraction behavior. So it's not that they wait until another model has distilled them and then check your performance — they can also see it in the process." [00:33:48.250]
The Legal Gray Zone: Distillation Has No Settled Law
Violating a Terms of Service agreement is not the same as legal infringement, and even the question of whether training on legitimately purchased data constitutes copyright infringement remains unsettled — with different signals emerging from U.S. courts.
"Violating the user agreement does not necessarily constitute legal infringement under law. This depends on other legal provisions. And also, their user agreements themselves restrict things very broadly." [00:39:40.430]
"There's a California court ruling — one precedent — that found this does not constitute infringement, as long as you paid for the content." [00:42:36.890]
Intelligence Supply May Be Outpacing Demand for Many Use Cases
A quietly emerging concern beneath the distillation debate: for many white-collar and productivity tasks, the current generation of models is already "good enough," and the commercial logic of spending hundreds of billions to train ever-stronger frontier models is in question.
"The founder of an application company told me: for users' ten tasks, eight of them can be solved by DeepSeek V4 Flash — and Flash is truly very cheap and relatively fast." [00:45:31.310]
"For many white-collar workers' jobs, the current models are truly already sufficient. You build a stronger model — it's like using a dragon-slaying sword when there are no dragons to slay." [00:46:57.890]
2. Contrarian Perspectives
A Student Model Can, in Theory, Surpass Its Teacher
The conventional wisdom is that distillation has a hard ceiling — you can approach but never exceed the teacher. The hosts push back: technically, there is no absolute ceiling, especially with multi-teacher distillation.
"Is it technically absolutely impossible for a student model to surpass the teacher? It may not be absolutely impossible. There are possibilities. This I think is actually a research topic. Distillation has many methods that can be optimized. For example, can I learn from not just one model, but more teacher models — learn from many teachers?" [00:18:14.210]
"If you learn from the best teacher model and you also surpass it, then you may become number one in the world. I think this is an open question technically." [00:36:15.690]
The Frontier Lab Business Model Is Already Being Disrupted, Even Before Distillation Is Definitively Proven
The commercial logic of spending tens of billions to train frontier models is already breaking down — not because distillation definitively "works," but because highly efficient open-weight models (e.g., DeepSeek V4, Grok) are killing the pricing power of expensive proprietary models.
"There is a famous diagram of the 'kill line' — the models below and to the right of V4 and Grok 6 in that zone are all being killed. There really is a problem now: for many tasks, the current batch of strongest models is already more than sufficient." [00:45:31.310]
Anthropic's Own Data Practices Are Morally Equivalent to the Distillation It Condemns
Anthropic (and OpenAI, Google DeepMind) trained their original models using unlicensed books and scraped content, and Anthropic specifically is now settling a $1.5 billion class-action lawsuit over downloading books from piracy sites. Their condemnation of distillation is, structurally, a double standard.
"Anthropic recently had a $1.5 billion class-action settlement. The specific case was that they downloaded many books from pirated book websites without paying for those books. And now many American authors collectively sued them for this." [00:42:08.410]
"Overall, that is definitely a double standard. It's just that these two behaviors are also not completely identical." [00:42:36.890]
ByteDance's Refusal to Distill May Actually Be Costing It Competitive Ground
Rather than being a principled advantage, ByteDance's abstention from distillation may be a real and material reason why Seed is not achieving state-of-the-art results — even against domestic Chinese competitors — and is causing top researchers to leave.
"I think not distilling is a fairly important reason among several. But you also can't attribute everything entirely to this one thing." [00:17:15.290]
"Since Zhang Yiming made this explicit statement, some people have wanted to leave Seed... They feel that they shouldn't give up this method. Because your career has a finite lifespan — I of course hope that during my years at Seed, I can produce my own representative work." [00:22:07.470]
3. Companies Identified
Anthropic Description: U.S. frontier AI lab, maker of the Claude series of models. Why mentioned: Named as the most aggressive public accuser of distillation by Chinese companies; built behavioral detection classifiers and fingerprinting systems to catch distillation; named DeepSeek, Kimi, and Minimax in February and Alibaba's Qwen in June; settling a $1.5B lawsuit for its own use of pirated training data; its Claude 4 series cited as a top "teacher model."
"Anthropic in February named three companies: DeepSeek, Kimi, and Minimax. Then in June it named another company: Alibaba's Qwen. It said this was the largest-scale distillation attack ever discovered — 28.8 million interactions." [00:25:32.130]
DeepSeek Description: Chinese AI lab, maker of the DeepSeek series of open-weight models. Why mentioned: DeepSeek R1's January 2025 release (with six openly released distilled sub-models) was the second major inflection point for the distillation boom; its V4 model (nearly 3T parameters) cited as a frontier open-weight benchmark; named by OpenAI and Anthropic as suspected of distillation.
"R1's technical report was released in January 2025. At the same time, it also released six smaller distilled models — four based on Qwen 2.5 and two based on Meta Llama 3. All six models learned from R1 as teacher. This is actually the most classic original purpose of distillation: compression." [00:10:20.950]
ByteDance / Seed Description: ByteDance's large AI research lab developing the "Doubao" / Seed model series. Why mentioned: Zhang Yiming's rare public (internal) stance against distillation; Seed reportedly restructured its data team in June to form a new top-level "AI Data and Security" department; Seed 2.1 reportedly did use some distillation methods; Seed losing some researchers over the no-distillation policy.
"ByteDance from June began restructuring its data team and is now forming a new AI Data and Security department at the same level as Seed and Flow — essentially elevating data as a critical capability." [00:23:05.350]
OpenAI Description: U.S. frontier AI lab, maker of the GPT and o-series models. Why mentioned: o1 cited as the pivotal model that unlocked distillation's ROI by introducing chain-of-thought and inference-time scaling; named DeepSeek as suspected of distillation; its user agreements prohibit using outputs to train competing models.
"o1 is the first model that truly opened up the reasoning model paradigm." [00:08:51.250]
Alibaba / Qwen (千问) Description: Alibaba's AI division and its Qwen series of models. Why mentioned: Named by Anthropic in June 2025 as having conducted 28.8 million interactions suspected to be distillation — described as "the largest-scale distillation attack ever discovered"; Alibaba subsequently reportedly banned internal use of Claude Code across the group.
"When Anthropic named Qwen in June, Alibaba reportedly banned Claude Code across the entire group. Anthropic told Congress it was 28.8 million interactions — the largest-scale distillation attack ever discovered." [00:25:32.130]
Kimi (Moonshot AI) Description: Chinese AI startup building the Kimi series of models. Why mentioned: Named by Anthropic in February as a suspected distillation actor; used as a benchmark example of a Chinese domestic model currently near the frontier.
"Anthropic in February named three companies: DeepSeek, Kimi, and Minimax." [00:25:32.130]
Minimax (海螺) Description: Chinese AI startup. Why mentioned: Named by Anthropic in February as a suspected distillation actor.
"Anthropic in February named three companies: DeepSeek, Kimi, and Minimax." [00:25:32.130]
Zhipu AI (智谱) Description: Chinese AI lab backed by Tsinghua University, maker of the GLM series. Why mentioned: Notably absent from any public accusation of distillation by U.S. frontier labs — despite being included in the discussion of top Chinese labs.
"Zhipu — no one has said anything. No American company has publicly stated that Zhipu is distilling." [00:25:02.310]
xAI / Grok Description: Elon Musk's AI lab, maker of the Grok series. Why mentioned: Grok 4 (released the day before recording) cited as a new model causing the frontier to feel competitive again; Musk quoted as saying Grok 7 will surpass all current models but acknowledged Anthropic's next model may be very strong.
"Yesterday Grok 4 was just released. And Musk, in a reply to another company's tweet, said Grok 7 will surpass all current models — but he added that Anthropic's next model may be very strong." [00:44:33.690]
Meta Description: U.S. tech giant, maker of the open-weight Llama series. Why mentioned: Llama 3 used as one of the base models for DeepSeek's six distilled R1 sub-models; V4 (1.6T parameters) cited as a large open-weight model that requires significant compute to deploy.
"Four of the six models used Qwen 2.5 as the base, and the other two used Meta Llama 3 as the base." [00:10:20.950]
Google DeepMind Description: Google's combined AI research division. Why mentioned: Google published an article in February noting that "distillation and other IP theft behaviors" had increased markedly since the prior year; Gemini cited as another teacher model.
"Google in February published an article saying: we have observed that since last year, distillation and other behaviors they call 'IP theft' have been increasing." [00:08:21.830]
Horizon Robotics / Haomo (豪末) Description: Chinese autonomous driving company. Why mentioned: CEO Guo Weihao was one of the first industry figures to explain distillation to the host in a real interview context — specifically the need to compress large cloud-trained models to fit on NVIDIA Orin chips in vehicles.
"The first time I had a long offline conversation about distillation was when interviewing Haomo's CEO Guo Weihao. They do autonomous driving. He talked extensively about distillation in autonomous driving — because you train a large model in the cloud and then have to put it on the car, which uses Orin chips with relatively limited compute." [00:11:20.490]
4. People Identified
Zhang Yiming (张一鸣) Description: Founder of ByteDance; one of the most powerful figures in global tech. Why mentioned: Made a rare internal-meeting statement explicitly prohibiting ByteDance's Seed lab from distillation; cited as someone whose understanding of human nature may be more important than his technical knowledge when assessing organizational risk; his decision is causing some talent attrition.
"Zhang Yiming may not be the most technically knowledgeable, but he is most likely the best at understanding human nature." [00:19:41.750]
Wang Yinglei (王英磊) / Adam Wang Description: Former TikTok executive (oversaw live streaming, reported to both Wen Jia and Alex Zhu); now leading ByteDance's new AI Data and Security department. Why mentioned: Appointed to lead the new top-level data organization at ByteDance — signaling how seriously the company is investing in non-distillation data capabilities as a strategic alternative.
"The new AI Data and Security department's head is Wang Yinglei, Adam Wang. He previously led TikTok's live streaming and reported to both Wen Jia and Alex. So he's worked closely with them before." [00:23:05.350]
Elon Musk Description: CEO of xAI; owner of X (Twitter). Why mentioned: Quoted from a July interview predicting Grok 7 will surpass all current models, while also acknowledging Anthropic's next model could be very strong — illustrating the volatility of frontier model rankings.
"Musk, in a reply to a tweet, said Grok 7 will surpass all current models. But he also added: Anthropic's next model may be very strong." [00:44:33.690]
Jensen Huang (黄仁勋) Description: CEO of NVIDIA. Why mentioned: In a late-July interview, was asked about the feasibility of distilling large open-weight models — implicitly noting that deploying a 2.8T parameter model like K3 requires enormous compute and electricity costs that make large-scale covert distillation harder to conceal.
"Huang Renxun in a late July interview was also asked this question — about distilling a very large open-weight model. First, deploying such a large open-weight model requires substantial compute. Then if you continue running various data through it, your electricity costs would also be quite high. In short, it would be relatively easy to monitor." [00:32:25.530]
Yao Shunyu (姚顺宇) Description: AI researcher; previously discussed in Episode 178 of this podcast. Why mentioned: Referenced in the context of the debate over whether intelligence demand (especially for white-collar work beyond coding) is large enough to justify ever-stronger models — his views helped frame the supply-demand mismatch concern.
"Last time we discussed the Yao Shunyu episode... when you expand from coding into the white-collar market, the demand for intelligence may not be that high." [00:46:29.370]
Guo Weihao (郭维浩) Description: CEO of Haomo, a Chinese autonomous driving company. Why mentioned: One of the first practitioners to explain distillation's real-world use case (compression for vehicle edge deployment) to the podcast hosts in a field interview.
"The first time I had a long offline conversation about distillation was when interviewing Haomo's CEO Guo Weihao." [00:11:20.490]
5. Operating Insights
Build High-Quality "True Demand" User Pipelines as a Strategic Moat
The most difficult and most valuable part of distillation is not the training process itself — it's sourcing authentic, high-quality prompts from real expert users (researchers, senior engineers, graduate students). The company that industrializes this user funnel most effectively wins the data war, regardless of which teacher model it uses.
"You build many relay stations, then you have people use the leading GPT or Claude models through these relay stations. And ideally these users are ones who can genuinely produce high-quality questions — maybe a highly capable STEM student, a researcher, or a senior programmer. This is actually an operations job, with a technical component." [00:29:57.330]
Data Pipelines Are the Durable Competitive Advantage, Not the Distillation Act Itself
The real proprietary asset in distillation is the end-to-end data pipeline: how you construct questions, filter real user data, use models to augment and rewrite data, handle error correction, and manage data format and mix ratios at scale and low cost. This system-level know-how is what separates leaders from fast followers.
"How do you further use models to augment and rewrite this data? There is also error correction involved, as well as the overall format and mix ratio of the data. This data pipeline has difficulty and at minimum contains experiential know-how. And whether your system is cost-efficient and stable will also affect your effectiveness and efficiency at large scale." [00:31:26.210]
Use Multi-Teacher Distillation to Break the Ceiling — Don't Rely on a Single Frontier Source
Depending on one teacher model replicates that model's errors and idiosyncrasies. Blending multiple teacher models (multi-teacher distillation) is both a technical strategy to approach or potentially surpass any single teacher, and a legal/compliance strategy to reduce attribution and detection risk.
"Can I learn not just from one model, but from more teacher models — learn broadly from many? Of course it's also possible you train badly — you could broadly fail too. But at least there are many spaces to explore." [00:18:14.210]
6. Overlooked Insights
The "Distilling Yourself" Failure Mode Is Real and Nearly Undetectable
One practitioner casually mentioned that a model company used a third-party relay station for distillation and discovered after the fact that it had been distilling itself — essentially because the relay was routing its queries back through its own model. This is not a hypothetical edge case; it is a documented failure mode in live production distillation pipelines that could waste enormous resources and go unnoticed without explicit provenance tracking.
"I know there is one model company that used a relay station and then discovered that what it had distilled was actually itself — it was distilling itself." [00:30:26.910]
This implies that any company running distillation at scale without end-to-end data provenance auditing is flying blind. For investors, this is a red flag about operational rigor at AI companies running aggressive data acquisition programs. For operators, it means the first infrastructure investment in any distillation program should be provenance tracking — before scaling data volume.
Zhipu AI's Absence From All Accusation Lists Is a Significant Signal Worth Investigating
Every other major Chinese frontier lab — DeepSeek, Kimi, Minimax, Qwen — has been publicly named by at least one U.S. lab as a suspected distillation actor. Zhipu AI has not been named by anyone. This was mentioned only in a single passing sentence but is analytically important: it could mean Zhipu is either (a) technically more sophisticated at concealment, (b) genuinely not distilling (suggesting a different strategic path), or (c) simply not yet on the radar because it is perceived as less threatening. If (b), it is the one Chinese lab building a potentially durable, non-distillation-dependent moat — which would make it the most interesting long-term bet on Chinese AI independence from U.S. frontier models.
"Zhipu — no one has said anything publicly. No American company has publicly stated that Zhipu is distilling or suspected of distilling." [00:25:02.310]