Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/STRATECHERY/Who’s Afraid of Chinese Models?…
NEWS
// NEWSLETTER ISSUE
STRATECHERY

Who’s Afraid of Chinese Models? (Stratechery Article 7-20-2026)

DATE July 20, 2026SOURCE STRATECHERYPARTICIPANTS BEN THOMPSON
// KEY TAKEAWAYS5 ITEMS
  1. 01Theme 1: Intelligence Is Becoming a Commodity
  2. 02Theme 2: Frontier Labs Will Survive
  3. 03Theme 3: China's Open-Weights Strategy Is Geopolitically Deliberate
  4. 04Theme 4: Distillation Is a Structural Advantage for Chinese Labs That U.S. Open-Weight Builders Cannot Match
  5. 05Theme 5: Cybersecurity Is the Real and Urgent Danger
// SUMMARY

Ben Thompson | July 20, 2026


1. Key Themes

Theme 1: Intelligence Is Becoming a Commodity — Cost Structure Is the New Moat

AI has transitioned from a software-like zero-marginal-cost business to one where COGS are real and decisive. In commodity markets, the lowest-cost producer wins — not the one with the best brand.

"We are rapidly approaching a state in which intelligence for many economically beneficial tasks is in fact a commodity. Anyone building a basic CRUD app, for example, can likely do so using models from multiple providers. And, in a commodity market, the route to profitability is not through charging higher prices... but rather through having a superior cost structure."

The five factors that determine COGS for intelligence — model footprint, inference efficiency, memory efficiency, serving efficiency, and token efficiency — are where competition will be fought.

"The COGS for intelligence is a function of a few different factors: Model footprint... Inference efficiency... Memory efficiency... Serving efficiency... Token efficiency: The fewer tokens required to reach a correct answer, the lower the inference cost."


Theme 2: Frontier Labs Will Survive — The Real Threat Is Being Displaced Up the Stack

Thompson argues the panic over Chinese models is overblown for frontier labs. Their structural advantages — months of head start on cost optimization and proprietary inference data flywheels — are durable.

"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence, thanks to model capability, serving scale, and token efficiency. They are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs."

The real threat is not Chinese models eating their lunch, but frontier labs themselves threatening software incumbents by moving up the stack into customer experience.

"It's striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with... this imperative to move up the stack does mean that frontier models are absolutely a threat to software providers, including Microsoft."


Theme 3: China's Open-Weights Strategy Is Geopolitically Deliberate — "Commoditize Your Complements"

China's push for open-source AI is not altruism — it is a calculated strategy to erode U.S. AI advantage while strengthening China's dominance in physical-world applications like robotics.

"The strategy for China is obvious: commoditize your complements. Note that Xi explicitly ties openness to AI 'moving from the digital world into the physical world'; the physical world is the world dominated by China, and the country's lead in areas like robotics is going to massively benefit from widely available AI models."

"China does not want the U.S. to gain an asymmetric advantage in AI; to the extent that China can weaken the U.S. frontier labs while strengthening any and all potential U.S. adversaries so much the better."


Theme 4: Distillation Is a Structural Advantage for Chinese Labs That U.S. Open-Weight Builders Cannot Match

Chinese labs can legally distill from U.S. frontier models, while U.S. open-weight builders cannot — forcing them into an absurd indirect path through Chinese intermediaries.

"Because U.S. open weight model makers must follow the frontier labs' terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs."

Dean Meyer and Konstantine Buhler's analysis: "Distillation compresses the costly final gap between a strong base and a near-frontier system... Every Western frontier advance therefore creates another teacher for Chinese labs. Western builders must either reproduce those capabilities independently or wait to learn from Chinese models. This gap gives Chinese labs a recurring structural advantage over Western companies."


Theme 5: Cybersecurity Is the Real and Urgent Danger — Current U.S. Policy Is Making It Worse

While economic fears about Chinese models are overblown, the cybersecurity dimension is genuinely alarming — and current Trump administration restrictions are actively backfiring.

"Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!"

The Hugging Face breach is the canary in the coal mine: defenders were forced to use Chinese open-source models because U.S. frontier model guardrails blocked legitimate incident response.

"Hugging Face's defenders turned instead to the open-source GLM 5.2 model from China's Z.ai lab... In an incident report, the company recommended that defenders 'have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.'"


2. Contrarian Perspectives

Perspective 1: Chinese Models Are Not Actually Cheaper on a Unit-of-Intelligence Basis

The consensus is that Chinese models like Kimi K3 are disrupting U.S. labs on price. Thompson argues this is a category error — token price ≠ intelligence price.

"Kimi, for example, reportedly uses significantly more tokens than Sol, rendering its price advantage moot... What is fungible is what is constructed from tokens, which is to say intelligence... The COGS for intelligence is a function of a few different factors."

Evidence: Kimi K3 costs $3/M input tokens vs. Sol's $5/M, but if Kimi requires substantially more tokens per task due to reasoning inefficiency, its effective cost-per-correct-answer may be higher or comparable.

"I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence."


Perspective 2: Distillation Should Be Legalized, Not Fought — It's Philosophically Indefensible to Prohibit It

The conventional view treats distillation as IP theft. Thompson inverts this: LLMs themselves are distillations of the open internet, making terms-of-service bans on distillation philosophically incoherent.

"Why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?"

Policy prescription: The U.S. should pass a law (1) making data collection for training fair use, and (2) barring ToS that prohibit distillation for U.S. companies — turning a vulnerability into a national innovation asset.


Perspective 3: The Agent Paradigm Unlocks a Volume-Over-Price Business Model That Frontier Labs Are Underestimating

Frontier labs remain anchored on a training-cost-dominated financial model. Thompson argues the inference market will grow so fast that lower prices + massive volume is a better strategy than defending high margins.

"Going forward, however, I expect the inference market to grow much faster than training costs (and that includes the assumption that training costs will continue to skyrocket), which means they really can make it up in volume. It wasn't clear this would be the case as recently as eight months ago, but the agent paradigm unlock is so massive that frontier labs should have more confidence that they can not just survive but thrive with lower prices."


3. Companies Identified

CompanyDescriptionWhy MentionedKey Quote
AnthropicU.S. frontier AI lab, maker of Claude/FablePositioned as a low-cost, high-capability incumbent that will survive commoditization — but criticized for ideological overreach and safety theater"The ideological angle of Anthropic in particular is impossible to ignore. This is a company that believes only it can be entrusted with AI."
OpenAIU.S. frontier AI labNamed alongside Anthropic as likely lowest-cost producers of frontier intelligence"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence."
Moonshot AI (Kimi)Chinese AI startup, maker of Kimi K3Central case study; its K3 model triggered market panic and the article's core analysis"Kimi K3... approaches the state-of-the-art in terms of capabilities... Moonshot was forced to pause taking on new subscriptions late on Sunday to manage overwhelming demand."
AlibabaChinese tech conglomerateReleased Qwen3.8 Max (2.4T parameters), reverting to open-weights after briefly going closed"Alibaba plans to make the model open-weight soon... Alibaba stopped releasing weights for its leading edge models earlier this year, but appears to have reverted that change."
SpaceXAIElon Musk's AI infrastructure companyCited as an example of compute resellers profiting from current supply scarcity"Nvidia's customers, like SpaceXAI, can turn around and resell compute at high margins as well to a company like Anthropic."
MicrosoftEnterprise software giantNamed as a key player trying to help enterprises run their own models to resist frontier lab encroachment"Companies like Microsoft are increasingly obsessed with helping companies run their own models... That is much more viable if Chinese models are a viable alternative."
Thinking MachinesWestern open-weight model makerCited as an example of Western labs forced to use Chinese models to bootstrap reinforcement learning"Thinking Machines, for example, which just released an open-weight model, relies on Chinese models to solve the cold start problem for reinforcement learning."
Hugging FaceAI model/dataset collaboration platformSuffered a breach where U.S. frontier model guardrails blocked defenders, forcing use of a Chinese model"Hugging Face said its production infrastructure was breached by an 'autonomous' AI agent system... So Hugging Face's defenders turned instead to the open-source GLM 5.2 model from China's Z.ai lab."
Z.aiChinese AI labMade GLM 5.2, the open-source model Hugging Face used for incident response when U.S. models were blocked"Hugging Face's defenders turned instead to the open-source GLM 5.2 model from China's Z.ai lab — running it on their own infrastructure to analyse the 17,000+ logs."
NvidiaSemiconductor companyFramed as the neutral infrastructure winner regardless of which model wins; Jensen Huang's "token factory" framing challenged"Nvidia GPUs are model agnostic: they generate tokens, and do so in the fastest and most efficient way possible."
MetaSocial media/AI companyMentioned as selling spare compute capacity to Anthropic, illustrating current supply dynamics"SpaceXAI and Meta selling capacity to Anthropic."

4. People Identified

PersonDescriptionWhy MentionedKey Quote
Jensen HuangCEO of NvidiaHis "token factory" framing is critiqued as outdated in the reasoning/agent era, where tokens are not fungible"Nvidia CEO Jensen Huang has described what Nvidia is building as 'token factories'... This is a framing that definitely made sense during the first paradigm of AI, the ChatGPT era, when tokens were delivered straight to the end user."
Satya NadellaCEO of MicrosoftCited (via X post) as a key voice pushing for enterprise model self-hosting as a defense against frontier lab encroachment"Companies like Microsoft are increasingly obsessed with helping companies run their own models."
Xi JinpingPresident of ChinaHis July 2026 speech explicitly mandating AI openness is identified as the strategic directive behind Alibaba's reversal to open weights"We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing."
Dean Meyer & Konstantine BuhlerAnalysts/writers on XCo-authored analysis explaining why distillation gives Chinese labs a structural, recurring advantage over Western open-weight builders"Distillation compresses the costly final gap between a strong base and a near-frontier system... Every Western frontier advance therefore creates another teacher for Chinese labs."

5. Operating Insights

Insight 1: When Evaluating AI Model Costs, Price Per Token Is the Wrong Metric — Measure Cost Per Correct Answer

For any operator building AI-powered products, the procurement decision should be based on end-task efficiency, not headline token pricing. A cheaper-per-token model that requires more reasoning steps may have higher effective COGS.

"Kimi, for example, reportedly uses significantly more tokens than Sol, rendering its price advantage moot... The COGS for intelligence is a function of a few different factors: [including] Token efficiency: The fewer tokens required to reach a correct answer, the lower the inference cost."

Insight 2: For Cybersecurity Use Cases, Pre-Vet and Self-Host an Open-Weight Model Before an Incident Occurs

The Hugging Face breach revealed a critical operational gap: enterprises relying solely on frontier model APIs for security operations may find those models guardrailed during an active incident, forcing the use of potentially adversarial alternatives.

"Hugging Face... recommended that defenders 'have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.'"

Insight 3: Sticky Developer Tooling (Harnesses) May Be the Most Durable Moat in the AI Stack

For software builders, the implication is that capturing the developer workflow early — before users form habits — is the most defensible position in an otherwise commoditizing model layer.

"It's striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users."


6. Overlooked Insights

Insight 1: Running Inference = Collecting Training Data — This Creates a Compounding Flywheel That Pure Cost Analysis Misses

Thompson briefly notes that inference is not just a revenue event — it is a data collection event that feeds the next training run. This means market share in inference today compounds into model quality advantages tomorrow, raising the stakes of the volume-vs-margin decision beyond simple unit economics.

"Intelligence isn't in fact a perfect commodity, in part because applied intelligence makes itself smarter. Specifically, whoever is running inference is also collecting data, and that data goes into making the next iteration of the model better. This is... all the more reason for the frontier labs to lower prices and increase usage as more compute comes online."

Insight 2: Current High Inference Prices Are a Policy Artifact of Compute Scarcity, Not Structural — Prices Will Compress Sharply as Supply Catches Up

The current "price umbrella" in AI inference is not a reflection of sustainable economics — it is a temporary artifact of GPU scarcity. Operators and investors building financial models around current inference pricing may be dramatically overestimating long-run margins.

"Right now there is a price umbrella that is downstream of the lack of compute... I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence."