Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/20VC/20VC: "Anti-Data Centres is a Ch…
POD
// EPISODE
20VC

20VC: "Anti-Data Centres is a Chinese Psyop" | How Many Planned Data Centers Will Actually Get Built? | Is Energy AI's Biggest Bottleneck? With Thomas Sohmers, Co-Founder @ Positron

DATE September 19, 2026SOURCE 20VCPARTICIPANTS HARRY STEBBINGS, THOMAS SOHMERS
// KEY TAKEAWAYS6 ITEMS
  1. 01Inference Is a Memory-Bound Problem, Not a Compute-Bound One
  2. 02The Memory Wall Is the Real Bottleneck of the AI Buildout
  3. 03Inference Economics Are Wildly More Profitable Than the Market Believes
  4. 04"Pacing the Frontier" Is Strategically and Geopolitically Naive
  5. 05Anti-Data-Center Sentiment Is Factually Wrong and Geopolitically Dangerous
  6. 06Token Price Collapse Masks Massive Value Increase Per Token

1. Key Themes

Inference Is a Memory-Bound Problem, Not a Compute-Bound One

Sohmers draws a sharp technical distinction that reframes the entire infrastructure debate: training is compute-bound (more flops = better models), but inference is fundamentally memory-bound because every generated token requires reading the full set of model weights sequentially. "That forward pass, that inference portion of it is heavily, heavily memory-bound due to the fact that basically for every single token that's generated... that requires going through the weights, the parameters" [00:05:24]. This is the core thesis behind Positron's entire product strategy.

The Memory Wall Is the Real Bottleneck of the AI Buildout

Compute has scaled roughly 7x faster than memory bandwidth over the last decade, creating a structural imbalance that gets worse every year. "Between 2014... till 2024, you had about 120x improvement in the flops of GPUs... The improvement in memory bandwidth is only 17x" [00:07:40]. This divergence stems from physical limits (SRAM cells haven't shrunk with Moore's Law in ~15 years) combined with a historical lack of incentive, since pre-transformer ML (CNNs) was purely compute-bound.

Inference Economics Are Wildly More Profitable Than the Market Believes

Cached tokens cost roughly 1/1000th as much to serve as freshly computed tokens, yet providers price them close to full rate — creating enormous hidden margin. "You make all of your money on selling cashed input and output tokens... There's a reason why Anthropic is being reported to have 80 points of gross margin right now on API business" [00:10:56]. Sohmers directly rebuts the popular narrative of AI labs burning unsustainable cash: "It's absurd to me that the meme of OpenAI, Anthropic, et cetera, are just burning cash... If they stopped training, they'd be massively profitable overnight" [00:12:16].

"Pacing the Frontier" Is Strategically and Geopolitically Naive

Sohmers is skeptical of the safety-driven push to slow down AI development, seeing it as both practically impossible (China won't pace) and dangerous (it concentrates power). "I don't see Putin signing up" is essentially validated: "Agreed. And I think this is a little bit the same naivety... thinking, oh, we're so great, we're so advanced, so far ahead that we can't get caught up to" [00:18:46]. He also floats a cynical but plausible commercial motive: pacing conveniently reduces training costs ahead of an Anthropic IPO — "a little bit of the pacing the frontier discussion is, oh, this is a great way to reduce costs ahead of IPO" [00:12:42].

Anti-Data-Center Sentiment Is Factually Wrong and Geopolitically Dangerous

Sohmers argues the bipartisan backlash against data centers in the US is built on false premises (water/electricity usage) and structurally advantages China, which faces no such constraints. "A single In-N-Out uses more water than the largest data centres in the United States... it's golf courses [that] are orders of magnitude more" [00:00:00]/[00:22:34]. He calls the trend "almost entirely a Chinese psyop" [00:00:00] because it hampers the one country (the US) capable of competing while China "bulldoze[s] all these people's homes and do[es] rolling blackouts... in order to serve the greater good of new training capacity" [00:23:29].

Token Price Collapse Masks Massive Value Increase Per Token

The headline stat — token prices falling from $60 to under $1 per million tokens in five years — understates what's actually happening, because today's tokens are categorically more capable. "A $60 token five years ago, no one would pay a cent for today. That was a complete garbage token" [01:00:32]. Sohmers estimates the real value increase is "closer to a thousandfold, not just the 60 fold" [01:01:53] price decline, suggesting cost-per-token is the wrong lens entirely — Greg Brockman is reportedly moving OpenAI toward "cost per useful result" pricing [01:02:39].

On-Device Small Models Will Increase, Not Decrease, Cloud Token Volume

Contrary to the assumption that local/small models cannibalize frontier model usage, Sohmers argues they act as force-multipliers that generate more cloud queries. "If they have a local LLM that is constantly checking their email, their calendar messages... deciding to do these lookups to cloud hosted models frequently, that's now on a per person basis, a massive increase in the number of tokens being consumed" [00:47:10]. He sees the next 1-2 orders of magnitude of token growth coming from autonomous local agents escalating tasks to smarter cloud models, not humans prompting directly.

Context Length Is the Real Constraint on Agentic Capability, Not Model Size

Sohmers argues that raw model capability has outpaced the system's ability to feed it relevant context. "A million token context length can only hold a portion of some of our internal companies' largest code repositories... the main limiter today isn't the model capabilities itself... It's on how much context can that model have" [00:43:11]. He also notes a critical distinction between advertised and usable context length, citing the RULER "needle in a haystack" benchmark where "GPT 5.6... could only do this about 70% of the time. GPT 6 Astra does it like over 95% of the time" [00:59:52].

Chinese Labs Are Out-Innovating on Algorithmic Efficiency Because of Export Controls

Export restrictions forced Chinese labs to innovate around memory/compute constraints rather than simply buying more hardware, producing techniques the US labs haven't adopted. "DeepSeek beginning of 2025 with DeepSeek v3... [got a] massive decrease in KV cache size with multi-head latent attention... none of the major US companies... are definitely not using MLA" [00:57:22]/[00:58:01]. This is a rare case of sanctions accelerating the sanctioned party's innovation rather than just slowing them down.

Sovereign Debt, Not AI Company Financials, Is the Real Systemic Risk

Sohmers turns the "AI bubble/debt" concern on its head — he's not worried about AI infrastructure debt but about the currency and bond system underneath it. "I believe in Oracle's business model and ability to execute... a whole lot more than United States government... my biggest economic concern... [is] a more acute specific crisis that arises out of the compounding of national debt leading to devaluation of the currency" [00:30:48]/[01:31:11].

2. Contrarian Perspectives

"Pacing the Frontier" Is a Cynical PR Move, Not Genuine Safety Policy

Sohmers suggests the loudest voices calling for AI safety pauses may be motivated by pre-IPO financial engineering rather than pure altruism, while conceding Dario Amodei himself is likely a true believer. "I think that is the biggest internal reason for everyone other than Dario... there's a lot of strategic reasons of saying, okay, by having these auditors, et cetera, that removes some potential responsibility, culpability from like legal perspectives" [00:15:41]. Harry adds the framing that it conveniently serves multiple parties' interests simultaneously: "Sam can have a reason not to IPO because his numbers aren't as good as Anthropic's... And Elon wants time to catch up as well" [00:15:23].

Government Regulation of AI Is "The Modern Road to Serfdom" — Worse Than Feared AI Risks

Rather than fearing runaway AI (Terminator scenarios), Sohmers argues the greater danger is regulatory capture concentrating AI capability among a handful of governments/companies. "Making legal to do matrix multiplications is like the thing that will set us back to... pre-enlightenment capabilities. That is like the biggest attack on classical liberal freedom concepts that I can think of" [00:14:40]. He predicts Dario Amodei will regret advocating for government oversight: "Dario basically... verbally begging for government governments to take over Anthropic... I would love to see his reaction if... he realizes, oh shit... it just becomes a bureaucracy that halts all progress" [00:16:41].

Higher Electricity Prices From Data Centers Is a Fabricated Concern

Against conventional wisdom that data centers strain grids and raise consumer costs, Sohmers argues the opposite is structurally true: data centers bring their own generation and are blocked from even helping lower grid prices. "There is absolutely zero cases where a data center could be potentially pulling power from anything that's already been allocated... we're just not allowing them to hook up to the grid where they could actually be lowering the prices for everyone" [00:26:03]. He also calls out incumbent utilities for lobbying against new capacity precisely because it would lower prices — a self-interested dynamic rarely discussed.

Model Sizes Will Keep Growing Even Though the Market Narrative Says "Smaller, Owned Models" Win

Pushing back on Harry's premise that enterprises will increasingly own smaller proprietary models, Sohmers cites concentration data: "Somewhere around 80, 85% of all tokens consumed and produced are done by just the top four model companies... I can totally buy... that 5% of all tokens consumed will be done by things on-prem" [00:45:19]. The "sovereign small model" narrative, in his view, is a minority use case, not the mainstream trend.

China's Open-Source Strategy Is a Temporary Tactic, Not Genuine Openness

Sohmers doesn't buy that China's embrace of open-weight models reflects genuine values around access. "As soon as they get into pole position, the ladder gets pulled up with them in some way... they definitely will not let the billion people that are not CCP party members benefit equally from technology" [00:17:57].

3. Companies Identified

Positron AI — Fabless semiconductor startup building inference-specific hardware (chips through full rack-scale systems) optimized for the memory-bound nature of LLM inference. Just raised an $875M Series C at a $5B valuation. Mentioned as the founder's own company solving the memory wall problem directly. "We can turn what you would have spent 500 megawatts with NVIDIA equipment and do that in 100 megawatt" [00:28:50]. Next-gen chips will have "eight times more memory capacity than the highest memory skew from NVIDIA" [00:56:04].

Anthropic — Frontier AI lab (Claude). Cited for extraordinary profitability hidden behind a cash-burn narrative and for Dario Amodei's "pacing the frontier" essay. "There's a reason why Anthropic is being reported to have 80 points of gross margin right now on API business" [00:11:22].

OpenAI — Frontier lab; cited for GPT-6 Astra's step-function leap to what Sohmers calls genuine AGI, and for in-house chip development (Jalapeno). "GPT-6 Astra, I do think is AGI" [00:49:45].

NVIDIA — Dominant GPU provider for inference workloads; cited as the benchmark Positron measures its memory capacity advances against, and noted as actually decreasing per-device memory due to market memory conditions. "NVIDIA is actually decreasing the amount of memory per device... based on the market memory conditions" [00:56:04].

Panthalassa — Startup building ocean-based data centers using a pumped-hydro cooling solution; a Positron partner company. "A company we're partnered with, and I'm good friends with the CEO... building ocean space data centers, basically a very interesting pumped hydro solution in the middle of the ocean" [00:27:44].

DeepSeek — Chinese AI lab credited with major algorithmic innovation (multi-head latent attention) driven by export-control constraints, materially reducing KV cache costs. "DeepSeek beginning of 2025 with DeepSeek v3 had made a lot of waves because they were able to get massive decrease in KV cache size with multi-head latent attention" [00:57:22]. Also reportedly developing its own chips.

Oracle — Cited favorably as a company Sohmers trusts more than the US government from a credit/execution standpoint. "I believe in Oracle's business model and ability to execute and do everything a whole lot more than United States government" [00:30:48].

Mercor and Surge — Data-labeling/data-economy companies cited as potentially becoming enormous businesses (~$200B) serving both frontier labs and enterprises, with Surge noted at $3B in revenue. Sohmers is only skeptical due to potential future vertical integration by the labs themselves.

SpaceX / Starlink (Elon Musk's ventures) — Referenced regarding space-based data centers as a long-term (not near-term) infrastructure alternative. "I would never, ever bet against Elon" [00:27:44].

Semi-Analysis — Cited as the source of the "AgentX" benchmark data showing ~96% cache-hit rates in real Claude Code agentic sessions. "About 96% of all the tokens that go through these entire sessions are cached" [00:39:17].

4. People Identified

Thomas Sohmers — Co-founder and Chairman of Positron AI; started in the semiconductor industry 13 years ago when "silicon was a dirty word in Silicon Valley" [00:54:02]. Central figure of the episode; demonstrates deep first-principles technical fluency across chip architecture, memory hierarchies, and quantization while also holding sharply-argued geopolitical/regulatory views.

Gavin Baker (Atreides) — Investor in Positron's Series C; cited twice for his views on data centers as "the greatest economic kind of needle mover for large parts of the country" [00:21:48] and for his framework on company quality: "I can't speak to a company that don't have numbers that are parabolically up and to the right" [01:03:58].

Dario Amodei (Anthropic CEO) — Discussed extensively regarding his "pacing the frontier" essay; portrayed as a genuine true believer in both AI's promise and risk, but naive about the consequences of inviting government regulation. "Dario and I would say the vast majority of people in Anthropic are true believers" [00:15:41].

Sam Altman (OpenAI CEO) — Referenced regarding IPO timing incentives and was present (with Ilya Sutskever) at the low-key ChatGPT launch event Sohmers personally attended at NeurIPS 2022. "It was so funny because... Sam and Ilia were there... at the end, they just said, hey, we launched this little fun experiment called ChatGPT, go check it out" [00:48:51].

Ilya Sutskever — Co-present at the NeurIPS 2022 ChatGPT launch, described by Sohmers firsthand.

Elon Musk — Cited regarding space-based data centers and general track record of high-risk bets paying off. "I would never, ever bet against Elon. I mean, I primarily bet for Elon" [00:27:44].

Greg Brockman (OpenAI) — Cited for signaling that OpenAI may move away from per-token pricing toward "cost per useful result" pricing. "He doesn't think that they're going to be pricing things in tokens much longer" [01:02:39].

Jensen Huang (NVIDIA CEO) — Mentioned as notably silent on "pacing the frontier," which Sohmers reads as meaningful.

Mark Zuckerberg (Meta) — Noted for explicitly opposing "pacing the frontier," saying Meta should continue as planned; Sohmers says he takes Zuckerberg's and Dario's commentary more seriously than politicians'.

5. Operating Insights

Price Cached and Uncached Tokens as Distinct Products With Distinct Margins

Sohmers reveals that the entire commercial model of API-based AI companies hinges on the spread between cached and uncached token pricing, since cached tokens cost ~1/1000th as much to serve. Any company building an LLM-based product should map its own cost structure the same way labs do internally, deliberately re-using cached context wherever possible to capture that margin rather than treating token costs as a flat, undifferentiated line item.

Treat "Advertised Capability" and "Usable Capability" as Separate Diligence Questions

Sohmers' point about context windows applies broadly to any vendor capability claim: "Just saying that something has this maximum context length is one thing. It's can it actually use that context length effectively is entirely different thing" [00:58:56]. Operators evaluating any AI vendor, chip, or infrastructure claim should demand benchmark performance at the claimed spec (e.g., RULER-style needle tests), not just the headline number.

When You Don't Care About Cost, That's a Signal to Push the Frontier Immediately

Sohmers describes Positron's own internal LLM usage policy as a heuristic for capital allocation: "If I can get 10 times the output value out of a model today, I very gladly pay 10 times more per token" [00:44:50]. This is a generalizable operating rule — for high-leverage internal workflows (e.g., chip design, core R&D), price sensitivity should be the last consideration, not the first.

Use Local/Small Models as an Escalation Filter, Not a Cost-Cutting Replacement

Rather than viewing on-device AI as a way to reduce cloud spend, Sohmers frames it as an amplifier: local models should be deployed to constantly monitor low-stakes data streams (email, calendar) and selectively escalate to frontier models only when genuinely needed. This architecture pattern — cheap constant local inference triggering expensive selective cloud inference — is a specific, actionable system design insight for any AI product team.

6. Overlooked Insights

The KV Cache Storage Problem Is Already Bigger Than the Model Itself

Buried in the technical discussion is a strikingly under-discussed fact: for frontier-scale models with long context, a single user's session cache can now exceed the size of the entire model's weights. "At these long context links for these size models, you have the individual user sessions being in the, let's say in the hundred gigabyte range. So with just 50 users on your service, the user context... end up being greater than the model weights that you're trying to store" [00:40:10]. This means the infrastructure economics of serving AI at scale are increasingly dominated by per-user state storage, not model hosting — a completely different (and far less discussed) infrastructure bottleneck than the GPU-shortage narrative that dominates most AI infrastructure conversations. This has direct implications for anyone building memory/storage infrastructure for AI, not just compute.

China's Algorithmic Innovation Under Sanctions May Be a Bigger Threat Than Its Compute Access

While the export-controls debate typically focuses on denying China raw compute/flops, Sohmers casually reveals that constraint has instead forced superior algorithmic efficiency (MLA, gated delta net) that US labs haven't even adopted yet: "None of the major US companies are... using MLA" [00:58:01]. This is a significant, under-examined strategic risk — the sanctions regime may be inadvertently accelerating exactly the kind of efficiency innovation that reduces China's dependence on the very chips being restricted, while US labs, flush with compute, have less incentive to develop the same techniques. This reframes the entire premise of chip export controls as potentially self-defeating.