Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/20VC/20VC: Are OpenAI and Anthropic O…
POD
// EPISODE
20VC

20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks

DATE July 20, 2026SOURCE 20VCPARTICIPANTS HARRY STEBBINGS, LIN QIAO
// KEY TAKEAWAYS6 ITEMS
  1. 01The Specialized Intelligence Thesis: Millions of Models, Not One AGI
  2. 02The Product-Market Fit vs. Durable Business Divergence
  3. 03Open Source as the Enabling Layer for Specialized Intelligence
  4. 0410x Token Cost Reduction → 100x Usage Explosion
  5. 052025 Is the Year of Co-Work
  6. 06Every Company Will Own Its Own Intelligence Stack

1. Key Themes

The Specialized Intelligence Thesis: Millions of Models, Not One AGI

Lin Qiao's core worldview is that the future of AI is not one dominant AGI model but millions of specialized models — one per application, per use case. The argument is grounded in data: the majority of the world's data is private, locked inside enterprises, and will never be used to train general-purpose models. The real frontier of intelligence is therefore private, specialized intelligence derived from that proprietary data.

"I really believe the future will not be a few small number of AGI models dominant world. I really believe the future will be — it may be scary but I think that's true — it will be millions of specialized models, one per application per use case." 00:20:23

The Product-Market Fit vs. Durable Business Divergence

Lin identified a structural shift that breaks a foundational SaaS-era assumption: in SaaS, product-market fit and a durable business were essentially the same thing. In AI, they are decoupled. Companies can achieve strong PMF but scale into bankruptcy because inference costs make serving their user base economically unviable at scale.

"Now, product market fit and durable business are two separate concepts. For startups, we have great companies that have product market fit... but they cannot scale because once they scale, they could scale into bankruptcy." 00:15:51

Open Source as the Enabling Layer for Specialized Intelligence

Fireworks made an early, high-conviction bet on open-source models when they were nascent. That bet paid off. Open models have now crossed a quality threshold where they can solve the majority of enterprise problems — and critically, they can be fine-tuned with small amounts of proprietary data to outperform general-purpose closed models on specific tasks.

"Both model categories cross a threshold solve so many variety of problems... open model cross a threshold is so much easier to tune... a small amount of unique data a particular company has and then we can hill climb towards your eval. And oftentimes, the end result of hill climbing is to solve your unique problem with your data, you are better than a general purpose model." 00:14:45

10x Token Cost Reduction → 100x Usage Explosion

Lin made a specific, quantified prediction: token costs will fall 10x in the next three years, driven by supply chain normalization, model efficiency improvements, and infrastructure optimization. This will in turn drive 100x usage growth because cost is the primary constraint on AI adoption at scale.

"I can imagine 10x cost reduction in the next three years and this 10x cost reduction will drive 100x usage." 00:44:15

2025 Is the Year of Co-Work — Broader and More Diverse Than Coding

Lin framed 2024 as the "year of coding" (highly concentrated around companies like Cursor) and 2025 as the "year of co-work," which is inherently more diversified across legal, finance, customer support, recruiting, sales, marketing, and healthcare. This is a structural shift in customer base composition and risk distribution for inference providers.

"I think last year is the year of coding, I think all major coding companies are us, and this year is the year of co-work and co-work is much more diversified by itself than coding because there are so many different categories of special purpose co-work." 00:00:00

Every Company Will Own Its Own Intelligence Stack — Just Like Its Own Software Stack

Lin drew a direct analogy to the SaaS era: every company eventually built its own software stack because standardized off-the-shelf software couldn't solve unique problems. The same is happening with AI intelligence. Owning intelligence will become a strategic necessity, not an option.

"There's a reason why every company builds their own software stack. There's no standardized software you just use off the shelf to solve your problem... Same, I think at a time every single company should own their own intelligence." 01:13:21

Distributed RL Training: Fireworks's Hidden Infrastructure Innovation

Working embedded with Cursor, Fireworks's CTO Dima pioneered a fully distributed reinforcement learning training architecture that decouples the trainer (weight updates) from RL rollout (environment interaction), running across five to six data center regions globally using scattered GPUs. This bypasses the need for a massive, expensive, tightly interconnected GPU cluster and solves fresh model weight synchronization across regions.

"We designed fully distributed system, we run across five, six data center regions globally and tapping to scattered GPUs and they're able to run massive jobs... we innovate a way we can distribute fresh model weights quickly, it's not too off, so numerically it's still sound while we are not limited by a very expensive deployment of GPU fleet." 00:33:15

The Hardware Depreciation Problem Changes Build vs. Buy Calculus

GPU hardware used to depreciate over six years with a three-year release cycle. Now, a single vendor releases three new SKUs per year. Models prefer the newest hardware, and older models have questionable long-term value. This dramatically changes the economics of owning vs. renting compute, and challenges the conventional wisdom about when to build data centers.

"In the past, it's six years, solid six years. And hardware release is usually three years... now within a year from one vendor alone, we have three SKUs... after three years, there are nine hardware SKUs in between. Do you still want to go back to nine generation older hardware running three years old model on that?" 00:53:03

Sovereign Intelligence: The National and Enterprise Independence Imperative

Drawing an analogy to electricity infrastructure, Lin argued that just as every country needs to own its power grid, every country and every company needs to own its intelligence infrastructure. The temporary TikTok ban was cited as the wake-up call — if a government can turn off your AI infrastructure, you are dangerously exposed.

"I definitely see that possibility. I also see if we think about the general intelligence model as the electricity layer, as a power line, every country should own their own power line." 00:58:19


2. Contrarian Perspectives

Anthropic and OpenAI May Be Dramatically Overvalued

The implicit argument Lin makes — that 90% of enterprise workflows can now be handled by open-source models, fine-tuned cheaply — means the total addressable market for closed frontier models is far smaller than assumed. Harry pushes this directly and Lin doesn't disagree.

"When 90% of enterprise workflows can be done... with open models, the usage for frontier models will not be as large as it was if it was needed for everything. So are these companies actually dramatically overvalued and overestimated if the majority can just go through open?" 00:15:29

"I think people start to realize it." 00:15:51

AGI Believers Are Building Power Lines, Not the Future of Intelligence

Lin reframes the entire AGI narrative: Sam Altman, Dario Amodei, Larry Page, and Sergey Brin are building important infrastructure (like power lines), but the actual value creation happens at the specialized layer. The power line doesn't replace the appliances — it enables them.

"I view Anthropic as a company fully believing AGI. The definition of AGI is there's this one model that can solve all the problems in the best way... the question is, is this power line going to replace everything we do? I don't think so." 00:08:43

Jensen Huang Is Building Open-Source Models Purely as a Supply Chain Fix, Not a Business

Most observers view NVIDIA's NemoTron model development as strategic market expansion or competitive positioning. Lin frames it entirely differently — Jensen is solving a supply chain problem. If the US lacks a US-native open model, that's a blockage in the five-layer AI stack, and Jensen's job is simply to unblock it.

"Why models? I think it's pure supply chain question. If US doesn't have a US native open model, it's a problem. It's a supply chain problem. He is solely there to solve the supply chain problem." 00:40:38

Gross Margin Optimization Is a Trap in Hyper-Growth AI

Conventional wisdom demands founders optimize gross margin as proof of business quality. Lin explicitly rejects this for the current phase — margin optimization imposes constraints that slow innovation, and in a winner-take-most market, sacrificing margin for speed and coverage is the rational strategy.

"Optimize for growth margin is optimized for differentiation... but I want to avoid over-optimizing for gross margin... one possible way to optimize gross margin, we do not grow at all. We just optimize the heck out of it. I know we can climb to a high number, but that's absolutely a disaster outcome." 00:54:54

Building Chips Is Premature for Almost Everyone Doing It

While OpenAI, Anthropic, DeepSeek, and Meta all announce chip programs, Lin argues this is only rational when workloads have fully stabilized — and AI workloads are still wildly dynamic. The companies building chips now are mostly doing it prematurely or out of necessity that doesn't apply broadly.

"Once your workload is stabilized, once your business is stabilized, it doesn't change too often, then that's the time to consider building a chip. I still see the whole AI world, especially models, customization, is very dynamic. Workload patterns are very dynamic." 01:00:01


3. Companies Identified

Fireworks AI

AI inference and specialized intelligence platform. Sits between chip providers and model providers, enabling enterprises to fine-tune open-source models and deploy them with workload-specific optimization. Processes 40 trillion tokens per day, majority from customized (not off-the-shelf) models. At approximately $800M ARR, targeting at least a double by year-end. 200 employees. Harry describes writing a $10M check after a 15-minute meeting, and calls it potentially a $500B company.

"We process more than 40 trillion tokens a day. Majority of those tokens are coming from a customized model, not from off-the-shelf models." 00:36:36

Cursor

AI coding tool, described as the pioneer in fine-tuning models for coding use cases. Fireworks's CTO was embedded there for months building distributed RL training infrastructure. Scaled to billions in revenue in a matter of years.

"In coding space, Cursor probably is one of the pioneers starting to tune their model and now almost all coding companies tune their own models." 00:25:18

Harvey

AI-native legal platform. Mentioned as one of two dominant players in a legal AI market that has consolidated from many companies to approximately two. Has committed to building its own model.

"It seems like there were a lot of those companies around two years ago, but now it's pretty much two." 00:56:00

Lagora

AI-native legal platform competing with Harvey. Contrasting approach — has not committed to building its own model.

"You have two companies that compete in the legal space, Harvey and Lagora, and Harvey have committed to building their own model and then Lagora have not." 00:21:55

Meta (Facebook)

Cited for long-term chip development (MTIA since 2018) and for being the exemplar of when building your own chip makes sense — after years of stable, massive-scale AI workloads (ranking and recommendation). Also cited for responsible build-vs-buy thinking in data centers.

"Meta has been building their chips for more than five years, way more than five years. And MTIA has been a project since 2018, maybe earlier." 00:59:35

NVIDIA

Praised for the five-layer AI stack framework (application → model → infrastructure → chips → energy) and for NemoTron open-source model development. Also noted for its acquisition of Groq (SRAM-based ASIC), which Lin sees as a smart architectural combination with GPU for pre-fill vs. generation workloads.

"NVIDIA recently acquired a company also called Groq with Q. It's a large SRAM-based ASIC accelerator." 00:50:08

Groq (acquired by NVIDIA)

SRAM-based ASIC accelerator. Lin explains the architectural logic — SRAM-intense chips excel at the generation phase of LLM inference, while GPU Flops-intense chips excel at pre-fill. The combination is architecturally powerful.

"It's a great combination between a Flops Intense GPU and SRAM-intense ASIC because Flop Intense is really good for first half of LN processing... and SRAM-intense is really good for generation." 00:50:30

Together AI

Inference provider mentioned as the price leader in the market. Positioned as the low-cost option vs. Fireworks's quality-differentiated positioning.

"Together is price king and I don't mean this disparaging but they're cheaper — if you want cheap you go there and respectfully if you want better quality product you go to you." 00:44:50

OpenRouter

Model routing aggregator. Mentioned in the context of the top six models being Chinese, raising national security concerns about dependencies on Chinese open-source models.

"When you look at OpenRouter I think the top six models today are Chinese models and they're incredible quality." 00:19:32

Capital One

Traditional enterprise customer that came to Fireworks inbound, without the company actively pursuing enterprise sales. Cited as validation of enterprise demand.

"We have customers like Geico, like Capital One, all these companies. So even without us pursuing enterprise, traditional enterprise, they come to us." 01:12:21

Geico

Traditional enterprise customer, same context as Capital One — inbound enterprise adoption without active pursuit.

"We have customers like Geico, like Capital One." 01:12:21

Anthropic

Cited as the prime example of an AGI-first company. Lin argues Anthropic's core belief in one dominant model is philosophically opposed to the specialization thesis. Also mentioned in the context of exploring Samsung chip manufacturing.

"I view Anthropic as a company fully believing AGI." 00:08:43

OpenAI

Referenced for releasing cheaper models and for Sam Altman's offer of a 5% stake to the US administration. Also noted for chip development initiatives.

"Is the next step actually we just see a massive reduction in price from the frontier models?" 00:17:24

DeepSeek

Chinese AI lab mentioned for building its own chips and for the quality of its open-source models. Also noted for China potentially restricting access to its models.

"DeepSeek building their own chips." 00:38:52

LinkedIn

Where Lin worked before Facebook, building large-scale data systems. Relevant to understanding her technical background.

Salesforce

George Hu's former employer, where he served as President. Cited as context for the caliber of operator Fireworks has recruited.


4. People Identified

Lin Qiao

Founder and CEO of Fireworks AI. Former Meta/Facebook engineering leader who spent seven years there intentionally to learn people management before founding. PhD in distributed systems and databases. Started the company at age 48. Deep background in large-scale recommendation systems. Co-founded Fireworks with Dima (CTO). Processed more than 40 trillion tokens per day, ~$800M ARR, 200 employees.

"I always want to have a tech business myself... I decided I want to go to a place I can learn the most of people. And the best company at that time is Facebook. It's a rising star in Silicon Valley." 00:05:39

George Hu

President of Fireworks AI (recently hired). Former President of Salesforce. Described by Harry as one of the most direct, no-BS operators he's ever met. Lin identified his rare trait: extremely experienced with high-altitude business vision, yet genuinely curious without assuming he knows the answers.

"I find a unique character about him is he's extremely experienced, has high altitude of business vision, but he's also very curious. He doesn't make assumptions... He didn't come with that attitude. He knows AI goes at insanely fast pace." 01:05:26

Dima

Co-founder and CTO of Fireworks AI. Described as having brilliant intellectual horsepower combined with extreme humility — a rare combination. Was embedded at Cursor for months building the distributed RL training infrastructure that powered Cursor's recent model launches.

"Dima... he is brilliant, high intellectual horsepower, but extremely humble. It's a weird combination." 01:06:22

Jensen Huang

CEO of NVIDIA. Praised for operating with radical information density — replying to emails within one minute, staying constantly close to ground-level detail. Lin frames this not as workaholism but as the strategic imperative of maintaining judgment quality in a high-velocity environment. Also credited for the five-layer AI stack framework.

"He's everywhere... I sent him an email, he will reply in one minute... Leadership is just judgment. It's not privilege. It's judgment. You basically have the context. You need to have the right context to make the right judgment for the team." 01:09:01

Jonathan (Groq CEO - Jonathan Ross)

CEO of Groq. Praised by Harry as excellent. Lin expressed mutual admiration. Groq was recently acquired by NVIDIA.

"I spoke to Jonathan before this show. Jonathan is excellent. He said what a fan he is of yours." 00:50:26

Eric Vishria (Eric Vishria, Benchmark)

Partner at Benchmark, who led the investment in Fireworks despite a personal rule against investing in big tech executives turned founders. He broke the rule for Lin, though he later confessed his own advisor questioned whether big tech executives succeed as startup founders.

"Eric Vischeria has a rule, don't invest in big tech directors. But he broke that rule with you, which is very special." 00:04:44

Alfred Lynn

Referenced as someone Harry spoke to as part of his diligence on Fireworks before the show. Sequoia Capital partner.

Matt Miller

Referenced as the person who introduced Harry to Lin. Appears to be associated with a venture or investment context.


5. Operating Insights

The "Earn the Right to Build" Capital Allocation Framework

Lin has an explicit rule about vertical integration: you only build down the stack when the business has scaled to the point where the economics of ownership are unambiguous. You don't pre-emptively build data centers or chips to reduce dependency — you earn that right through scale. Until then, agility and focus on differentiation are the priority.

"We need to earn the rights of building anything, so focus is everything for us, and we want to focus on where we add the biggest amount of value based on our strength... When makes sense to build, you earn the right to build for your own giant traffic and if it saves like five times more cost, then you should go do it." 00:38:52

Hire for Extreme Ownership, Not Competence First

Lin's primary hiring signal is not technical competence but a disposition to claim end-to-end ownership of problems without being assigned. People who automatically see the full scope of a problem, work laterally across the org to solve it, and deliver regardless of obstacles have the highest long-term output and growth curve.

"We want people with the high competence, but more importantly, the strong indicator whether they will do well in this wave... is whether they are really built for taking extreme ownership. Extreme ownership as in, people just automatically claim, 'Hey, this is an end-to-end problem. I'm going to see through the whole thing and work with a bunch of people to make it happen.'" 01:08:13

Build Relationships with Executives Before You Need Them

Lin met George Hu a year before hiring him — at which point she told him directly the company was too small for him. She kept him engaged as an advisor reviewing executive candidates. When the company hit the right inflection point, the relationship was already deep, making the hire natural and fast.

"I told him, hey, we're probably too small for you. But I would like to work with you at some capacity. So he helped me actually build out the team, interview a lot of executives... And that early relationship paid off." 01:04:59

Decouple Optimization from Experimentation Phases

Lin's engineering philosophy: never optimize a system until you're certain it's the system you want to scale. During high experimentation phases, over-optimization imposes constraints that slow down iteration velocity and kill discovery. Only once a system is proven do you spend the engineering capital to squeeze efficiency from it.

"During system development and in a high velocity system expanding phase, we don't want to overbuild because we're in kind of high experimentation... Optimization doesn't make any sense. Once we know this is a system that we want to build 100% and we are going to scale this a thousand times bigger, then we go optimize the heck out of it." 00:48:14

Zero KLD as a Quality Standard Across the Training-Inference Boundary

Fireworks enforces bit-equivalent accuracy between their training and inference systems (Zero KLD — zero Kullback-Leibler divergence). This prevents quality loss when a fine-tuned model moves from training to production. Most inference providers don't do this. For customers investing heavily in fine-tuning, quality loss at the training-inference boundary effectively discounts their training investment.

"Between the training system and the inference system when model move over we have bit equivalents. So the numerics are fully the same. We do not lose a bit of accuracy. That's really hard to achieve... because we know primary business is in model customization and inference of customized model. And we want our customers every single dollar investing training maximize it." 00:45:54


6. Overlooked Insights

Consumer Recommendation Systems Are the Next Major Inflection Point for AI Inference

Buried in a single throwaway sentence, Lin mentioned that consumer-facing companies are beginning to rethink their traditional recommendation systems using generative AI — and that Fireworks, with its deep roots in Meta's large-scale recommendation infrastructure, is uniquely positioned to serve this transition. This is a massive, largely unspoken opportunity. Recommendation systems power the core monetization engines of every major consumer internet company (social, e-commerce, streaming). If GenAI rewrites how recommendations work, the inference demand spike would dwarf anything seen in the coding or co-work waves.

"More interestingly, we start to see an uptick of consumer facing companies are all starting to look into GenAI technology. They are changing how they are thinking about their traditional business of doing recommendation, for example. And that's very interesting to me because we have obviously worked at a huge recommendation system in the world at Meta, and we are very eager to see how that transforms into a new economy for us." 00:35:32

ROI Attribution and AI Monitoring Is the Next Under-Invested Infrastructure Layer

Lin identified that the entire industry is currently in "token maximization" mode — measuring success by usage volume. The next phase, already beginning, is "ROI maximization" — measuring the actual business return from AI spend. The tooling to do this (monitoring, attribution, cost-per-outcome measurement) barely exists today. This is an implicit signal for a major infrastructure category that will be needed as AI matures into production across enterprises, yet almost no capital is currently flowing there.

"Monitoring the ROI, I think the industry started to kind of pay attention to it. But eventually that's what matters. Not how much spend is, what is the return? And what is the cost? And what is the attribution?... The token maxing is just a thing in time, but we're quickly moving to ROI maxing, which is all about running a business." 01:11:36