Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/20VC/20VC: Mercor CPO on Revenue Conc…
POD
// EPISODE
20VC

20VC: Mercor CPO on Revenue Concentration from Frontier Labs | Why Large Enterprise is Scared to Partner with Frontier Labs | Why Small Specialised Models is the Future with Osvald Nitski

DATE July 25, 2026SOURCE 20VCPARTICIPANTS HARRY STEBBINGS, OSVALD NITSKI
// KEY TAKEAWAYS6 ITEMS
  1. 01Open Source Raises the Floor, Not the Ceiling
  2. 02Latent Demand Is Vastly Underestimated
  3. 03The Future Is Specialized Models for Every Enterprise
  4. 04Token Spend Will Exceed Salary Spend
  5. 05RL Environments Are the New Frontier Data Type
  6. 06Product Management Is Being Fundamentally Restructured

1. Key Themes

Open Source Raises the Floor, Not the Ceiling

Open source model improvements don't cannibalize the core data business because the real value lies at the frontier of model performance — the capabilities that no existing model can yet achieve. As capabilities become commoditized, new frontiers emerge.

"Open source models just raise the floor of what people are interested in. As long as customers still have new capabilities that they want to get better at, our business still continues to grow. Open source models just mean that nobody is buying anything that Kimi K3 can already do." — Osvald Nitski 00:04:20

Latent Demand Is Vastly Underestimated

The common framing that "90% of enterprise workflows can be handled by AI" is based on currently imagined use cases. The more significant opportunity lies in long-horizon tasks that enterprises aren't even trying yet.

"There's a whole category of latent demand that people aren't even trying to do with models yet. Most commonly we think these are like long horizon tasks like setting up a procurement agent to fully automate your procurement team for months on end... We think that's just not even captured in these calculations when someone says enterprise workflows are being handled because nobody's trying to do these things yet." — Osvald Nitski 00:05:14

The Future Is Specialized Models for Every Enterprise

The era of one-size-fits-all models is transitioning to enterprise-specific specialized models, each requiring proprietary eval and training data. This creates a durable, expanding market for companies like Mercor.

"I buy it. I think it's also self-serving towards Mercor in that we think that every specialized model will need enterprise-specific eval and training data to show the model how to perform in its setting." — Osvald Nitski 00:09:05

Token Spend Will Exceed Salary Spend — And That's Fine

Mercor already spends more on tokens than on salaries, and Osvald sees this as not just acceptable but necessary for hyper-growth companies servicing insatiable demand. He projects this ratio will increase industry-wide.

"We do. 100% sounds reasonable. So as I said, we're — I've only worked at hyper growth companies and that's what Mercor is and continues to be more so every day as the growth just accelerates. For us, it makes sense because the demand that we have is so high... We can't spend money fast enough to service all of the demand that we have." — Osvald Nitski 00:13:04

RL Environments Are the New Frontier Data Type

The fastest-growing data type at Mercor is reinforcement learning environments — simulations of apps and real-world states that train agents to behave in deployment-like conditions. This is where the frontier is moving, and most people haven't caught up.

"The data type that's growing the fastest for us is environments... these simulations of apps that you might want your agent to use... the shift here is that the data that the agents are being evaled and trained on looks a lot closer to what they see in deployment." — Osvald Nitski 00:39:03

Product Management Is Being Fundamentally Restructured

The PM-to-engineer ratio is inverting. As engineering velocity increases with coding agents, understanding business needs — not building — becomes the bottleneck. PMs must now think at the business level, not the tool level.

"The ratio of PM to end will change over time to have fewer engineers per PM as engineering velocity increases with better coding agents... We need to be very careful as a hyper growth company to grow the teams in lockstep... The trend will be higher PM to end ratio though." — Osvald Nitski 00:29:47

AI Services Are a Knowledge Dissemination Problem, Not a Permanent Business Model

The rise of AI services at companies like Palantir and Microsoft is temporary — driven by the concentration of AI deployment knowledge in San Francisco. As that knowledge spreads, services become commoditized.

"We have basically a concentration of a bunch of people in San Francisco who really know how to deploy agents, eval agents, be AI first in engineering and in other areas. That knowledge just isn't out there yet. And eventually it will be... This is, I think, like a decade-long change." — Osvald Nitski 00:19:39

Cybersecurity Is an Uncapped, Adversarial Data Market

Unlike most enterprise workflows that approach sufficiency, cybersecurity is permanently adversarial — the goalposts always move. This makes it structurally immune to the "90% solvable" framing, and a massive growth category for training data.

"There's never going to be that 90% for security because the goalposts are always going to move. So most cyber as a category is growing. And the nature of the data types is much more kind of like uncapped, evolving, adversarial in terms of where the goalposts are." — Osvald Nitski 00:48:10

Robotics Data Is the Next Generational Opportunity

Physical-world data for robotics is nascent compared to Gen AI and even autonomous vehicles. Osvald sees this as a major new revenue line within three years, with a Waymo-style inflection rather than a ChatGPT-style overnight moment.

"I think that real world, like physical data is going to grow significantly over the next three years. Robotics is an interesting area for us. The data market for robotics is nascent relative to Gen AI, relative to autonomous vehicles as well. And we think that's going to grow a lot." — Osvald Nitski 00:54:14


2. Contrarian Perspectives

The "90% of Enterprise Workflows Are Solved" Claim Is Fundamentally Miscalculated

The widely-cited statistic that current models handle 90% of enterprise workflows ignores an enormous category of latent demand — tasks no one is even attempting yet. Mercor's own Apex benchmarks suggest top models are only at ~50% on long-horizon workflows.

"In our Apex benchmarks, we're getting closer to around 50% of long horizon workflows. Top models are scoring around that much... There's a class of workflows that we shouldn't even be thinking about in terms of binary, like can the models do it or not?... We need to be thinking more about continuous uncapped rewards." — Osvald Nitski 00:07:44

Large Enterprises Put Sensitive Core Data on Open (Often Chinese) Models

There is a deep irony in enterprise security behavior: companies put sensitive, core IP on open-weight models (often from Chinese labs) because they don't want it on closed proprietary models — yet those open models may carry greater actual risk. The conventional wisdom about proprietary vs. open data sensitivity is backwards in practice.

"Am I the only one who sees the irony — we put the sensitive, sensitive data on open source, most likely Chinese models, and we put the HR and procurement data on the closed model?" — Harry Stebbings 00:07:07

"Well, it depends on where you run the open models, right? Whether or not that's a bad idea. So the beauty about open weights models is that the inference can happen in multiple places. So you could make mistakes using them, but you have more control." — Osvald Nitski 00:07:22

VC-Subsidized Founder-Led Annotation Is a Structural Competitor That Can't Scale — But Labs Love It

A cottage industry of founders doing annotation themselves, funded by VC capital, is undercutting structured vendors on price. This is a real near-term competitive threat to Mercor, even though it is fundamentally unscalable.

"We're facing what looks like a cottage industry of founders doing annotation themselves... Labs love this because it's just like totally mispriced... they're smart people, they're founders... but they're running the projects themselves. This is just like VC-subsidized work that labs love. The problem is scaling it beyond a few data points." — Osvald Nitski 00:41:52

Services Revenue Is Not a Sign of Business Strength — It's Temporary Arbitrage on Knowledge

While Palantir and Microsoft's surging services revenue is celebrated by markets, Osvald argues it's structurally temporary: it exists only because AI deployment knowledge is still concentrated in SF. Once that knowledge disseminates, it becomes a commodity.

"As the knowledge of how to use AI gets disseminated throughout industry... that knowledge just isn't out there yet. And eventually it will be. And maybe you won't need at that point teams to go and set things up... it'll become more of like a job function similar to software engineering." — Osvald Nitski 00:19:39

Frontier Labs Building Application Products Is History Repeating — They Will Lose to Focused Competitors

Fears that Anthropic or OpenAI will destroy vertical AI companies (like Ironclad in legal) are overstated. Historical precedent from Google+ and Microsoft's many failed diversification attempts suggests focused competitors nearly always win against unfocused giants.

"I think if you look to precedents here, large companies often try to, you know, make new bets, diversify. But they lose to companies that have intense focus on their market... I would wonder if there's anything to learn from history with Google and Microsoft having many business units, many efforts, but a core business that has driven all of their revenue." — Osvald Nitski 00:45:27


3. Companies Identified

Mercor

AI-powered marketplace and data platform that matches expert annotators with AI labs and enterprises to produce eval and training datasets. Mentioned as the subject company — growing 10x in headcount, spending more on tokens than salaries, moving toward RL environments and enterprise self-serve.

"We end every week with millions more in the bank... the business is very healthy and we can't spend money fast enough." — Osvald Nitski 00:35:43

Anthropic

Leading closed frontier AI lab, one of Mercor's primary customers for eval and training data. Also highlighted as Salesforce's major AI spend target ($300M/year) and cited as a potential competitive threat to vertical AI application companies.

"Mr. Benioff from Salesforce said that he spends 300 million a year on Anthropic, which works out to be about 3.8% of developer salaries if you average the salaries." — Harry Stebbings 00:11:50

Waymo

Autonomous vehicle company cited as the best analogy for how robotics will scale — not a sudden ChatGPT moment, but a slow, city-by-city rollout that eventually reaches undeniable utility.

"I take Waymo more than I take Uber... technically it works and it might be, you know, there might be like regulatory challenges or other challenges with scaling." — Osvald Nitski 00:56:27

Palantir

Enterprise AI and data analytics company cited as evidence of the rising "services" deployment model for AI. Harry noted Palantir is "skyrocketing" as services become a growing part of AI enterprise revenue.

"We obviously seeing Palantir skyrocket and services becoming an increasing part of everyone's business. Is that the future of AI enterprise deployment?" — Harry Stebbings 00:19:18

Fireworks AI

AI inference and fine-tuning company. Founder Lin Qiao cited as a proponent of the specialized-model-per-company thesis, which Osvald endorses as consistent with Mercor's own trajectory.

"I had Lin Qiao, the founder of Fireworks, on the show the other day. And she was like, exactly that is why we'll have specialized models for every single company." — Harry Stebbings 00:08:27

ClickHouse

Open-source analytical database company. CEO Aaron cited for aggressively 6x-ing AI spend to stay at the frontier, presented as one model for how companies should think about AI investment.

"We saw Aaron from ClickHouse say that he's 6x'd spend and that's what they need to do because we need to be at the frontier." — Harry Stebbings 00:10:27

Factory

AI coding and software development company. Founder Matan cited for the provocative take that "services are just an excuse for crap product" — a counterpoint to the services-as-deployment thesis.

"Matan from Factory said to me, you know what? Services, they're just an excuse for crap product." — Harry Stebbings 00:20:17

Ironclad (referred to as "Lagora")

AI-powered legal contract platform. Harry mentioned it as his portfolio investment and used it as the test case for whether Anthropic would eventually come after vertical AI legal applications.

"I'm an investor in Lagora. And everyone's like, no, your real competition is actually Anthropic." — Harry Stebbings 00:44:15

Surge (Scale AI's human labeling division)

Human data labeling competitor to Mercor. Referenced in context of price competition and discounting on data projects.

"Are they price sensitive? Are they haggling going, oh, well, Edwin at Surge gave me a 10% discount. Can I have that?" — Harry Stebbings 00:40:26

Claude (Anthropic's AI assistant)

Referenced as an internal tool displacing Figma at Mercor via Claude's built-in design capabilities ("Cloud Design"), and also as a risk when employees over-delegate judgment to models.

"Cloud Design has done a great job. People really like using it. It's easy to use. And we've just had a natural movement towards it." — Osvald Nitski 00:18:45

Salesforce

Enterprise software company. CEO Marc Benioff cited for $300M annual Anthropic spend and for building out an AI services division — used as a data point in the AI ROI and enterprise spend debate.

"Mr. Benioff from Salesforce said that he spends 300 million a year on Anthropic." — Harry Stebbings 00:11:50


4. People Identified

Brandon (Mercor CEO/Co-founder)

Co-founder and CEO of Mercor. Referenced multiple times as having exceptional market judgment, deep relationships with frontier labs, and the ability to validate data hypotheses faster than anyone. Also cited for predicting that AI token spend will hit 100% of salary costs.

"Brandon said on the show that it would hit 100% and he said that you already spend more today than you do on salaries." — Harry Stebbings 00:12:55 "I think Brandon does it very well. I think our operations team does it very well. But ultimately, it's kind of like a guess." — Osvald Nitski 00:18:00

Lin Qiao

Founder of Fireworks AI. Cited for the thesis that every company will have specialized models tailored to their specific business priorities, which Osvald strongly endorsed.

"I had Lin Qiao, the founder of Fireworks, on the show the other day. And she was like, exactly that is why we'll have specialized models for every single company." — Harry Stebbings 00:08:27

Alex Karp

CEO of Palantir. Referenced for two specific public arguments: enterprise skepticism toward sharing sensitive data with frontier model providers, and questionability of AI ROI for large enterprises.

"I had Alex Karp and... the way he said about the incredible skepticism we see from large enterprises towards data and sharing data with the frontier model providers." — Harry Stebbings 00:06:02

Matan (Factory)

Founder of Factory AI. Cited for the contrarian view that services are a product failure signal, not a deployment feature.

"Matan from Factory said to me, you know what? Services, they're just an excuse for crap product." — Harry Stebbings 00:20:17

Marc Benioff

CEO of Salesforce. Referenced for the concrete data point of $300M annual Anthropic spend, and for Salesforce building out a dedicated AI services department — framed as a bellwether for enterprise AI budget direction.

"Mr. Benioff from Salesforce said that he spends 300 million a year on Anthropic, which works out to be about 3.8% of developer salaries if you average the salaries." — Harry Stebbings 00:11:50

Aaron (ClickHouse CEO)

CEO of ClickHouse. Cited as a concrete example of aggressive AI-first investment philosophy — 6x-ing spend on the premise that frontier access compounds.

"We saw Aaron from ClickHouse say that he's 6x'd spend and that's what they need to do because we need to be at the frontier." — Harry Stebbings 00:10:27


5. Operating Insights

Fight Product Surface Area Expansion Relentlessly — Especially When Engineering Is Cheap

When coding agents make building fast and cheap, the natural temptation is to add features indiscriminately. Osvald's hard-won lesson: this creates operational chaos. The PM's primary job becomes reduction, not addition.

"We're constantly in this battle to try to simplify our product surface area and find the interactions and the workflows that are most scalable... as a product team, constantly fighting to reduce surface area and simplify things." — Osvald Nitski 00:13:49

"We supported, I think, too many workflows for human data projects... We tried to serve every ask. We made a tool that's maximally flexible... That's just chaos to manage. And what we needed to do sooner was to put guardrails on the type of services that we support." — Osvald Nitski 00:16:00

Hiring Interviews: Replace Take-Homes With an AI Fluency Screen + Whiteboard Judgment Test

Mercor has moved to a two-stage process: one assignment using an AI agent to produce an artifact (confirms AI fluency), then whiteboarding sessions focused on experiment design, statistics, and systems design — skills that are hard to fake and hard to outsource to models.

"We do one take home assignment, which is like, can you just use an agent... produce this artifact for me and we'll look at it. Do that once, you know that the person's AI fluent. And then we move towards a lot of whiteboarding... we care a lot about being able to set up good experiments and understanding statistics, having good judgment and systems design as well." — Osvald Nitski 00:24:25

Never Delegate Judgment or Decision-Making to Models — Draw the Line There

Osvald's personal rule, which he enforces culturally across his team: AI tools are for execution, not for the judgment calls that define your actual job. Delegating decisions to models degrades the capability over time.

"I want to be very careful never to delegate judgment or decision making to models... don't delegate your decision making, like your actual job to a model because you're going to lose that ability." — Osvald Nitski 00:26:37

Hold Friday Afternoon Product Meetings to Enforce Accountability

A small but deliberately chosen tactical choice: product meetings on Friday afternoon prevent early weekend departures and enforce a hard accountability checkpoint at the end of the week.

"We need to have weekly meetings to maintain accountability. Do them on Friday a bit later in the day, make sure no one's leaving early on the weekend." — Osvald Nitski 00:31:28

Scale Supply Ahead of Demand Only for Exceptional Talent — Not Broadly

Mercor's selective pre-buying of talent supply: retain exceptional experts who might be needed for future high-value work, and create off-the-shelf data assets during low-demand periods for later resale. This is not a general policy — it's targeted at the highest-skill tier.

"At times, we retain exceptional talent to do work that might be valuable in the future. And we can do off-the-shelf data creation to basically make use of supply when demand is low and then resell that data later. In that case, we do. Otherwise, we don't." — Osvald Nitski 00:54:37


6. Overlooked Insights

Mercor's Apex Benchmark Is a Proprietary Evaluation Framework — and a Potential Business in Itself

Osvald casually mentions that Mercor runs its own "Apex benchmarks" that measure long-horizon workflow completion rates for top models. This is a non-trivial capability: a proprietary, real-world benchmark for agent performance that no existing public benchmark captures. If Mercor controls the definition of what "done" looks like for long-horizon enterprise tasks, they control the optimization objective for every lab buying their training data. This is a strategic moat hiding inside a throwaway sentence.

"In our Apex benchmarks, we're getting closer to around 50% of long horizon workflows. Top models are scoring around that much." — Osvald Nitski 00:07:44

The "Saj" Reference Reveals a Secretive, Non-Imitative Competitor Worth Watching

In an otherwise dismissive review of competitors ("every time I look at one of their websites, they're just doing something we did like a week or a month ago"), Osvald carves out a specific exception for a competitor referred to as "Saj" — describing them as secretive, not copying Mercor, and genuinely differentiated. This is the only competitor Mercor's CPO treats with real intellectual respect, and they named them on a public podcast while simultaneously admitting they don't know why they're different. That combination — secretive, differentiated, and surprising a well-informed insider — is a strong signal of a company worth independent research.

"Would you say that about Saj? It's happened before. Yeah. It's happened before. They're a bit out there. Honestly, I don't spend that much time thinking about them because I spend more time thinking about customers. We've seen it. They're a bit out there in that they don't copy us as much, and they do seem a bit different from others in the field. Hard to say why. They're very secretive." — Osvald Nitski 00:52:45