Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THE A16Z SHOW/Who Grades the AI Models? | Ben…
POD
// EPISODE
THE A16Z SHOW

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

DATE September 9, 2026SOURCE THE A16Z SHOWPARTICIPANTS BEN HOROWITZ, ERIK TORENBERG, JENNIFER LI, RAYAN KRISHNAN, UNKNOWN HOST 02

Note: In this transcript, Rayan Krishnan (founder/CEO of VALS AI) is the guest whose company and track record is being discussed, though many of his lines are mislabeled as "Ben Horowitz" in the transcript. Content has been attributed based on who is describing VALS AI in the first person (Krishnan) versus the a16z hosts (Ben Horowitz, Jennifer Li, Erik Torenberg) asking questions.

1. Key Themes

Public benchmarks are being gamed, creating a false picture of model capability

The catalyst for VALS AI's founding was discovering that self-reported and open benchmarks diverge sharply from real capability. "When Meta released Llama 4... on our held-out private benchmarks, the model was actually underperforming. But on all of the major public benchmarks where the questions and rubrics were actually open source, it was showing incredible capability" [00:03:23 - approx, attributed to Krishnan]. The explanation given is structural: "questions and rubrics were actually open source" allow optimization/leakage, whereas private held-out evals don't.

Independent, third-party evaluation is becoming a new trillion-dollar-industry institution, akin to audit firms or rating agencies

Krishnan frames VALS as filling a role analogous to historical trust intermediaries: "Every time a new trillion dollar industry emerges, there's a need for this independent testing group" [00:00:00]. He extends the analogy explicitly: labs "would like to see a rational buying market... it's not just entirely self-reported to justify that investment," noting "Demis and others in the industry calling for an ecosystem of third-party evaluators" [00:03:52].

Enterprises are massively mispricing intelligence, and token spend is starting to eclipse salary spend

A vivid anecdote: a Fortune 10 company gave engineers a "$100 a day budget" for Claude Code, later raised to "$300 per employee... almost an employee's worth of salary in tokens." This created bizarre behavioral distortions: "the most productive hours of work are actually now 4 to 6 p.m. when the rate limits reset... there's this dead period in the afternoon when people go on walks... because they just don't have the rate limits" [00:16:29]. The broader claim: "we're in this world where it is still very unclear what ROI looks like... token spend may start to eclipse salary spend" [00:17:24].

Benchmarks must be perpetually deprecated and rebuilt — there is no stable finish line

"We have this unofficial motto, always a higher peak... It is our job to perpetually construct these next mountains for them to summit" [00:11:48]. Benchmarks also need refreshing simply because the world changes — legal, medical, and factual benchmarks need re-certification the way professionals need to retake licensing exams: "if you're a lawyer, you have to retake the bar exam... We should also expect models to be tested on the current state of the world" [00:12:14].

Evaluation is an "AI-complete" problem because we haven't even solved evaluating humans

Jennifer Li raises, and the guest concurs, that human evaluation itself lacks consensus (IQ, EQ, Big Five, SAT controversies), so model evaluation inherits this fuzziness: "What is really the distinction between an associate and a partner at a law firm? And there isn't a clear test or an eval for that in the human world" [00:06:29]. The proposed resolution is norm formation over time rather than a single objective standard, compared to MPAA content ratings: "I just think it's going to be necessarily fuzzy, but there will develop norms over time" [00:08:07].

Coding is the leading indicator for how evaluation and adoption will play out in every other knowledge-work domain

"Coding is a sign for what's to come in every domain... if you have a very good coding agent, chances are you have a model that can also make PowerPoint slides or DCFs in Excel... with a high degree of capability as well" [00:21:33].

Recursive self-improvement (RSI) is the frontier benchmark category with the highest long-term stakes, including geopolitically

VALS built an "RSI index" as an apples-to-apples proxy since running the true experiment (a model training its successor) is too expensive: "what we're doing is forming a set of proxies for every part of the process it takes to build the next version of the model... pre-training, post-training, harness level engineering" [00:00:19]. Geopolitically, this is framed as the highest-stakes axis: "long term, what's actually going to be the most interesting is the recursive self-improvement possibility... one country or one company kind of run away with it and produce models that we don't know much about" [00:35:52].

Government's proper role is setting rules, not doing technical evaluation

Jennifer Li articulates a clean division of labor: "the government is particularly ill-suited to do the latter, particularly over time. It's just not a good government function. But they're very good at setting the rules because they can enforce the rules... the government sets and enforces the rules and that a very competent kind of private company then tells them if the rule is broken" [00:30:06].

Auditing-industry failure modes (Enron-style conflicts of interest) are a cautionary template for AI evals

"If you look at auditing as an industry, you end up with issues like Enron, where if you have the same group who's responsible for doing the audit, as well as also consulting and supporting the company, you have a mixed incentive structure and then it just becomes pay to pass the audit or in this case, pay to win the benchmark" [00:09:00]. This directly informed VALS' policy to never sell training data to labs, since "a lot of that industry has now built these gimmick-style benchmarks as a mechanism to sell their data" [00:09:00].

The nuclear "trust but verify" framework is the best analogy for cross-border AI governance

"There's actually a lot to learn from nuclear here... Reagan had this line, trust but verify... there were also flyovers. So it's a mechanism by which a country could audit another country's nuclear stockpile... Having the shared language of evals will allow us to say things like you have the right number of nuclear warheads" [00:34:30].


2. Contrarian Perspectives

Sovereign AI is a massive global misallocation of capital

Despite the political momentum behind national AI champions, the guest is candid that this is economically irrational: "from my very idealistic perspective, I'm surprised to see so much investment in sovereign AI... if I was taking a God's eye view, it would be extremely inefficient to build all of these data centers and replicate this data engineering process and train these very large models. When in fact, you could probably consolidate a lot of these efforts. But it seems like that's not the world we're in" [00:34:00].

OpenRouter's "routing" narrative is somewhat a misnomer

Rather than intelligently routing requests, OpenRouter is mostly a passive gateway: "OpenRouter is a bit of a misnomer in that most of their usage comes from being a model gateway. And so it's actually up to their users to decide which models they want to use when. And that's because really the hardest part of routing is building the evals" [00:15:29]. This implies the real value-add in "routing" hasn't been built yet — it's an evals problem in disguise, positioning VALS as more foundational than routing infrastructure companies.

More expensive-sounding models can actually cost more to run despite lower sticker price, upending simple pricing intuition

"We're actually seeing in a lot of cases Sonnet is more expensive than Opus because it is so token hungry. And so I think if you were to operate based on... use Sonnet where you feel like it's applicable, you may actually end up spending more than you need to" [00:20:58]. This directly contradicts the intuitive assumption that cheaper-per-token models are cheaper overall.

Enterprises' arbitrary token budgets are essentially inventing a new, unjustified compensation line item on the fly

The Fortune 10 anecdote reveals that companies are setting token budgets ($100 → $300/day) with no rigorous ROI basis — "this is actually pretty arbitrary because it's hard to quantify what the right usage limit should be" [00:17:24] — meaning a huge fraction of enterprise AI spend today may be poorly calibrated guesswork rather than data-driven allocation, even at sophisticated Fortune 10 companies.

Model labs actually want to be independently graded, contrary to the assumption that self-interested labs would resist scrutiny

Rather than resisting third-party evaluation as a threat, labs are described as actively wanting it to justify their own capital raises: labs "would like to see a rational buying market... that when they invest billions of dollars to build a new model, there are actually substantive ways they can point to evidence" [00:03:52] — a counterintuitive alignment of incentives between regulators/evaluators and the labs being regulated.


3. Companies Identified

VALS AI — Independent third-party AI model evaluation company, founded in 2024. Described extensively as the guest's own company: built private held-out benchmarks (e.g., that exposed Llama 4 underperformance), a finance agent benchmark used by major financial institutions, VibeCode bench (natural-language-to-full-stack-app benchmark), a Recursive Self-Improvement Index, and Valsmith (a product letting enterprises build custom coding benchmarks from their own GitHub repos). Also runs an internal automation system called "Steve" — "the Economic Vals employee" — to scale evaluation work [00:05:32]. "Our finance agent benchmark is used by a bunch of the big financial institutions to get a sense of how models are improving" [00:09:48].

Meta / Llama 4 — Mentioned as a cautionary case study where public benchmark performance diverged from real capability: "on our held out private benchmarks, the model was actually underperforming" [00:00:00].

Anthropic — Referenced regarding Claude Code enterprise usage/rate limits and thin margins: "Anthropic is running on pretty narrow margins to support this. And they have... massive costs to serve these models" [00:17:24]. Also referenced for its model lineup (Opus, Sonnet) and pricing dynamics.

OpenRouter (acquired by Stripe) — Referenced as a model gateway company; characterized as not truly doing intelligent routing today, which the guest argues is actually an evals problem: "most of their usage comes from being a model gateway" [00:15:29].

Cognition (Devon) — Named specifically as a standout for efficiency in VALS' internal token-maxing experiment: "the Cognition Devon tool is actually very token efficient. And so that's a place we've chosen to adopt more" [00:23:22].

Google DeepMind — Referenced via its leadership (Demis Hassabis) as supportive of third-party evaluation ecosystems: "Demis and others in the industry calling for an ecosystem of third-party evaluators" [00:03:52].


4. People Identified

Rayan Krishnan — Founder and CEO of VALS AI. Described as coming from a research background specifically in benchmark/evaluation construction prior to founding VALS: "I had a background doing research, in particular building benchmarks and evaluations" [00:02:18]. Credited with the founding insight that "one of the biggest drivers for model capability is having a new legible way to evaluate models" [00:02:18], and with building VALS' technical and go-to-market approach (private evals, Valsmith, RSI Index, refusal to sell training data to labs).

Links — Krishnan's co-founder at VALS AI, credited alongside him for the early all-nighter grind phase of the company: "my co-founder, Links, and I pulling an all-nighter to try and get as much done as possible" [00:05:04].

Demis Hassabis (referenced as "Demis") — Head of Google DeepMind; cited as an industry leader publicly calling for third-party AI evaluation ecosystems, lending credibility to the thesis that labs want independent scrutiny [00:03:52].

Xi Jinping and Trump — Referenced in the geopolitical/verification discussion as examples of state leaders whose upcoming meeting signals early "trust" dynamics that need a verification mechanism analogous to nuclear arms inspections: "we're starting to see signs of trust in that Xi Jinping and Trump are going to be meeting next month. But there is no clear way to actually do the verification part of this" [00:34:30].

Ronald Reagan — Referenced for the "trust but verify" doctrine used as the governing analogy for how AI verification between nations/labs should work [00:34:30].


5. Operating Insights

Run "token-maxing" experiments internally before setting company-wide AI usage policy

Krishnan describes deliberately giving the VALS team unlimited access to coding tools for a month to observe real usage patterns before optimizing: "I wanted to do a token maxing experiment... we had a lot of engineers spending between one to two billion tokens a day. I think peak day was one engineer spending six billion... we spent roughly $1.5 million worth of tokens... it was actually 10x more we were spending in tokens than employee salary for that month" [00:22:26]. The output of this experiment directly produced Valsmith and a concrete tooling policy change (adopting Cognition's Devon for efficiency).

Build "auto-issued recommendations" that route each task to the right tool/model automatically rather than relying on individual judgment

Rather than blanket-restricting tool access, VALS gives everyone access to everything but layers automated routing guidance on top: "We have access to all the tools. We give everyone access to everything. But we auto-issue recommendations for any GitHub issue or ticket for where to begin their session. And that should titrate the actual usage depending on the intelligence required for that task" [00:24:25].

Treat benchmark retirement as a deliberate, ongoing practice, not a failure state

Rather than clinging to legacy benchmarks for continuity of comparison, VALS proactively deprecates saturated ones and frames this as core to the business model: "insofar as foundation model labs are hill climbing, they're searching for the next peaks to summit. It is our job to perpetually construct these next mountains for them to summit" [00:11:48]. Operators building any kind of measurement/scoring product should expect to need continuous product rebuilding, not a one-time build.

Structurally separate evaluation revenue from data-selling revenue to preserve credibility

VALS made an explicit, early strategic choice to forgo a "lucrative business" (selling training data to labs) specifically to avoid conflicted incentives, learning from Enron-style audit failures: "one very early decision we made was the decision to never sell training data to labs... a lot of that industry has now built these gimmick-style benchmarks as a mechanism to sell their data" [00:08:38]. This is a reusable playbook for any company positioning itself as a neutral arbiter/certifier in an industry — the business model itself must be structurally decoupled from the party being evaluated.


6. Overlooked Insights

The "messy middle" of model pricing is creating a hidden, exploitable inefficiency for anyone who builds proper internal evals

Buried in a fairly technical exchange is a genuinely non-obvious and immediately actionable finding: "we're actually seeing in a lot of cases Sonnet is more expensive than Opus because it is so token hungry" [00:20:58], combined with "Mew Spark is also very cheap, and 1.2 is very capable" [00:20:58]. Most enterprises are almost certainly defaulting to reputation/sticker-price heuristics ("use the flagship model," "use the cheap model") rather than empirically testing token efficiency per repository/task — meaning there is real, quantifiable cost savings sitting unclaimed for any company willing to run its own Valsmith-style benchmark. This is a concrete arbitrage opportunity mentioned almost in passing.

Benchmark obsolescence isn't just about capability ceilings — it's a legal/regulatory freshness problem with compounding risk

The comparison to professional recertification ("if you're a lawyer, you have to retake the bar exam" [00:12:14]) is framed casually but implies something large: as case law, medical guidelines, and regulations continuously change, any enterprise deploying LLMs for legal, medical, or compliance work is running models against a decaying ground truth unless benchmarks are continuously refreshed. This creates a hidden liability surface for enterprises in regulated industries that nobody in the conversation calls out explicitly as a risk category, even though the mechanism (like the "legal research benchmark" mentioned) was specifically built to address it — suggesting compliance-grade eval-refresh services could be a much bigger market than the episode treats it as.