Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
- 01Frontier training costs ride a 4x-per-generation escalator, and the catch-up curve is steep
- 02RL is the new scaling axis, and the plots "never stop going up"
- 03Models are accelerating their own training, which bends the capex asymptote
- 04Open vs. closed mirrors the Linux vs. Windows/macOS outcome: open wins volume, closed still wins big
- 05The commercial model is "rental vs. ownership" of intelligence
- 06Enterprise adoption will customize systems before models
1. Key Themes
Frontier training costs ride a 4x-per-generation escalator, and the catch-up curve is steep
Laskin gives an unusually concrete cost curve for reaching the frontier. He frames capital needs as a function of chip generations: "for every generation of model, there's a 4x multiplier in compute, roughly." He walks the progression from H100 to Blackwell to Vera Rubin: "maybe a couple of years ago was an H100 and maybe it was 100,000 H100s... 100,000 Blackwells, which is roughly a 4x multiplier on an H100. Now the next generation is going to be around 100,000 Vera Rubens, I believe. So that kind of gives you the scale and moves from hundreds of millions to billions to tens of billions." 00:09:27 On catching up, he says "a year ago or maybe 18 months ago, it would have been... hundreds of millions of dollars. I think that now, or let's say even six months ago is probably order billion... Going into next year, I think order 10." 00:09:02 The implication is that each year of delay raises the entry price, which is why open frontier labs are structurally rare.
RL is the new scaling axis, and the plots "never stop going up"
Beam's training split shows where compute is shifting. Laskin: "To train Beam, which is a 500 billion parameter model total 23B active, it was 6,000 GB300s. We ran it, I think, for a few weeks... Reinforcement learning was a little over 10,000 GB300s for four weeks. So actually there were more flops spent on reinforcement learning." 00:12:19 His conviction comes from AlphaGo: "there was a plot in AlphaGo where it just never stopped improving... if you start to figure out how to apply that recipe to stuff that is economically valuable, then... it becomes an economic question of how much money do I want to put in." 00:13:18 He says Reflection's RL curves behave the same way: "If you look at our plots, they just keep going up. And it's just a matter of compute, basically." 00:13:46 The bottleneck moves from "can you train these things" to "where can you get the data and what is economically valuable."
Models are accelerating their own training, which bends the capex asymptote
Laskin argues that capital needs won't compound forever, because efficiency gains are accelerating. "When it was just human researchers doing the work, there's probably a 7x improvement each year... in reinforcement learning. We're probably at a point where it's, you know, 30x like efficiency, depending how good your model is for actually improving itself." 00:10:58 He adds: "the speed is probably four or more times faster than researchers just doing it alone. And that'll probably accelerate. So the amount of intelligence you can extract per training flop is increasing." 00:11:52 Revenue per flop rises as intelligence density rises, so the capex-to-revenue ratio may self-correct.
Open vs. closed mirrors the Linux vs. Windows/macOS outcome: open wins volume, closed still wins big
On token share: "maybe six months ago, it was majority closed, minority open. When you go to any gateway like open router or Vercel, and it's flipped almost exactly from 70/30 closed open to 70/30 open closed now." 00:25:15 His long-run analogy is operating systems: "95% plus of servers, computers in the world run on an open source operating system like Linux. That doesn't mean that the closed stuff is very valuable. Microsoft and Apple, these are very extremely valuable companies." 00:25:37 He also expects economic value, not just tokens, to flow to open because "everything else you need to run it is still expensive" and open models need a different kind of cloud built around them. 00:26:07
The commercial model is "rental vs. ownership" of intelligence
Laskin frames open-model monetization as a housing analogy: "when you're buying a token, you're renting like a piece of a whole stack, which is the harness... the model, the inference software, the cluster management software, the GPUs... I liken it to we start typically off renting our apartments. But then as you grow up, you want to own a house." 00:22:51 Reflection's role is providing "cluster management software... inference software... the harness" plus services, which he describes as a "demand driver" for inference. 00:24:41 The trigger for the switch is spend: "if you're spending, let's say, $100 million plus, which is not uncommon at all on closed models a year, then you start thinking about... a more optimal way to do this." 00:29:42
Enterprise adoption will customize systems before models
Laskin offers a distinct view of how enterprise open-model consumption will unfold. AI natives (Cursor, Cognition, Abridge, Harvey) moved closed → open → fine-tuned quickly because "they set up native products that are data collectors." 00:28:29 Enterprises will differ: "the majority of enterprise token consumption is going to come not from customized models, but from customized systems. Meaning you took an open model. You didn't actually fine-tune it yet. You just customized a system around it, like your agentic harness around it to make it work for some KYC flow." 00:28:03 They also don't go straight to ownership: "I have not seen kind of a straight shot kind of sprint to an ownership market... you have to go through the rental stage." 00:29:10
The win formula: intelligence density × compute × trust
His competitive framework is simple: "what intelligence density are you able to offer times how much compute do you have times how much trust do you have with organizations that they would want to work with you?" 00:31:52 He argues compute scarcity prevents the usual commoditization death spiral from producing a single winner: "everything is getting commoditized. Like everything across the whole stack is so hypercompetitive... the model margins are going to be compressed." 00:32:41 Elad Gil confirms the pressure from the buyer side: "we see a lot of the AI native companies negotiate their closed deals very aggressively because the open ecosystem exists." 00:33:29
Chinese open models as subsidy, Trojan horse, and geopolitical leverage
Laskin sees Chinese open models as a net gift to Western builders, while recognizing the strategic play: "open models are Trojan horses for the infrastructure that they bring with them. So you have an open model but that alone is not very useful. So you buy into the entire software ecosystem and infrastructure ecosystem of a given country." 00:41:48 He warns of full-stack lock-in: "Chinese companies like Huawei are going to be coming in and offering a full stack solution and then locking into that country's supply." 00:42:15 He also notes demand for Western alternatives: "most of Fortune 500, the open model footprint is actually pretty low today because of sort of resistance or aversion to Chinese models." 00:39:17 His historical frame is that exporting open infrastructure "has been actually the American playbook for a long time." 00:43:57
Safety through openness: Linus's Law applied to models
Laskin's central safety argument is historical and empirical. On encryption: "the decision after some catastrophic failures on the closed side where effectively a small handful of engineers designed certain systems that had unintended consequences that they couldn't predict and got easily hacked... strong encryption protocols became open. And that actually gave birth to the whole field of cybersecurity." 00:45:56 On models: "we have a few hundred safety researchers within closed labs that understand how these things work. And despite their best intentions, it is impossible to cover the long tail of vulnerabilities or unintended consequences." 00:00:00 He argues alignment is "a whack-a-mole thing" that is "deeply boring and unsatisfying" and benefits from many contributors: "shouldn't there be 100,000 like researchers and computer scientists looking at and patching up these bugs?" 00:52:44
Research team shape: ~100 people, pods per capability, and a new "forward deployed researcher" role
Laskin describes a stable organizational shape: "There are pods, right, of like five to 10 people that go and target a particular capability. And within coding, there are five to 10 capabilities." 01:08:57 For deployment, "you basically need a pod for any real-world capability. And each enterprise has many, many of them... this new type of forward deployed engineer that is a bit more scientific and evaluations-oriented." 01:09:24 He calls it "more of a training gap" than a demand gap: "we would take as many as possible today." 01:09:51
2. Contrarian Perspectives
Chinese open models are a gift to the West, not just a threat
The prevailing frame is that Chinese open weights are a competitive threat. Laskin argues the opposite on net: "the fact that a great open model ecosystem came from China is actually a massive benefit for the world. Like, there are so many companies in the West... that have been able to build more durable businesses as a result of that." 00:35:25 He highlights the "tall poppy syndrome," where application builders get subsumed by closed providers, and open models are the counterweight. Sarah Guo agrees with the subsidy framing: "I always felt that Chinese open source was basically a subsidy by the Chinese government to U.S. enterprise or to Western enterprise." 00:38:06 This sits uneasily with his simultaneous argument that the West needs its own competitive open ecosystem.
Restricting cyber-offensive capability in closed models makes the world less safe
Most safety orthodoxy says to remove dangerous capabilities from frontier models. Laskin says capability removal cuts both ways: "When you remove cyber offensive capabilities, you also remove cyber defensive capabilities. And as a result, the players who would want to help defending are incapable of doing so." 00:00:00 His empirical evidence: "a very powerful closed model had unintended consequences where it went and hacked into another company. And the only way that company could remediate itself was by using open models to protect itself." 00:47:16 Sarah Guo adds that a recent incident saw a lab revert to open source "because they weren't able to use the existing state-of-the-art labs because of the guardrails." 00:55:26
Frontier AI safety discourse has become "dogmatic," and the closed-lab safety posture rhymes with Leninist centralization
Laskin challenges the safety establishment's framing, and draws a sharp analogy from his Soviet childhood: "in Lenin's Communist Manifesto, there's a statement around, hey, we're building this like socialist state that's going to benefit everyone. But we do need this temporary state of dictatorship where everything is centralized... that is kind of what that thinking reminds me of. It's like, listen, we're going to take care of everyone." 00:56:33 He also targets doomer probabilities: "when leaders of companies say like on the theoretical side, there's a 10% chance that we all die. You know, that doesn't help." 00:50:10
Enterprises will consume open models via customized systems, not fine-tuned weights
The industry narrative is that open models win because they can be fine-tuned. Laskin predicts the opposite for the enterprise: harness and system customization first, fine-tuning later, because enterprises architect around existing products and face friction. "I think there are a lot of frictions there to just do a straight shot to fine-tuning." 01:00:00 The 90%-plus customization stat he cites applies to AI-native customers, not enterprises.
AI research feels like engineering now, not a Wild West of ideas
Laskin concedes a point many researchers dislike: "it feels more of an engineering discipline... like the rocket ship analogy." 00:14:45 The axes of scaling are set (pre-training, synthetic data, RL), and "it doesn't feel like a Wild West the way I felt five years ago," partly because of the "hardware lottery": "even if you have a great new idea, if it's not really a good fit for the hardware, it doesn't make sense." 00:16:11 The implication is that execution, capital, and talent density beat blue-sky exploration for now.
3. Companies Identified
ReflectionAI
Open-weight frontier model lab co-founded by Misha Laskin and Yanis. Released Beam, "Reflection's first open model." 00:01:40 Why mentioned: the subject of the episode, scaled from ~30 to ~300 people in a year, with Beam trained on 6,000 GB300s for pre-training plus 10,000+ GB300s for four weeks of RL. Laskin: "we're now at around 300 people and... assembled all the teams on pre-training, mid-training, reinforcement learning, scaled up our... first models end to end." 00:01:40
Beam (Reflection's model)
A 500B-parameter total, 23B-active open model focused on coding and agentic tasks. Laskin: "Beam tends to be three to four times more efficient than models of the same capability class and much more efficient when it comes to... models that are larger out there. Where the efficiency gains... end up being something like 10x." 00:20:27 He claims "the largest scale that's ever been done in open source" for RL: "I've not seen a 10,000 GB300 for four weeks run documented yet." 00:20:47
Google DeepMind
Laskin and co-founder Yanis's prior employer, where they led RL on the first Gemini models. Why mentioned: origin of the founding thesis and a contrast on infrastructure maturity. "A lot of the tools that we just take for granted or I took for granted as a researcher because they just worked, you have to build." 00:03:30
Gemini
Google's frontier model family. Laskin: "My co-founder, Yanis, and I had been working on the first series of Gemini models. We were working on the reinforcement learning team... we had just shipped the Gemini 1 and 1.5 model." 00:04:04
AlphaGo
DeepMind's Go-playing system, a foundational RL success. Why mentioned: both the reason Laskin entered AI and the template for RL scaling: "the first AlphaGo systems were trained on expert amateur human games. So they did this imitation learning first and then reinforcement learning." 00:22:04
OpenAI (including O1)
Frontier closed lab. Laskin cites O1 as evidence RL moved faster than expected: "reinforcement learning started working faster than we thought it would... O1 came out at that time." 00:05:55 Sarah Guo also notes OpenAI's recent math theorem results 00:59:18 and the early coding bet via Copilot 00:17:21.
Anthropic
Frontier closed lab. Sarah Guo: "Anthropic placed a pretty early bet on code." 00:16:51 Why mentioned: example of focusing on a vertical to generate revenue that funds scaling.
Mistral
European open-model lab. Laskin: "There was a small model from Mistral at the time." 00:05:46 Sarah Guo also cites Mistral as an early commercialization pioneer for open models. 00:22:35
Meta (Llama 2 / Llama 3)
Early Western open-model provider. Laskin: "LAMA 2 had just released and we saw LAMA 3 coming." 00:05:46 Why mentioned: the open base that Reflection originally planned to build on before concluding it had to build its own.
Cursor, Cognition, Abridge, Harvey
AI-native companies with products built around customized open models. Laskin: "a product that's built around something that's big and customized." 00:27:41 Why mentioned: archetype of "digital natives" who are the biggest dedicated-inference customers today.
OpenRouter and Vercel
Model gateways. Why mentioned: the data source for Laskin's claim that token share has flipped to roughly 70/30 open over closed. 00:25:15
Linux
Open-source operating system. Why mentioned: the template for open-vs-closed market structure. "95% plus of servers, computers in the world run on an open source operating system like Linux." 00:25:37
Microsoft and Apple
Closed-OS incumbents. Why mentioned: proof that closed players can still be extremely valuable alongside a dominant open standard. 00:25:37
Dell
Infrastructure provider. Laskin describes enterprises going "with an infrastructure provider like Dell and say, I just want to set up the bare metal, but I want to serve stuff into my enterprise." 00:30:04 Why mentioned: part of the on-prem resurgence driven by compute shortages at hyperscalers.
Huawei
Chinese chip and infrastructure company. Why mentioned: example of a full-stack offering that could lock countries into Chinese supply. "Chinese companies like Huawei are going to be coming in and offering a full stack solution." 00:42:15
SpaceX (compute cluster)
Laskin: "We trained the final model on, you know, the SpaceX cluster that we received in, in July." 00:62:50 Why mentioned: the source of Beam's final training compute, a notable and unusual compute-provider relationship.
Stargate
Large-scale data center project Laskin toured. "The thing that struck me was the amount of cars in the parking lots." 00:60:30 Why mentioned: evidence of job creation from data centers.
NVIDIA (H100, Blackwell, GB300, Vera Rubin)
Chip provider across the compute generations Laskin uses as his cost yardstick. 00:09:54
4. People Identified
Misha Laskin
Co-founder and CEO of ReflectionAI; former Google DeepMind researcher; physics PhD. Why mentioned: the guest. Notable for the RL-first thesis and for building a ~300-person frontier lab in a year. "Everything has been hard... you actually have to get 30 things right." 00:02:35
Yanis (Ioannis Antonoglou, Reflection co-founder)
Co-founder of ReflectionAI; leads research and technology. Laskin: "Yanis, my co-founder, was one of the founding engineers at DeepMind and was a key contributor to all the big RL projects that came out of that lab, including AlphaGo." 00:04:52 Why mentioned: the research leader behind Beam and the original reason Laskin entered AI.
Lee Sedol
Go world champion defeated by AlphaGo. Laskin uses the "Lee Sedol level" as the benchmark when describing how RL agents become efficient: "by the time you got it to Lee Sedol level, it was just very smart in its search." 00:21:10
Linus Torvalds (creator of Linux)
Source of Linus's Law. Laskin: "Linus' law, like creator of Linux, is that with enough eyeballs, all bugs become shallow." 00:46:54 Why mentioned: foundation of the open-safety argument.
Reid Hoffman
Mentioned for the aphorism about "assembling a plane as you're flying it" that Laskin uses to describe building the lab. 00:01:40
Sarah Guo and Elad Gil (hosts)
Hosts of No Priors; investors. Why mentioned: Gil's confirmation that AI natives use open-model alternatives to negotiate aggressively with closed providers 00:33:34; Guo's reframing of Chinese open source as a subsidy to Western enterprise 00:38:06.
Lenin
Cited by Laskin as a historical analogue for centralization-justified-by-future-benefit arguments. 00:56:33
5. Operating Insights
Build your hiring around a "capability pod" structure, then redeploy it as a forward-deployed unit
Laskin's org design is instructive for any applied-AI company: "There are pods, right, of like five to 10 people that go and target a particular capability. And within coding, there are five to 10 capabilities." 01:08:57 The same structure exports to customers: "you basically need a pod for any real-world capability. And each enterprise has many." The skill set (evals, data generation, harness intuition) is identical between core research and deployment, so the pod becomes a reusable unit. 00:09:24
Sell the demand driver (services), not the commodity (inference)
Laskin's direct experience: "It's very hard to go in and just say, we're going to give you inference. I've not heard that work." 00:31:06 His alternative is to "go in, unlock really valuable use cases that drive a lot of compute demand" so services act as the front door to an inference business. 00:24:41 Applied generally: when the product layer is commoditizing, sell the outcome that pulls the commodity.
Use evals as the lever for generating training data in a customer's domain
Laskin explains how to unlock economically valuable verticals without touching customer data: "It's not even that you're going to be training on the customer's data. It's more that you're setting up evaluations. And if you can set up a good evaluation for their tasks, then you can generate synthetic data that approximated and you get good generalization." 00:19:10 The same pattern applies in finance KYC and compliance, cyber defense, and legal. Build the eval first.
Pick the point of customer adoption by spend threshold
Time sales conversations to the rental-to-ownership inflection: customers spending "$100 million plus" annually on closed models, or facing hyperscaler compute scarcity, are the real open-model buyers. 00:29:42 The pitch before that point is premature.
Protect time-to-ship by cutting exploration once you hire senior talent
Laskin's resourcing logic: "Part of why getting great talent in matters a lot is because you get to cut down on exploration and then you can focus on execution of the things that work." 00:08:36 On the Beam timeline, he notes that "a lot of stuff... didn't make it in," and recommends exhausting high-impact known bets before expanding risk. 00:63:12
6. Overlooked Insights
Open-model gateway data shows a six-month flip, and it's the strongest quantitative claim in the episode
Said in passing, but significant: "maybe six months ago, it was majority closed, minority open... it's flipped almost exactly from 70/30 closed open to 70/30 open closed now." 00:25:15 A reversal of that magnitude in half a year, on gateway traffic, suggests the demand curve for open weights is steeper than most narratives acknowledge. Caveat: gateways skew toward developers and AI natives, which Laskin himself notes elsewhere, so it may overstate enterprise adoption.
The scarce input is moving from model weights to "everything else you need to run it," which hints at a new infrastructure layer to invest in
Laskin mentions almost offhandedly that open models "need a different kind of cloud built around them" because GPU accelerators are a more expensive compute substrate than CPUs. 00:26:37 Paired with his point that enterprises going on-prem with providers like Dell ask "who helps them with that layer between the bare metal and actually making them successful?" 00:30:32, this points to an underbuilt middleware category: cluster management, inference serving, and harness tooling for owned intelligence. The model is free-ish; the operating stack is where the unmet demand and the money sit.