Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/LIGHTCONE/Open Models Change The Economics…
POD
// EPISODE
LIGHTCONE

Open Models Change The Economics of AI

DATE September 12, 2026SOURCE LIGHTCONEPARTICIPANTS JEFFREY MORGAN, SPEAKER 02, SPEAKER 03, SPEAKER 04
// KEY TAKEAWAYS6 ITEMS
  1. 01The Enterprise Shift From "Cost Play" to "Control Play"
  2. 02Chinese Open Models Dominate Cloud-Hosted Coding Agent Usage
  3. 03AT&T as a Live Case Study of Enterprise Open Model Adoption
  4. 04Explosive, Measurable Token Growth Driven by Agentic Coding and "Co-Work"
  5. 05The Stack Is Unbundling Into New Layers of Value (Knowledge, Coordination, Execution)
  6. 06The Rise of Ultra-Cheap "Flash"-Class Models as the Next Frontier
In this episode

Key Themes

The Enterprise Shift From "Cost Play" to "Control Play"

Enterprises initially adopt open models for cost savings, but the deeper motivation is customization and control over AI systems. Cost is the entry point, not the destination.

"Cost is by far the largest pain point that open models can jump in and solve. But, you know, every business has a vision of getting better control over AI and customizing it for their business. And that's really their North Star." 00:00:00 - Jeffrey Morgan

Chinese Open Models Dominate Cloud-Hosted Coding Agent Usage

There's a striking bifurcation: in cloud-hosted use cases (especially coding agents), Chinese-origin models have taken overwhelming share, while local/on-device usage remains a balanced mix of US, European, and Chinese models.

"These two graphs are pretty stunning in comparison. Like basically for local models, the U.S. and Chinese models are neck and neck... And for cloud-hosted models, like the U.S. is like recoloring the x-axis. It's like 100% Chinese models. Basically, we need more U.S. labs to make large models." 00:23:11 - one of the hosts

AT&T as a Live Case Study of Enterprise Open Model Adoption

A concrete enterprise data point: AT&T has already moved a large share of production token volume to open models, driven primarily by coding agents.

"I think there was a great article in the information yesterday from AT&T. And it ends up they've already shifted 40 percent of their token consumption to open models." 00:02:44 - Jeffrey Morgan

Explosive, Measurable Token Growth Driven by Agentic Coding and "Co-Work"

Olama's own usage data shows two step-function inflections in per-developer token consumption in 2025 — one from coding agents (Kimi, GLM, Minimax launches) and a second, larger one from OpenClaw-style autonomous agents that extended AI use beyond developers into finance, support, sales, and marketing.

"As a whole through Olamas cloud, we saw 150X since the start of the year... whereas open models... were mostly being served as custom models [before]... seeing out-of-the-box open models being served, that really only took off at the start of this year." 00:04:48 - Jeffrey Morgan

The Stack Is Unbundling Into New Layers of Value (Knowledge, Coordination, Execution)

As raw open-model tokens become abundant and cheap, the scarce/valuable problems move up the stack — into orchestration, memory, sandboxes/execution, and security — creating room for new best-of-breed startups rather than one bundled "God stack."

"The new scarcity, the problems now are what's above the tokens, right? You know, how do you orchestrate an agent from A to B? These are problems that are... just really hard to solve for an individual dev. Like, there's no way they're going to build all those layers." 00:15:27 - Jeffrey Morgan

The Rise of Ultra-Cheap "Flash"-Class Models as the Next Frontier

After closing the intelligence gap with closed frontier models, the new competitive axis for open models is extreme cost/latency efficiency — enabling "unlimited token" usage patterns reminiscent of early ChatGPT.

"We're maybe like less than three months behind between the frontier closed models and the open models. But the next problem to solve is extreme efficiency... this new class of Flash models where they're good enough for 80% of the tasks, they're really fast and they're ultra cheap." 00:29:40 - Jeffrey Morgan

Hybrid Execution: Local + Cloud, Open + Closed, Working Together (Not Winner-Take-All)

Rather than one model type winning, the steady state is a blend — easy tasks run locally or on cheap open models, hard tasks route to frontier closed models, mediated by routers, echoing prior cloud-computing patterns (e.g., DynamoDB + Postgres coexistence).

"I think the steady state is that most of the software's, most of the models are open... for the hardest tasks that's reserved for these frontier labs... And then from there, there's a whole bunch of problems in the middle... maybe it's a combination of open and closed models working together." 00:18:43 - Jeffrey Morgan

Model Release Cadence Is Accelerating, Compressing the Value of Custom Fine-Tuning

The open-source model release cycle has gone from roughly six months to iterating multiple times in a single season, which changes the calculus on whether it's worth fine-tuning your own model versus waiting for the next release.

"This summer, we've already seen three iterations of the DeepSeq Flash model as an example. What used to be more of a six-month cycle... the gap is closing, essentially. And I think that makes it even harder to custom-train models." 00:05:54 - Jeffrey Morgan

Next-Gen Personal/Desktop Hardware (NVIDIA DGX, Apple Silicon) Is Resetting What's Possible Locally

New hardware (DGX Spark/Station, Apple Silicon) is enabling frontier-adjacent model performance on desks, which the team believes will pull the "fastest coding loop" back from cloud to local over time.

"You can run more than a frontier model at high speeds at a price point that isn't very far off what you can buy from a classic workstation computer." 00:24:38 - Jeffrey Morgan

Security/Safety Tooling Gaps Are the Real Adoption Blocker, Not Capability

Open models lack the out-of-the-box safety/guardrail tooling that closed providers bundle, and this — not raw capability or even geopolitics — is cited as the primary blocker for enterprise adoption, creating a startup opportunity.

"Open models don't have all the safety tooling that closed model providers give you out of the box, but that's super important, especially for businesses to adopt." 00:17:30 - Jeffrey Morgan (attributed content, Speaker 03)

Contrarian Perspectives

Being "Just a Layer" on Top of Infrastructure Is Not a Weakness in AI — It's an Advantage

Conventional wisdom from the Platform-as-a-Service era (Heroku, early Docker) held that being a thin layer on top of someone else's infrastructure made you vulnerable to being subsumed. Olama's founders explicitly reject this lesson for the AI era.

"There's this concept that if you're a layer on top of something else, that you're in kind of a vulnerable position as a startup, which is absolutely not true in the AI world. And in fact, going up the stack can sometimes be even better because you're closer to the customer." 00:53:37 - Jeffrey Morgan

Non-Determinism Is a Feature, Not a Bug — Overturning a Core Systems Engineering Principle

Decades of infrastructure/DevOps discipline was built around determinism — systems must run "exactly as designed." LLM-based systems invert this: imperfection and variability are inherent and desirable.

"These LMs are never perfect. And like in the systems world, you want everything to be exactly as it's designed to run. It's tested. It's validated. But LMs by definition are not... It's a feature, not a bug." 00:54:06 - Jeffrey Morgan

The "God Model" Thesis Has Lost — Orchestrated Small Models Are Winning in Practice

Early AGI discourse assumed one giant frontier model would dominate every use case. In practice, the market is bifurcating toward networks of smaller, cheaper, task-specific models coordinated together, which are more trustworthy, repeatable, and affordable.

"When we first started talking about AGI... a lot of AI researchers would come out and say, like, there's just going to be a giant God model and it's going to do everything. But... it hasn't quite worked out that way... if it was going to be a God model versus lots of smaller special purpose or even just simpler models, it's turning out to be the latter so far." 00:31:35 - Jeffrey Morgan

Geopolitical Origin of a Model Matters Less to Many Customers Than Where and How It's Run

Despite intense public discourse on US vs. China model geopolitics, Olama's customer conversations reveal a segmented reality — many enterprises genuinely don't care about model origin as long as execution environment and security are controlled, undercutting the assumption that origin is a universal dealbreaker.

"A lot of the geopolitical angles around this start with, you know, where the model's from. And the more we spend time with customers and users, a lot of it's actually how the models run, where it's run, how it's run... it starts to matter a lot more [than origin]." 00:33:41 - Jeffrey Morgan

Enterprises Already Solve "Manchurian Candidate" Security Fears Via Existing Supply-Chain Discipline

Rather than treating Chinese open-weight models as a novel unmanageable security risk, the guest argues Fortune 500 IT/security teams already have mature practices from decades of open-source dependency management that directly apply.

"What you don't see a lot on some of the press articles is how robust some of the IT and security teams are at the businesses that we know of... This isn't a new problem... It's been a thing for decades." 00:35:30 - Jeffrey Morgan

Companies Identified

Olama — Platform for running open-source AI models locally and in the cloud; used by 9 million developers, 178,000 GitHub stars, 85% of Fortune 500. Mentioned as the guest's own company and central case study of open-model economics and distribution. "Olama is used by 9 million developers, has 178,000 GitHub stars, and is used by 85% of the Fortune 500." 00:00:43 - one of the hosts

AT&T — Large telecom enterprise cited as concrete proof point of enterprise open-model adoption at scale. "They've already shifted 40 percent of their token consumption to open models." 00:02:44 - Jeffrey Morgan

DeepSeek — Chinese AI lab; models cited repeatedly for rapid iteration (three versions of "DeepSeq Flash" in one summer) and driving the highest growth on Olama Cloud, plus a new multimodal LLM launch. "If we go back to the model breakdown on Alamas Cloud, the highest growth area is definitely the DeepSeq model." 00:30:08 - Jeffrey Morgan

Kimi (Moonshot AI) — Chinese model noted for becoming the best model for web development, intensifying head-to-head competition with closed frontier labs. "We all saw with the Kimi model how for web development it became the best model. And that sent this new shockwave across the market." 00:32:57 - Jeffrey Morgan

GLM / Zipu AI — Chinese model family (GLM-5.3 highlighted for impressive cybersecurity capability) seen as approaching frontier-tier performance in specific domains. "We saw the announcement and the release of the GLM-5.3 model and its capabilities from a cybersecurity standpoint, you know, being extremely impressive." 00:06:51 - Jeffrey Morgan

Minimax — Chinese model provider cited among the first wave that made open models viable for coding agents at the start of the year. "We saw Kimmy, the GLM models, Minimax launch. Finally, we had open models that could power coding agents." 00:03:56 - Jeffrey Morgan

OpenClaw — Agent framework/product credited with driving a massive step-change in token consumption in April by opening automation to non-developers. "In April, we saw this incredible growth from OpenClaw, really, which was then not just developers, but the rest of the world could take a hard problem, give it to an open model and let it go complete the task." 00:04:14 - Jeffrey Morgan

Hermes (agent project) — Follow-on agent project after OpenClaw that continued driving growth. "Subsequently the Hermes project, Hermes agent project take off." 00:03:03 - Jeffrey Morgan

NVIDIA — Chipmaker praised for enabling the open model ecosystem via hardware (DGX Spark/Station) and open releases (Nemotron models), described as strategically strengthening its hardware moat by empowering open source rather than competing on software/tokens. "NVIDIA as a company is so interesting because their moat is not like trying to start new software businesses or sell tokens. They seem to be quite interested in just releasing a lot of open source and helping the ecosystem." 00:23:38 - Jeffrey Morgan

Nemotron (NVIDIA) — Model family cited as the "first wave" of a strong US open-model response, and praised for transparency ("you can go and introspect what made this model"). "With the launch of the Nemotron Ultra model, we're seeing kind of the first wave of that. And it's really exciting." 00:23:31 - Jeffrey Morgan

Apple / Apple Silicon / MLX — Hardware and software stack praised for enabling strong local LLM performance on MacBooks/Mac Studio, described as having a mature tech stack for local inference. "There's an incredibly mature tech stack through the MLX project with Apple where they've done some amazing work to run LLMs on the Mac Studio." 00:26:15 - Jeffrey Morgan

Open Router — Cited as a best-of-breed "router"/aggregator product giving developers wide model selection without negotiating with dozens of providers individually. "Open Router obviously is a good example of that from a wide model selection. So the developer doesn't have to sign up for, you know, 100 different providers." 00:56:24 - Jeffrey Morgan

Open Code — Open-source coding harness described as one of the most popular harnesses among Olama users, integrating with any model. "Open code is a great one and one of the most popular harnesses from Olami users." 00:10:21 - Jeffrey Morgan

Codex (harness) — Cited as an example of an open-source harness that Olama integrates with. "The Codex harness is open source." 00:10:21 - Jeffrey Morgan

Hugging Face — Referenced as a repository of specialized/uncensored security-research models, and noted for needing to use open-weight models to detect a hack from a frontier lab. "There's sort of this interesting moment right now where a hugging face had to use open weight models to actually even detect the hack from the Frontier." 00:07:27 - Jeffrey Morgan

Docker — Founders' prior company (Docker Desktop); referenced repeatedly as the source of the team's engineering culture and as a cautionary/instructive tale about monetization timelines and layer-position strategy. "My co-founder and I previously built Docker Desktop while at Docker. So we really got an understanding of what makes a great developer experience." 00:36:51 - Jeffrey Morgan

Sakana AI — Mentioned as a project combining open and closed models with good results in routing/orchestration. "I think we've seen a ton of projects, whether it's from Sakana AI or Open Router that have combined the two and seen really good results." 00:19:20 - Jeffrey Morgan

Cursor — Cited as a famous example of a company that took an off-the-shelf open model (like DeepSeek or Kimi) and fine-tuned it for its own use case at scale. "You'd take an off-the-shelf model like DeepSeq or Kimmy, and you'd fine-tune it for your use case. Like, for example, you know, Cursor had famously done." 00:05:15 - Jeffrey Morgan

Anthropic (Claude/Opus) — Referenced multiple times as a leading closed frontier lab and comparison point (Opus 4.6 vs. Qwen 3.8 for coding; Claude refusing pen-testing tasks; Anthropic platform team's talk on knowledge/coordination/execution). "Yeah, QIN 3.838B is now as good as Opus 4.6 for coding." 00:21:32 - Jeffrey Morgan

Gemma (Google DeepMind) — Cited as a strong choice for local model deployment. "We have the incredible models from the original Lama models, of course, but also the Gemma models from DeepMind. These are great choices for local." 00:21:56 - Jeffrey Morgan

GBT/GPT Luna (OpenAI) — Referenced as an example of extreme price effectiveness driving widespread customer adoption. "Seeing, for example, the GBT Luna model become very, very price effective for customers has been a huge boom." 00:29:40 - Jeffrey Morgan

MongoDB — Cited historically as an example of a developer-first tool that quickly moved into enterprise use, paralleling Olama's own trajectory. "For example, you know, databases, we saw some of this too, where a database that started for devs like MongoDB very quickly also moved to enterprise." 00:46:32 - Jeffrey Morgan

AWS / DynamoDB — Referenced as a historical example of proprietary cloud infrastructure being used alongside open alternatives (Postgres), paralleling the open/closed model blend today. "AWS had DynamoDB, which was kind of their proprietary scale out database. But then a lot of customers use that in conjunction with PostgresDB." 00:19:57 - Jeffrey Morgan

VMware — Cited as source of engineering talent/culture for the Olama team. "A lot of our team, you know, we aren't AI researchers by background. We're from VMware and Docker and from, you know, other networking companies." 00:12:34 - Jeffrey Morgan

Heroku / Google App Engine — Cited as historical examples of bundled cloud platforms that developers ultimately moved away from in favor of best-of-breed products. "If you look back to the original generation cloud products, you have, like, the Herokus of the world. You have Google App Engine. All those things were bundled together." 00:14:35 - Jeffrey Morgan

People Identified

Jeffrey (Jeff) Morgan — Co-founder and CEO of Olama; former co-founder at a company acquired by Docker; built Docker Desktop. Central guest of the episode, described as having deep insight into state-of-the-art AI usage due to Olama's position in the token flow. "Jeff knows a lot about the state of the art of AI, what models score highest on benchmarks, and what developers actually download and keep using." 00:00:43 - one of the hosts

Michael (Olama co-founder) — Jeff Morgan's co-founder, former college roommate at University of Waterloo, and co-founder of Jeff's prior company (acquired by Docker). Cited for steering the pivot to Olama and sticking with the founding vision through years of pre-product-market-fit searching. "Michael, my co-founder was the co-founder of my first company. He was my college roommate at University of Waterloo." 00:49:51 - Jeffrey Morgan

Peter (Benchmark partner) — Series A investor in Olama and previously Series A investor in Docker; cited as backing the founders based on trust in the people rather than the specific point-in-time product. "We had known Peter, the partner at Benchmark, from our previous lives building at Docker because he was the Series A investor in Docker... a lot of it was weighted on, you know, the people and also why we exist." 00:39:10 - Jeffrey Morgan

Jared (YC partner) — Referenced as the YC group partner who worked with Olama's founders through their pivots and remained a valued advisor. "Just talking to Jared and like five other groups of founders every week really helped you feel less lonely." 00:49:51 - Jeffrey Morgan

Jensen (Huang, NVIDIA CEO) — Referenced for framing the "five-layer cake" model of AI infrastructure (apps, model, infrastructure/inference, chips, energy) that Olama is trying to help orchestrate for open models. "The classic five-layer cake, right, that Jensen mentioned, which is like the apps... the model... the infrastructure and inference... the chips, and then you've got the energy." 00:13:04 - Jeffrey Morgan

Elon (Musk) — Mentioned as among the first recipients of the NVIDIA DGX Spark hardware alongside Olama. "We were one of the first people, along with Elon and a few others, to receive the DDX Spark as well." 00:25:12 - Jeffrey Morgan

Operating Insights

Build a Repeatable "Day-Zero Launch" Playbook for Fast-Moving Ecosystems

Olama developed a systematized three-part playbook (harness readiness, hardware/inference provider optimization, and model packaging) to execute successful launches within a compressed, often 24-hour window before a new model drops — treating what looks chaotic as an operational discipline.

"A lot of this stuff comes together in the last 24 hours before the model gets released. And so it's generally a fire drill." 00:11:17 - Jeffrey Morgan

Give Yourself a Hard Deadline to Break Analysis Paralysis

After two years of overthinking their product direction, the founders gave themselves two weeks to ship the first version of Olama — a forcing function that, combined with the Llama 2 launch, immediately produced more traction than years of prior work.

"Let's give ourselves two weeks to launch the first version of Ollama... in two weeks, all of that happened, going from idea to shipping it to getting to more users than we had ever had with our previous stuff. And before that was two years of just, frankly, overthinking." 00:43:35 - Jeffrey Morgan

Don't Treat an Open-Source User Base as an Anonymous "Blob" — Stay in Direct Contact

A specific self-identified mistake: having a wildly popular open-source project can create a false sense of customer understanding. The team explicitly flags that failing to maintain direct relationships with users/customers (versus just watching download/star metrics) was a gap they are now actively correcting.

"One of the biggest risks of having an open source project that takes off is you consider your user base and customer base, your customer, just a blob on the internet. Which is a really risky way to think about customers." 00:49:22 - Jeffrey Morgan

Match Team Composition to the Actual Problem You're Solving, Not the One You Started With

The founders realized their team's DNA (developer-tools/security backgrounds) was mismatched to their original enterprise-security pivot idea, and explicitly restructured their thinking around what their team was naturally suited to build.

"We kind of tried to really introspect our team, which I wish we had done sooner because security is a very different team and sale than developer tools." 00:43:06 - Jeffrey Morgan

Recognize Which "Systems Engineering" Instincts From Prior Eras No Longer Apply in AI

Explicit warning for operators bringing infrastructure/DevOps backgrounds into AI: old assumptions (determinism required, being a thin layer is risky, headcount needs mirror pre-AI ops complexity) actively need to be unlearned rather than reused.

"There are a lot of lessons we learned in the previous generation of DevOps and infrastructure that aren't valid anymore in the AI world." 00:53:19 - Jeffrey Morgan

Overlooked Insights

The Token-Cost Collapse Enables an Entirely New Category: Cheap Model Orchestration as a Business Model

Buried in a discussion about "Flash" models is a much bigger structural point: once per-token costs fall low enough, the economic strategy shifts from "pick the best single model" to "chain many cheap models together" to solve problems previously requiring one expensive model — this is a new, underappreciated compute/product design pattern that could define a wave of startups.

"By having these cheaper models, not only are they more accessible, they can run faster and you can access them in higher volume, but you can start to chain them together and build new problems that are solved by orchestration on top. And that's a really exciting area for new startups." 00:31:09 - Jeffrey Morgan

Local Inference Is About to Make a Comeback via Hardware, Reversing the "Cloud Migration" Narrative Everyone Assumes Is One-Directional

Most of the conversation (and the industry narrative) assumes agentic AI is inevitably moving to the cloud. But buried in the hardware discussion is a contrarian trajectory: the team believes the fastest coding-agent experience will eventually pull back to local devices once desktop hardware (DGX Spark/Station-class) is powerful enough — essentially predicting a reversal of the local-to-cloud migration that almost no one else in the ecosystem is discussing.

"We started local. Clearly, the coding agent demand is in the cloud. But that's going to come back locally in our minds because the hardware will catch up when you have a GB300 on your desk and you want the fastest coding loop... that's been a journey of starting local, going to the cloud. And then we think that'll come back local." 00:26:40 - Jeffrey Morgan