Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THE A16Z SHOW/World Models, Robotics, and the…
POD
// EPISODE
THE A16Z SHOW

World Models, Robotics, and the Future of 3D AI

DATE September 13, 2026SOURCE THE A16Z SHOWPARTICIPANTS JUSTIN JOHNSON, SOFIA PUCCINI, THEO JAFFEE
In this episode

Note: The transcript labels are inconsistent — content indicates Justin Johnson (World Labs co-founder) is the guest being interviewed, while Theo Jaffee and Sofia Puccini are the interviewing hosts. Quotes are attributed based on who is describing World Labs/Atlas in the first person versus who is asking questions.

1. Key Themes

World Models as a New Horizontal Platform, Analogous to LLMs

The central thesis of World Labs is that just as language models became a general-purpose engine for text, "world models" — grounded in visual and physical understanding — represent an entirely new horizontal category of AI infrastructure applicable across many verticals. "Our thesis is that there exists another category of model called world models that should be based in visual understanding, should be based in physical understanding, that can be used to generate, simulate, reconstruct worlds. And if we can build these with the right generality, they should be applicable to tons of different industries, right? From entertainment to VR to construction to robotics." [00:00:00]

Atlas's Three Core Capabilities: Generation, Reconstruction, Simulation

Atlas isn't a single-purpose tool — it's built to generate entirely new worlds from text/image prompts, reconstruct real spaces from as few as one photo, and simulate robotic behavior within those reconstructed spaces. "It does basically three different kinds of things. Generation, reconstruction, and simulation... I can take just a few photos of a space, as few as one up to 100 or more. And the more photos you give it, the more you can reconstruct that existing real world space accurately." [00:01:37]

Solving the "Slot Machine" Problem in Long-Horizon Generation

A key technical differentiator is spatial grounding of reference images in 3D space combined with pixel-precise camera control, which prevents the drift/hallucination that plagues typical video generation models at longer time horizons. "So the idea here is that even as you go to really long generations, it's not a slot machine. You're generating and directing and really controlling what the model outputs." [00:04:43]

Decoupling from Gaussian Splats: The Architectural Leap from Marble to Atlas

World Labs' first product, Marble, bottlenecked all outputs (video, 3D scenes, meshes) through Gaussian splat representations. Atlas represents a full architectural rebuild that lets 2D and 3D generation branch earlier, avoiding unnecessary computation and quality loss. "Marble really put Gaussian splats really at the center of many things that it did... whenever we output videos or images from Marble, you were always generating a Gaussian scene, then rendering those Gaussians back to images or videos... Atlas... was built fully from the ground up to both rethink the architecture, to make it more unified, more easy to scale." [00:06:59]

Real-to-Sim-to-Real as a Fast Path for Robotics Adaptation

Rather than only relying on massive in-context learning or teleoperation datasets, World Labs sees a path where a handful of casual phone photos/videos can rapidly build a custom simulation environment to fine-tune a general robotics foundation model for a specific space. "You could come in here and take just a couple of casual videos with your phone or a couple of images with your phone, use Atlas to reconstruct the space and now stage all kinds of robotics interactions... And now you've got a robot maybe in the span of a couple minutes even that could come in and is perfectly adapted to this space." [00:12:36]

3D as an Input Control Surface, Not Just an Output Format

A subtle but important theme: 3D isn't only valuable as a final deliverable (game engine assets, meshes) — it's equally valuable as an input/control interface, even when the final output is flat 2D video, because it gives directors precise control that text prompts cannot. "I don't just want to write in text like pan in, pan out truck left... I want to be able to grab the camera and make it like zoom exactly as I like... Atlas has camera control as a native 3D... input type." [00:20:04]

Creative Control as a Design Philosophy

World Labs is explicitly positioning itself against "generative slot machine" products, emphasizing directorial control as a differentiator from typical diffusion/video-gen tools. "We never wanted to be building these generative slot machines where you just like pull the thing and hope you get out a good generation... We want to build these tools that have really deep control... you feel like you're a director in the thing, guiding this model." [00:17:36]

2. Contrarian Perspectives

OpenAI's Retreat from World Models as "Innovator's Dilemma"

Rather than assuming OpenAI abandoned Sora's "world simulator" research direction because it wasn't valuable, Justin Johnson frames it as a strategic blind spot from incumbents overinvested in their winning platform (LLMs). "I think there is maybe a sense of a bit of innovators dilemma, right? Like you're sitting on LLMs and you've got this as a really powerful thing and you're leading the pack there. It doesn't necessarily make sense to try to invest in that next thing. But... we think that there is another category of models out here. There's a ton of opportunities out here and we want to be the ones to go and grab them." [00:14:03]

Explicit 3D Representations (Gaussians/Meshes) Won't Be Fully Obsoleted by Pure Pixel Streaming

Despite the industry excitement around "just stream pixels directly from the model," Johnson argues explicit 3D representations remain structurally necessary for real-time, embedded, and mobile/VR use cases for the foreseeable future, and adoption of AI-native workflows will lag far behind what's technically possible because existing industries (gaming, VFX, architecture) have entrenched pipelines. "I think we all believe in the future of technology... But the reality is that it takes a lot longer for these things to permeate than we expect... instead of trying to switch them over all at once to a fully AI-native thing, if you can get with AI into their workflows and meet them where they are, then that's another great use of splats." [00:08:47]

Fully Hand-Authored AAA Games May Be Ending — But Human Creative Direction Won't Disappear

Justin Johnson floats the provocative idea that a game like GTA VI could be "the last great piece of human written software" before generative world models take over asset creation — while Sofia Puccini pushes back that games remain fundamentally an art form requiring human direction, drawing a parallel to how past creative technology shifts (hand-painted Disney cells, synced audio/video, CG/Pixar) each transformed production without eliminating the human director role. "I do wonder if GTA six is just the last great piece of like human written software before like all future video games are kind of generated on the fly." [00:16:07] / "I think there will always be like human in the loop with, cause video games are an art." [00:16:21]

Black Hole / Quantum Physics Simulation Exposes a Real Architectural Limit

Rather than claiming Atlas is a universal physics engine, Johnson candidly states the model likely fails outside roughly Newtonian regimes and would require rethinking both architecture and training data for non-Newtonian or quantum-scale physics — a notable admission of current limitation from the person building the product. "Anything that's like roughly Newtonian... that doesn't work at is probably just a quick SFT away. If you want something that's sort of non-Newtonian, right? Like black hole physics or... quantum mechanics realm, I think then you'd have to rethink a bit the architecture and the data side." [00:20:48]

3. Companies Identified

World Labs — Startup building spatial intelligence products and world models, founded by Justin Johnson. Mentioned as the maker of Marble (first product) and Atlas (newly announced multimodal world model). Described as pursuing generation, reconstruction, and simulation as a unified horizontal AI platform. "They announced Atlas, the world's first multimodal world model that generates image and video frames with pixel perfect camera control and reconstructs them in 3D." [00:01:05]

OpenAI — Referenced for its 2023/2024 Sora research paper, "video generation models as world simulators," and its pivot of Sora toward a consumer TikTok-clone product rather than continuing world-model research. Mentioned as a cautionary example of innovator's dilemma. "OpenAI's paper on it was called like, video generation models as world simulators, which seems like very similar to the vision of world labs. And then... they did Sora 2 basically as like a TikTok clone." [00:13:21]

4. People Identified

Justin Johnson — Co-founder of World Labs, guest on the episode, driving force behind Atlas and Marble. Identified as the technical/strategic architect of the company's world-model thesis and product roadmap. "We are live with Justin Johnson, who is a co-founder of World Labs, which is a startup that builds spatial intelligence products." [00:01:05]

Tice — Referenced as a friend at OpenAI who built a viral project reconstructing San Francisco as a giant Gaussian splat "GTA-style" walkable/drivable city simulation, mentioned as an early proof point for the explicit-3D approach to world generation. "Our friend Tice from OpenAI did this project that went kind of viral that was like GTA, but it's San Francisco... a giant Gaussian splat of the whole city." [00:05:08]

5. Operating Insights

Ship the Base Model Separately From the Product

World Labs deliberately separated the model announcement (Atlas) from product launch, signaling a platform-first go-to-market strategy rather than bundling capability with a single application. "This is not a product launch. This is a model announcement. So the products are coming later. This is the base model. It's going to be used to power our future products from World Labs." [00:01:37]

Meet Legacy Workflows Where They Are Instead of Forcing Migration

Rather than pushing customers to abandon existing 3D pipelines (gaming, VFX, architecture) for a fully AI-native approach, World Labs' strategy is to integrate AI capabilities into existing tools and formats (meshes, splats) that plug into current production workflows — reducing adoption friction. "Instead of trying to switch them over all at once to a fully AI-native thing, if you can get with AI into their workflows and meet them where they are, then that's another great use of splats and other explicit 3D representations." [00:09:14]

Build One Unified Architecture Rather Than Bolting On Modalities

The lesson from Marble to Atlas: rather than iterating incrementally on a bottlenecked architecture (everything routed through Gaussian splats), World Labs did a full architectural rebuild to let 2D and 3D outputs branch natively and early — a costly but structurally necessary decision to unlock scale and quality. "Atlas... was built fully from the ground up to both rethink the architecture, to make it more unified, more easy to scale, and to have it have that sort of that branch between 2D and 3D happen earlier." [00:07:29]

Design for Both Human and Agent Consumers From Day One

World Labs is explicitly building Atlas to be API-consumable by coding agents as well as humans, anticipating that future creative software will be assembled by agents stitching together model capabilities rather than only humans using GUIs. "If we can provide a model that has all these general capabilities... and then we can have an agent, a coding agent go and stitch all these capabilities together. Then we can let people really quickly iterate on really cool experiences." [00:15:05]

6. Overlooked Insights

The Recursive Chip-Design Flywheel

Almost as a throwaway joke, Justin Johnson mentions using a 3D world model to help design the semiconductor chips that will train the next generation of world models — a potentially significant compounding-returns insight (AI-designed chips accelerating AI training) that wasn't explored further but hints at a genuine long-term infrastructure moat/strategy. "One of our dreams we joked when starting the company is you want to build these 3D world models and how cool would it be to use a 3D world model to design a chip that we then use to train the next generation world model." [00:21:23]

Bullet-Time Reconstruction From Just Three iPhones Is a Bigger Deal Than It Sounds

The offhand mention that three synchronized iPhone cameras can produce Matrix-style "bullet time" freeze-frame flythroughs — and that this worked far better than the team expected — is understated but signals that Atlas's few-shot 3D reconstruction is already production-viable for consumer-grade hardware, with major implications for indie VFX, social content, and sports/event capture without expensive multi-camera rigs. "We can do this thing... you take like a couple of regular iPhones and stick them on tripods, just like three iPhones... use Atlas to do these bullet time reframing shots... this worked better than any of us were expecting. It's completely insane." [00:18:22]