Dwarkesh Patel
Independent podcast host who has become the go-to media platform for top AI researchers and executives.
“If a true human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That's 15x today's spot prices.”
“Math, of course, is the exception. And I feel like this is actually an important driver of progress in this domain and also in coding. It's not just verifiability. It has to be grindable.”
Source→“I think Karpathy said this when he came on my podcast, is that for humans, many billions of years of evolution had to go into basically pre-training us. And so we're being unfair when we're comparing how little data we see within our lifetimes to what these cold-started LLMs, who are just starting off with a totally random initialization, have to learn from.”
Source→“Dwarkesh argues that the dominant driver of AI improvement is data quality and quantity — not architectural cleverness, hyperparameter tuning, or training tricks. The speed at which open-source models catch up to frontier models is itself evidence: distillable data flows through public APIs, but proprietary training recipes do not.”
Source→“Is Nigeria own a lot of SK Hynix and like Anthropic? I'm guessing not right. It's not enough for them to just own the S and P 500.”
Source→“There aren't that many institutions that have thought as hard as Jane Street about how to turn smart people into some of the most competent researchers and engineers in the world.”
Source→“Omni is the next step towards more accurate world models. Because in order to predict the next frame of a video, you have to have a deep understanding of physics and spatial dynamics.”
Source→“After Cursor injects these hint tokens they run another forward pass — the trajectory itself doesn't change but the hint causes the model to assign lower probability to the error tokens. Cursor then trains the original model to match those probabilities, basically teaching it to downweight these specific mistakes.”
Source→“Your colleague, Chad Jones, has a very interesting result about how the share of the economy that is going towards paying for computing, basically paying for the transistors, has been decreasing.”
Source→“Crusoe was one of the first clouds to adopt Envy Sentinel, NVIDIA's own GPU monitoring and self-healing software for enhanced GPU uptime utilization and reliability... Crusoe can swap in a healthy node in less than 10 minutes.”
Source→“A 10-layer neural network pass... 10 steps of reasoning... is able to amortize and approximate to a very high fidelity a nearly intractable search problem.”
Source→“Jensen Huang: 'I didn't deeply internalize how difficult it would be to build a foundation AI lab like OpenAI and Anthropic… I'm not going to make that same mistake again.'”
“There's a talk by Ilya where he says today we know not to do pipeline parallelism.”
Source→“Last week Horace was kind enough to give me and my friends a great lecture on large scale pre-training systems and there were some concepts that I wanted to animate for a write-up on my blog.”
Source→“There's a talk by Ilya where he says today we know not to do pipeline parallelism.”
Source→“All of a sudden, Dwarkesh Patel's podcast has become must-listen content among top AI researchers and executives.”
AI-extracted from podcast / newsletter / paper summaries. May contain errors.