Daniel Litt: The Mathematician's Guide to AI
- 01AI Math Results Are Impressive but Narrow
- 02The Erdős Unit Distance Problem as a Benchmark for Genuine AI Creativity
- 03AI Reasoning in Math Is Surprisingly Human-Like, Not Alien
- 04The Fact That Models Scale on Natural Language (Not Lean) Has Major Implications for Generalization
- 05Incentive Structures in Academic Math Are Breaking Down, Creating a "Slot Machine" Dynamic
- 06Mathematical Diversity Is Existentially Important
A16Z Podcast | Participants: Daniel Litt (Professor of Mathematics, University of Toronto), Leisha Li (A16Z Infra Partner)
1. Key Themes
AI Math Results Are Impressive but Narrow — "Applying Known Techniques" Is Not the Same as Understanding
Daniel Litt draws a sharp distinction between what current models are actually doing versus what mathematicians do. The models are excellent at pulling together technical ideas from across many papers and completing long computations, but they lack the big-picture intuition that drives real mathematical discovery.
"They're more like they're very, very good at applying some known techniques, which is to be clear, that's like a very powerful thing to do... but it's like some kind of fairly narrow band of what mathematicians care about so far." 00:10:47
The Erdős Unit Distance Problem as a Benchmark for Genuine AI Creativity
Litt identifies the AI-produced solution to the Erdős unit distance problem (announced mid-May) as the best current example of a fully autonomous AI result that went beyond mere technique application — bringing in ideas from another area and spawning further human discoveries.
"The Erdős unit distance problem... it brought in some techniques from another area. I think those techniques were like not especially like deep or new. They were sort of classical ideas from the 60s, but they were new to this area of studying, you know, point configurations in the plane. And so that was pretty cool." 00:03:14
"A bunch of mathematicians took those ideas and used them to find counter examples to a bunch of other interesting open questions. So, for example, like the sum product conjecture over the real numbers." 00:04:39
AI Reasoning in Math Is Surprisingly Human-Like, Not Alien
Against the narrative of AI doing "inhuman" symbol-pushing, Litt notes that the chain-of-thought released by OpenAI for its math results looks very much like how a human mathematician actually thinks.
"The argument was very human... It was very recognizable. It was like, if I tried to imagine like my chain of thought trying to solve a problem, like it might look kind of like that." 00:06:38
The Fact That Models Scale on Natural Language (Not Lean) Has Major Implications for Generalization
Leisha Li surfaces a non-obvious point: the scaling success in math has come from informal natural language reasoning, not from verified formal proof systems. Litt draws a direct inference that this means the techniques will generalize far beyond math.
"I really like your point, by the way, that they're mostly doing natural language reasoning rather than Lean. Like I think that that suggests to me that like these — you hear a lot of people say like math is a verifiable domain, like that explains the progress or whatever — like my sense is that because they're primarily scaling informal reasoning, like probably the techniques are going to generalize to other domains pretty well." 00:09:07
Incentive Structures in Academic Math Are Breaking Down, Creating a "Slot Machine" Dynamic
Litt describes a concrete and alarming failure mode: postdocs are now running AI systems to generate papers at scale, producing correct but intellectually hollow results, with multiple groups outputting identical proofs of the same theorem within days of each other.
"Here's an experiment you can do. You can take Codex. You can say, go online and find five recent conjectures in algebraic geometry and prove them... I've run this experiment. And with some back and forth, I was able to, you know, in an hour get, like, three, you know, quite bad papers. But correct papers." 00:36:17
"We've seen examples where, like, three or four or five papers with the exact same proof of the exact same theorem have come out within a couple days of each other. Which is clearly, you know, some situation where someone's playing the slot machine." 00:37:15
Mathematical Diversity Is Existentially Important — and AI Monoculture Is a Real Risk
Litt warns that if mathematical exploration becomes subordinated to what AI models naturally pursue, you lose the cognitive diversity that has historically driven breakthroughs. The models appear to be mode-collapsing onto similar reasoning paths.
"A lot of progress in mathematics comes from, like, letting, you know, a thousand different flowers bloom. And people pursue their own curiosity... it's really high dimensional space of mathematics. And then, like, opportunistically, you suddenly get some applications or, like, answers to old questions we found. And it's, like, not clear to me that if, you know, we kind of subordinate mathematical exploration to what the models want to pursue... if what you're getting is, like, one mathematician duplicated a thousand times." 00:00:16
AI Cannot Yet Do the "Fuzzy" Structural Checking That Experts Actually Use to Catch Errors
Litt reveals that expert mathematicians check papers not by reading line-by-line but by stress-testing global structure — asking whether an argument implies something known to be false, or checking special cases. Current models cannot do this.
"Well, he was proving something that was too strong to be true... me and a bunch of other experts, like, immediately, like, realized it was wrong... And so far the models seem not to have been able to do this. Like, the specific error is quite subtle and — but they're also not able to do this kind of overall kind of big picture tech checking, which is kind of, like, how people in practice check papers." 00:57:53
The Bimodal Learning Split Is Already Happening in Classrooms
Litt notes that educators are already observing a bifurcation: students who leverage AI to deepen their learning versus those who use it to skip thinking entirely and then fail everything else.
"I've heard people start talking about a bimodal distribution in their classes where there's some people who are really, like, figuring out how to take advantage of new tools and other people who are just, like, letting them do their homework and then bombing everything else." 00:45:38
2. Contrarian Perspectives
Even Fully Superhuman AI Still Requires Human Mathematicians — Not for Capability, but for Direction
Most people assume that once AI surpasses humans at math, human mathematicians become obsolete. Litt argues the opposite: even in a world of robustly superhuman models, you need humans to decide what mathematics gets pursued, because there is no guarantee that uninstructed models will pursue broad, diverse, fundamental research.
"Let's suppose the models become, like, really robustly superhuman... I claim, like, still, actually, we still want human mathematicians... if we kind of instrumentalize what we want them to do, like, we want to say, like, oh, you know, make our life better or whatever, it might not be the case, like, that, like, what they decide to do is pursue a wide variety of interesting research. Right? Like, they might just try to take the direct path." 00:40:38
Mathematical "Ugliness" Is a Failure Mode, Not a Guide — Win by Any Means Necessary
Against the common mathematician's instinct to abandon a proof that feels ugly, Litt argues this is a limiting bias, and that the orientation should be scientific (what is fundamental?) rather than aesthetic.
"Win by any means necessary, in my opinion. Like, I like to think of what I'm doing, like doing kind of physics except with concepts. So, you know, instead of beauty, I try to think about maybe like what is kind of fundamental. What's going to open up further understanding most." 00:21:43
"One failure mode I see among young mathematicians sometimes is like, you have something and you kind of think you know how to prove it. And then like the proof feels really ugly. And you decide. But like, okay, I mean, what if you're wrong and it's not ugly? Like, why limit yourself?" 00:21:16
AI Math Proofs Being Short Is Not a Sign of Elegance — It's a Sign of Capability Limits
The widely celebrated observation that AI-produced math proofs are refreshingly short and clever is reframed by Litt as simply a reflection of what the models can currently verify, not genuine mathematical taste.
"My sense is that the reason they're not producing long, complicated proofs is that they cannot. Just, like, the ability to check correctness is not yet there... I think the problem with producing a very long thing is they might not know they're wrong." 00:52:36
The Goal of Mathematics Is Not to Produce Papers — Understanding That Lives Only in Model Weights Is "Pretty Unsatisfying"
Against any framing where AI-generated proofs represent mathematical progress in themselves, Litt insists the entire point is producing human understanding — and that knowledge residing only in model weights fails that test.
"The goal of mathematics is not to produce mathematics papers, it's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me that's like pretty unsatisfying." 00:00:00
3. Companies Identified
OpenAI Leading AI lab; maker of ChatGPT and the o-series reasoning models. Mentioned for producing the leading math reasoning results, releasing chain-of-thought traces that look recognizably human, and releasing a list of 10 math problems formalized in Lean as a benchmark. Daniel Litt uses ChatGPT (specifically GPT-4.5/o-series "Sol" and "5.6 Pro") as his primary daily research tool.
"OpenAI released some chain of thought. It was very recognizable. It was like, if I tried to imagine like my chain of thought trying to solve a problem, like it might look kind of like that." 00:06:45
"With this recent list of 10 problems, released by OpenAI, those were all formalized in Lean. Which is, of course, very good evidence that they're true." 00:53:36
Anthropic AI lab; maker of Claude (specifically Claude Opus 4/Fable). Mentioned as neck-and-neck with OpenAI on math benchmarks after catching up around the Opus 4 generation. Claude was used by Levent-Poge to find an elliptic curve of rank 30.
"Around Opus 4.5 or Opus 4.6, they like more or less caught up... some experimentations suggest to me they're pretty neck and neck." 00:12:31
"It's due to, I guess, Claude Cable prompted by Levent-Poge and a collaborator whose name I unfortunately forget." 00:49:24
Google DeepMind / Gemini AI lab. Gemini DeepThink (described as formerly frontier-level) was used by Daniel Litt in producing his own paper where AI proved several lemmas.
"I like worked with, actually this was with Gemini DeepThink, which at the time was also on the frontier, no longer." 00:29:42
Cursor AI coding tool. Cited by Leisha Li for a published study attempting to reproduce SQLite in Rust as a long-horizon coding benchmark, illustrating the parallels between AI limitations in code and in math.
"Cursor put out about testing their long horizon harness. In this case, they were trying to reproduce SQLite in Rust, and it was just so telling, like, how far we are from — you know, and that seems like a very comparable task." 00:54:46
4. People Identified
Daniel Litt Professor of Mathematics, University of Toronto. Algebraic geometer and problem-solver. Publicly vocal about his evolving views on AI and mathematics. Has co-authored a paper where Gemini DeepThink proved lemmas. Has run experiments with Codex generating correct algebraic geometry papers in an hour.
"There was some lemma I wanted to prove and the models couldn't do it. So none of the frontier models could do it. And so I like worked out a ton of examples on my own and I realized, oh, well, maybe like here's some reason why it could be true... once I had that statement, the models were able to very quickly prove that sort of better statement." 00:00:36
Levent-Poge (Levent Alpoge) Mathematician. Mentioned for being behind several of the notable recent semi-autonomous AI math results, including an elliptic curve of rank 30 produced with Claude. Noted as prolific on Twitter and releasing PDF write-ups of results.
"A lot of these nice recent results have come from Levent with unclear amounts of autonomy. So my sense is that some of them are semi-autonomous rather than fully autonomous." 00:49:44
Noam Elkies Mathematician at Harvard, described as perhaps the most famous person who pursues elliptic curve rank records. Mentioned as context for the significance of the rank-30 elliptic curve result.
"There are a few very, very talented mathematicians who like these kinds of questions. So Noam Elkies being maybe the most famous example. So, Elkies and Klagsbrun are the ones who kind of have been pushing this record for a while. And they recently found a rank 29 example." 00:50:27
Zev Klagsbrun Mathematician; collaborator of Noam Elkies. Mentioned as one of the people who held the previous elliptic curve rank record (rank 29) before the AI-assisted rank-30 result.
"Elkies and Klagsbrun are the ones who kind of have been pushing this record for a while. And they recently found a rank 29 example." 00:50:27
Ava Howell Mentioned as a possible third collaborator on the rank-30 elliptic curve result alongside Levent-Poge and Claude.
"Levent and Claude and I think Ava Howell, maybe, is the third collaborator." 00:51:16
Mark Zelke and Metab Swani Researchers at OpenAI. Mentioned for an appearance on a recent episode with Leisha Li in which they noted that the AI-produced proofs have been notably short — a point Litt subsequently reframes as a capability constraint rather than elegance.
"I had Mark Zelke and Metab Swani on from OpenAI recently. And they were saying how, you know, what's kind of been the most charming or delightful is just that the proofs have been relatively short." 00:51:45
Carl Simpson Mathematician. Mentioned as one of the architects of the deep analogy between cohomology of algebraic varieties and representations of fundamental groups — a 30–40 year research program that exemplifies theory-building as mathematical activity.
"That analogy has like led to a huge amount of developments for the last like 30 or 40 years, by people like Carl Simpson and Jokuro Machizuki and others." 00:16:43
Shinichi Mochizuki (Jokuro Machizuki) Mathematician. Mentioned alongside Carl Simpson as a key figure in the cohomology/fundamental group analogy program.
"Trying to realize that dream has led to a lot of beautiful mathematics for the last like 30 or 40 years, by people like Carl Simpson and Jokuro Machizuki and others." 00:16:14
Bryan Birch and Peter Swinnerton-Dyer Mathematicians. Mentioned as the originators of the Birch and Swinnerton-Dyer Conjecture (a Millennium Prize Problem), which Litt cites as a landmark example of data-driven conjecture-making — one of the first ever computer-aided bits of mathematics.
"The Birch-Swinnerton-Dyer conjecture... was discovered by like, it was like the first big data conjecture. So Birch-Swinnerton-Dyer had like found all these statistics on elliptic curves in the 60s. Like it was one of the first ever computer-aided bits of mathematics." 00:17:47
Balazs Szegedy Mathematician. Mentioned by Leisha Li as a Hungarian mathematician she has collaborated with, in the context of Budapest's tradition of teaching group theory in primary school.
"I used to collaborate a bunch with some Hungarian mathematicians, Balazs Szegedy among them. And I heard that in Budapest, they would just teach group theory when you're in primary school." 00:43:45
William Thurston Legendary mathematician. Mentioned by Leisha Li for his observation that mathematical progress is in part a sociological phenomenon — communities of mathematicians converging on interesting structures.
"I think like Thurston made some comment or point about this in the 70s, that it is a sociological phenomenon of more and more mathematicians start examining something, then you'll maybe, you know, converge on some interesting structures." 00:20:33
5. Operating Insights
Use AI to Unlock Projects You Were Procrastinating on Due to Skill Gaps, Not Just to Accelerate Existing Work
Litt's most practically useful AI adoption insight is not that AI sped up his existing research, but that it unlocked an entirely new category of projects — specifically computational/coding work he had been deferring for years due to his own weak coding skills. This is a reframe for any operator or knowledge worker: identify the skill-gap bottlenecks that have been creating years of procrastination and attack those first.
"I suck at coding. And so now I know, you know, my good friend who's really good at coding. And now I have all these coding projects. Cause like suddenly I — if I had a question where coding would have been really useful, I would have procrastinated on it for six months... The models are like very good for kind of massively parallel things. Like if you want to find an example of something, you can just ask it to find, you know, work through a thousand examples in parallel." 00:19:06
Use Inability to Grind as a Discovery Signal — Don't Let AI Remove the Productive Friction
Litt illustrates a concrete case where his own resistance to an ugly brute-force calculation forced him to search for a better argument, which led to a genuinely superior conceptual proof. The operating lesson: don't use AI to instantly resolve every difficult or unpleasant step. The friction of not being able to do something easily is often the signal that a better approach exists.
"I realized like this horrible grind proof would work. And I like could not bring myself to do it. And so I looked for another argument... You can now put it into ChatGPT 5.6 Pro. And it will output like the worst proof you've ever seen, like 10 pages of like just brutal calculation with no insight whatsoever. And so, okay, this would have been a perfectly fine proof, but it would not have led to this discovery of like, I think, a kind of beautiful conceptual explanation." 00:31:06
The Correct AI Workflow: Use It to Understand, Not to Produce — Stay in the Loop
Litt is explicit that he deliberately does not use autonomous harnesses for his core research because he wants to remain the agent of understanding. This is an operating principle for any knowledge profession: design your AI workflow so that you are using the model to deepen your own comprehension, not to outsource the comprehension itself.
"I personally do not enjoy autonomous mathematics very much, so I mostly do not use the harness. I mostly try to use it to help me understand stuff... You don't want to automate your job away because that does involve you being in the loop to understand it, which is, you know, necessary to participate." 00:58:59
6. Overlooked Insights
The 800-Page AI-Generated "Proof" on ArXiv Is a Canary — and It Points to a Coming Verification Crisis
Litt very briefly mentions, almost in passing, that someone recently posted an 800-page AI-generated claimed proof of resolution of singularities in positive characteristic — one of the deepest open problems in algebraic geometry. He dismisses it as certainly wrong given current model capabilities. But the deeper signal is enormous: we are already in a world where AI can generate plausible-sounding, technically dense, multi-hundred-page mathematical documents that no human will read and no AI can reliably verify. This is not just a math problem — it is a preview of a coming verification and trust crisis across all long-form technical knowledge production (legal briefs, clinical trial write-ups, engineering specifications, financial models). The institutions that build robust verification infrastructure for long-horizon AI outputs will have an enormous structural advantage.
"Someone recently posted a claimed proof of resolution of singularities in positive characteristic, which was 800 AI-generated pages. It's, like, definitely — I mean, I'm sorry, I haven't read it. I haven't done an error, but there's no way it's correct... The capacities of the current models, which are reasonably well-calculated. And it's definitely no human has read it. Definitely the models are not able to check this kind of thing." 00:54:00
Mode Collapse in AI Reasoning Is Already Empirically Observable — and Nobody Is Raising the Alarm
Litt casually mentions — and Leisha Li almost immediately moves past — that multiple groups have produced the exact same proof of the exact same theorem within days of each other, because the frontier models are consistently converging on the same reasoning paths. This is not just an academic nuisance. It is direct empirical evidence of mode collapse in frontier reasoning models across independent users. For investors and operators: any domain where AI reasoning is being used for differentiated competitive outputs (investment theses, drug candidate selection, strategic analysis) is quietly experiencing the same convergence. The teams that detect this and deliberately inject diversity into their AI-assisted reasoning pipelines — through prompt engineering, model diversity, or structured human dissent — will have a genuine edge over competitors who assume AI-assisted work is inherently differentiated.
"Three or four or five papers with the exact same proof of the exact same theorem have come out within a couple days of each other... ChatGPT is kind of consistently finding the same. But that's also interesting... It's kind of mode collapsed on, like, certain paths of reasoning." 00:37:15