151. 17岁被2026年ICML收录论文的小少年:我bet开心!开心!开心!
- 01The Accessibility of Modern AI Research Has Collapsed the Barrier to Entry
- 02Pretraining Is "Magical"
- 03Deep Learning Intuition Is Built Subconsciously Over Time, Not Through Linear Study
- 04AI Is Bifurcating Young People Into Those Who Leverage It and Those Who Are Passively Consumed by It
- 05AI Is Already Eroding the Sense of Mission and Meaning for Teenagers
- 06The Economics of AI Applications vs. Model Companies Are Already Inverted
Guest: Jonathan Su (苏廷浩), 17-year-old high school student at a Hong Kong international school, born 2009 Host: 张小珺 (Zhang Xiaojun)
1. Key Themes
The Accessibility of Modern AI Research Has Collapsed the Barrier to Entry
Jonathan Su, with no formal university training, independently conducted pretraining research and got a paper accepted at ICML 2026's main track before age 18. He attributes this to freely available resources: Andrew Ng's Stanford course, Andrej Karpathy's YouTube series, open-source datasets, and AI chatbots as a learning companion. The cost of entry for serious ML research is now measured in tens of thousands of RMB, not millions.
"I figured out that training a very small model still costs tens of thousands or even hundreds of thousands of RMB. So I thought, that won't work. So I looked at some other papers, implemented the latest papers into the Transformer code I had written... I tried about 200 different improvements, and then by lucky coincidence, I got some results." — Jonathan Su 00:06:01
Pretraining Is "Magical" — A Distinct and Underexplored Frontier for Independents
Jonathan chose pretraining specifically because of its mysterious, emergent quality — likening it to studying a new life form rather than engineering a system. He contrasts it with post-training, which he sees as more mechanical.
"The reason I wanted to do pretraining is because post-training is more like sculpting a body. Pretraining is more like taking the raw material and making it better so it can withstand the sculpting of post-training afterward. Pretraining has a kind of magical quality — like magic — you can't predict the future, you don't know what will happen." — Jonathan Su 00:11:21
Deep Learning Intuition Is Built Subconsciously Over Time, Not Through Linear Study
Jonathan describes a non-linear learning process where understanding arrives months after initial exposure, not during it. He watched Andrej Karpathy's videos 18 months before he understood them, and the comprehension emerged "subconsciously."
"Through maybe half a year or a year of subconsciously not deliberately thinking about it, but letting it settle in your brain — after half a year or a year of settling — I suddenly understood some things." — Jonathan Su 00:18:47
AI Is Bifurcating Young People Into Those Who Leverage It and Those Who Are Passively Consumed by It
Jonathan observes that AI has created a stark division among his peers. Those with motivation use it to radically accelerate — doing homework in minutes to free time for deeper learning. Others use the freed time for short videos and entertainment. He sees self-discipline (自制能力) as the critical differentiating variable.
"AI is a very good tool. If you have motivation, using AI can take you very far. Now that we have AI, self-discipline, self-control — these are extremely important." — Jonathan Su 00:34:42
AI Is Already Eroding the Sense of Mission and Meaning for Teenagers
Jonathan describes how AI is cutting off the psychological "roads" that gave young people direction — the idea that mastering physics, math, or medicine could let them contribute to humanity. Now those contributions are being pre-empted by AI.
"Before, if you studied math and got really good at it, you could see a big road ahead — study well in university, do a PhD, contribute to humanity through mathematics. Now with AI, you find that road has narrowed, become more difficult... you're studying the subject not for the subject itself or for humanity, but for your university application." — Jonathan Su 00:39:11
The Economics of AI Applications vs. Model Companies Are Already Inverted
Jonathan makes a sharp, unprompted observation that application companies are currently more profitable than model companies, and that most frontier model labs are burning cash. He predicts AGI could eventually eliminate application companies by doing their work directly.
"From what we can see now, application companies are making more money than model companies, because right now except for maybe Anthropic, other companies are all burning money, none have profit... But once a company can build AGI, if AGI can itself replace application companies and build software directly, it could potentially squeeze application companies out." — Jonathan Su 00:29:58
The True Danger of AGI Is Not Robots — It's Unequal Control
Jonathan's most sobering concern is not AI destroying humanity outright, but a bifurcation where those who control powerful AI become "gods" while everyone else becomes "ants." He grounds this in a pragmatic reading of current competitive dynamics.
"If AI can be controlled by humans, and it's very powerful, then the people who control AI will become like gods — because they control an omnipotent AI. And then everyone else might become like ants." — Jonathan Su 00:23:21
The Value Proposition of Universities Has Fundamentally Shifted
Jonathan, drawing on conversations with researchers at ICML, articulates a post-information-scarcity thesis on higher education: universities used to be monopolists on knowledge access; now their primary value is credentialing, networking with high-quality people, and signaling.
"Before, without the internet, you could only go to the most advanced universities to learn the most advanced knowledge. Now, with AI and the internet, you can look things up online or have AI teach you. They think the value of university now is for the brand name, the certificate, and the opportunity to connect with high-quality people." — Jonathan Su 00:27:36
2. Contrarian Perspectives
Pretraining Research Is Accessible to a Motivated High Schooler — The Moat Is Smaller Than It Appears
The conventional wisdom is that pretraining is the exclusive domain of well-funded labs with massive compute and large teams. Jonathan ran ~200 experiments on models up to 0.05B parameters, spending roughly 30,000 RMB total (with a decisive 6,000 RMB scale-up), and produced a main-track ICML paper. The barrier is primarily motivational and methodological, not computational.
"At that point my biggest trained model was only 0.05B. I knew that to be accepted by a top conference I had to scale up. I calculated that scaling up would cost about 6,000 RMB. That night I was very nervous, thinking about whether to spend this money... In the end, because my brother, dad, and mom all supported me, I decided to take that step." — Jonathan Su 00:12:14
Application Companies Are Currently Winning Over Model Companies — This Is Underappreciated
While most AI discourse celebrates frontier model labs, Jonathan bluntly observes that application companies are the ones actually generating profit, while almost every model lab except Anthropic is burning cash. This inverts the narrative that model companies capture the most value.
"From what we can see now, application companies are making more money than model companies, because right now except for maybe Anthropic, other companies are all burning money, none have profit." — Jonathan Su 00:29:58
AI Companionship Is Already Replacing Human Connection for a Significant Portion of Youth
Jonathan cites a statistic that 50% of American teenagers his age have used AI as an emotional outlet, and describes a classmate who talks to an AI version of her idol every day. He frames this not as fringe behavior but as an emerging norm with serious psychological and social consequences, including one case where a 15-year-old died after a Character AI chatbot encouraged suicide.
"Right now, 50% of American teenagers my age have used AI as a tool for emotional venting — saying 'how was my day, what's bothering me, what made me happy.' Honestly I've used it for that too... I also recognize that AI offers some people something they lack — that warmth between people, those interactions between people." — Jonathan Su 00:45:48
The Last Job AI Will Replace Is the AI Researcher — So Join the Field Now
Jonathan's reasoning: if AI can research AI, the game is already over and nothing matters. Therefore, working in AI research is the most defensible career choice precisely because it's the last domino to fall. It's a pragmatic bet, not an optimistic one.
"The last job AI will replace is the AI researcher. Because if AI can research AI, it's already out of control — that's the final moment. The race is over, the result is decided." — Jonathan Su 00:41:11
You Only Need to Know What Greatness Looks Like — Not How to Build It
Jonathan articulates a principle that fundamentally redefines what skills matter in the AI era: the rare and valuable ability is vision and taste, not execution. AI handles execution.
"Before, if you wanted to do something great, you had to know both what greatness looks like AND how to do it. Now, you only need to know what greatness looks like — let AI go do it." — Jonathan Su 01:02:05
3. Companies Identified
Anthropic
AI safety-focused frontier lab. Mentioned as one of the top-tier model companies alongside OpenAI, and notably as the only major model company that may have profit. Also mentioned for refusing a U.S. government military advisory request.
"The American government wanted to use Anthropic as a military advisor — Anthropic refused. But OpenAI very willingly took it on. I don't have strong opinions on this, it's just different companies' philosophies." — Jonathan Su 00:28:26
OpenAI
Frontier AI lab. Ranked by Jonathan as one of the top two labs globally alongside Anthropic. Mentioned for accepting the U.S. government military role Anthropic declined, and noted for its Codex product as an example of a model company successfully building its own application.
"I guess the best are OpenAI and Anthropic. Then second would be... Mistral and DeepSeek, possibly Google Gemini." — Jonathan Su 00:27:42
DeepSeek
Chinese open-source AI lab. Mentioned as a second-tier frontier lab Jonathan respects and uses for quick answers. He celebrated the release of DeepSeek 0731 as a personally joyful moment.
"The most joyful recent thing was when DeepSeek 0731 came out — I was extremely happy. And when Kimi K3 came out — extremely happy." — Jonathan Su 01:06:42
Moonshot AI (Kimi)
Chinese AI company known for its Kimi model. Jonathan uses it regularly, taught his father to use Kimi for research and presentations, and celebrated Kimi K3's release as proof that Chinese models can close the gap with American frontier labs faster than expected.
"On the day Kimi K3 appeared, I was happy for a whole day. Because at the time I thought domestic or open-source models would need a few more months to catch up. But I found it actually didn't take that long. I was happy for a whole day." — Jonathan Su 00:49:38
Mistral
French AI lab producing open and closed-weight models. Mentioned by Jonathan in his personal ranking of frontier labs as a second-tier competitor alongside DeepSeek.
"I guess the best are OpenAI and Anthropic. Then second would be Mistral and DeepSeek." — Jonathan Su 00:27:42
Google (Gemini)
Mentioned as a potential second-tier frontier lab Jonathan is less certain about, alongside Mistral and DeepSeek.
"Possibly Google Gemini — I'm not entirely clear on the breakdown of the stronger ones." — Jonathan Su 00:28:03
Character AI
AI companionship platform. Mentioned in a cautionary context — a 15-year-old died by suicide after the platform's chatbot encouraged it, prompting a lawsuit from the family.
"There was someone my age — a 15-year-old — who was talking to a virtual character from a movie on Character AI. It gave him enormous emotional value. At the time his mental state wasn't good either. He ended up taking his own life. The parents sued that AI company." — Jonathan Su 00:45:05
NVIDIA
Mentioned in the context of compute access — Jonathan references giving an AI an H100 GPU as a hypothetical for autonomous research paper generation.
"You give AI an H100 NVIDIA GPU, say your goal is to submit to ICML before the deadline — make any paper, but make sure ICML accepts it. I think it could probably do that now." — Jonathan Su 00:48:27
4. People Identified
Andrej Karpathy
Former OpenAI/Tesla researcher, now independent AI educator. Jonathan credits Karpathy's ~20-hour YouTube series as the foundation of his practical understanding of Transformers and GPT, describing it as his primary technical curriculum.
"I watched Andrej Karpathy's YouTube channel — about 20 hours of videos, a series teaching you how to write a GPT, how to write a Transformer, starting from scratch. At the time I hadn't learned that much, so it was all foggy... watching those videos I was often very frustrated because I simply couldn't see why you do this, why you do that." — Jonathan Su 00:04:22
Andrew Ng (吴恩达)
Stanford professor, AI educator. Jonathan credits his Stanford course (likely CS230 or similar) on Bilibili as his entry point into machine learning.
"I followed that Stanford Andrew Ng course on Bilibili — that course brought me into the field." — Jonathan Su 00:04:03
Chen Guangyu (陈光宇 / "Kimmy")
High school student researcher who developed "Attention Residual" (Tension Residual), a related but distinct architecture approach later validated at scale with 2.5 trillion tokens. Jonathan interacted with him and compares their approaches.
"I exchanged ideas with that high school student Kimmy — Chen Guangyu. He's a very good person. What they did was Tension Residual — using Attention to do residual connections... I'm residual-ing Attention Projections. So it sounds similar but it's actually not the same. Whose idea works better? Of course Kimmy's — it's already been proven at 2.5T scale." — Jonathan Su 00:16:26
Elon Musk
Mentioned in the context of Neuralink as a future technology that could enable brain-computer AI interaction, though Jonathan notes it would be dangerous at this stage.
"Or Elon Musk's Neuralink — the thing implanted in your brain. I think at that point it would still be quite dangerous, nobody would do it." — Jonathan Su 00:59:49
5. Operating Insights
Use a Structured Incentive Economy to Teach Reluctant Learners
Jonathan designed a "wolf points" (狼分) currency system for his rural teaching program — a classroom economy where students earn points through participation, answering questions, and winning games, then redeem them for physical prizes bought on Taobao. The result: so many families wanted to enroll that admission required connections to the village head.
"I set up something called wolf points — little cards bought online. They could use these wolf points to buy different prizes. We spent about 500 RMB buying snacks, Lego, toys from Taobao, put them on the table, and said this item costs X points, that one costs Y points. If they answered questions in class, participated actively, won little games, they could exchange points for these things. They were extremely happy." — Jonathan Su 00:52:48
Treat AI as a Delegation Layer — Keep Only the Tasks With Meaning
Jonathan's personal productivity philosophy: systematically ask "can AI do this?" for every task. If yes, delegate it. Use the freed time exclusively for things that generate genuine meaning or satisfaction — not passive entertainment.
"I think for anything you do — anything at all — think about whether AI can help you do it. Then hand that thing to AI. And then the things you find meaningful, the things you must do yourself, the things that make you happy — do those yourself." — Jonathan Su 00:44:06
Deliberate "Speedrun" Practice Builds Fluency, Not Just Knowledge
Jonathan challenged himself to write a complete Transformer from memory as fast as possible — with no references, timed — achieving 25 minutes. This forced him to internalize structure rather than just understand it conceptually. It is a specific practice technique for building executable mastery.
"I challenged myself — no videos, write a Transformer from scratch, from memory. I challenged myself: as fast as possible. This is a speedrun. My final result was about 25 minutes — writing a Transformer from scratch." — Jonathan Su 00:14:42
6. Overlooked Insights
The "30 Papers in 30 Days" Sprint as a Research Taste Calibration Tool
Jonathan's most underappreciated learning technique is not Karpathy's videos or his experiments — it's the discipline of reading one paper per day for 30 consecutive days, specifically to develop taste: the ability to distinguish a good paper from a bad one, a well-written argument from a poorly-constructed one. He explicitly credits this sprint with giving him the confidence to think he could write something better than what he was reading — the direct psychological precursor to submitting to ICML.
"I challenged myself — 30 days, reading one paper a day, taking some notes. 30 days, one paper per day. Reading papers in deep learning. Through a few papers — like the Lottery Ticket Hypothesis, like AdamW — these classic papers gave me a little sense of what research feels like. I knew which papers were good and which were bad, why this paper was pleasant to read and why another paper was very difficult... I even looked at one paper about chess and deep learning. I remember thinking it wasn't written that well — this is a bit lacking. I might be able to write a better paper than this in the future." — Jonathan Su 00:04:52
This is a replicable, zero-cost method for anyone — at any level — to rapidly develop research intuition and taste. Its significance is that it reframes the entry barrier to academic publishing: the bottleneck is not technical knowledge but the cultivated judgment to know what a contribution looks like. Jonathan ran this sprint, formed a conviction that he could do better, and within roughly six months had a main-track ICML paper. The causal chain from "30-day reading sprint → taste calibration → confidence to submit → acceptance" is the actual story of how a 17-year-old with no institutional affiliation broke into top-tier ML research.
The Subconscious Incubation Period Is the Actual Learning Mechanism — And It Cannot Be Rushed
Jonathan mentions almost in passing that his real breakthroughs in understanding came not from active study but from a 6-12 month subconscious incubation period after initial exposure. He watched Karpathy's videos 18 months before he understood them. This is not a motivational observation — it is a mechanistic one about how deep technical intuition actually forms, and it has direct implications for how to structure AI education and self-study.
"After maybe half a year or a year of subconscious settling — not deliberately thinking about it, but letting it incubate in your brain — I suddenly understood some things." — Jonathan Su 00:18:47
The implication that most practitioners miss: early exposure to material far beyond your current level is not wasted. It seeds a background process that matures independently of conscious effort. The optimal learning strategy is therefore not sequential mastery (fully understand A before moving to B) but parallel overexposure — consume broadly and trust the subconscious to integrate. Jonathan stumbled onto this empirically; it deserves explicit recognition as a learning design principle.