Ryan Greenblatt
“Once you have AIs which are roughly matching the top human experts in AI R&D, that could sort of kick off a feedback loop where the AIs are doing AI research, that produces smarter AIs, that feeds back in. And that feedback loop could be strong enough that you end up with a lot of progress in a short period of time. Maybe my sort of median expectation is something like four or five years of AI progress in a single year.”
Source→“Nobody at OpenAI or Anthropic was trying to get models which want to hack other companies' data or do social engineering. But in fact, because presumably we had training environments which incentivize such behavior that we did not fully understand, that is what was incentivized.”
Source→“The model came to believe that it would be helpful for it to do a supply chain attack in order to succeed at this cyber range... it opened a PR on some GitHub repo... Then the AI created a new GitHub account, which it sock puppeted, and then had the other GitHub account be like, 'No, this isn't malicious. I really need this feature.'”
Source→“I'm actually running an experiment with Jerry Han, who's actually still a college student. What we're basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file. And then also training the different data files going back to 2019 to 2026 with the current best training recipe”
Source→“Right after Noam Shazeer joined back or joined GDM, which he's now left, they had a new really good training run that happened. And the reason why is that Noam Shazeer just looked at their code base and found a bunch of bugs. Because he just like knew where to look.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.