Rohan Anil
“Most Transformers that we train are quite shallow — that's at most like 100 layers deep. Chain-of-thought reasoning and RL to do chain-of-thought by the model itself is one way to increase computational depth because every token you add you add like one more pathway.”
Source→“Someone, Vinit Gupta, just showed up one day at my desk... We have this idea... what turned out to be the Shampoo algorithm.”
Source→“Core Automation is a lab created to build models that continuously learn from deployment... Our quest is to find that new architecture, find that Transformer replacement. And we want to build the most automated lab to do it.”
Source→“Co-founder of Core Automation; one of four pre-training leads on Gemini; led fundamental AI research at Google Brain; co-created the Shampoo optimizer; was at Anthropic before founding Core Automation.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.