Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
New episode with my good friends Sholto Douglas & Trenton Bricken. Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic.
We talk through what’s changed in the last year of AI research; the new RL regime and how far it can scale; how to trace a model’s thoughts; and how countries, workers, and students should prepare for AGI.
See you next year for v3. Here’s last year’s episode, btw. Enjoy!
Watch on YouTube; listen on Apple Podcasts or Spotify.
----------
SPONSORS
* WorkOS ensures that AI companies like OpenAI and Anthropic don't have to spend engineering time building enterprise features like access controls or SSO. It’s not that they don't need these features; it's just that WorkOS gives them battle-tested APIs that they can use for auth, provisioning, and more. Start building today at workos.com.
* Scale is building the infrastructure for safer, smarter AI. Scale’s Data Foundry gives major AI labs access to high-quality data to fuel post-training, while their public leaderboards help assess model capabilities. They also just released Scale Evaluation, a new tool that diagnoses model limitations. If you’re an AI researcher or engineer, learn how Scale can help you push the frontier at scale.com/dwarkesh.
* Lighthouse is THE fastest immigration solution for the technology industry. They specialize in expert visas like the O-1A and EB-1A, and they’ve already helped companies like Cursor, Notion, and Replit navigate U.S. immigration. Explore which visa is right for you at lighthousehq.com/ref/Dwarkesh.
To sponsor a future episode, visit dwarkesh.com/advertise.
----------
TIMESTAMPS
(00:00:00) – How far can RL scale?
(00:16:27) – Is continual learning a key bottleneck?
(00:31:59) – Model self-awareness
(00:50:32) – Taste and slop
(01:00:51) – How soon to fully autonomous agents?
(01:15:17) – Neuralese
(01:18:55) – Inference compute will bottleneck AGI
(01:23:01) – DeepSeek algorithmic improvements
(01:37:42) – Why are LLMs ‘baby AGI’ but not AlphaZero?
(01:45:38) – Mech interp
(01:56:15) – How countries should prepare for AGI
(02:10:26) – Automating white collar work
(02:15:35) – Advice for students
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
AMA: career advice given AGI, how I research ft. Sholto & Trenton
I recorded an AMA! I had a blast chatting with my friends Trenton Bricken and Sholto Douglas. We discussed my new book, career advice given AGI, how I pick guests, how I research for the show, and some other nonsense.
My book, “The Scaling Era: An Oral History of AI, 2019-2025” is available in digital format now. Preorders for the print version are also open!
Watch on YouTube; listen on Apple Podcasts or Spotify.
Timestamps
(0:00:00) - Book launch announcement
(0:04:57) - AI models not making connections across fields
(0:10:52) - Career advice given AGI
(0:15:20) - Guest selection criteria
(0:17:19) - Choosing to pursue the podcast long-term
(0:25:12) - Reading habits
(0:31:10) - Beard deepdive
(0:33:02) - Who is best suited for running an AI lab?
(0:35:16) - Preparing for fast AGI timelines
(0:40:50) - Growing the podcast
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
Sholto Douglas & Trenton Bricken — How LLMs actually think
Had so much fun chatting with my good friends Trenton Bricken and Sholto Douglas on the podcast.
No way to summarize it, except:
This is the best context dump out there on how LLMs are trained, what capabilities they're likely to soon have, and what exactly is going on inside them.
You would be shocked how much of what I know about this field, I've learned just from talking with them.
To the extent that you've enjoyed my other AI interviews, now you know why.
So excited to put this out. Enjoy! I certainly did :)
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform.
There's a transcript with links to all the papers the boys were throwing down - may help you follow along.
Follow Trenton and Sholto on Twitter.
Timestamps
(00:00:00) - Long contexts
(00:16:12) - Intelligence is just associations
(00:32:35) - Intelligence explosion & great researchers
(01:06:52) - Superposition & secret communication
(01:22:34) - Agents & true reasoning
(01:34:40) - How Sholto & Trenton got into AI research
(02:07:16) - Are feature spaces the wrong way to think about intelligence?
(02:21:12) - Will interp actually work on superhuman models
(02:45:05) - Sholto’s technical challenge for the audience
(03:03:57) - Rapid fire
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe