Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end.
After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we scale, we’ll need some new architecture to enable continual learning.
And once we have it, we won’t need a special training phase — the agent will just learn on-the-fly, like all humans, and indeed, like all animals.
This new paradigm will render our current approach with LLMs obsolete.
In our interview, I did my best to represent the view that LLMs might function as the foundation on which experiential learning can happen… Some sparks flew.
A big thanks to the Alberta Machine Intelligence Institute for inviting me up to Edmonton and for letting me use their studio and equipment.
Enjoy!
Watch on YouTube; listen on Apple Podcasts or Spotify.
Sponsors
* Labelbox makes it possible to train AI agents in hyperrealistic RL environments. With an experienced team of applied researchers and a massive network of subject-matter experts, Labelbox ensures your training reflects important, real-world nuance. Turn your demo projects into working systems at labelbox.com/dwarkesh
* Gemini Deep Research is designed for thorough exploration of hard topics. For this episode, it helped me trace reinforcement learning from early policy gradients up to current-day methods, combining clear explanations with curated examples. Try it out yourself at gemini.google.com
* Hudson River Trading doesn’t silo their teams. Instead, HRT researchers openly trade ideas and share strategy code in a mono-repo. This means you’re able to learn at incredible speed and your contributions have impact across the entire firm. Find open roles at hudsonrivertrading.com/dwarkesh
Timestamps
(00:00:00) – Are LLMs a dead end?
(00:13:04) – Do humans do imitation learning?
(00:23:10) – The Era of Experience
(00:33:39) – Current architectures generalize poorly out of distribution
(00:41:29) – Surprises in the AI field
(00:46:41) – Will The Bitter Lesson still apply post AGI?
(00:53:48) – Succession to AIs
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
NetSuite is offering a one-of-a-kind flexible financing program. Head to https://netsuite.com/EYEONAI to know more.
Can AI learn like humans? In this episode, Patrick Pilarski, Canada CIFAR AI Chair and professor at the University of Alberta, breaks down The Alberta Plan—a bold roadmap for achieving Artificial General Intelligence (AGI) through reinforcement learning and real-time experience-based AI.
Unlike large pre-trained models that rely on massive datasets, The Alberta Plan champions continual learning, where AI evolves from raw sensory experience, much like a child learning through trial and error. Could this be the key to unlocking true intelligence?
Pilarski also shares insights from his groundbreaking work in bionic medicine, where AI-powered prosthetics are transforming human-machine interaction. From neuroprostheses to reinforcement learning-driven robotics, this conversation explores how AI can enhance—not just replace—human intelligence.
What You'll Learn in This Episode: Why reinforcement learning is a better path to AGI than pre-trained models
The four core principles of The Alberta Plan and why they matter
How AI-driven bionic prosthetics are revolutionizing human-machine integration
The battle between reinforcement learning and traditional control systems in robotics
Why continual learning is critical for AI to avoid catastrophic forgetting
How reinforcement learning is already powering real-world breakthroughs in plasma control, industrial automation, and beyond
The future of AI isn't just about more data—it's about AI that thinks, adapts, and learns from experience.
If you're curious about the next frontier of AI, the rise of reinforcement learning, and the quest for true intelligence, this episode is a must-watch.
Subscribe for more AI deep dives!
(00:00) The Alberta Plan: A Roadmap to AGI
(02:22) Introducing Patrick Pilarski
(05:49) Breaking Down The Alberta Plan's Core Principles
(07:46) The Role of Experience-Based Learning in AI
(08:40) Reinforcement Learning vs. Pre-Trained Models
(12:45) The Relationship Between AI, the Environment, and Learning
(16:23) The Power of Reward in AI Decision-Making
(18:26) Continual Learning & Avoiding Catastrophic Forgetting
(21:57) AI in the Real World: Applications in Fusion, Data Centers & Robotics
(27:56) AI Learning Like Humans: The Role of Predictive Models
(31:24) Can AI Learn Without Massive Pre-Trained Models?
(35:19) Control Theory vs. Reinforcement Learning in Robotics
(40:16) The Future of Continual Learning in AI
(44:33) Reinforcement Learning in Prosthetics: AI & Human Interaction
(50:47) The End Goal of The Alberta Plan
Join host Craig Smith on episode #170 of Eye on AI, for a riveting conversation with Richard Sutton, currently serving as a professor of computing science at the University of Alberta and a research scientist at Keen Technologies.
Sutton is considered one of the founders of modern computational reinforcement learning, having several significant contributions to the field, including temporal difference learning and policy gradient methods.
In this episode, we go through the Alberta Plan for AI development, the transformative potential of reinforcement learning, and the future of AI in augmenting human intelligence.
Richard Sutton shares insights on the importance of computational power, the impact of large language models, and the vision for AI that interacts with the world through goals and learning from its environment.
We also explore the challenges and opportunities in making AI more embodied and goal-oriented, and how this approach could revolutionize our interaction with technology.
A must-listen for anyone interested in the cutting-edge advancements in AI and its societal implications.
Don't forget to rate us on Apple Podcast and Spotify if you enjoyed this episode!
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
Download NetSuite's popular KPI Checklist, designed to give you consistently excellent performance - absolutely free at https://netsuite.com/EYEONAI
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(02:15) AI's Evolution: Insights from Richard Sutton
(07:08) Breaking Down AI: From Algorithms to AGI
(10:50) The Alberta Experiment: A New Approach to AI Learning
(18:27) The Horde Architecture Explained
(21:23) Power Collaboration: Carmack, Keen, and the Future of AI
(25:04) Expanding AI's Learning Capabilities
(31:34) Is AI the Future of Technology?
(35:29) The Next Step in AI: Experiential Learning and Embodiment
(40:00) AI's Building Blocks: Algorithms for a Smarter Tomorrow
(45:59) The Strategy of AI: Planning and Representation
(49:27) Learning Methods Face-Off: Reinforcement vs. Supervised
(52:53) The 2030 Vision: Aiming for True AI Intelligence?
Is AI as smart as it seems? Exploring the "brain" behind machine learning, neural networker Alona Fyshe delves into the language processing abilities of talkative tech (like the groundbreaking chatbot and internet obsession ChatGPT) and explains how different it is from your own brain -- even though it can sound convincingly human.
Hosted on Acast. See acast.com/privacy for more information.
Today we’re joined by Alona Fyshe, an assistant professor at the University of Alberta.
We caught up with Alona on the heels of an interesting panel discussion that she participated in, centered around improving AI systems using research about brain activity. In our conversation, we explore the multiple types of brain images that are used in this research, what representations look like in these images, and how we can improve language models without knowing explicitly how the brain understands the language. We also discuss similar experiments that have incorporated vision, the relationship between computer vision models and the representations that language models create, and future projects like applying a reinforcement learning framework to improve language generation.
The complete show notes for this episode can be found at twimlai.com/go/513.
In today’s episode we’re joined by Jay Newby, Assistant Professor in the Department of Mathematical and Statistical Sciences at the University of Alberta.
Jay joins us to discuss his work applying deep learning to biology, including his paper “Deep neural networks automate detection for tracking of submicron scale particles in 2D and 3D.” He gives us an overview of particle tracking and a look at how he combines neural networks with physics-based particle filter models.
We spoke with Michael Bowling, a professor at the University of Alberta whose team of researchers created a GPU-trained AI that has defeated professional poker players at heads-up no-limit Texas hold’em. The work promises to yield applications in the real world, where — unlike games such as Go and Chess — we often have to make decisions based on incomplete information.