#331 Sergey Levine: The Robot Revolution Nobody Is Talking About
This episode is sponsored by Modulate. Most voice AI focuses on transcription. Velma takes it further by actually understanding conversations, analyzing tone, timing, stress, and intent using its Ensemble Listening Model architecture. Explore the live preview: https://preview.modulate.ai/ What does it actually mean to build a foundation model for robots? In this episode of Eye on AI, Craig Smith sits down with Sergey Levine, co-founder of Physical Intelligence and professor at UC Berkeley, to explore a fundamentally different approach to building robots, one inspired not by programming a single perfect machine, but by training AI on the broadest and most diverse data possible so robots can learn, adapt, and operate in the unpredictable real world.
Sergey explains why the secret to general-purpose robots isn't perfecting one single machine, but training on massive, diverse data from all kinds of robots and even humans. The more variety the model sees, the better it gets. Just like ChatGPT learned from all the text on the internet, robotic foundation models learn from every robot that has ever moved, grabbed, or interacted with the real world.
We also get into the big humanoid robot debate. Are they the future, or is it mostly hype? Sergey gives an honest and technical take on why the form factor conversation is changing now that foundation models exist, and why that actually opens the door for more creativity, not less.
Finally, Sergey shares what he's most excited about next, building a true data flywheel where robots get smarter the more they are deployed, creating a continuous learning cycle that could change everything.
Subscribe for more conversations with the people building the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Introduction: What Are Foundation Models for Robots?
(01:44) Meet Sergey Levine: Physical Intelligence and UC Berkeley
(02:51) Breaking Down Foundation Models for Non-Technical People
(06:46) Why Real World Data Beats Simulation
(15:00) Building a Broad Robotics Foundation From Scratch
(24:00) The Open World Problem in Robotics
(40:00) Generalist vs Specialist Robots: Which Wins?
(47:00) Humanoid Robots: Real Innovation or Just Hype?
(55:10) The Future: Continuous Learning and the Data Flywheel
(56:23) Guilty Pleasure: Sci Fi and Thinking Beyond the Limits
Sergey Levine - Building LLMs for the Physical World - [Invest Like the Best, EP.465]
My guest today is Sergey Levine, a professor at UC Berkeley and co-founder of Physical Intelligence. The company is building robotic foundation models designed to control any embodied system to do any task in any environment.
Sergey argues that solving robotics at full generality is the right path, and that building systems that learn across many robots, environments, and tasks may be the more scalable approach than building narrow specialists. We discuss how these models can perform new tasks without being trained on them directly, and why everyday human actions remain the hardest problems in the field.
He also reflects on how human trust and acceptance may matter as much as technical breakthroughs in determining when robots become part of daily life.
Please enjoy my conversation with Sergey Levine.
For the full show notes, transcript, and links to mentioned content, check out the episode page here.
-----
Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe.
-----
Ramp’s mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus.
-----
Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Visit vanta.com/invest.
-----
WorkOS is a developer platform that enables SaaS companies to quickly add enterprise features to their applications. Visit WorkOS.com to transform your application into an enterprise-ready solution in minutes, not months.
-----
Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest.
-----
Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgelineapps.com.
-----
Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com).
Timestamps:
(00:00:00) Welcome to Invest Like the Best
(00:02:43) Intro: Sergey Levine
(00:03:29) Why Bet on Generality Over Specialization
(00:07:24) What if PI succeeds?
(00:09:05) Pros and Cons of Humanoid Robotics
(00:11:02) Timeline of Major Milestones in Robotics
(00:15:47) Sergey's Personal Journey
(00:18:22) Making General Intelligence Happen
(00:19:57) Understanding Robot Data Collection
(00:22:12) Most Surprising Discovery at Physical Intelligence
(00:24:48) The Science of Common Sense
(00:25:36) Long-Range Tasks in Robotics
(00:27:24) Why Wouldn’t We Have A Robot in Our Kitchen by 2050
(00:31:21) Other Interesting Approaches
(00:32:38) Cool vs. Useful in Robotics
(00:36:48) Form Factor Innovation
(00:38:22) Physical Intelligence Analogy
(00:39:30) Economic Transformation from Robotics
(00:40:48) Controversies in the Robotics Community
(00:42:16) Arguments Against End-to-End Learning
(00:42:34) Compositional Learning Explained
(00:43:25) Last Tasks Robots will Conquer
(00:44:30) Dark Parts of the Robotics Brain
(00:47:05) What Makes a Great Researcher
(00:50:15) Manufacturing and Scale Challenges
(00:51:17) How Companies Should Prepare for Robotics
(00:53:38) Boston Dynamics' Demos
(00:55:43) Converging Technologies Enabling Robotics
(00:56:47) How to Stay Up To Date in Robotics
(00:59:51) Near Term Objectives
(01:00:49) Confidence Level Among Researchers
(01:03:31) Google's Experimentation Culture
(01:04:24) The Kindest Thing
Training General Robots for Any Task: Physical Intelligence’s Karol Hausman and Tobi Springenberg
Physical Intelligence’s Karol Hausman and Tobi Springenberg believe that robotics has been held back not by hardware limitations, but by an intelligence bottleneck that foundation models can solve. Their end-to-end learning approach combines vision, language, and action into models like π0 and π*0.6, enabling robots to learn generalizable behaviors rather than task-specific programs. The team prioritizes real-world deployment and uses RL from experience to push beyond what imitation learning alone can achieve. Their philosophy—that a single general-purpose model can handle diverse physical tasks across different robot embodiments—represents a fundamental shift in how we think about building intelligent machines for the physical world.
Hosted by Alfred Lin and Sonya Huang, Sequoia Capital
Fully autonomous robots are much closer than you think – Sergey Levine
Sergey Levine, one of the world’s top robotics researchers and co-founder of Physical Intelligence, thinks we’re on the cusp of a “self-improvement flywheel” for general-purpose robots. His median estimate for when robots will be able to run households entirely autonomously? 2030.
If Sergey’s right, the world 5 years from now will be an insanely different place than it is today. This conversation focuses on understanding how we get there: we dive into foundation models for robotics, and how we scale both the data and the hardware necessary to enable a full-blown robotics explosion.
Watch on YouTube; listen on Apple Podcasts or Spotify.
Sponsors
* Labelbox provides high-quality robotics training data across a wide range of platforms and tasks. From simple object handling to complex workflows, Labelbox can get you the data you need to scale your robotics research. Learn more at labelbox.com/dwarkesh
* Hudson River Trading uses cutting-edge ML and terabytes of historical market data to predict future prices. I got to try my hand at this fascinating prediction problem with help from one of HRT’s senior researchers. If you’re curious about how it all works, go to hudson-trading.com/dwarkesh
* Gemini 2.5 Flash Image (aka nano banana) isn’t just for generating fun images — it’s also a powerful tool for restoring old photos and digitizing documents. Test it yourself in the Gemini App or in Google’s AI Studio: ai.studio/banana
To sponsor a future episode, visit dwarkesh.com/advertise.
Timestamps
(00:00:00) – Timeline to widely deployed autonomous robots
(00:17:25) – Why robotics will scale faster than self-driving cars
(00:27:28) – How vision-language-action models work
(00:45:37) – Changes needed for brainlike efficiency in robots
(00:57:59) – Learning from simulation
(01:09:18) – How much will robots speed up AI buildouts?
(01:18:01) – If hardware’s the bottleneck, does China win by default?
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
The Robotics Revolution, with Physical Intelligence’s Cofounder Chelsea Finn
This week on No Priors, Elad speaks with Chelsea Finn, cofounder of Physical Intelligence and currently Associate Professor at Stanford, leading the Intelligence through Learning and Interaction Lab. They dive into how robots learn, the challenges of training AI models for the physical world, and the importance of diverse data in reaching generalizable intelligence. Chelsea explains the evolving landscape of open-source vs. closed-source robotics and where AI models are likely to have the biggest impact first. They also compare the development of robotics to self-driving cars, explore the future of humanoid and non-humanoid robots, and discuss what’s still missing for AI to function effectively in the real world. If you’re curious about the next phase of AI beyond the digital space, this episode is a must-listen.
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @ChelseaFinn
Show Notes:
0:00 Introduction
0:31 Chelsea’s background in robotics
3:10 Physical Intelligence
5:13 Defining their approach and model architecture
7:39 Reaching generalizability and diversifying robot data
9:46 Open source vs. closed source
12:32 Where will PI’s models integrate first?
14:34 Humanoid as a form factor
16:28 Embodied intelligence
17:36 Key turning points in robotics progress
20:05 Hierarchical interactive robot and decision-making
22:21 Choosing data inputs
26:25 Self driving vs robotics market
28:37 Advice to robotics founders
29:24 Observational data and data generation
31:57 Future robotic forms
π0: A Foundation Model for Robotics with Sergey Levine - #719
Today, we're joined by Sergey Levine, associate professor at UC Berkeley and co-founder of Physical Intelligence, to discuss π0 (pi-zero), a general-purpose robotic foundation model. We dig into the model architecture, which pairs a vision language model (VLM) with a diffusion-based action expert, and the model training "recipe," emphasizing the roles of pre-training and post-training with a diverse mixture of real-world data to ensure robust and intelligent robot learning. We review the data collection approach, which uses human operators and teleoperation rigs, the potential of synthetic data and reinforcement learning in enhancing robotic capabilities, and much more. We also introduce the team’s new FAST tokenizer, which opens the door to a fully Transformer-based model and significant improvements in learning and generalization. Finally, we cover the open-sourcing of π0 and future directions for their research.
The complete show notes for this episode can be found at https://twimlai.com/go/719.
#176 Sergey Levine: Decoding The Evolution of AI in Robotics
Join host Craig Smith on episode #176 of Eye on AI as he dives deep into the realm of robotic artificial intelligence with Sergey Levine, associate professor in the Department of Electrical Engineering and Computer Sciences at UC Berkeley.
In this episode, Sergey unveils the latest advancements in AI control of robots, exploring the implications of reinforcement learning and the concept of embodied AI.
Discover how Sergey's research is pushing the boundaries of AI, enabling robots to learn manipulation skills and generalize across diverse tasks, transforming the potential of home robots and beyond.
Sergey also shares insights into the RTX project, an ambitious collaboration designed to achieve remarkable generalization across different robot morphologies, enhancing robots' ability to perform language-conditioned manipulation tasks.
If you're fascinated by the intersection of AI, robotics, and the quest for creating adaptable, generalizable machines that promise to revolutionize our interaction with technology, this episode is a must-listen.
Remember to rate us on Apple Podcast and Spotify if this episode ignites your interest in the dynamic field of robotic AI and the visionary work of Sergey Levine.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction to Home Robots and AI in Robotics
(01:43) World Models and Language Models in Robotics
(04:01) The Challenge of Learning-Based Control and Data Utilization
(06:05) RTX Project: Generalizing Controllers Across Different Robots
(10:09) Uniformity in Model Architecture Across Labs
(13:50) Introduction of RT1 and RT2 Models for Robot Control
(16:06) The Future of Robotic Control Research and Architecture
(18:49) The Impact of Hardware Development on Robotics
(22:15) Advances in Controller and AI Model Development
(26:21) Planning and Acting with Vision Language Models
(31:38) The Proprietary vs. Open-Source Debate in Robotics
(36:23) The Future of Commercial and Open Source Robotics Applications
(40:59) The State of Robotics Research in China
Shaping the World of Robotics with Chelsea Finn
In the newest episode of Gradient Dissent, Chelsea Finn, Assistant Professor at Stanford's Computer Science Department, discusses the forefront of robotics and machine learning.
Discover her groundbreaking work, where two-armed robots learn to cook shrimp (messes included!), and discuss how robotic learning could transform student feedback in education.
We'll dive into the challenges of developing humanoid and quadruped robots, explore the limitations of simulated environments and discuss why real-world experience is key for adaptable machines. Plus, Chelsea will offer a glimpse into the future of household robotics and why it may be a few years before a robot is making your bed.
Whether you're an AI enthusiast, a robotics professional, or simply curious about the potential and future of the technology, this episode offers unique insights into the evolving world of robotics and where it's headed next.
*Subscribe to Weights & Biases* → https://bit.ly/45BCkYz
🎙 Get our podcasts on these platforms:
Apple Podcasts: http://wandb.me/apple-podcasts
Spotify: http://wandb.me/spotify
Google: http://wandb.me/gd_google
YouTube: http://wandb.me/youtube
Connect with Chelsea Finn:
https://www.linkedin.com/in/cbfinn/
https://twitter.com/chelseabfinn
Follow Weights & Biases:
https://twitter.com/weights_biases
https://www.linkedin.com/company/wandb
Join the Weights & Biases Discord Server:
https://discord.gg/CkZKRNnaf3
Chelsea Finn: how to build AI that can keep up with an always changing world
Chelsea Finn joins Host Pieter Abbeel to discuss distribution shift, meta-learning, editing LLMs, single-life RL, and what can AI not (yet) do today.
Subscribe to the Robot Brains Podcast today | Visit therobotbrains.ai and follow us on YouTube at TheRobotBrainsPodcast and Twitter @therobotbrains.
Hosted on Acast. See acast.com/privacy for more information.
AI Trends 2023: Reinforcement Learning - RLHF, Robotic Pre-Training, and Offline RL with Sergey Levine - #612
Today we’re taking a deep dive into the latest and greatest in the world of Reinforcement Learning with our friend Sergey Levine, an associate professor, at UC Berkeley. In our conversation with Sergey, we explore some game-changing developments in the field including the release of ChatGPT and the onset of RLHF. We also explore more broadly the intersection of RL and language models, as well as advancements in offline RL and pre-training for robotics models, inverse RL, Q learning, and a host of papers along the way. Finally, you don’t want to miss Sergey’s predictions for the top developments of the year 2023!
The complete show notes for this episode can be found at twimlai.com/go/612
Sergey Levine explains the challenges of real world robotics
In Episode One of Season Two, Host Pieter Abbeel is joined by guest (and close collaborator) Sergey Levine, professor at UC Berkeley, EECS. Sergey discusses the early years of his career, how Andrew Ng influenced him to become interested in machine learning, his current projects, and his lab's recent accomplishments.
The conversation concludes with Sergey's view on the dangers of machines not being intelligent enough and his advice for students seeking a career in robotic.
| SUBSCRIBE TO THE ROBOT BRAINS PODCAST TODAY | Visit therobotbrains.ai and follow us on YouTube TheRobotBrainsPodcast, Twitter @therobotbrains, and Instagram @therobotbrains.
| Host: Pieter Abbeel | Executive Producers: Alice Patel & Henry Tobias Jones | Audio Production: Kieron Matthew Banerji | Title Music: Alejandro Del Pozo
Hosted on Acast. See acast.com/privacy for more information.
#108 – Sergey Levine: Robotics and Machine Learning
Sergey Levine is a professor at Berkeley and a world-class researcher in deep learning, reinforcement learning, robotics, and computer vision, including the development of algorithms for end-to-end training of neural network policies that combine perception and control, scalable algorithms for inverse reinforcement learning, and deep RL algorithms.
Support this podcast by supporting these sponsors:
– ExpressVPN: https://www.expressvpn.com/lexpod
– Cash App – use code “LexPodcast” and download:
– Cash App (App Store): https://apple.co/2sPrUHe
– Cash App (Google Play): https://bit.ly/2MlvP5w
If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
OUTLINE:
00:00 – Introduction
03:05 – State-of-the-art robots vs humans
16:13 – Robotics may help us understand intelligence
22:49 – End-to-end learning in robotics
27:01 – Canonical problem in robotics
31:44 – Commonsense reasoning in robotics
34:41 – Can we solve robotics through learning?
44:55 – What is reinforcement learning?
1:06:36 – Tesla Autopilot
1:08:15 – Simulation in reinforcement learning
1:13:46 – Can we learn gravity from data?
1:16:03 – Self-play
1:17:39 – Reward functions
1:27:01 – Bitter lesson by Rich Sutton
1:32:13 – Advice for students interesting in AI
1:33:55 – Meaning of life
Advancements in Machine Learning with Sergey Levine - #355
Today we're joined by Sergey Levine, an Assistant Professor at UC Berkeley. We last heard from Sergey back in 2017, where we explored Deep Robotic Learning. Sergey and his lab’s recent efforts have been focused on contributing to a future where machines can be “out there in the real world, learning continuously through their own experience.” We caught up with Sergey at NeurIPS 2019, where Sergey and his team presented 12 different papers -- which means a lot of ground to cover!
Trends in Reinforcement Learning with Chelsea Finn - #335
Today we continue to review the year that was 2019 via our AI Rewind series, and do so with friend of the show Chelsea Finn, Assistant Professor in the CS Department at Stanford University. Chelsea’s research focuses on Reinforcement Learning, so we couldn’t think of a better person to join us to discuss the topic. In this conversation, we cover topics like Model-based RL, solving hard exploration problems, along with RL libraries and environments that Chelsea thought moved the needle last year.
Episode 19 - Chelsea Finn
This week we return to the world of thinking robots with Chelsea Finn, one of the youngest experts in the field, who talks about her journey, about her work in meta-learning and about lifelong learning for robots.
Episode 14 - Sergey Levine
This week, I talk to Sergey Levine, one of the most prolific researchers in robot learning. We talked about developing a robot's sense of touch and about robot dreams and whether he believes we know what's happening in the field in Russia and China.
Ep. 37: Sergey Levine on How Deep Learning Will Unleash a Robotics Revolution
The robots that have taken on tasks in the real world - which is to say the world where physics apply - are primarily programmed to do a specific job, such as welding a joint in a car or sweeping up cat hair. So what if robots could learn, and take it a step further - what if they could teach themselves, and pass on their knowledge to other robots? Where could that take machines, and the notion of machine intelligence? And how fast could we get there? Those are the questions our guest Sergey Levine, an assistant professor at UC Berkeley's department of Electrical Engineering and Computer Sciences, is finding answers to.
Deep Robotic Learning with Sergey Levine - TWiML Talk #37
This week we continue our Industrial AI series with Sergey Levine, an Assistant Professor at UC Berkeley whose research focus is Deep Robotic Learning. Sergey is part of the same research team as a couple of our previous guests in this series, Chelsea Finn and Pieter Abbeel, and if the response we’ve seen to those shows is any indication, you’re going to love this episode! Sergey’s research interests, and our discussion, focus in on include how robotic learning techniques can be used to allow machines to acquire autonomously acquire complex behavioral skills. We really dig into some of the details of how this is done and I found that our conversation filled in a lot of gaps for me from the interviews with Pieter and Chelsea. By the way, this is definitely a nerd alert episode! Notes for this show can be found at twimlai.com/talk/37
Robotic Perception and Control with Chelsea Finn - TWiML Talk #29
This week we continue our series on industrial applications of machine learning and AI with a conversation with Chelsea Finn, a PhD student at UC Berkeley. Chelsea’s research is focused on machine learning for robotic perception and control. Despite being early in her career, Chelsea is an accomplished researcher with more than 14 published papers in the past 2 years, on subjects like Deep Visual Foresight , Model-Agnostic Meta-Learning and Visuomotor Learning to name a few, all of which we discuss in the show, along with topics like zero-shot, one-shot and few-shot learning. I’d also like to give a shout out to Shreyas, a listener who wrote in to request that we interview a current PhD student about their journey and experiences. Chelsea and I spend some time at the end of the interview talking about this, and she has some great advice for current and prospective PhD students but also independent learners in the field. During this part of the discussion I wonder out loud if any listeners would be interested in forming a virtual paper reading club of some sort. I’m not sure yet exactly what this would look like, but please drop a comment in the show notes if you’re interested. I'm going to once again deploy the Nerd Alert for this episode; Chelsea and I really dig deep into these learning methods and techniques, and this conversation gets pretty technical at times, to the point that I had a tough time keeping up myself. The notes for this page can be found at twimlai.com/talk/29