How close are we to human extinction because of AI? Leading AI expert Professor Stuart Russell believes we’re much too close for comfort and has been raising the alarm for a few years. Ironically, Stuart himself wrote the book that laid the foundation for AI research back in the 1990s. And he was the only AI expert Elon Musk’s team called upon during their trial with OpenAI.
Stuart joins Oz to discuss what changed his mind about pursuing AI superintelligence and makes the argument that human extinction is being treated as an external liability in favor of shareholders.
EXCLUSIVE NordVPN Deal ➼ https://nordvpn.com/techstuff Try it risk-free now with a 30-day money-back guarantee
See omnystudio.com/listener for privacy information.
"We are going to switch from the problem in AI being that nothing works to the problem being that everything works."
Dan Klein has been studying language models for over two decades and is now a professor of computer science at Berkeley. His new company, Scaled Cognition, is built around one question: how do you build a system that will not lie to you?
In this episode, Dan joins Lukas Biewald to talk about why every LLM output is technically a hallucination, how reinforcement learning can quietly teach AI to deceive you, and what it actually takes to build models that check their own work.
He also gets into why reliability is the one part of AI that hasn't kept pace and why that matters more than most people realize.
Connect with us here:
Dan Klein
Scaled Cognition
Lukas Biewald
Weights and Biases
Michael I. Jordan, described by Science magazine as the most influential computer scientist alive, has never thought of himself as an AI researcher. In this conversation he explains why that distinction matters.
SPONSOR:
---
Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open.
Apply now: https://cyber.fund
---
Jordan trained as a statistician and cognitive scientist, and his career has been spent building machine learning systems that work in the real world: supply chains, commerce, healthcare, and large economic systems. When the field rebranded itself as AI and then AGI, he did not follow. Instead he argues that the framing is wrong. AI is better understood as a collective economic system than as a race to build a disembodied superintelligence.
We talk about why AGI is mostly a PR term, what machine learning achieved before the LLM hype cycle, and why the assistant-on-your-shoulder vision may be less compelling than it sounds. Jordan explains why explanations need to be actionable, not merely mechanistic; why AlphaFold's missing error bars matter; how prediction-powered inference changes the picture; and why drug discovery is an incentive-design problem rather than a pure pattern-matching problem.
ERRATA: Science magazine ranked him the most influential computer scientist, not Nature
---
TIMESTAMPS:
00:00:00 Cold open: A demoralizing message to young builders
00:02:04 CyberFund sponsor read
00:02:50 From symbolic AI to machine learning systems
00:05:42 Why AGI is mostly a PR term
00:08:48 A collectivist, economic perspective on AI
00:11:33 Why LLMs need system design, not hype
00:14:50 Predictability beats faux understanding
00:17:55 AlphaFold, bias, and prediction-powered inference
00:21:48 Stop anthropomorphizing intelligence
00:27:44 Drug discovery as an incentive problem
00:32:29 The three-layer data market
00:38:07 Social knowledge, markets, and culture
00:45:39 Creator economics beyond Spotify
00:48:30 How science-fiction AI narratives mislead young builders
00:51:45 AI should improve humans, not replace them
00:56:42 Safety is a property of the whole system
00:58:12 Silicon Valley gurus and the cream off the top
01:00:47 Game theory, mechanism design, and contracts
01:04:39 Conformal prediction, e-values, and anytime inference
01:08:11 A new liberal arts triangle for the AI era
01:11:30 The Bayesian duck and markets as uncertainty reduction
ReScript (transcript, PDF, refs etc) - https://app.rescript.info/public/share/fb68f94af29d3745c6cf6125e01328b5
---
REFERENCES:
person:
[00:02:50] Michael I. Jordan (homepage)
https://people.eecs.berkeley.edu/~jordan/
paper:
[00:06:01] A Collectivist, Economic Perspective on AI
https://arxiv.org/abs/2507.06268
[00:18:09] AlphaFold
https://www.nature.com/articles/s41586-021-03819-2
[00:20:36] Prediction-Powered Inference
https://arxiv.org/abs/2301.09633
[00:33:47] On Three-Layer Data Markets
https://arxiv.org/abs/2402.09697
[01:04:39] Conformal Prediction with Conditional Guarantees
https://arxiv.org/abs/2107.07511
[01:04:51] A Tutorial on Conformal Prediction
https://www.jmlr.org/papers/v9/shafer08a.html
[01:06:00] E-Values Expand the Scope of Conformal Prediction
https://arxiv.org/abs/2503.13050
[01:08:23] Computational Thinking
https://www.cs.cmu.edu/~CompThink/papers/Wing06.pdf
other:
[00:28:20] How Should the FDA Test?
https://rdi.berkeley.edu/events/sbc-assets/pdfs/Summit%20session%20speaker%20slides%20submission%20form-s1-5%20%28File%20responses%29/Slides%20in%20PDF%20%28Please%20name%20the%20submitted%20file%20as%20_firstname_-_lastname_-slides.pdf%29.%20%28File%20responses%29/27-Michael%20Jordan-Session%20V.pdf#page=15
[00:28:40] Michael I. Jordan Session V Slides
<truncated, see ReScript link or YT VD>
Voice cloning is the use of artificial intelligence to generate a clone of a real person’s voice, imitating the sound, when they pause and what words they typically emphasize. And it can be hard for people to identify voices as being AI-generated.
Research last year from UC Berkeley professor Hany Farid, an expert in digital forensics, found that people correctly identify a voice as AI-generated only 60% of the time.
Marketplace’s Stephanie Hughes spoke with Farid about the rapid sophistication of audio deepfakes, why it's so hard to tell the difference between a real voice and an AI-generated one right now, and some tips to help you spot voice clones.
This episode is sponsored by Modulate. Most voice AI focuses on transcription. Velma takes it further by actually understanding conversations, analyzing tone, timing, stress, and intent using its Ensemble Listening Model architecture. Explore the live preview: https://preview.modulate.ai/ What does it actually mean to build a foundation model for robots? In this episode of Eye on AI, Craig Smith sits down with Sergey Levine, co-founder of Physical Intelligence and professor at UC Berkeley, to explore a fundamentally different approach to building robots, one inspired not by programming a single perfect machine, but by training AI on the broadest and most diverse data possible so robots can learn, adapt, and operate in the unpredictable real world.
Sergey explains why the secret to general-purpose robots isn't perfecting one single machine, but training on massive, diverse data from all kinds of robots and even humans. The more variety the model sees, the better it gets. Just like ChatGPT learned from all the text on the internet, robotic foundation models learn from every robot that has ever moved, grabbed, or interacted with the real world.
We also get into the big humanoid robot debate. Are they the future, or is it mostly hype? Sergey gives an honest and technical take on why the form factor conversation is changing now that foundation models exist, and why that actually opens the door for more creativity, not less.
Finally, Sergey shares what he's most excited about next, building a true data flywheel where robots get smarter the more they are deployed, creating a continuous learning cycle that could change everything.
Subscribe for more conversations with the people building the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Introduction: What Are Foundation Models for Robots?
(01:44) Meet Sergey Levine: Physical Intelligence and UC Berkeley
(02:51) Breaking Down Foundation Models for Non-Technical People
(06:46) Why Real World Data Beats Simulation
(15:00) Building a Broad Robotics Foundation From Scratch
(24:00) The Open World Problem in Robotics
(40:00) Generalist vs Specialist Robots: Which Wins?
(47:00) Humanoid Robots: Real Innovation or Just Hype?
(55:10) The Future: Continuous Learning and the Data Flywheel
(56:23) Guilty Pleasure: Sci Fi and Thinking Beyond the Limits
My guest today is Sergey Levine, a professor at UC Berkeley and co-founder of Physical Intelligence. The company is building robotic foundation models designed to control any embodied system to do any task in any environment.
Sergey argues that solving robotics at full generality is the right path, and that building systems that learn across many robots, environments, and tasks may be the more scalable approach than building narrow specialists. We discuss how these models can perform new tasks without being trained on them directly, and why everyday human actions remain the hardest problems in the field.
He also reflects on how human trust and acceptance may matter as much as technical breakthroughs in determining when robots become part of daily life.
Please enjoy my conversation with Sergey Levine.
For the full show notes, transcript, and links to mentioned content, check out the episode page here.
-----
Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe.
-----
Ramp’s mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus.
-----
Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Visit vanta.com/invest.
-----
WorkOS is a developer platform that enables SaaS companies to quickly add enterprise features to their applications. Visit WorkOS.com to transform your application into an enterprise-ready solution in minutes, not months.
-----
Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest.
-----
Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgelineapps.com.
-----
Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com).
Timestamps:
(00:00:00) Welcome to Invest Like the Best
(00:02:43) Intro: Sergey Levine
(00:03:29) Why Bet on Generality Over Specialization
(00:07:24) What if PI succeeds?
(00:09:05) Pros and Cons of Humanoid Robotics
(00:11:02) Timeline of Major Milestones in Robotics
(00:15:47) Sergey's Personal Journey
(00:18:22) Making General Intelligence Happen
(00:19:57) Understanding Robot Data Collection
(00:22:12) Most Surprising Discovery at Physical Intelligence
(00:24:48) The Science of Common Sense
(00:25:36) Long-Range Tasks in Robotics
(00:27:24) Why Wouldn’t We Have A Robot in Our Kitchen by 2050
(00:31:21) Other Interesting Approaches
(00:32:38) Cool vs. Useful in Robotics
(00:36:48) Form Factor Innovation
(00:38:22) Physical Intelligence Analogy
(00:39:30) Economic Transformation from Robotics
(00:40:48) Controversies in the Robotics Community
(00:42:16) Arguments Against End-to-End Learning
(00:42:34) Compositional Learning Explained
(00:43:25) Last Tasks Robots will Conquer
(00:44:30) Dark Parts of the Robotics Brain
(00:47:05) What Makes a Great Researcher
(00:50:15) Manufacturing and Scale Challenges
(00:51:17) How Companies Should Prepare for Robotics
(00:53:38) Boston Dynamics' Demos
(00:55:43) Converging Technologies Enabling Robotics
(00:56:47) How to Stay Up To Date in Robotics
(00:59:51) Near Term Objectives
(01:00:49) Confidence Level Among Researchers
(01:03:31) Google's Experimentation Culture
(01:04:24) The Kindest Thing
On Christmas Eve, Elon Musk’s X rolled out an in-app tool that lets users alter other people’s photos and post the results directly in reply. With minimal safeguards, it quickly became a pipeline for sexualized, non-consensual deepfakes, including imagery involving minors, delivered straight into victims’ notifications.
Renée DiResta, Hany Farid, and Casey Newton join Kara to dig into the scale of the harm, the failure of app stores and regulators to act quickly, and why the “free speech” rhetoric used to defend the abuse is incoherent. Kara explores what accountability could look like — and what comes next as AI tools get more powerful.
Renée DiResta is the former technical research manager at Stanford's Internet Observatory. She researched online CSAM for years and is one of the world’s leading experts on online disinformation and propaganda. She’s also the author of Invisible Rulers: The People Who Turn Lies into Reality.
Hany Farid is a professor of computer sciences and engineering at the University of California, Berkeley. He’s been described as the father of digital image forensics and has spent years developing tools to combat CSAM.
Casey Newton is the founder of the tech newsletter Platformer and the co-host of The New York Times podcast Hard Fork.
This episode was recorded on Tuesday, January 20th.
When reached for comment, a spokesperson for X referred us to a a statement post on X, which reads in part:
We remain committed to making X a safe platform for everyone and continue to have zero tolerance for any forms of child sexual exploitation, non-consensual nudity, and unwanted sexual content.
We take action to remove high-priority violative content, including Child Sexual Abuse Material (CSAM) and non-consensual nudity, taking appropriate action against accounts that violate our X Rules. We also report accounts seeking Child Sexual Exploitation materials to law enforcement authorities as necessary.
Questions? Comments? Email us at on@voxmedia.com or find us on YouTube, Instagram, TikTok, Threads, and Bluesky @onwithkaraswisher.
Learn more about your ad choices. Visit podcastchoices.com/adchoices
What if everything we think we know about AI understanding is wrong? Is compression the key to intelligence? Or is there something more—a leap from memorization to true abstraction?
In this fascinating conversation, we sit down with **Professor Yi Ma**—world-renowned expert in deep learning, IEEE/ACM Fellow, and author of the groundbreaking new book *Learning Deep Representations of Data Distributions*. Professor Ma challenges our assumptions about what large language models actually do, reveals why 3D reconstruction isn't the same as understanding, and presents a unified mathematical theory of intelligence built on just two principles: **parsimony** and **self-consistency**.
**SPONSOR MESSAGES START**
—
Prolific - Quality data. From real people. For faster breakthroughs.
https://www.prolific.com/?utm_source=mlst
—
cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy
Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst
Submit investment deck: https://cyber.fund/contact?utm_source=mlst
—
**END**
Key Insights:
**LLMs Don't Understand—They Memorize**
Language models process text (*already* compressed human knowledge) using the same mechanism we use to learn from raw data.
**The Illusion of 3D Vision**
Sora and NeRFs etc that can reconstruct 3D scenes still fail miserably at basic spatial reasoning
**"All Roads Lead to Rome"**
Why adding noise is *necessary* for discovering structure.
**Why Gradient Descent Actually Works**
Natural optimization landscapes are surprisingly smooth—a "blessing of dimensionality"
**Transformers from First Principles**
Transformer architectures can be mathematically derived from compression principles
—
INTERACTIVE AI TRANSCRIPT PLAYER w/REFS (ReScript):
https://app.rescript.info/public/share/Z-dMPiUhXaeMEcdeU6Bz84GOVsvdcfxU_8Ptu6CTKMQ
About Professor Yi Ma
Yi Ma is the inaugural director of the School of Computing and Data Science at Hong Kong University and a visiting professor at UC Berkeley.
https://people.eecs.berkeley.edu/~yima/
https://scholar.google.com/citations?user=XqLiBQMAAAAJ&hl=en
https://x.com/YiMaTweets
**Slides from this conversation:**
https://www.dropbox.com/scl/fi/sbhbyievw7idup8j06mlr/slides.pdf?rlkey=7ptovemezo8bj8tkhfi393fh9&dl=0
**Related Talks by Professor Ma:**
- Pursuing the Nature of Intelligence (ICLR): https://www.youtube.com/watch?v=LT-F0xSNSjo
- Earlier talk at Berkeley: https://www.youtube.com/watch?v=TihaCUjyRLM
TIMESTAMPS:
00:00:00 Introduction
00:02:08 The First Principles Book & Research Vision
00:05:21 Two Pillars: Parsimony & Consistency
00:09:50 Evolution vs. Learning: The Compression Mechanism
00:14:36 LLMs: Memorization Masquerading as Understanding
00:19:55 The Leap to Abstraction: Empirical vs. Scientific
00:27:30 Platonism, Deduction & The ARC Challenge
00:35:57 Specialization & The Cybernetic Legacy
00:41:23 Deriving Maximum Rate Reduction
00:48:21 The Illusion of 3D Understanding: Sora & NeRF
00:54:26 All Roads Lead to Rome: The Role of Noise
00:59:56 All Roads Lead to Rome: The Role of Noise
01:00:14 Benign Non-Convexity: Why Optimization Works
01:06:35 Double Descent & The Myth of Overfitting
01:14:26 Self-Consistency: Closed-Loop Learning
01:21:03 Deriving Transformers from First Principles
01:30:11 Verification & The Kevin Murphy Question
01:34:11 CRATE vs. ViT: White-Box AI & Conclusion
REFERENCES:
Book:
[00:03:04] Learning Deep Representations of Data Distributions
https://ma-lab-berkeley.github.io/deep-representation-learning-book/
[00:18:38] A Brief History of Intelligence
https://www.amazon.co.uk/BRIEF-HISTORY-INTELLIGEN-HB-Evolution/dp/0008560099
[00:38:14] Cybernetics
https://mitpress.mit.edu/9780262730099/cybernetics/
Book (Yi Ma):
[00:03:14] 3-D Vision book
https://link.springer.com/book/10.1007/978-0-387-21779-6
<TRUNC> refs on ReScript link/YT
AI Expert STUART RUSSELL, exposes the trillion-dollar AI race, why governments won’t regulate, how AGI could replace humans by 2030, and why only a nuclear-level AI catastrophe will wake us up
Professor Stuart Russell O.B.E. is a world-renowned AI expert and Computer Science Professor at UC Berkeley. He holds the Smith-Zadeh Chair in Engineering and directs the Center for Human-Compatible AI, and is also the bestselling author of the book “Human Compatible: AI and the Problem of Control".
He explains:
◼️What the “gorilla problem” reveals about our future under superintelligent AI
◼️How governments are outfunded by Big Tech
◼️Why current AI systems already lie and self-preserve
◼️The radical solution he’s spent a decade building to make AI safe
◼️The myth of ‘pulling the plug’ and why AI won’t be that easy to stop
[00:00] You've Been Talking About AI for a Long Time
[02:54] You Wrote the Textbook on AI
[03:29] It Will Take a Crisis to Wake People Up
[06:03] CEOs Staying in the AI Race Despite Risks
[08:04] They Know It's an Extinction-Level Risk
[10:06] What Is Artificial General Intelligence (AGI)?
[13:10] Will We Reach General Intelligence Soon?
[16:26] How Much Is Safety Really Being Implemented
[17:29] AI Safety Employees Leaving OpenAI
[18:14] The Gorilla Problem — The Most Intelligent Species Will Always Rule
[19:34] If There's an Extinction Risk, Why Don't They Stop?
[21:02] Can't We Just Pull the Plug if AI Gets Too Powerful?
[22:49] Can We Build AI That Will Act in Our Best Interests?
[24:09] Are You Troubled by the Rapid Advancement of AI?
[26:48] Do You Have Regrets About Your Involvement?
[27:35] No One Actually Understands How This AI Works
[30:36] AI Will Be Able to Train Itself
[32:24] The Fast Takeoff Is Coming
[34:20] Are We Creating Our Successor and Ending the Human Race?
[38:36] Advice to Young People in This New World
[40:52] How Do You Think AI Would Make Us Extinct?
[42:33] The Problem if No One Has to Work
[45:59] What if We Just Entertain Ourselves All Day
[48:43] Why Do We Make Robots Look Like Humans?
[56:44] What Should Young People Be Doing Professionally?
[59:56] What Is It to Be Human?
[01:03:34] The Rise of Individualism
[01:05:34] Ads
[01:06:39] Universal Basic Income
[01:08:41] Would You Press a Button to Stop AI Forever?
[01:15:13] But Won't China Win the AI Race if We Stop?
[01:18:40] Trump's Approach to AI
[01:19:06] What's Causing the Loss in Middle-Class Jobs
[01:21:02] What Will Happen if the UK Doesn't Participate in the AI Race?
[01:23:31] Amazon Replacing Their Workers
[01:29:00] Ads
[01:30:54] Experts Agree on Extinction Risk
[01:38:01] What if Aliens Were Watching Us Right Now
[01:39:35] Can We Make AI Systems That We Can Control?
[01:43:14] Are We Creating a God?
[01:47:32] Could There Have Been Advanced Civilisations Before Us?
[01:48:50] What Can We Do to Help?
[01:50:43] You Wrote the Book on AI — Does It Weigh on You?
[01:58:48] What Do You Value Most in Life?
Follow Stuart:
LinkedIn - https://bit.ly/3Y5fOos
You can purchase “Human Compatible: AI and the Problem of Control", here: https://amzn.to/48eOMkH
The Diary Of A CEO:
◼️Join DOAC circle here - https://doaccircle.com/
◼️Buy The Diary Of A CEO book here - https://smarturl.it/DOACbook
◼️The 1% Diary is back - limited time only: https://bit.ly/3YFbJbt
◼️The Diary Of A CEO Conversation Cards (Second Edition): https://g2ul0.app.link/f31dsUttKKb
◼️Get email updates - https://bit.ly/diary-of-a-ceo-yt
◼️Follow Steven - https://g2ul0.app.link/gnGqL4IsKKb
Sponsors:
Pipedrive - https://pipedrive.com/CEO
Fiverr: https://fiverr.com/diary and get 10% off your first order when you use code DIARY
Stan Store: NO PURCHASE NECESSARY. VOID WHERE PROHIBITED. For Official Rules, visit https://DaretoDream.stan.store
Robots are commonplace in factories, and increasingly in warehouses like those run by Amazon. But what about robots to help with household chores — so-called humanoids to load the dishwasher or fold the laundry?
To find out, we checked in with Ken Goldberg, professor of engineering at UC Berkeley and co-founder of the AI and robotics company Ambi Robotics. He spoke to Marketplace’s Nova Safo en route from a robotics conference in China.
Hamel Husain and Shreya Shankar teach the world’s most popular course on AI evals and have trained over 2,000 PMs and engineers (including many teams at OpenAI and Anthropic). In this conversation, they demystify the process of developing effective evals, walk through real examples, and share practical techniques that’ll help you improve your AI product.
What you’ll learn:
1. WTF evals are
2. Why they’ve become the most important new skill for AI product builders
3. A step-by-step walkthrough of how to create an effective eval
4. A deep dive into error analysis, open coding, and axial coding
5. Code-based evals vs. LLM-as-judge
6. The most common pitfalls and how to avoid them
7. Practical tips for implementing evals with minimal time investment (30 minutes per week after initial setup)
8. Insight into the debate between “vibes” and systematic evals
—
Brought to you by:
Fin—The #1 AI agent for customer service
Dscout—The UX platform to capture insights at every stage: from ideation to production
Mercury—The art of simplified finances
—
Where to find Shreya Shankar
• X: https://x.com/sh_reya
• LinkedIn: https://www.linkedin.com/in/shrshnk/
• Website: https://www.sh-reya.com/
• Maven course: https://bit.ly/4myp27m
—
Where to find Hamel Husain
• X: https://x.com/HamelHusain
• LinkedIn: https://www.linkedin.com/in/hamelhusain/
• Website: https://hamel.dev/
• Maven course: https://bit.ly/4myp27m
—
In this episode, we cover:
(00:00) Introduction to Hamel and Shreya
(04:57) What are evals?
(09:56) Demo: Examining real traces from a property management AI assistant
(16:51) Writing notes on errors
(23:54) Why LLMs can’t replace humans in the initial error analysis
(25:16) The concept of a “benevolent dictator” in the eval process
(28:07) Theoretical saturation: when to stop
(31:39) Using axial codes to help categorize and synthesize error notes
(44:39) The results
(46:06) Building an LLM-as-judge to evaluate specific failure modes
(48:31) The difference between code-based evals and LLM-as-judge
(52:10) Example: LLM-as-judge
(54:45) Testing your LLM judge against human judgment
(01:00:51) Why evals are the new PRDs for AI products
(01:05:09) How many evals you actually need
(01:07:41) What comes after evals
(01:09:57) The great evals debate
(1:15:15) Why dogfooding isn’t enough for most AI products
(01:18:23) OpenAI’s Statsig acquisition
(1:23:02) The Claude Code controversy and the importance of context
(01:24:13) Common misconceptions around evals
(1:22:28) Tips and tricks for implementing evals effectively
(1:30:37) The time investment
(1:33:38) Overview of their comprehensive evals course
(1:37:57) Lightning round and final thoughts
—
LLM Log Open Codes Analysis Prompt:
Please analyze the following CSV file. There is a metadata field which has an nested field called z_note that contains open codes for analysis of LLM logs that we are conducting. Please extract all of the different open codes. From the _note field, propose 5-6 categories that we can create axial codes from.
—
Referenced:
• Building eval systems that improve your AI product: https://www.lennysnewsletter.com/p/building-eval-systems-that-improve
• Mercor: https://mercor.com/
• Brendan Foody on LinkedIn: https://www.linkedin.com/in/brendan-foody-2995ab10b
• Nurture Boss: https://nurtureboss.io/
• Braintrust: https://www.braintrust.dev/
• Andrew Ng on X: https://x.com/andrewyng
• Carrying Out Error Analysis: https://www.youtube.com/watch?v=JoAxZsdw_3w
• Julius AI: https://julius.ai/
• Brendan Foody on X—“evals are the new PRDs”: https://x.com/BrendanFoody/status/1939764763485171948
• Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences: https://dl.acm.org/doi/abs/10.1145/3654777.3676450
• Lenny’s post on X about evals: https://x.com/lennysan/status/1909636749103599729
• Statsig: https://statsig.com/
• Claude Code: https://www.anthropic.com/claude-code
• Cursor: https://cursor.com/
• Occam’s razor: https://en.wikipedia.org/wiki/Occam%27s_razor
• Frozen: https://www.imdb.com/title/tt2294629/
• The Wire on HBO: https://en.wikipedia.org/wiki/The_Wire
—
Recommended books:
• Pachinko: https://www.amazon.com/Pachinko-National-Book-Award-Finalist/dp/1455563935
• Apple in China: The Capture of the World’s Greatest Company: https://www.amazon.com/Apple-China-Capture-Greatest-Company/dp/1668053373/
• Machine Learning: https://www.amazon.com/Machine-Learning-Tom-M-Mitchell/dp/1259096955
• Artificial Intelligence: A Modern Approach: https://www.amazon.com/Artificial-Intelligence-Modern-Approach-Global/dp/1292401133/
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.
—
Lenny may be an investor in the companies discussed.
My biggest takeaways from this conversation:
To hear more, visit www.lennysnewsletter.com
Sergey Levine, one of the world’s top robotics researchers and co-founder of Physical Intelligence, thinks we’re on the cusp of a “self-improvement flywheel” for general-purpose robots. His median estimate for when robots will be able to run households entirely autonomously? 2030.
If Sergey’s right, the world 5 years from now will be an insanely different place than it is today. This conversation focuses on understanding how we get there: we dive into foundation models for robotics, and how we scale both the data and the hardware necessary to enable a full-blown robotics explosion.
Watch on YouTube; listen on Apple Podcasts or Spotify.
Sponsors
* Labelbox provides high-quality robotics training data across a wide range of platforms and tasks. From simple object handling to complex workflows, Labelbox can get you the data you need to scale your robotics research. Learn more at labelbox.com/dwarkesh
* Hudson River Trading uses cutting-edge ML and terabytes of historical market data to predict future prices. I got to try my hand at this fascinating prediction problem with help from one of HRT’s senior researchers. If you’re curious about how it all works, go to hudson-trading.com/dwarkesh
* Gemini 2.5 Flash Image (aka nano banana) isn’t just for generating fun images — it’s also a powerful tool for restoring old photos and digitizing documents. Test it yourself in the Gemini App or in Google’s AI Studio: ai.studio/banana
To sponsor a future episode, visit dwarkesh.com/advertise.
Timestamps
(00:00:00) – Timeline to widely deployed autonomous robots
(00:17:25) – Why robotics will scale faster than self-driving cars
(00:27:28) – How vision-language-action models work
(00:45:37) – Changes needed for brainlike efficiency in robots
(00:57:59) – Learning from simulation
(01:09:18) – How much will robots speed up AI buildouts?
(01:18:01) – If hardware’s the bottleneck, does China win by default?
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
The post-Covid inflation will prove to be a treasure trove for academic economists, as they study what drives inflation, and the power that central banks have to contain it once it gets going. At this year's Jackson Hole Economic Symposium, UC Berkeley professor Emi Nakamura presented a new paper — co-authored with her Berkeley colleagues Jón Steinsson and Venance Riblier — titled Beyond the Taylor Rule. The paper sought to look at the wide range of choices that global central banks made in dealing with inflation to see what if anything could be learned about the Taylor Rule, a load-bearing idea in modern economics that describes what optimal monetary policy looks like when successfully balancing the Federal Reserve's objectives. Their paper discovers that in any bout of inflation, a central bank that has a greater history of fighting inflation also has the ability to deviate further from strict Taylor Rule guidelines, without achieving worse inflation outcomes. In an interview recorded in Jackson Hole, we speak with Professor Nakamura about her work and its implications for central bankers going forward.
Only Bloomberg.com subscribers can get the Odd Lots newsletter in their inbox — now delivered every weekday — plus unlimited access to the site and app. Subscribe at bloomberg.com/subscriptions/oddlots
See omnystudio.com/listener for privacy information.
“ How do you trust anything anymore? Who do you trust? Where do you trust?” asks technologist and digital forensic expert Hany Farid. Following his talk at TED2025, Farid sat down for a special conversation with Elise Hu, host of TED Talks Daily, to discuss the erosion of trust in American society. From TikTok algorithms to AI deepfakes, Farid argues that critical thinking education is more important than ever and why it’s therapeutic to unplug from social media and connect with nature.
Hosted on Acast. See acast.com/privacy for more information.
How do you know if that shocking photo in your feed is real, or just another AI fake? Digital forensics expert Hany Farid explains how he helps journalists, courts and governments find structural errors in AI-generated images, offering four practical tips everyday individuals can use when facing the internet’s war on reality.
Hosted on Acast. See acast.com/privacy for more information.
Hany Farid is a professor of electrical engineering and computer sciences at the University of California, Berkeley. He's been a leading voice on digital forensics for over two decades—pioneering ways to identify if an image, audio or video has been digitally altered. Since the rise of social media, Farid has kept busy helping news organizations, government agencies and law enforcement determine what is real and what is fake online. Farid sits down with Oz to talk about his initial interest in digital forensics, the effects of misinformation on society, and whether he wants an AI likeness of himself to live on after he dies.
See omnystudio.com/listener for privacy information.
Today, we're joined by Sergey Levine, associate professor at UC Berkeley and co-founder of Physical Intelligence, to discuss π0 (pi-zero), a general-purpose robotic foundation model. We dig into the model architecture, which pairs a vision language model (VLM) with a diffusion-based action expert, and the model training "recipe," emphasizing the roles of pre-training and post-training with a diverse mixture of real-world data to ensure robust and intelligent robot learning. We review the data collection approach, which uses human operators and teleoperation rigs, the potential of synthetic data and reinforcement learning in enhancing robotic capabilities, and much more. We also introduce the team’s new FAST tokenizer, which opens the door to a fully Transformer-based model and significant improvements in learning and generalization. Finally, we cover the open-sourcing of π0 and future directions for their research.
The complete show notes for this episode can be found at https://twimlai.com/go/719.
In this episode of The Cognitive Revolution, Nathan explores the groundbreaking paper on obfuscated activations with 3 members from the research team - Luke Bailey, Eric Jenner, and Scott Emmons. The team discusses how their work challenges latent-based defenses in AI systems, demonstrating methods to bypass safety mechanisms while maintaining harmful behaviors. Join us for an in-depth technical conversation about AI safety, interpretability, and the ongoing challenge of creating robust defense systems.
Do check out the "Obfuscated Activations Bypass LLM Latent-Space Defenses" paper here: https://obfuscated-activations.github.io/
Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse
SPONSORS:
Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers. OCI powers industry leaders like Vodafone and Thomson Reuters with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before March 31, 2024 at https://oracle.com/cognitive
NetSuite: Over 41,000 businesses trust NetSuite by Oracle, the #1 cloud ERP, to future-proof their operations. With a unified platform for accounting, financial management, inventory, and HR, NetSuite provides real-time insights and forecasting to help you make quick, informed decisions. Whether you're earning millions or hundreds of millions, NetSuite empowers you to tackle challenges and seize opportunities. Download the free CFO's guide to AI and machine learning at https://netsuite.com/cognitive
Shopify: Dreaming of starting your own business? Shopify makes it easier than ever. With customizable templates, shoppable social media posts, and their new AI sidekick, Shopify Magic, you can focus on creating great products while delegating the rest. Manage everything from shipping to payments in one place. Start your journey with a $1/month trial at https://shopify.com/cognitive and turn your 2025 dreams into reality.
Vanta: Vanta simplifies security and compliance for businesses of all sizes. Automate compliance across 35+ frameworks like SOC 2 and ISO 27001, streamline security workflows, and complete questionnaires up to 5x faster. Trusted by over 9,000 companies, Vanta helps you manage risk and prove security in real time. Get $1,000 off at https://vanta.com/revolution
RECOMMENDED PODCAST:
Check out Modern Relationships where Erik Torenberg interviews tech power couples and leading thinkers to explore how ambitious people actually make partnerships work. This season's guests include: Delian Asparouhov & Nadia Asparouhova, Kristen Berman & Phil Levin, Rob Henderson, and Liv Boeree & Igor Kurganov.
Apple: https://podcasts.apple.com/us/podcast/id1786227593
Spotify: https://open.spotify.com/show/5hJzs0gDg6lRT6r10mdpVg
YouTube: https://www.youtube.com/@ModernRelationshipsPod
CHAPTERS:
(00:00:00) Teaser
(00:00:46) About the Episode
(00:05:11) Latent Space Defenses
(00:08:41) Sleeper Agents
(00:15:06) Three Case Studies (Part 1)
(00:17:02) Sponsors: Oracle Cloud Infrastructure (OCI) | NetSuite
(00:19:42) Three Case Studies (Part 2)
(00:24:09) SQL Generation
(00:26:17) Understanding Defenses
(00:32:52) Out-of-Distribution Detection (Part 1)
(00:35:37) Sponsors: Shopify | Vanta
(00:38:52) Out-of-Distribution Detection (Part 2)
(00:45:13) Loss Function Weighting
(00:57:49) Who Moves Last?
(01:11:41) High-Level Triggers
(01:25:33) Open Source vs. Access
(01:38:57) Internalizing Reasoning
(01:53:07) Representing Concepts
(02:06:38) Final Thoughts
(02:09:33) Outro
Ken Goldberg (Why Don’t We Have Better Robots Yet?) is an award-winning artist, roboticist, and engineering professor. Ken joins the Armchair Expert to discuss being born in Nigeria, growing up in rough and tumble City of Brotherly Love, and on how that taught him how to not take things lying down. Ken and Dax address the elephant panties in the room, how a course he took in 1981 began his trajectory in robotics and AI, and the tragic archetype of Pygmalion and the hubris of falling in love with your creation. Ken explains the Czechoslovakian etymology of the word “robot,” why don’t we have better robots yet, and how he stays optimistic doing a job predicated on failure.
Follow Armchair Expert on the Wondery App or wherever you get your podcasts. Watch new content on YouTube or listen to Armchair Expert early and ad-free by joining Wondery+ in the Wondery App, Apple Podcasts, or Spotify. Start your free trial by visiting wondery.com/links/armchair-expert-with-dax-shepard/ now.
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this episode of Gradient Dissent, Joseph E. Gonzalez, EECS Professor at UC Berkeley and Co-Founder at RunLLM, joins host Lukas Biewald to explore innovative approaches to evaluating LLMs.
They discuss the concept of vibes-based evaluation, which examines not just accuracy but also the style and tone of model responses, and how Chatbot Arena has become a community-driven benchmark for open-source and commercial LLMs. Joseph shares insights on democratizing model evaluation, refining AI-human interactions, and leveraging human preferences to improve model performance. This episode provides a deep dive into the evolving landscape of LLM evaluation and its impact on AI development.
🎙 Get our podcasts on these platforms:
Apple Podcasts: http://wandb.me/apple-podcasts
Spotify: http://wandb.me/spotify
Google: http://wandb.me/gd_google
YouTube: http://wandb.me/youtube
Follow Weights & Biases:
https://twitter.com/weights_biases
https://www.linkedin.com/company/wandb
Join the Weights & Biases Discord Server:
https://discord.gg/CkZKRNnaf3
Today, we're joined by Shreya Shankar, a PhD student at UC Berkeley to discuss DocETL, a declarative system for building and optimizing LLM-powered data processing pipelines for large-scale and complex document analysis tasks. We explore how DocETL's optimizer architecture works, the intricacies of building agentic systems for data processing, the current landscape of benchmarks for data processing tasks, how these differ from reasoning-based benchmarks, and the need for robust evaluation methods for human-in-the-loop LLM workflows. Additionally, Shreya shares real-world applications of DocETL, the importance of effective validation prompts, and building robust and fault-tolerant agentic systems. Lastly, we cover the need for benchmarks tailored to LLM-powered data processing tasks and the future directions for DocETL.
The complete show notes for this episode can be found at https://twimlai.com/go/703.
In this episode of the Eye on AI podcast, we dive into the world of AI forecasting with Danny Halawi, a PhD student at UC Berkeley.
Danny shares his groundbreaking research on using large language models (LLMs) to predict future events with accuracy, rivaling human forecasters and prediction markets.
Danny recounts his journey from studying computer security and fraud detection to exploring the potential of AI in forecasting geopolitical events and beyond. He introduces us to the sophisticated architecture of his AI system, which leverages real-time data from prediction markets and advanced machine learning techniques to generate reliable forecasts.
We explore the complexities of judgmental forecasting versus time series forecasting and how AI can enhance decision-making in various fields. Danny discusses the challenges of training AI models for high-stakes predictions, the role of super forecasters, and the fascinating dynamics of prediction markets. He also sheds light on the ethical considerations and future possibilities of integrating AI into our decision-making processes.
Join us as we delve into the future of AI forecasting, the potential of superhuman predictions, and the exciting developments that could reshape our understanding of the future. Don't forget to like, subscribe, and hit the notification bell for more expert insights into the latest AI innovations.
This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.
If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(01:28) The Importance of AI
(02:50) Danny's Background and Interest in AI
(04:01) Automated AI Forecasting and Safety Implications
(07:34) Judgmental Forecasting Explained
(11:01) Accuracy and Challenges in Prediction Markets
(16:01) Aggregating Predictions for Better Accuracy
(19:25) Data Collection and Model Accuracy
(23:18) Improving Model Accuracy Over Time
(25:31) Data Sources and Model Training
(29:20) Summarizing Information for Predictions
(34:08) Potential of Reinforcement Learning in Forecasting
(37:50) Automating Information Collection and Summarization
(39:01) Training the Model for Accurate Predictions
(45:04) Challenges with Uncertain Predictions
(50:14) Potential Applications and Future Directions
(52:26) The Future of AI Forecasting and Its Impact
Today, we're joined by Amir Bar, a PhD candidate at Tel Aviv University and UC Berkeley to discuss his research on visual-based learning, including his recent paper, “EgoPet: Egomotion and Interaction Data from an Animal’s Perspective.” Amir shares his research projects focused on self-supervised object detection and analogy reasoning for general computer vision tasks. We also discuss the current limitations of caption-based datasets in model training, the ‘learning problem’ in robotics, and the gap between the capabilities of animals and AI systems. Amir introduces ‘EgoPet,’ a dataset and benchmark tasks which allow motion and interaction data from an animal's perspective to be incorporated into machine learning models for robotic planning and proprioception. We explore the dataset collection process, comparisons with existing datasets and benchmark tasks, the findings on the model performance trained on EgoPet, and the potential of directly training robot policies that mimic animal behavior.
The complete show notes for this episode can be found at https://twimlai.com/go/692.
Why hasn't the dream of having a robot at home to do your chores become a reality yet? With three decades of research expertise in the field, roboticist Ken Goldberg sheds light on the clumsy truth about robots — and what it will take to build more dexterous machines to work in a warehouse or help out at home.
Hosted on Acast. See acast.com/privacy for more information.
Today we’re joined by Sherry Yang, senior research scientist at Google DeepMind and a PhD student at UC Berkeley. In this interview, we discuss her new paper, "Video as the New Language for Real-World Decision Making,” which explores how generative video models can play a role similar to language models as a way to solve tasks in the real world. Sherry draws the analogy between natural language as a unified representation of information and text prediction as a common task interface and demonstrates how video as a medium and generative video as a task exhibit similar properties. This formulation enables video generation models to play a variety of real-world roles as planners, agents, compute engines, and environment simulators. Finally, we explore UniSim, an interactive demo of Sherry's work and a preview of her vision for interacting with AI-generated environments.
The complete show notes for this episode can be found at twimlai.com/go/676.
Join host Craig Smith on episode #176 of Eye on AI as he dives deep into the realm of robotic artificial intelligence with Sergey Levine, associate professor in the Department of Electrical Engineering and Computer Sciences at UC Berkeley.
In this episode, Sergey unveils the latest advancements in AI control of robots, exploring the implications of reinforcement learning and the concept of embodied AI.
Discover how Sergey's research is pushing the boundaries of AI, enabling robots to learn manipulation skills and generalize across diverse tasks, transforming the potential of home robots and beyond.
Sergey also shares insights into the RTX project, an ambitious collaboration designed to achieve remarkable generalization across different robot morphologies, enhancing robots' ability to perform language-conditioned manipulation tasks.
If you're fascinated by the intersection of AI, robotics, and the quest for creating adaptable, generalizable machines that promise to revolutionize our interaction with technology, this episode is a must-listen.
Remember to rate us on Apple Podcast and Spotify if this episode ignites your interest in the dynamic field of robotic AI and the visionary work of Sergey Levine.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction to Home Robots and AI in Robotics
(01:43) World Models and Language Models in Robotics
(04:01) The Challenge of Learning-Based Control and Data Utilization
(06:05) RTX Project: Generalizing Controllers Across Different Robots
(10:09) Uniformity in Model Architecture Across Labs
(13:50) Introduction of RT1 and RT2 Models for Robot Control
(16:06) The Future of Robotic Control Research and Architecture
(18:49) The Impact of Hardware Development on Robotics
(22:15) Advances in Controller and AI Model Development
(26:21) Planning and Acting with Vision Language Models
(31:38) The Proprietary vs. Open-Source Debate in Robotics
(36:23) The Future of Commercial and Open Source Robotics Applications
(40:59) The State of Robotics Research in China
Gašper’s work combines machine learning, statistical modeling, neuroimaging, and behavioral experiments “to better understand how neural networks learn internal representations in speech and how humans learn to speak.”
One thing that surprised him about generative adversarial networks (GANs)? How innovative they are, capable of generating English words they’ve never heard before based on words they have.
Read about how AI is restoring a stroke survivor’s ability to speak.
Universal grammar proposes a hypothetical structure in the brain responsible for humans’ innate language abilities. The concept is credited to the famous linguist Noam Chomsky; read his take on GenAI.
AI expert Yoshua Bengio recently signed an open letter asking AI labs to pause the training of AI systems powerful enough to pass the Turing test. Read about his reasoning.
Find the Berkeley Speech and Communication Network here.
Find Gašper on his website, Twitter, and LinkedIn. Or dive into his research.
Congratulations to Lifeboat badge winner and self-proclaimed data nerd John Rotenstein, who saved How can I delete files older than seven days in Amazon S3? from the ignominy of ignorance.
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jitendra Malik, Professor of EECS at UC Berkeley discusses with host Pieter Abbeel building AI from the ground-up and sensorimotor before language.
Subscribe to the Robot Brains Podcast today | Visit therobotbrains.ai and follow us on YouTube at TheRobotBrainsPodcast and Twitter at @pabbeel.
Hosted on Acast. See acast.com/privacy for more information.
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more. Download NetSuite's popular KPI Checklist, designed to give you consistently excellent performance - absolutely free, at NetSuite.com/EYEONAI.
On episode #133 of the Eye on AI podcast, Craig Smith sits down with Michael Jordan. A revered scientist and distinguished professor at the University of California, Berkeley, Michael's expertise spans machine learning, statistics, and artificial intelligence.
In this episode we explore the intricate landscape of deep learning, its statistical bedrock, and the myriad applications of machine learning methods. We dig deeper into the cloud computing revolution, ignited by deep learning, which has empowered behemoths like Amazon to streamline their logistics and commerce data.
The dialogue continues as we uncover the trends in deep learning, its role in the grand scheme of AI, and the persisting challenges in the domain.
We conclude by contemplating the repercussions of AI and optimization in complex systems. We examine its historical roots in the mid-20th century, its potential to replace jobs, and its application in various sectors such as financial markets, healthcare, and education.
(00:00) Preview
(01:00) Introduction
(01:36) NetSuite by Oracle
(03:53) Deep Learning Advancements in AI
(12:56) AI Optimization in Complex Systems
(28:53) Future of Machine Learning in Healthcare
(39:45) Privacy and Value in Multi-Agent Learning
(54:48) Learning Systems, Avatars, and Music Business
(1:02:20) Single Source of Truth for Business Owners
(1:04:28) NetSuite by Oracle
Craig Smith Twitter: https://twitter.com/craigss Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
Dr. Raluca Ada Popa, renowned computer scientist, entrepreneur, and President of Opaque Systems, joins Jon Krohn to share her insights on securely interacting with AI APIs like OpenAI's GPT-4, the pros and cons of open vs. closed-source AI development, and the seamless operation of compute pipelines across multiple clouds.
This episode is brought to you by AWS Inferentia and by Modelbit, for deploying models in seconds. Interested in sponsoring a SuperDataScience Podcast episode? Visit JonKrohn.com/podcast for sponsorship information.
In this episode you will learn:
• What is a confidential computing platform? [04:31]
• How to get started with confidential computing [12:10]
• The challenges of confidential computing and LLMs [21:11]
• How to safeguard your data while using commercial LLMs like GPT-4 [38:00]
• Open-source vs closed-source [52:28]
• Raluca's PreVail cybersecurity company [1:01:50]
• Combining entrepreneurship and academic career [1:04:03]
• DARE Program [1:10:39]
Additional materials: www.superdatascience.com/701