Asimov: Building An Omniscient RL Oracle with ReflectionAI’s Misha Laskin
Superintelligence, at least in an academic sense, has already been achieved. But Misha Laskin thinks that the next step towards artificial superintelligence, or ASI, should look both more user and problem-focused. ReflectionAI co-founder and CEO Misha Laskin joins Sarah Guo to introduce Asimov, their new code comprehension agent built on reinforcement learning (RL). Misha talks about creating tools and designing AI agents based on customer needs, and how that influences eval development and the scope of the agent’s memory. The two also discuss the challenges in solving scaling for RL, the future of ASI, and the implications for Google’s “non-acquisition” of Windsurf.
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @MishaLaskin | @reflection_ai
Chapters:
00:00 – Misha Laskin Introduction
00:44 – Superintelligence vs. Super Intelligent Autonomous Systems
03:26 – Misha’s Journey from Physics to AI
07:48 – Asimov Product Release
11:52 – What Differentiates Asimov from Other Agents
16:15 – Asimov’s Eval Philosophy
21:52 – The Types of Queries Where Asimov Shines
24:35 – Designing a Team-Wide Memory for Asimov
28:38 – Leveraging Pre-Trained Models
32:47 – The Challenges of Solving Scaling in RL
37:21 – Training Agents in Copycat Software Environments
38:25 – When Will We See ASI?
44:27 – Thoughts on Windsurf’s Non-Acquisition
48:10 – Exploring Non-RL Datasets
55:12 – Tackling Problems Beyond Engineering and Coding
57:54 – Where We’re At in Deploying ASI in Different Fields
01:02:30 – Conclusion
Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs
LLMs are democratizing digital intelligence, but we’re all waiting for AI agents to take this to the next level by planning tasks and executing actions to actually transform the way we work and live our lives.
Yet despite incredible hype around AI agents, we’re still far from that “tipping point” with best in class models today. As one measure: coding agents are now scoring in the high-teens % on the SWE-bench benchmark for resolving GitHub issues, which far exceeds the previous unassisted baseline of 2% and the assisted baseline of 5%, but we’ve still got a long way to go.
Why is that? What do we need to truly unlock agentic capability for LLMs? What can we learn from researchers who have built both the most powerful agents in the world, like AlphaGo, and the most powerful LLMs in the world?
To find out, we’re talking to Misha Laskin, former research scientist at DeepMind. Misha is embarking on his vision to build the best agent models by bringing the search capabilities of RL together with LLMs at his new company, Reflection AI. He and his cofounder Ioannis Antonoglou, co-creator of AlphaGo and AlphaZero and RLHF lead for Gemini, are leveraging their unique insights to train the most reliable models for developers building agentic workflows.
Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital
00:00 Introduction
01:11 Leaving Russia, discovering science
10:01 Getting into AI with Ioannis Antonoglou
15:54 Reflection AI and agents
25:41 The current state of Ai agents
29:17 AlphaGo, AlphaZero and Gemini
32:58 LLMs don’t have a ground truth reward
37:53 The importance of post-training
44:12 Task categories for agents
45:54 Attracting talent
50:52 How far away are capable agents?
56:01 Lightning round
Mentioned:
The Feynman Lectures on Physics: The classic text that got Misha interested in science.
Mastering the game of Go with deep neural networks and tree search: The original 2016 AlphaGo paper.
Mastering the game of Go without human knowledge: 2017 AlphaGo Zero paper
Scaling Laws for Reward Model Overoptimization: OpenAI paper on how reward models can be gamed at all scales for all algorithms.
Mapping the Mind of a Large Language Model: Article about Anthropic mechanistic interpretability paper that identifies how millions of concepts are represented inside Claude Sonnet
Pieter Abeel: Berkeley professor and founder of Covariant who Misha studied with
A2C and A3C: Advantage Actor Critic and Asynchronous Advantage Actor Critic, the two algorithms developed by Misha’s manager at DeepMind, Volodymyr Mnih, that defined reinforcement learning and deep reinforcement learning