Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated. Noam explains how, if properly scaffolded, today’s models can reason for weeks or even months on complex tasks. He also discusses real-world implications of test-time compute, from building poker solver bots to disproving legendary math conjectures. Together, they also unpack the large gaps in current AI safety frameworks, explore the bottlenecks for recursive self-improvement, and look ahead at the future of multi-agent collaboration and global knowledge sharing.
Read more: Implications of Large-Scale Test-Time Compute
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @polynoamial | @OpenAI
Chapters:
00:00 – Cold Open
00:43 – Noam Brown Introduction
01:23 – Why Benchmarks Are Broken
04:19 – Compute Budgets and Projections
05:34 – How Long Should Models Think?
06:47 – Benchmark-Maxxing
08:34 – Using Poker Bots as Evals
11:26 – Safety Evals When Model Capability Scales With Budget
14:41 – Release Cycle vs. Agent Runtime
17:06 – Latent Model Capability
20:59 – Limits on Recursive Self-Improvement
27:09 – Large-Scale Multi-Agent Coordination
29:11 – Competition at the Frontier
31:51 – Breaking the Benchmark Grid Equilibrium
33:29 – Why Benchmarks Should be Evaluated by Cost
36:18 – Conclusion
OpenAI’s IMO Team on Why Models Are Finally Solving Elite-Level Math
In just two months, a scrappy three-person team at OpenAI sprinted to fulfill what the entire AI field has been chasing for years—gold-level performance on the International Mathematical Olympiad problems. Alex Wei, Sheryl Hsu and Noam Brown discuss their unique approach using general-purpose reinforcement learning techniques on hard-to-verify tasks rather than formal verification tools. The model showed surprising self-awareness by admitting it couldn’t solve problem six, and revealed the humbling gap between solving competition problems and genuine mathematical research breakthroughs.
Hosted by Sonya Huang, Sequoia Capital
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Solving Poker and Diplomacy, Debating RL+Reasoning with Ilya, what’s *wrong* with the System 1/2 analogy, and where Test-Time Compute hits a wall
Full Video Episode
Timestamps
00:00 Intro – Diplomacy, Cicero & World Championship 02:00 Reverse Centaur: How AI Improved Noam’s Human Play 05:00 Turing Test Failures in Chat: Hallucinations & Steerability 07:30 Reasoning Models & Fast vs. Slow Thinking Paradigm 11:00 System 1 vs. System 2 in Visual Tasks (GeoGuessr, Tic-Tac-Toe) 14:00 The Deep Research Existence Proof for Unverifiable Domains 17:30 Harnesses, Tool Use, and Fragility in AI Agents 21:00 The Case Against Over-Reliance on Scaffolds and Routers 24:00 Reinforcement Fine-Tuning and Long-Term Model Adaptability 28:00 Ilya’s Bet on Reasoning and the O-Series Breakthrough 34:00 Noam’s Dev Stack: Codex, Windsurf & AGI Moments 38:00 Building Better AI Developers: Memory, Reuse, and PR Reviews 41:00 Multi-Agent Intelligence and the “AI Civilization” Hypothesis 44:30 Implicit World Models and Theory of Mind Through Scaling 48:00 Why Self-Play Breaks Down Beyond Go and Chess 54:00 Designing Better Benchmarks for Fuzzy Tasks 57:30 The Real Limits of Test-Time Compute: Cost vs. Time 1:00:30 Data Efficiency Gaps Between Humans and LLMs 1:03:00 Training Pipeline: Pretraining, Midtraining, Posttraining 1:05:00 Games as Research Proving Grounds: Poker, MTG, Stratego 1:10:00 Closing Thoughts – Five-Year View and Open Research Directions
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
AI won't plateau — if we give it time to think | Noam Brown
To get smarter, traditional AI models rely on exponential increases in the scale of data and computing power. Noam Brown, a leading research scientist at OpenAI, presents a potentially transformative shift in this paradigm. He reveals his work on OpenAI's new o1 model, which focuses on slower, more deliberate reasoning — much like how humans think — in order to solve complex problems.
Hosted on Acast. See acast.com/privacy for more information.
OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason Better
Combining LLMs with AlphaGo-style deep reinforcement learning has been a holy grail for many leading AI labs, and with o1 (aka Strawberry) we are seeing the most general merging of the two modes to date. o1 is admittedly better at math than essay writing, but it has already achieved SOTA on a number of math, coding and reasoning benchmarks.
Deep RL legend and now OpenAI researcher Noam Brown and teammates Ilge Akkaya and Hunter Lightman discuss the ah-ha moments on the way to the release of o1, how it uses chains of thought and backtracking to think through problems, the discovery of strong test-time compute scaling laws and what to expect as the model gets better.
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital
Mentioned in this episode:
Learning to Reason with LLMs: Technical report accompanying the launch of OpenAI o1.
Generator verifier gap: Concept Noam explains in terms of what kinds of problems benefit from more inference-time compute.
Agent57: Outperforming the human Atari benchmark, 2020 paper where DeepMind demonstrated “the first deep reinforcement learning agent to obtain a score that is above the human baseline on all 57 Atari 2600 games.”
Move 37: Pivotal move in AlphaGo’s second game against Lee Sedol where it made a move so surprising that Sedol thought it must be a mistake, and only later discovered he had lost the game to a superhuman move.
IOI competition: OpenAI entered o1 into the International Olympiad in Informatics and received a Silver Medal.
System 1, System 2: The thesis if Danial Khaneman’s pivotal book of behavioral economics, Thinking, Fast and Slow, that positied two distinct modes of thought, with System 1 being fast and instinctive and System 2 being slow and rational.
AlphaZero: The predecessor to AlphaGo which learned a variety of games completely from scratch through self-play. Interestingly, self-play doesn’t seem to have a role in o1.
Solving Rubik’s Cube with a robot hand: Early OpenAI robotics paper that Ilge Akkaya worked on.
The Last Question: Science fiction story by Isaac Asimov with interesting parallels to scaling inference-time compute.
Strawberry: Why?
O1-mini: A smaller, more efficient version of 1 for applications that require reasoning without broad world knowledge.
00:00 - Introduction
01:33 - Conviction in o1
04:24 - How o1 works
05:04 - What is reasoning?
07:02 - Lessons from gameplay
09:14 - Generation vs verification
10:31 - What is surprising about o1 so far
11:37 - The trough of disillusionment
14:03 - Applying deep RL
14:45 - o1’s AlphaGo moment?
17:38 - A-ha moments
21:10 - Why is o1 good at STEM?
24:10 - Capabilities vs usefulness
25:29 - Defining AGI
26:13 - The importance of reasoning
28:39 - Chain of thought
30:41 - Implication of inference-time scaling laws
35:10 - Bottlenecks to scaling test-time compute
38:46 - Biggest misunderstanding about o1?
41:13 - o1-mini
42:15 - How should founders think about o1?
Noam Brown: from Open AI on solving Poker and Diplomacy with AI
Noam Brown joins host Pieter Abbeel to discuss solving poker and Diplomacy with AI.
Subscribe to the Robot Brains Podcast today | Visit therobotbrains.ai and follow us on YouTube at TheRobotBrainsPodcast and Twitter @therobotbrains.
Hosted on Acast. See acast.com/privacy for more information.
The bot Cicero can collaborate, scheme and build trust with humans. What does this mean for the next frontier of AI? With Noam Brown, Research Scientist at Meta
AGI can beat top players in chess, poker, and, now, Diplomacy. In November 2022, a bot named Cicero demonstrated mastery in this game, which requires natural language negotiation and cooperation with humans. In short, Cicero can lie, scheme, build trust, pass as human, and ally with humans. So what does that mean for the future of AGI?
This week’s guest is research scientist Noam Brown. He co-created Cicero on the Meta Fundamental AI Research Team, and is considered one of the smartest engineers and researchers working in AI today.
Co-hosts Sarah Guo and Elad Gil talk to Noam about why all research should be high risk, high reward, the timeline until we have AGI agents negotiating with humans, why scaling isn’t the only path to breakthroughs in AI, and if the Turing Test is still relevant.
Show Links:
More about Noam Brown
Read the research article about Cicero (diplomacy) published in Science.
Read the research article about Liberatus (heads-up poker) published in Science.
Read the research article about Pluribus (multiplayer poker) published in Science.
Watch the AlphaGo Documentary.
Read “How Smart Are the Robots Getting?” by New York Times reporter Cade Metz
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @Polynoamial
Show Notes:
[01:43] - What sparked Noam’s interest in researching AI that could defeat games
[6:00] - How the AlexaNET and AlphaGo changed the landscape of AI research
[8:09] - Why Noam chose Diplomacy as the next game to work on after poker
[9:51] - What Diplomacy is and why the game was so challenging for an AI bot
[14:50] - Algorithmic breakthroughs and significance of AI bots that win in No-Limit Texas Hold'em poker
[23:29] - The Nash Equilibrium and optimal play in poker
[24:53] - How Cicero interacted with humans
[27:58] - The relevance and usefulness of the Turing Test
[31:05] - The data set used to train Cicero
[31:54] - Bottlenecks to AI researchers and challenges with scaling
[40:10] - The next frontier in researching games for AI
[42:55] - Domains that humans will still dominate and applications for AI bots in the real world
[48:13] - Reasoning challenges with AI
#344 – Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation
Noam Brown is a research scientist at FAIR, Meta AI, co-creator of AI that achieved superhuman level performance in games of No-Limit Texas Hold’em and Diplomacy. Please support this podcast by checking out our sponsors:
– True Classic Tees: https://trueclassictees.com/lex and use code LEX to get 25% off
– Audible: https://audible.com/lex to get 30-day free trial
– InsideTracker: https://insidetracker.com/lex to get 20% off
– ExpressVPN: https://expressvpn.com/lexpod to get 3 months free
EPISODE LINKS:
Noam’s Twitter: https://twitter.com/polynoamial
Noam’s LinkedIn: https://www.linkedin.com/in/noam-brown-8b785b62/
webDiplomacy: https://webdiplomacy.net/
Noam’s papers:
Superhuman AI for multiplayer poker: https://par.nsf.gov/servlets/purl/10119653
Superhuman AI for heads-up no-limit poker: https://par.nsf.gov/servlets/purl/10077416
Human-level play in the game of Diplomacy: https://www.science.org/doi/10.1126/science.ade9097
PODCAST INFO:
Podcast website: https://lexfridman.com/podcast
Apple Podcasts: https://apple.co/2lwqZIr
Spotify: https://spoti.fi/2nEwCF8
RSS: https://lexfridman.com/feed/podcast/
YouTube Full Episodes: https://youtube.com/lexfridman
YouTube Clips: https://youtube.com/lexclips
SUPPORT & CONNECT:
– Check out the sponsors above, it’s the best way to support this podcast
– Support on Patreon: https://www.patreon.com/lexfridman
– Twitter: https://twitter.com/lexfridman
– Instagram: https://www.instagram.com/lexfridman
– LinkedIn: https://www.linkedin.com/in/lexfridman
– Facebook: https://www.facebook.com/lexfridman
– Medium: https://medium.com/@lexfridman
OUTLINE:
Here’s the timestamps for the episode. On some podcast players you should be able to click the timestamp to jump to that time.
(00:00) – Introduction
(05:37) – No Limit Texas Hold ’em
(09:30) – Solving poker
(22:40) – Poker vs Chess
(29:18) – AI playing poker
(1:02:46) – Heads-up vs Multi-way poker
(1:13:37) – Greatest poker player of all time
(1:17:10) – Diplomacy game
(1:27:01) – AI negotiating with humans
(2:09:26) – AI in geopolitics
(2:14:11) – Human-like AI for games
(2:20:12) – Ethics of AI
(2:24:26) – AGI
(2:28:25) – Advice to beginners
569: A.I. For Crushing Humans at Poker and Board Games
Research Scientist at Meta AI, Dr. Noam Brown, joins Jon Krohn to discuss his award-winning no-limit poker-playing algorithms and the real-world implications of his game-playing A.I. breakthroughs.
In this episode you will learn:
What Meta A.I. is and how it fits into Meta, the company [3:01]
Noam's award-winning no-limit poker-playing algorithms, Libratus and Pluribus algorithms. [4:33]
What game theory is and how does Noam integrate it into his models? [8:45]
The real-world implications of Noam’s game-playing A.I. breakthroughs [25:24]
Why Noam elected to become a researcher at a big tech firm instead of in academia [27:06]
The main barriers to getting AI game theory techniques beyond games to self-driving cars [30:16]
Recommendations for people who want to break into poker AI [37:45]
Additional materials: www.superdatascience.com/569