ARC Prize v2 Launch! (Francois Chollet and Mike Knoop)
We are joined by Francois Chollet and Mike Knoop, to launch the new version of the ARC prize! In version 2, the challenges have been calibrated with humans such that at least 2 humans could solve each task in a reasonable task, but also adversarially selected so that frontier reasoning models can't solve them. The best LLMs today get negligible performance on this challenge.
https://arcprize.org/
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Goto https://tufalabs.ai/
***
TRANSCRIPT:
https://www.dropbox.com/scl/fi/0v9o8xcpppdwnkntj59oi/ARCv2.pdf?rlkey=luqb6f141976vra6zdtptv5uj&dl=0
TOC:
1. ARC v2 Core Design & Objectives
[00:00:00] 1.1 ARC v2 Launch and Benchmark Architecture
[00:03:16] 1.2 Test-Time Optimization and AGI Assessment
[00:06:24] 1.3 Human-AI Capability Analysis
[00:13:02] 1.4 OpenAI o3 Initial Performance Results
2. ARC Technical Evolution
[00:17:20] 2.1 ARC-v1 to ARC-v2 Design Improvements
[00:21:12] 2.2 Human Validation Methodology
[00:26:05] 2.3 Task Design and Gaming Prevention
[00:29:11] 2.4 Intelligence Measurement Framework
3. O3 Performance & Future Challenges
[00:38:50] 3.1 O3 Comprehensive Performance Analysis
[00:43:40] 3.2 System Limitations and Failure Modes
[00:49:30] 3.3 Program Synthesis Applications
[00:53:00] 3.4 Future Development Roadmap
REFS:
[00:00:15] On the Measure of Intelligence, François Chollet
https://arxiv.org/abs/1911.01547
[00:06:45] ARC Prize Foundation, François Chollet, Mike Knoop
https://arcprize.org/
[00:12:50] OpenAI o3 model performance on ARC v1, ARC Prize Team
https://arcprize.org/blog/oai-o3-pub-breakthrough
[00:18:30] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al.
https://arxiv.org/abs/2201.11903
[00:21:45] ARC-v2 benchmark tasks, Mike Knoop
https://arcprize.org/blog/introducing-arc-agi-public-leaderboard
[00:26:05] ARC Prize 2024: Technical Report, Francois Chollet et al.
https://arxiv.org/html/2412.04604v2
[00:32:45] ARC Prize 2024 Technical Report, Francois Chollet, Mike Knoop, Gregory Kamradt
https://arxiv.org/abs/2412.04604
[00:48:55] The Bitter Lesson, Rich Sutton
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
[00:53:30] Decoding strategies in neural text generation, Sina Zarrieß
https://www.mdpi.com/2078-2489/12/9/355/pdf
R1, OpenAI’s o3, and the ARC-AGI Benchmark: Insights from Mike Knoop
In this episode of Gradient Dissent, host Lukas Biewald sits down with Mike Knoop, Co-founder and CEO of Ndea, a cutting-edge AI research lab. Mike shares his journey from building Zapier into a major automation platform to diving into the frontiers of AI research. They discuss DeepSeek’s R1, OpenAI’s O-series models, and the ARC Prize, a competition aimed at advancing AI’s reasoning capabilities. Mike explains how program synthesis and deep learning must merge to create true AGI, and why he believes AI reliability is the biggest hurdle for automation adoption.
This conversation covers AGI timelines, research breakthroughs, and the future of intelligent systems, making it essential listening for AI enthusiasts, researchers, and entrepreneurs.
Mentioned Show Notes:
https://ndea.com
https://arcprize.org/blog/r1-zero-r1-results-analysis
https://arcprize.org/blog/oai-o3-pub-breakthrough
🎙 Get our podcasts on these platforms:
Apple Podcasts: http://wandb.me/apple-podcasts
Spotify: http://wandb.me/spotify
Google: http://wandb.me/gd_google
YouTube: http://wandb.me/youtube
Connect with Mike Knoop"
@mikeknoop
Follow Weights & Biases:
https://twitter.com/weights_biases
https://www.linkedin.com/company/wandb
Join the Weights & Biases Discord Server:
https://discord.gg/CkZKRNnaf3
The ARC Prize: Efficiency, Intuition, and AGI, with Mike Knoop, co-founder of Zapier
Nathan interviews Mike Knoop, co-founder of Zapier and co-creator of the ARC Prize, about the $1 million competition for more efficient AI architectures. They discuss the ARC AGI benchmark, its implications for general intelligence, and the potential impact on AI safety. Nathan reflects on the challenges of intuitive problem-solving in AI and considers hybrid approaches to AGI development.
Apply to join over 400 founders and execs in the Turpentine Network: https://hmplogxqz0y.typeform.com/to/JCkphVqj
RECOMMENDED PODCAST
🎙️ Second Opinion - A new podcast for health-tech insiders from Christina Farr of the Second Opinion newsletter. Join Christina Farr, Luba Greenwood, and Ash Zenooz every week as they challenge industry experts with tough questions about the best bets in health-tech.
Apple Podcasts: https://podcasts.apple.com/us/podcast/id1759267211
Spotify: https://open.spotify.com/show/0A8NwQE976s32zdBbZw6bv
-
🎙️ History 102 with WhatifAltHist
Every week, creator of WhatifAltHist Rudyard Lynch and Erik Torenberg cover a major topic in history in depth -- in under an hour. This season will cover classical Greece, early America, the Vikings, medieval Islam, ancient China, the fall of the Roman Empire, and more.
Subscribe on Spotify: https://open.spotify.com/show/36Kqo3BMMUBGTDo1IEYihm
Apple: https://podcasts.apple.com/us/podcast/history-102-with-whatifalthists-rudyard-lynch-and/id1730633913
YouTube: https://www.youtube.com/@History102-qg5oj
SPONSORS:
Oracle Cloud Infrastructure (OCI) is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds; offers one consistent price, and nobody does data better than Oracle. If you want to do more and spend less, take a free test drive of OCI at https://oracle.com/cognitive
The Brave search API can be used to assemble a data set to train your AI models and help with retrieval augmentation at the time of inference. All while remaining affordable with developer first pricing, integrating the Brave search API into your workflow translates to more ethical data sourcing and more human representative data sets. Try the Brave search API for free for up to 2000 queries per month at https://bit.ly/BraveTCR
Omneky is an omnichannel creative generation platform that lets you launch hundreds of thousands of ad iterations that actually work customized across all platforms, with a click of a button. Omneky combines generative AI and real-time advertising data. Mention "Cog Rev" for 10% off https://www.omneky.com/
Head to Squad to access global engineering without the headache and at a fraction of the cost: head to https://choosesquad.com/ and mention “Turpentine” to skip the waitlist.
CHAPTERS:
(00:00:00) About the Show
(00:06:06) The ARC Benchmark
(00:09:34) Other Benchmarks
(00:10:58) Definition of AGI
(00:14:38) The rules of the contest
(00:18:16) ARC test set (Part 1)
(00:18:23) Sponsors: Oracle | Brave
(00:20:31) ARC test set (Part 2)
(00:22:50) Stair-stepping benchmarks
(00:26:17) ARC Prize
(00:28:34) The rules of the ARC Prize
(00:31:12) Compute costs (Part 1)
(00:34:47) Sponsors: Omneky | Squad
(00:36:34) Compute costs (Part 2)
(00:51:20) Intuition
(00:54:32) Human Intelligence
(00:56:06) Current Frontier Language Models
(00:57:44) Program Synthesis
(01:04:10) Is the model learning or memorizing?
(01:15:02) Exploring Solutions
(01:17:02) Non-backpropagation evolutionary architecture search
(01:19:49) Expectations for an AGI world
(01:24:11) Reliability and out of domain generalization
(01:28:35) What a person would do
(01:29:51) What is the right generalization
(01:35:32) The ARC AGI Challenge
(01:48:32) FunSearch
(01:50:41) Kolmogorov-Arnold-Networks
(01:54:18) Grokking
(01:55:42) Outro
Zapier’s Mike Knoop launches ARC Prize to Jumpstart New Ideas for AGI
As impressive as LLMs are, the growing consensus is that language, scale and compute won’t get us to AGI. Although many AI benchmarks have quickly achieved human-level performance, there is one eval that has barely budged since it was created in 2019.
Google researcher François Chollet wrote a paper that year defining intelligence as skill-acquisition efficiency—the ability to learn new skills as humans do, from a small number of examples. To make it testable he proposed a new benchmark, the Abstraction and Reasoning Corpus (ARC), designed to be easy for humans, but hard for AI. Notably, it doesn’t rely on language.
Zapier co-founder Mike Knoop read Chollet’s paper as the LLM wave was rising. He worked quickly to integrate generative AI into Zapier’s product, but kept coming back to the lack of progress on the ARC benchmark. In June, Knoop and Chollet launched the ARC Prize, a public competition offering more than $1M to beat and open-source a solution to the ARC-AGI eval.
In this episode Mike talks about the new ideas required to solve ARC, shares updates from the first two weeks of the competition, and shares why he’s excited for AGI systems that can innovate alongside humans.
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital
Mentioned:
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models: The 2019 paper that first caught Mike’s attention about the capabilities of LLMs
On the Measure of Intelligence: 2019 paper by Google researcher François Chollet that introduced the ARC benchmark, which remains unbeaten
ARC Prize 2024: The $1M+ competition Mike and François have launched to drive interest in solving the ARC-AGI eval
Sequence to Sequence Learning with Neural Networks: Ilya Sutskever paper from 2014 that influenced the direction of machine translation with deep neural networks.
Etched: Luke Miles on LessWrong wrote about the first ASIC chip that accelerates transformers on silicon
Kaggle: The leading data science competition platform and online community, acquired by Google in 2017
Lab42: Swiss AU lab that hosted ARCathon precursor to ARC Prize
Jack Cole: Researcher on team that was #1 on the leaderboard for ARCathon
Ryan Greenblatt: Researcher with current high score (50%) on ARC public leaderboard
(00:00) Introduction
(01:51) AI at Zapier
(08:31) What is ARC AGI?
(13:25) What does it mean to efficiently acquire a new skill?
(19:03) What approaches will succeed?
(21:11) A little bit of a different shape
(25:59) The role of code generation and program synthesis
(29:11) What types of people are working on this?
(31:45) Trying to prove you wrong
(34:50) Where are the big labs?
(38:21) The world post-AGI
(42:51) When will we cross 85% on ARC AGI?
(46:12) Will LLMs be part of the solution?
(50:13) Lightning round
Francois Chollet — Why the biggest AI models can't solve simple puzzles
Here is my conversation with Francois Chollet and Mike Knoop on the $1 million ARC-AGI Prize they're launching today.
I did a bunch of socratic grilling throughout, but Francois’s arguments about why LLMs won’t lead to AGI are very interesting and worth thinking through.
It was really fun discussing/debating the cruxes. Enjoy!
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here.
Timestamps
(00:00:00) – The ARC benchmark
(00:11:10) – Why LLMs struggle with ARC
(00:19:00) – Skill vs intelligence
(00:27:55) - Do we need “AGI” to automate most jobs?
(00:48:28) – Future of AI progress: deep learning + program synthesis
(01:00:40) – How Mike Knoop got nerd-sniped by ARC
(01:08:37) – Million $ ARC Prize
(01:10:33) – Resisting benchmark saturation
(01:18:08) – ARC scores on frontier vs open source models
(01:26:19) – Possible solutions to ARC Prize
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
How the ARC Prize is democratizing the race to AGI with Mike Knoop from Zapier
The first step in achieving AGI is nailing down a concise definition and Mike Knoop, the co-founder and Head of AI at Zapier, believes François Chollet got it right when he defined general intelligence as a system that can efficiently acquire new skills. This week on No Priors, Miked joins Elad to discuss ARC Prize which is a multi-million dollar non-profit public challenge that is looking for someone to beat the Abstraction and Reasoning Corpus (ARC) evaluation.
In this episode, they also get into why Mike thinks LLMs will not get us to AGI, how Zapier is incorporating AI into their products and the power of agents, and why it’s dangerous to regulate AGI before discovering its full potential.
Show Links:
About the Abstraction and Reasoning Corpus
Zapier Central
ARC Prize
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @mikeknoop
Show Notes:
(0:00) Introduction
(1:10) Redefining AGI
(2:16) Introducing ARC Prize
(3:08) Definition of AGI
(5:14) LLMs and AGI
(8:20) Promising techniques to developing AGI
(11:0) Sentience and intelligence
(13:51) Prize model vs investing
(16:28) Zapier AI innovations
(19:08) Economic value of agents
(21:48) Open source to achieve AGI
(24:20) Regulating AI and AGI
Zapier Co-Founder Mike Knoop on category creation, API evolution & AI architecture | E1769
This Week in Startups is presented by:
Embroker. The Embroker Startup Insurance Program helps startups secure the most important types of insurance at a lower cost and with less hassle. Save up to 20% off of traditional insurance today at Embroker.com/twist. While you’re there, get an extra 10% off using offer code TWIST.
Lemon.io - Hire pre-vetted remote developers, get 15% off your first 4 weeks of developer time at https://Lemon.io/twist
Eight Sleep. Good sleep is the ultimate game changer. Now you can add the Pod Pro Cover to any mattress! Go to eightsleep.com/twist to check out the Pod Pro Cover and get $150 off at checkout!
*
Today’s show:
Zapier’s Mike Knoop joins Jason to discuss the early days of Zapier before breaking down the evolution of app integrations and API usage (1:20). They dive into reducing friction for Zapier users, regulating AI, the limitations of present-day AI architecture, and more (43:06).
*
Check out Zapier: https://zapier.com/
Follow Mike: https://twitter.com/mikeknoop
*
Time stamps:
(0:00) Mike Knoop joins Jason
(1:20) Zapier’s origin story
(8:40) Zapier’s key inflection point and its profit-sharing model
(14:05) Zapier’s business model
(15:48) Embroker - Use code TWIST to get an extra 10% off insurance at https://Embroker.com/twist
(17:03) The evolution of app integrations and API usage
(23:41) Lemon.io - Get 15% off your first 4 weeks of developer time at https://Lemon.io/twist
(25:00) Zapier demo + incorporating AI into your workflow
(36:43) Eight Sleep - Go to https://eightsleep.com/twist to check out the Pod Cover and get $150 off at checkout!
(38:14) Linkedin tightening the belt on its API
(41:21) Zapier’s enterprise customers
(43:06) Reducing friction for Zapier users
(51:41) Regulating AI
(55:19) The limitations of present-day AI architecture
*
Read LAUNCH Fund 4 Deal Memo: https://www.launch.co/four
Apply for Funding: https://www.launch.co/apply
Buy ANGEL: https://www.angelthebook.com
Great recent interviews: Steve Huffman, Brian Chesky, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland, PrayingForExits, Jenny Lefcourt
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
*
Follow Jason:
Twitter: https://twitter.com/jason
Instagram: https://www.instagram.com/jason
LinkedIn: https://www.linkedin.com/in/jasoncalacanis
*
Follow TWiST:
Substack: https://twistartups.substack.com
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
*
Subscribe to the Founder University Podcast: https://www.founder.university/podcast