ARC Prize v2 Launch! (Francois Chollet and Mike Knoop)
We are joined by Francois Chollet and Mike Knoop, to launch the new version of the ARC prize! In version 2, the challenges have been calibrated with humans such that at least 2 humans could solve each task in a reasonable task, but also adversarially selected so that frontier reasoning models can't solve them. The best LLMs today get negligible performance on this challenge.
https://arcprize.org/
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Goto https://tufalabs.ai/
***
TRANSCRIPT:
https://www.dropbox.com/scl/fi/0v9o8xcpppdwnkntj59oi/ARCv2.pdf?rlkey=luqb6f141976vra6zdtptv5uj&dl=0
TOC:
1. ARC v2 Core Design & Objectives
[00:00:00] 1.1 ARC v2 Launch and Benchmark Architecture
[00:03:16] 1.2 Test-Time Optimization and AGI Assessment
[00:06:24] 1.3 Human-AI Capability Analysis
[00:13:02] 1.4 OpenAI o3 Initial Performance Results
2. ARC Technical Evolution
[00:17:20] 2.1 ARC-v1 to ARC-v2 Design Improvements
[00:21:12] 2.2 Human Validation Methodology
[00:26:05] 2.3 Task Design and Gaming Prevention
[00:29:11] 2.4 Intelligence Measurement Framework
3. O3 Performance & Future Challenges
[00:38:50] 3.1 O3 Comprehensive Performance Analysis
[00:43:40] 3.2 System Limitations and Failure Modes
[00:49:30] 3.3 Program Synthesis Applications
[00:53:00] 3.4 Future Development Roadmap
REFS:
[00:00:15] On the Measure of Intelligence, François Chollet
https://arxiv.org/abs/1911.01547
[00:06:45] ARC Prize Foundation, François Chollet, Mike Knoop
https://arcprize.org/
[00:12:50] OpenAI o3 model performance on ARC v1, ARC Prize Team
https://arcprize.org/blog/oai-o3-pub-breakthrough
[00:18:30] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al.
https://arxiv.org/abs/2201.11903
[00:21:45] ARC-v2 benchmark tasks, Mike Knoop
https://arcprize.org/blog/introducing-arc-agi-public-leaderboard
[00:26:05] ARC Prize 2024: Technical Report, Francois Chollet et al.
https://arxiv.org/html/2412.04604v2
[00:32:45] ARC Prize 2024 Technical Report, Francois Chollet, Mike Knoop, Gregory Kamradt
https://arxiv.org/abs/2412.04604
[00:48:55] The Bitter Lesson, Rich Sutton
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
[00:53:30] Decoding strategies in neural text generation, Sina Zarrieß
https://www.mdpi.com/2078-2489/12/9/355/pdf
Francois Chollet - ARC reflections - NeurIPS 2024
François Chollet discusses the outcomes of the ARC-AGI (Abstraction and Reasoning Corpus) Prize competition in 2024, where accuracy rose from 33% to 55.5% on a private evaluation set.
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. Are you interested in working on reasoning, or getting involved in their events?
They are hosting an event in Zurich on January 9th with the ARChitects, join if you can.
Goto https://tufalabs.ai/
***
Read about the recent result on o3 with ARC here (Chollet knew about it at the time of the interview but wasn't allowed to say):
https://arcprize.org/blog/oai-o3-pub-breakthrough
TOC:
1. Introduction and Opening
[00:00:00] 1.1 Deep Learning vs. Symbolic Reasoning: François’s Long-Standing Hybrid View
[00:00:48] 1.2 “Why Do They Call You a Symbolist?” – Addressing Misconceptions
[00:01:31] 1.3 Defining Reasoning
3. ARC Competition 2024 Results and Evolution
[00:07:26] 3.1 ARC Prize 2024: Reflecting on the Narrative Shift Toward System 2
[00:10:29] 3.2 Comparing Private Leaderboard vs. Public Leaderboard Solutions
[00:13:17] 3.3 Two Winning Approaches: Deep Learning–Guided Program Synthesis and Test-Time Training
4. Transduction vs. Induction in ARC
[00:16:04] 4.1 Test-Time Training, Overfitting Concerns, and Developer-Aware Generalization
[00:19:35] 4.2 Gradient Descent Adaptation vs. Discrete Program Search
5. ARC-2 Development and Future Directions
[00:23:51] 5.1 Ensemble Methods, Benchmark Flaws, and the Need for ARC-2
[00:25:35] 5.2 Human-Level Performance Metrics and Private Test Sets
[00:29:44] 5.3 Task Diversity, Redundancy Issues, and Expanded Evaluation Methodology
6. Program Synthesis Approaches
[00:30:18] 6.1 Induction vs. Transduction
[00:32:11] 6.2 Challenges of Writing Algorithms for Perceptual vs. Algorithmic Tasks
[00:34:23] 6.3 Combining Induction and Transduction
[00:37:05] 6.4 Multi-View Insight and Overfitting Regulation
7. Latent Space and Graph-Based Synthesis
[00:38:17] 7.1 Clément Bonnet’s Latent Program Search Approach
[00:40:10] 7.2 Decoding to Symbolic Form and Local Discrete Search
[00:41:15] 7.3 Graph of Operators vs. Token-by-Token Code Generation
[00:45:50] 7.4 Iterative Program Graph Modifications and Reusable Functions
8. Compute Efficiency and Lifelong Learning
[00:48:05] 8.1 Symbolic Process for Architecture Generation
[00:50:33] 8.2 Logarithmic Relationship of Compute and Accuracy
[00:52:20] 8.3 Learning New Building Blocks for Future Tasks
9. AI Reasoning and Future Development
[00:53:15] 9.1 Consciousness as a Self-Consistency Mechanism in Iterative Reasoning
[00:56:30] 9.2 Reconciling Symbolic and Connectionist Views
[01:00:13] 9.3 System 2 Reasoning - Awareness and Consistency
[01:03:05] 9.4 Novel Problem Solving, Abstraction, and Reusability
10. Program Synthesis and Research Lab
[01:05:53] 10.1 François Leaving Google to Focus on Program Synthesis
[01:09:55] 10.2 Democratizing Programming and Natural Language Instruction
11. Frontier Models and O1 Architecture
[01:14:38] 11.1 Search-Based Chain of Thought vs. Standard Forward Pass
[01:16:55] 11.2 o1’s Natural Language Program Generation and Test-Time Compute Scaling
[01:19:35] 11.3 Logarithmic Gains with Deeper Search
12. ARC Evaluation and Human Intelligence
[01:22:55] 12.1 LLMs as Guessing Machines and Agent Reliability Issues
[01:25:02] 12.2 ARC-2 Human Testing and Correlation with g-Factor
[01:26:16] 12.3 Closing Remarks and Future Directions
SHOWNOTES PDF:
https://www.dropbox.com/scl/fi/ujaai0ewpdnsosc5mc30k/CholletNeurips.pdf?rlkey=s68dp432vefpj2z0dp5wmzqz6&st=hazphyx5&dl=0
Pattern Recognition vs True Intelligence - Francois Chollet
Francois Chollet, a prominent AI expert and creator of ARC-AGI, discusses intelligence, consciousness, and artificial intelligence.
Chollet explains that real intelligence isn't about memorizing information or having lots of knowledge - it's about being able to handle new situations effectively. This is why he believes current large language models (LLMs) have "near-zero intelligence" despite their impressive abilities. They're more like sophisticated memory and pattern-matching systems than truly intelligent beings.
***
MLST IS SPONSORED BY TUFA AI LABS!
The current winners of the ARC challenge, MindsAI are part of Tufa AI Labs. They are hiring ML engineers. Are you interested?! Please goto https://tufalabs.ai/
***
He introduced his "Kaleidoscope Hypothesis," which suggests that while the world seems infinitely complex, it's actually made up of simpler patterns that repeat and combine in different ways. True intelligence, he argues, involves identifying these basic patterns and using them to understand new situations.
Chollet also talked about consciousness, suggesting it develops gradually in children rather than appearing all at once. He believes consciousness exists in degrees - animals have it to some extent, and even human consciousness varies with age and circumstances (like being more conscious when learning something new versus doing routine tasks).
On AI safety, Chollet takes a notably different stance from many in Silicon Valley. He views AGI development as a scientific challenge rather than a religious quest, and doesn't share the apocalyptic concerns of some AI researchers. He argues that intelligence itself isn't dangerous - it's just a tool for turning information into useful models. What matters is how we choose to use it.
ARC-AGI Prize:
https://arcprize.org/
Francois Chollet:
https://x.com/fchollet
Shownotes:
https://www.dropbox.com/scl/fi/j2068j3hlj8br96pfa7bi/CHOLLET_FINAL.pdf?rlkey=xkbr7tbnrjdl66m246w26uc8k&st=0a4ec4na&dl=0
TOC:
1. Intelligence and Model Building
[00:00:00] 1.1 Intelligence Definition and ARC Benchmark
[00:05:40] 1.2 LLMs as Program Memorization Systems
[00:09:36] 1.3 Kaleidoscope Hypothesis and Abstract Building Blocks
[00:13:39] 1.4 Deep Learning Limitations and System 2 Reasoning
[00:29:38] 1.5 Intelligence vs. Skill in LLMs and Model Building
2. ARC Benchmark and Program Synthesis
[00:37:36] 2.1 Intelligence Definition and LLM Limitations
[00:41:33] 2.2 Meta-Learning System Architecture
[00:56:21] 2.3 Program Search and Occam's Razor
[00:59:42] 2.4 Developer-Aware Generalization
[01:06:49] 2.5 Task Generation and Benchmark Design
3. Cognitive Systems and Program Generation
[01:14:38] 3.1 System 1/2 Thinking Fundamentals
[01:22:17] 3.2 Program Synthesis and Combinatorial Challenges
[01:31:18] 3.3 Test-Time Fine-Tuning Strategies
[01:36:10] 3.4 Evaluation and Leakage Problems
[01:43:22] 3.5 ARC Implementation Approaches
4. Intelligence and Language Systems
[01:50:06] 4.1 Intelligence as Tool vs Agent
[01:53:53] 4.2 Cultural Knowledge Integration
[01:58:42] 4.3 Language and Abstraction Generation
[02:02:41] 4.4 Embodiment in Cognitive Systems
[02:09:02] 4.5 Language as Cognitive Operating System
5. Consciousness and AI Safety
[02:14:05] 5.1 Consciousness and Intelligence Relationship
[02:20:25] 5.2 Development of Machine Consciousness
[02:28:40] 5.3 Consciousness Prerequisites and Indicators
[02:36:36] 5.4 AGI Safety Considerations
[02:40:29] 5.5 AI Regulation Framework
It's Not About Scale, It's About Abstraction - Francois Chollet
François Chollet discusses the limitations of Large Language Models (LLMs) and proposes a new approach to advancing artificial intelligence. He argues that current AI systems excel at pattern recognition but struggle with logical reasoning and true generalization.
This was Chollet's keynote talk at AGI-24, filmed in high-quality. We will be releasing a full interview with him shortly. A teaser clip from that is played in the intro!
Chollet introduces the Abstraction and Reasoning Corpus (ARC) as a benchmark for measuring AI progress towards human-like intelligence. He explains the concept of abstraction in AI systems and proposes combining deep learning with program synthesis to overcome current limitations. Chollet suggests that breakthroughs in AI might come from outside major tech labs and encourages researchers to explore new ideas in the pursuit of artificial general intelligence.
TOC
1. LLM Limitations and Intelligence Concepts
[00:00:00] 1.1 LLM Limitations and Composition
[00:12:05] 1.2 Intelligence as Process vs. Skill
[00:17:15] 1.3 Generalization as Key to AI Progress
2. ARC-AGI Benchmark and LLM Performance
[00:19:59] 2.1 Introduction to ARC-AGI Benchmark
[00:20:05] 2.2 Introduction to ARC-AGI and the ARC Prize
[00:23:35] 2.3 Performance of LLMs and Humans on ARC-AGI
3. Abstraction in AI Systems
[00:26:10] 3.1 The Kaleidoscope Hypothesis and Abstraction Spectrum
[00:30:05] 3.2 LLM Capabilities and Limitations in Abstraction
[00:32:10] 3.3 Value-Centric vs Program-Centric Abstraction
[00:33:25] 3.4 Types of Abstraction in AI Systems
4. Advancing AI: Combining Deep Learning and Program Synthesis
[00:34:05] 4.1 Limitations of Transformers and Need for Program Synthesis
[00:36:45] 4.2 Combining Deep Learning and Program Synthesis
[00:39:59] 4.3 Applying Combined Approaches to ARC Tasks
[00:44:20] 4.4 State-of-the-Art Solutions for ARC
Shownotes (new!): https://www.dropbox.com/scl/fi/i7nsyoahuei6np95lbjxw/CholletKeynote.pdf?rlkey=t3502kbov5exsdxhderq70b9i&st=1ca91ewz&dl=0
[0:01:15] Abstraction and Reasoning Corpus (ARC): AI benchmark (François Chollet)
https://arxiv.org/abs/1911.01547
[0:05:30] Monty Hall problem: Probability puzzle (Steve Selvin)
https://www.tandfonline.com/doi/abs/10.1080/00031305.1975.10479121
[0:06:20] LLM training dynamics analysis (Tirumala et al.)
https://arxiv.org/abs/2205.10770
[0:10:20] Transformer limitations on compositionality (Dziri et al.)
https://arxiv.org/abs/2305.18654
[0:10:25] Reversal Curse in LLMs (Berglund et al.)
https://arxiv.org/abs/2309.12288
[0:19:25] Measure of intelligence using algorithmic information theory (François Chollet)
https://arxiv.org/abs/1911.01547
[0:20:10] ARC-AGI: GitHub repository (François Chollet)
https://github.com/fchollet/ARC-AGI
[0:22:15] ARC Prize: $1,000,000+ competition (François Chollet)
https://arcprize.org/
[0:33:30] System 1 and System 2 thinking (Daniel Kahneman)
https://www.amazon.com/Thinking-Fast-Slow-Daniel-Kahneman/dp/0374533555
[0:34:00] Core knowledge in infants (Elizabeth Spelke)
https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf
[0:34:30] Embedding interpretive spaces in ML (Tennenholtz et al.)
https://arxiv.org/abs/2310.04475
[0:44:20] Hypothesis Search with LLMs for ARC (Wang et al.)
https://arxiv.org/abs/2309.05660
[0:44:50] Ryan Greenblatt's high score on ARC public leaderboard
https://arcprize.org/
Francois Chollet — Why the biggest AI models can't solve simple puzzles
Here is my conversation with Francois Chollet and Mike Knoop on the $1 million ARC-AGI Prize they're launching today.
I did a bunch of socratic grilling throughout, but Francois’s arguments about why LLMs won’t lead to AGI are very interesting and worth thinking through.
It was really fun discussing/debating the cruxes. Enjoy!
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here.
Timestamps
(00:00:00) – The ARC benchmark
(00:11:10) – Why LLMs struggle with ARC
(00:19:00) – Skill vs intelligence
(00:27:55) - Do we need “AGI” to automate most jobs?
(00:48:28) – Future of AI progress: deep learning + program synthesis
(01:00:40) – How Mike Knoop got nerd-sniped by ARC
(01:08:37) – Million $ ARC Prize
(01:10:33) – Resisting benchmark saturation
(01:18:08) – ARC scores on frontier vs open source models
(01:26:19) – Possible solutions to ARC Prize
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
#79 Consciousness and the Chinese Room [Special Edition] (CHOLLET, BISHOP, CHALMERS, BACH)
This video is demonetised on music copyright so we would appreciate support on our Patreon! https://www.patreon.com/mlst
We would also appreciate it if you rated us on your podcast platform.
YT: https://youtu.be/_KVAzAzO5HU
Panel: Dr. Tim Scarfe, Dr. Keith Duggar
Guests: Prof. J. Mark Bishop, Francois Chollet, Prof. David Chalmers, Dr. Joscha Bach, Prof. Karl Friston, Alexander Mattick, Sam Roffey
The Chinese Room Argument was first proposed by philosopher John Searle in 1980. It is an argument against the possibility of artificial intelligence (AI) – that is, the idea that a machine could ever be truly intelligent, as opposed to just imitating intelligence.
The argument goes like this:
Imagine a room in which a person sits at a desk, with a book of rules in front of them. This person does not understand Chinese.
Someone outside the room passes a piece of paper through a slot in the door. On this paper is a Chinese character. The person in the room consults the book of rules and, following these rules, writes down another Chinese character and passes it back out through the slot.
To someone outside the room, it appears that the person in the room is engaging in a conversation in Chinese. In reality, they have no idea what they are doing – they are just following the rules in the book.
The Chinese Room Argument is an argument against the idea that a machine could ever be truly intelligent. It is based on the idea that intelligence requires understanding, and that following rules is not the same as understanding.
in this detailed investigation into the Chinese Room, Consciousness and Syntax vs Semantics, we interview luminaries J.Mark Bishop and Francois Chollet and use unreleased footage from our interviews with David Chalmers, Joscha Bach and Karl Friston. We also cover material from Walid Saba and interview Alex Mattick from Yannic's Discord.
This is probably my favourite ever episode of MLST. I hope you enjoy it! With Keith Duggar.
Note that we are using clips from our unreleased interviews from David Chalmers and Joscha Bach -- we will release those shows properly in the coming weeks. We apologise for delay releasing our backlog, we have been busy building a startup company in the background.
TOC:
[00:00:00] Kick off
[00:00:46] Searle
[00:05:09] Bishop introduces CRA
[00:00:00] Stevan Hardad take on CRA
[00:14:03] Francois Chollet dissects CRA
[00:34:16] Chalmers on consciousness
[00:36:27] Joscha Bach on consciousness
[00:42:01] Bishop introduction
[00:51:51] Karl Friston on consciousness
[00:55:19] Bishop on consciousness and comments on Chalmers
[01:21:37] Private language games (including clip with Sam Roffey)
[01:27:27] Dr. Walid Saba on the chinese room (gofai/systematicity take)
[00:34:36] Bishop: on agency / teleology
[01:36:38] Bishop: back to CRA
[01:40:53] Noam Chomsky on mysteries
[01:45:56] Eric Curiel on math does not represent
[01:48:14] Alexander Mattick on syntax vs semantics
Thanks to: Mark MC on Discord for stimulating conversation, Alexander Mattick, Dr. Keith Duggar, Sam Roffey. Sam's YouTube channel is https://www.youtube.com/channel/UCjRNMsglFYFwNsnOWIOgt1Q
#51 Francois Chollet - Intelligence and Generalisation
In today's show we are joined by Francois Chollet, I have been inspired by Francois ever since I read his Deep Learning with Python book and started using the Keras library which he invented many, many years ago. Francois has a clarity of thought that I've never seen in any other human being! He has extremely interesting views on intelligence as generalisation, abstraction and an information conversation ratio. He wrote on the measure of intelligence at the end of 2019 and it had a huge impact on my thinking. He thinks that NNs can only model continuous problems, which have a smooth learnable manifold and that many "type 2" problems which involve reasoning and/or planning are not suitable for NNs. He thinks that many problems have type 1 and type 2 enmeshed together. He thinks that the future of AI must include program synthesis to allow us to generalise broadly from a few examples, but the search could be guided by neural networks because the search space is interpolative to some extent.
https://youtu.be/J0p_thJJnoo
Tim's Whimsical notes; https://whimsical.com/chollet-show-QQ2atZUoRR3yFDsxKVzCbj
#120 – François Chollet: Measures of Intelligence
François Chollet is an AI researcher at Google and creator of Keras.
Support this podcast by supporting our sponsors (and get discount):
– Babbel: https://babbel.com and use code LEX
– MasterClass: https://masterclass.com/lex
– Cash App: download app & use code “LexPodcast”
Episode links:
Francois’s Twitter: https://twitter.com/fchollet
Francois’s Website: https://fchollet.com/
On the Measure of Intelligence (paper): https://arxiv.org/abs/1911.01547
If you would like to get more information about this podcast go to https://lexfridman.com/podcast or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
OUTLINE:
00:00 – Introduction
05:04 – Early influence
06:23 – Language
12:50 – Thinking with mind maps
23:42 – Definition of intelligence
42:24 – GPT-3
53:07 – Semantic web
57:22 – Autonomous driving
1:09:30 – Tests of intelligence
1:13:59 – Tests of human intelligence
1:27:18 – IQ tests
1:35:59 – ARC Challenge
1:59:11 – Generalization
2:09:50 – Turing Test
2:20:44 – Hutter prize
2:27:44 – Meaning of life
François Chollet: Keras, Deep Learning, and the Progress of AI
François Chollet is the creator of Keras, which is an open source deep learning library that is designed to enable fast, user-friendly experimentation with deep neural networks. It serves as an interface to several deep learning libraries, most popular of which is TensorFlow, and it was integrated into TensorFlow main codebase a while back. Aside from creating an exceptionally useful and popular library, François is also a world-class AI researcher and software engineer at Google, and is definitely an outspoken, if not controversial, personality in the AI world, especially in the realm of ideas around the future of artificial intelligence. This conversation is part of the Artificial Intelligence podcast. If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on iTunes or support it on Patreon.