Prof. Randall Balestriero - LLMs without pretraining and SSL
Randall Balestriero joins the show to discuss some counterintuitive findings in AI. He shares research showing that huge language models, even when started from scratch (randomly initialized) without massive pre-training, can learn specific tasks like sentiment analysis surprisingly well, train stably, and avoid severe overfitting, sometimes matching the performance of costly pre-trained models. This raises questions about when giant pre-training efforts are truly worth it.
He also talks about how self-supervised learning (where models learn from data structure itself) and traditional supervised learning (using labeled data) are fundamentally similar, allowing researchers to apply decades of supervised learning theory to improve newer self-supervised methods.
Finally, Randall touches on fairness in AI models used for Earth data (like climate prediction), revealing that these models can be biased, performing poorly in specific locations like islands or coastlines even if they seem accurate overall, which has important implications for policy decisions based on this data.
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.
Goto https://tufalabs.ai/
***
TRANSCRIPT + SHOWNOTES:
https://www.dropbox.com/scl/fi/n7yev71nsjso71jyjz1fy/RANDALLNEURIPS.pdf?rlkey=0dn4injp1sc4ts8njwf3wfmxv&dl=0
TOC:
1. Model Training Efficiency and Scale
[00:00:00] 1.1 Training Stability of Large Models on Small Datasets
[00:04:09] 1.2 Pre-training vs Random Initialization Performance Comparison
[00:07:58] 1.3 Task-Specific Models vs General LLMs Efficiency
2. Learning Paradigms and Data Distribution
[00:10:35] 2.1 Fair Language Model Paradox and Token Frequency Issues
[00:12:02] 2.2 Pre-training vs Single-task Learning Spectrum
[00:16:04] 2.3 Theoretical Equivalence of Supervised and Self-supervised Learning
[00:19:40] 2.4 Self-Supervised Learning and Supervised Learning Relationships
[00:21:25] 2.5 SSL Objectives and Heavy-tailed Data Distribution Challenges
3. Geographic Representation in ML Systems
[00:25:20] 3.1 Geographic Bias in Earth Data Models and Neural Representations
[00:28:10] 3.2 Mathematical Limitations and Model Improvements
[00:30:24] 3.3 Data Quality and Geographic Bias in ML Datasets
REFS:
[00:01:40] Research on training large language models from scratch on small datasets, Randall Balestriero et al.
https://openreview.net/forum?id=wYGBWOjq1Q
[00:10:35] The Fair Language Model Paradox (2024), Andrea Pinto, Tomer Galanti, Randall Balestriero
https://arxiv.org/abs/2410.11985
[00:12:20] Muppet: Massive Multi-task Representations with Pre-Finetuning (2021), Armen Aghajanyan et al.
https://arxiv.org/abs/2101.11038
[00:14:30] Dissociating language and thought in large language models (2023), Kyle Mahowald et al.
https://arxiv.org/abs/2301.06627
[00:16:05] The Birth of Self-Supervised Learning: A Supervised Theory, Randall Balestriero et al.
https://openreview.net/forum?id=NhYAjAAdQT
[00:21:25] VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning, Adrien Bardes, Jean Ponce, Yann LeCun
https://arxiv.org/abs/2105.04906
[00:25:20] No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth Data (2025), Daniel Cai, Randall Balestriero, et al.
https://arxiv.org/abs/2502.06831
[00:33:45] Mark Ibrahim et al.'s work on geographic bias in computer vision datasets, Mark Ibrahim
https://arxiv.org/pdf/2304.12210
Want to Understand Neural Networks? Think Elastic Origami! - Prof. Randall Balestriero
Professor Randall Balestriero joins us to discuss neural network geometry, spline theory, and emerging phenomena in deep learning, based on research presented at ICML. Topics include the delayed emergence of adversarial robustness in neural networks ("grokking"), geometric interpretations of neural networks via spline theory, and challenges in reconstruction learning. We also cover geometric analysis of Large Language Models (LLMs) for toxicity detection and the relationship between intrinsic dimensionality and model control in RLHF.
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. Are you interested in working on reasoning, or getting involved in their events?
Goto https://tufalabs.ai/
***
Randall Balestriero
https://x.com/randall_balestr
https://randallbalestriero.github.io/
Show notes and transcript: https://www.dropbox.com/scl/fi/3lufge4upq5gy0ug75j4a/RANDALLSHOW.pdf?rlkey=nbemgpa0jhawt1e86rx7372e4&dl=0
TOC:
- Introduction
- 00:00:00: Introduction
- Neural Network Geometry and Spline Theory
- 00:01:41: Neural Network Geometry and Spline Theory
- 00:07:41: Deep Networks Always Grok
- 00:11:39: Grokking and Adversarial Robustness
- 00:16:09: Double Descent and Catastrophic Forgetting
- Reconstruction Learning
- 00:18:49: Reconstruction Learning
- 00:24:15: Frequency Bias in Neural Networks
- Geometric Analysis of Neural Networks
- 00:29:02: Geometric Analysis of Neural Networks
- 00:34:41: Adversarial Examples and Region Concentration
- LLM Safety and Geometric Analysis
- 00:40:05: LLM Safety and Geometric Analysis
- 00:46:11: Toxicity Detection in LLMs
- 00:52:24: Intrinsic Dimensionality and Model Control
- 00:58:07: RLHF and High-Dimensional Spaces
- Conclusion
- 01:02:13: Neural Tangent Kernel
- 01:08:07: Conclusion
REFS:
[00:01:35] Humayun – Deep network geometry & input space partitioning
https://arxiv.org/html/2408.04809v1
[00:03:55] Balestriero & Paris – Linking deep networks to adaptive spline operators
https://proceedings.mlr.press/v80/balestriero18b/balestriero18b.pdf
[00:13:55] Song et al. – Gradient-based white-box adversarial attacks
https://arxiv.org/abs/2012.14965
[00:16:05] Humayun, Balestriero & Baraniuk – Grokking phenomenon & emergent robustness
https://arxiv.org/abs/2402.15555
[00:18:25] Humayun – Training dynamics & double descent via linear region evolution
https://arxiv.org/abs/2310.12977
[00:20:15] Balestriero – Power diagram partitions in DNN decision boundaries
https://arxiv.org/abs/1905.08443
[00:23:00] Frankle & Carbin – Lottery Ticket Hypothesis for network pruning
https://arxiv.org/abs/1803.03635
[00:24:00] Belkin et al. – Double descent phenomenon in modern ML
https://arxiv.org/abs/1812.11118
[00:25:55] Balestriero et al. – Batch normalization’s regularization effects
https://arxiv.org/pdf/2209.14778
[00:29:35] EU – EU AI Act 2024 with compute restrictions
https://www.lw.com/admin/upload/SiteAttachments/EU-AI-Act-Navigating-a-Brave-New-World.pdf
[00:39:30] Humayun, Balestriero & Baraniuk – SplineCam: Visualizing deep network geometry
https://openaccess.thecvf.com/content/CVPR2023/papers/Humayun_SplineCam_Exact_Visualization_and_Characterization_of_Deep_Network_Geometry_and_CVPR_2023_paper.pdf
[00:40:40] Carlini – Trade-offs between adversarial robustness and accuracy
https://arxiv.org/pdf/2407.20099
[00:44:55] Balestriero & LeCun – Limitations of reconstruction-based learning methods
https://openreview.net/forum?id=ez7w0Ss4g9
(truncated, see shownotes PDF)
#86 - Prof. YANN LECUN and Dr. RANDALL BALESTRIERO - SSL, Data Augmentation, Reward isn't enough [NEURIPS2022]
Yann LeCun is a French computer scientist known for his pioneering work on convolutional neural networks, optical character recognition and computer vision. He is a Silver Professor at New York University and Vice President, Chief AI Scientist at Meta. Along with Yoshua Bengio and Geoffrey Hinton, he was awarded the 2018 Turing Award for their work on deep learning, earning them the nickname of the "Godfathers of Deep Learning".
Dr. Randall Balestriero has been researching learnable signal processing since 2013, with a focus on learnable parametrized wavelets and deep wavelet transforms. His research has been used by NASA, leading to applications such as Marsquake detection. During his PhD at Rice University, Randall explored deep networks from a theoretical perspective and improved state-of-the-art methods such as batch-normalization and generative networks. Later, when joining Meta AI Research (FAIR) as a postdoc with Prof. Yann LeCun, Randall further broadened his research interests to include self-supervised learning and the biases emerging from data-augmentation and regularization, resulting in numerous publications.
Episode recorded live at NeurIPS.
YT: https://youtu.be/9dLd6n9yT8U (references are there)
Support us! https://www.patreon.com/mlst
Host: Dr. Tim Scarfe
TOC:
[00:00:00] LeCun interview
[00:18:25] Randall Balestriero interview (mostly on spectral SSL paper, first ref)
061: Interpolation, Extrapolation and Linearisation (Prof. Yann LeCun, Dr. Randall Balestriero)
We are now sponsored by Weights and Biases! Please visit our sponsor link: http://wandb.me/MLST
Patreon: https://www.patreon.com/mlst
Yann LeCun thinks that it's specious to say neural network models are interpolating because in high dimensions, everything is extrapolation. Recently Dr. Randall Balestriero, Dr. Jerome Pesente and prof. Yann LeCun released their paper learning in high dimensions always amounts to extrapolation. This discussion has completely changed how we think about neural networks and their behaviour.
[00:00:00] Pre-intro
[00:11:58] Intro Part 1: On linearisation in NNs
[00:28:17] Intro Part 2: On interpolation in NNs
[00:47:45] Intro Part 3: On the curse
[00:48:19] LeCun
[01:40:51] Randall B
YouTube version: https://youtu.be/86ib0sfdFtw