Jack Morris on Finding the Next Big AI Breakthrough
We know that the top-tier AI labs are spending unbelievable amounts of money on talent. But what are these researchers actually working on? And how do we know that they're making progress? And furthermore, how can we even measure that progress? On this episode, we speak with Jack Morris, an AI researcher and Ph.D. candidate at Cornell University, who is also a part-time researcher at Meta. We talk about what he does, and why breakthroughs seem to be lumpy and unpredictable. We also talk about the battle between open- and closed-source approaches, US vs. Chinese labs, and how an individual talent thinks about where they want to spend their time, balancing the desire for research and prestige with a big fat paycheck.
Only Bloomberg.com subscribers can get the Odd Lots newsletter in their inbox — now delivered every weekday — plus unlimited access to the site and app. Subscribe at bloomberg.com/subscriptions/oddlots
See omnystudio.com/listener for privacy information.
Information Theory for Language Models: Jack Morris
Our last AI PhD grad student feature was Shunyu Yao, who happened to focus on Language Agents for his thesis and immediately went to work on them for OpenAI. Our pick this year is Jack Morris, who bucks the “hot” trends by -not- working on agents, benchmarks, or VS Code forks, but is rather known for his work on the information theoretic understanding of LLMs, starting from embedding models and latent space representations (always close to our heart).
Jack is an unusual combination of doing underrated research but somehow still being to explain them well to a mass audience, so we felt this was a good opportunity to do a different kind of episode going through the greatest hits of a high profile AI PhD, and relate them to questions from AI Engineering.
Papers and References made
* AI grad school:
* A new type of information theory:
* Embeddings
* Text Embeddings Reveal (Almost) As Much As Text: https://arxiv.org/abs/2310.06816
* Contextual document embeddings https://arxiv.org/abs/2410.02525
Harnessing the Universal Geometry of Embeddings: https://arxiv.org/abs/2505.12540
* Language models
* GPT-style language models memorize 3.6 bits per param:
* Approximating Language Model Training Data from Weights: https://arxiv.org/abs/2506.15553
* LLM Inversion
* “There Are No New Ideas In AI.... Only New Datasets”
* misc reference: https://junyanz.github.io/CycleGAN/
—
for others hiring AI PhDs, Jack also wanted to shout out his coauthor
Zach Nussbaum, his coauthor on Nomic Embed: Training a Reproducible Long Context Text Embedder.
Full Video Episode
Timestamps
00:00 Introduction to Jack Morris01:18 Career in AI03:29 The Shift to AI Companies03:57 The Impact of ChatGPT04:26 The Role of Academia in AI05:49 The Emergence of Reasoning Models07:07 Challenges in Academia: GPUs and HPC Training11:04 The Value of GPU Knowledge14:24 Introduction to Jack's Research15:28 Information Theory17:10 Understanding Deep Learning Systems19:00 The "Bit" in Deep Learning20:25 Wikipedia and Information Storage23:50 Text Embeddings and Information Compression27:08 The Research Journey of Embedding Inversion31:22 Harnessing the Universal Geometry of Embeddings34:54 Implications of Embedding Inversion36:02 Limitations of Embedding Inversion38:08 The Capacity of Language Models40:23 The Cognitive Core and Model Efficiency50:40 The Future of AI and Model Scaling52:47 Approximating Language Model Training Data from Weights01:06:50 The "No New Ideas, Only New Datasets" Thesis
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Attack of the C̶l̶o̶n̶e̶s̶ Text!
Come hang with the bad boys of natural language processing (NLP)! Jack Morris joins Daniel and Chris to talk about TextAttack, a Python framework for adversarial attacks, data augmentation, and model training in NLP. TextAttack will improve your understanding of your NLP models, so come prepared to rumble with your own adversarial attacks!
Sponsors:
Linode – Our cloud of choice and the home of Changelog.com. Deploy a fast, efficient, native SSD cloud server for only $5/month. Get 4 months free using the code changelog2019 OR changelog2020. To learn more and get started head to linode.com/changelog.
Changelog++ – You love our content and you want to take it to the next level by showing your support. We’ll take you closer to the metal with no ads, extended episodes, outtakes, bonus content, a deep discount in our merch store (soon), and more to come. Let’s do this!
Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com.
Rollbar – We move fast and fix things because of Rollbar. Resolve errors in minutes. Deploy with confidence. Learn more at rollbar.com/changelog.
Featuring:
Jack Morris – Website, GitHub, X
Chris Benson – Website, GitHub, LinkedIn, X
Daniel Whitenack – Website, GitHub, X
Show Notes:
TextAttack
Attacking Machine Learning with Adversarial Examples
Upcoming Events:
Register for upcoming webinars here!