We’re back with our final installment from Hard Fork Live, recorded at the Yerba Buena Center for the Arts in San Francisco. In this episode, we’re joined by Sayash Kapoor and Daniel Kokotajlo to talk about their differing visions of A.I. transformation: why Sayash thinks A.I. will diffuse throughout society like a “normal” technology, and why Daniel thinks an unprecedented acceleration is just around the corner. Then we’re joined by George Ekas from Toborlife AI, along with his dancing robot Toby. Finally, the podcaster Dwarkesh Patel drops by, and we take a few questions from the live audience.
Guests:
Sayash Kapoor, an A.I. researcher at Princeton University and a co-author of the newsletter “AI as Normal Technology”
Daniel Kokotajlo, the executive director of the AI Futures Project and a co-author of “AI 2027”
George Ekas, the director of engineering at Toberlife AI
Dwarkesh Patel, a tech podcaster
Additional Reading:
This A.I. Forecast Predicts Storms Ahead
AI as Normal Technology
Common Ground Between AI 2027 & AI as Normal Technology
Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Jon Krohn recaps the month of February in this episode of In Case You Missed It. Across four interviews with Will Falcon (Episode 965), Tom Griffiths (Episode 969), Antje Barth (Episode 963), and Praveen Murugesan (Episode 967), Jon questions the brains behind some of the AI industry’s most innovative companies about launching a startup, developing a popular product, what artificial intelligence can still learn from human intelligence, and how AI might finally start to think on its own.
Additional materials: www.superdatascience.com/972
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
Princeton Professor Tom Griffiths talks to Jon Krohn about his new book, The Laws of Thought, which grapples with the mathematical models behind biological and artificial intelligence, and what makes the human brain so fascinating for psychologists and computer scientists to study. In this episode, he details how the mathematical principles governing the external world can also be used to explore cognitive science, or “the internal world.”
This episode is brought to you by the Dell, by Intel, by Cisco and by Acceldata.
Additional materials: www.superdatascience.com/969
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(01:18) Tom Griffiths’ current research
(21:23) On mathematical inference in LLMs
(35:19) How to engineer inductive bias
(52:00) How to model curiosity into AI systems
Ryan is joined by Professor Tom Griffiths, the head of Princeton University’s AI Lab, to dive into findings from his new book The Laws of Thought, which explores the history of the philosophy, mathematics, and logic that underlie artificial intelligence, and scientists' efforts to describe our minds using mathematics. They discuss the challenges of understanding human cognition, the implications of probabilistic AI “thinking,” and where Aristotle fits into the philosophical discussions we’re having on consciousness and sentience in AI.
Episode notes:
The Laws of Thought details our quest to use mathematics to describe the ways we think, from its origins three hundred years ago to the ideas behind modern AI systems and how our human minds differ from the neural networks of AI.
Connect with Tom on LinkedIn and find more of his work at the Princeton website.
Congrats to user Andreas Rayo Kniep for winning a Populist badge for their answer to Is there a difference between the UTC and Etc/UTC time zones?.
We want to know what you're using to upskill and learn in the age of AI. Take this five minute survey on learning and AI to have your voice heard in our next Stack Overflow Knows Pulse Survey.
TRANSCRIPT
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this bonus episode, Princeton University professor and artificial intelligence researcher Tom Griffiths joins Sam to unpack The Laws of
Thought, his new book exploring how math has been used for
centuries to understand how minds — human and machine — actually work.
Tom walks through three main frameworks shaping intelligence today — rules and symbols, neural networks, and probability — and he explains why modern AI only makes sense when you see how those pieces fit together.
The conversation connects cognitive science, large language models, and the limits of human versus machine intelligence. Along the way, Tom and Sam dig into language, learning, and what humans still do better — like judgment, curation, and metacognition. Read the episode transcript here.
*Please take our listener survey: mitsmr.com/podcastsurvey
It's short — we promise! — and all respondents will receive a free MIT SMR article collection, "Maximizing the Value of Generative AI."
Me, Myself, and AI is a podcast produced by MIT Sloan Management Review and hosted by Sam Ransbotham. It is engineered by David Lishansky and produced by Allison Ryder.
We encourage you to rate and review our show. Your comments may be used in Me, Myself, and AI materials.
ME, MYSELF, AND AI® is a federally registered trademark of Massachusetts Institute of Technology. All rights reserved.
From undergraduate research seminars at Princeton to winning Best Paper award at NeurIPS 2025, Kevin Wang, Ishaan Javali, Michał Bortkiewicz, Tomasz Trzcinski, Benjamin Eysenbach defied conventional wisdom by scaling reinforcement learning networks to 1,000 layers deep—unlocking performance gains that the RL community thought impossible. We caught up with the team live at NeurIPS to dig into the story behind RL1000: why deep networks have worked in language and vision but failed in RL for over a decade (spoiler: it’s not just about depth, it’s about the objective), how they discovered that self-supervised RL (learning representations of states, actions, and future states via contrastive learning) scales where value-based methods collapse, the critical architectural tricks that made it work (residual connections, layer normalization, and a shift from regression to classification), why scaling depth is more parameter-efficient than scaling width (linear vs. quadratic growth), how Jax and GPU-accelerated environments let them collect hundreds of millions of transitions in hours (the data abundance that unlocked scaling in the first place), the “critical depth” phenomenon where performance doesn’t just improve—it multiplies once you cross 15M+ transitions and add the right architectural components, why this isn’t just “make networks bigger” but a fundamental shift in RL objectives (their code doesn’t have a line saying “maximize rewards”—it’s pure self-supervised representation learning), how deep teacher, shallow student distillation could unlock deployment at scale (train frontier capabilities with 1000 layers, distill down to efficient inference models), the robotics implications (goal-conditioned RL without human supervision or demonstrations, scaling architecture instead of scaling manual data collection), and their thesis that RL is finally ready to scale like language and vision—not by throwing compute at value functions, but by borrowing the self-supervised, representation-learning paradigms that made the rest of deep learning work.
We discuss:
* The self-supervised RL objective: instead of learning value functions (noisy, biased, spurious), they learn representations where states along the same trajectory are pushed together, states along different trajectories are pushed apart—turning RL into a classification problem
* Why naive scaling failed: doubling depth degraded performance, doubling again with residual connections and layer norm suddenly skyrocketed performance in one environment—unlocking the “critical depth” phenomenon
* Scaling depth vs. width: depth grows parameters linearly, width grows quadratically—depth is more parameter-efficient and sample-efficient for the same performance
* The Jax + GPU-accelerated environments unlock: collecting thousands of trajectories in parallel meant data wasn’t the bottleneck, and crossing 15M+ transitions was when deep networks really paid off
* The blurring of RL and self-supervised learning: their code doesn’t maximize rewards directly, it’s an actor-critic goal-conditioned RL algorithm, but the learning burden shifts to classification (cross-entropy loss, representation learning) instead of TD error regression
* Why scaling batch size unlocks at depth: traditional RL doesn’t benefit from larger batches because networks are too small to exploit the signal, but once you scale depth, batch size becomes another effective scaling dimension
—
RL1000 Team (Princeton)
* 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities: https://openreview.net/forum?id=s0JVsx3bx1
Full Video Episode
Timestamps
00:00:00 Introduction: Best Paper Award and NeurIPS Poster Experience00:01:11 Team Introductions and Princeton Research Origins00:03:35 The Deep Learning Anomaly: Why RL Stayed Shallow00:04:35 Self-Supervised RL: A Different Approach to Scaling00:05:13 The Breakthrough Moment: Residual Connections and Critical Depth00:07:15 Architectural Choices: Borrowing from ResNets and Avoiding Vanishing Gradients00:07:50 Clarifying the Paper: Not Just Big Networks, But Different Objectives00:08:46 Blurring the Lines: RL Meets Self-Supervised Learning00:09:44 From TD Errors to Classification: Why This Objective Scales00:11:06 Architecture Details: Building on Braw and SymbaFowl00:12:05 Robotics Applications: Goal-Conditioned RL Without Human Supervision00:13:15 Efficiency Trade-offs: Depth vs Width and Parameter Scaling00:15:48 JAX and GPU-Accelerated Environments: The Data Infrastructure00:18:05 World Models and Next State Classification00:22:37 Unlocking Batch Size Scaling Through Network Capacity00:24:10 Compute Requirements: State-of-the-Art on a Single GPU00:21:02 Future Directions: Distillation, VLMs, and Hierarchical Planning00:27:15 Closing Thoughts: Challenging Conventional Wisdom in RL Scaling
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab racing to solve software engineering at scale. We caught up with John live at NeurIPS 2025 to dig into the state of code evals heading into 2026: why SWE-bench went from ignored (October 2023) to the industry standard after Devin’s launch (and how Walden emailed him two weeks before the big reveal), how the benchmark evolved from Django-heavy to nine languages across 40 repos (JavaScript, Rust, Java, C, Ruby), why unit tests as verification are limiting and long-running agent tournaments might be the future (CodeClash: agents maintain codebases, compete in arenas, and iterate over multiple rounds), the proliferation of SWE-bench variants (SWE-bench Pro, SWE-bench Live, SWE-Efficiency, AlgoTune, SciCode) and how benchmark authors are now justifying their splits with curation techniques instead of just “more repos,” why Tau-bench’s “impossible tasks” controversy is actually a feature not a bug (intentionally including impossible tasks flags cheating), the tension between long autonomy (5-hour runs) vs. interactivity (Cognition’s emphasis on fast back-and-forth), how Terminal-bench unlocked creativity by letting PhD students and non-coders design environments beyond GitHub issues and PRs, the academic data problem (companies like Cognition and Cursor have rich user interaction data, academics need user simulators or compelling products like LMArena to get similar signal), and his vision for CodeClash as a testbed for human-AI collaboration—freeze model capability, vary the collaboration setup (solo agent, multi-agent, human+agent), and measure how interaction patterns change as models climb the ladder from code completion to full codebase reasoning.
We discuss:
* John’s path: Princeton → SWE-bench (October 2023) → Stanford PhD with Diyi Yang and the Iris Group, focusing on code evals, human-AI collaboration, and long-running agent benchmarks
* The SWE-bench origin story: released October 2023, mostly ignored until Cognition’s Devin launch kicked off the arms race (Walden emailed John two weeks before: “we have a good number”)
* SWE-bench Verified: the curated, high-quality split that became the standard for serious evals
* SWE-bench Multimodal and Multilingual: nine languages (JavaScript, Rust, Java, C, Ruby) across 40 repos, moving beyond the Django-heavy original distribution
* The SWE-bench Pro controversy: independent authors used the “SWE-bench” name without John’s blessing, but he’s okay with it (”congrats to them, it’s a great benchmark”)
* CodeClash: John’s new benchmark for long-horizon development—agents maintain their own codebases, edit and improve them each round, then compete in arenas (programming games like Halite, economic tasks like GDP optimization)
* SWE-Efficiency (Jeffrey Maugh, John’s high school classmate): optimize code for speed without changing behavior (parallelization, SIMD operations)
* AlgoTune, SciCode, Terminal-bench, Tau-bench, SecBench, SRE-bench: the Cambrian explosion of code evals, each diving into different domains (security, SRE, science, user simulation)
* The Tau-bench “impossible tasks” debate: some tasks are underspecified or impossible, but John thinks that’s actually a feature (flags cheating if you score above 75%)
* Cognition’s research focus: codebase understanding (retrieval++), helping humans understand their own codebases, and automatic context engineering for LLMs (research sub-agents)
* The vision: CodeClash as a testbed for human-AI collaboration—vary the setup (solo agent, multi-agent, human+agent), freeze model capability, and measure how interaction changes as models improve
—
John Yang
* SWE-bench: https://www.swebench.com
* X: https://x.com/jyangballin
Full Video Episode
Timestamps
00:00:00 Introduction: John Yang on SWE-bench and Code Evaluations00:00:31 SWE-bench Origins and Devon's Impact on the Coding Agent Arms Race00:01:09 SWE-bench Ecosystem: Verified, Pro, Multimodal, and Multilingual Variants00:02:17 Moving Beyond Django: Diversifying Code Evaluation Repositories00:03:08 Code Clash: Long-Horizon Development Through Programming Tournaments00:04:41 From Halite to Economic Value: Designing Competitive Coding Arenas00:06:04 Ofir's Lab: SWE-ficiency, AlgoTune, and SciCode for Scientific Computing00:07:52 The Benchmark Landscape: TAU-bench, Terminal-bench, and User Simulation00:09:20 The Impossible Task Debate: Refusals, Ambiguity, and Benchmark Integrity00:12:32 The Future of Code Evals: Long Autonomy vs Human-AI Collaboration00:14:37 Call to Action: User Interaction Data and Codebase Understanding Research
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Universities have been in the crosshairs of the White House since President Trump took office — and Princeton University president Christopher Eisgruber is one of a handful of college administrators who have spoken out against it.
Kara speaks to the Eisgruber about his new book, Terms of Respect: How Colleges Get Free Speech Right, and right-wing attacks on universities that come under the guise of free speech, including from the late conservative activist Charlie Kirk and his organization Turning Point USA. They discuss why some campus leaders have fought against (and others have complied) with the Trump administration’s investigations into allegations of antisemitism and demands to overhaul diversity programs in college admissions and hiring. And they talk about the long-term impacts of losing academic freedom on the reputation and success of US higher education, the economy and society as a whole.
Please note: this interview was recorded on Monday September 29th, before President Trump said his administration was nearing a deal with Harvard while it also began a process called debarment that could allow it to bar the university from future federal grants.
Want to see Kara (and Scott Galloway) live on the Pivot Tour November 8th - 14th? Find tickets and details at PivotTour.com.
Questions? Comments? Email us at on@voxmedia.com or find us on YouTube, Instagram, TikTok, Threads, and Bluesky @onwithkaraswisher.
Learn more about your ad choices. Visit podcastchoices.com/adchoices
As more companies push AI in their workplaces, the technology is rapidly reshaping the way many of us do our jobs. But a lot of people — from entry-level employees to the C-Suite — are still in the dark about the limits of AI, its best uses, and how to make it work for them.
We called in a panel of AI experts to answer some our listeners’ burning questions about how to use it at work: Sayash Kapoor, co-author of the book AI Snake Oil: What Artificial Can Do, What it Can’t, and How to Tell the Difference and the Substack AI as Normal Technology; Rajeev Kapur, CEO of 1105 Media and author of the book AI Made Simple: A Beginner's Guide to Generative Intelligence; and futurist and author Amy Webb, founder and CEO of the consulting firm Future Today Strategy Group.
Kara, Sayash, Rajeev and Amy break down everything from how vibe coding works to thornier questions around privacy and regulation. They talk about how young people can prepare themselves to enter the workforce, and how all of us can develop skills to stay relevant. And, of course, they weigh in on the question so many of us are asking right now: Is AI coming for my job?
Questions? Comments? Email us at on@voxmedia.com or find us on YouTube, Instagram, TikTok, and Bluesky @onwithkaraswisher.
Learn more about your ad choices. Visit podcastchoices.com/adchoices
This week, we check in on the state of artificial intelligence in education. We talk with a co-founder of Alpha Schools, MacKenzie Price, about how her private K-12 schools are using A.I. to generate personalized lesson plans and enabling teachers to spend their time motivating rather than teaching students. Then, the Princeton historian D. Graham Burnett joins us to discuss the existential threat that A.I. poses to the traditional humanities degree and why he believes we’ll see thousands of new schools emerge outside the university system to carry on the exploration of what it means to be a person in the world. And finally, we hear directly from students who are on the front line of technological change.
Guests:
MacKenzie Price, co-founder of Alpha Schools
D. Graham Burnett, historian of science and technology at Princeton University
Additional Reading:
A.I.-Driven Education: Founded in Texas and Coming to a School Near You
Will the Humanities Survive Artificial Intelligence?
OpenAI and Microsoft Bankroll New A.I. Training for Teachers
Welcome to Campus. Here’s Your ChatGPT.
We want to hear from you. Email us at hardfork@nytimes.com. Find “Hard Fork” on YouTube and TikTok.
Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Two visions for the future of AI clash in this debate between Daniel Kokotajlo and Arvind Narayanan.
Is AI a revolutionary new species destined for runaway superintelligence, or just another step in humanity’s technological evolution—like electricity or the internet?
Daniel, a former OpenAI researcher and author of AI 2027, argues for a fast-approaching intelligence explosion. Arvind, a Princeton professor and co-author of AI Snake Oil, contends that AI is powerful but ultimately controllable and slow to reshape society. Moderated by Ryan and David, this conversation dives into the crux of capability vs. power, economic transformation, and the future of democratic agency in an AI-driven world.
------
💫 LIMITLESS | SUBSCRIBE & FOLLOW
https://limitless.bankless.com/
https://x.com/LimitlessFT
------
BANKLESS SPONSOR TOOLS:
🪙FRAX | SELF SUFFICIENT DeFi
https://bankless.cc/Frax
🦄UNISWAP | SWAP ON UNICHAIN
https://bankless.cc/unichain
🛞MANTLE | MODULAR LAYER 2 NETWORK
https://bankless.cc/Mantle
🌐SELF | PROVE YOUR SELF
https://bankless.cc/Self
🟠BINANCE | THE WORLDS #1 CRYPTO EXCHANGE
https://bankless.cc/binance
------
TIMESTAMPS
0:00 Intro
4:44 Is AI Normal Tech?
30:20 Capability & Power
44:52 Manageable or Existential?
52:36 AGI Milestone
1:00:31 AI by 2030
1:12:07 Making Sense
1:16:08 Closing & Disclaimers
------
RESOURCES
Daniel Kokotajlo
https://x.com/dkokotajlo
Arvind Narayanan
https://x.com/random_walker
Two Paths For AI - Daniel
https://www.newyorker.com/culture/open-questions/two-paths-for-ai
AI as Normal Technology - Arvind
https://knightcolumbia.org/content/ai-as-normal-technology
AI 2027 - Daniel
https://ai-2027.com/
AI Snake Oil - Arvind
https://www.aisnakeoil.com/
https://www.amazon.com/Snake-Oil-Artificial-Intelligence-Difference/dp/069124913X
Pause AI
https://pauseai.info/pdoom
------
Not financial or tax advice. See our investment disclosures here:
https://www.bankless.com/disclosures
This week Meta is on trial, in a landmark case over whether it illegally snuffed out competition when it acquired Instagram and WhatsApp. We discuss some of the most surprising revelations from old email messages made public as evidence in the case, and explain why we think the F.T.C.’s argument has gotten weaker in the years since the lawsuit was filed. Then we hear from Princeton computer scientist Arvind Narayanan on why he believes it will take decades, not years, for A.I. to transform society in the ways the big A.I. labs predict. And finally, what do dolphins, Katy Perry and A1 steak sauce have in common? They’re all important characters in our latest round of HatGPT.
Tickets to Hard Fork live are on sale now! See us June 24 at SFJAZZ.
Guest:
Arvind Narayanan, director of the Center for Information Technology at Princeton and co-author of “AI Snake Oil: What Artificial Intelligence Can Do, What it Can’t, and How to Tell the Difference.”
Additional Reading:
What if Mark Zuckerberg Had Not Bought Instagram and WhatsApp?
AI as Normal Technology
One Giant Stunt for Womankind
We want to hear from you. Email us at hardfork@nytimes.com. Find “Hard Fork” on YouTube and TikTok.
Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Today, we're joined by Arvind Narayanan, professor of Computer Science at Princeton University to discuss his recent works, AI Agents That Matter and AI Snake Oil. In “AI Agents That Matter”, we explore the range of agentic behaviors, the challenges in benchmarking agents, and the ‘capability and reliability gap’, which creates risks when deploying AI agents in real-world applications. We also discuss the importance of verifiers as a technique for safeguarding agent behavior. We then dig into the AI Snake Oil book, which uncovers examples of problematic and overhyped claims in AI. Arvind shares various use cases of failed applications of AI, outlines a taxonomy of AI risks, and shares his insights on AI’s catastrophic risks. Additionally, we also touched on different approaches to LLM-based reasoning, his views on tech policy and regulation, and his work on CORE-Bench, a benchmark designed to measure AI agents' accuracy in computational reproducibility tasks.
The complete show notes for this episode can be found at https://twimlai.com/go/704.
Arvind Narayanan is a professor of Computer Science at Princeton and the director of the Center for Information Technology Policy. He is a co-author of the book AI Snake Oil and a big proponent of the AI scaling myths around the importance of just adding more compute. He is also the lead author of a textbook on the computer science of cryptocurrencies which has been used in over 150 courses around the world, and an accompanying Coursera course that has had over 700,000 learners.
In Today's Episode with Arvind Narayanan We Discuss:
1. Compute, Data, Algorithms: What is the Bottleneck:
Why does Arvind disagree with the commonly held notion that more compute will result in an equal and continuous level of model performance improvement?
Will we continue to see players move into the compute layer in the need to internalise the margin? What does that mean for Nvidia?
Why does Arvind not believe that data is the bottleneck? How does Arvind analyse the future of synthetic data? Where is it useful? Where is it not?
2. The Future of Models:
Does Arvind agree that this is the fastest commoditization of a technology he has seen?
How does Arvind analyse the future of the model landscape? Will we see a world of few very large models or a world of many unbundled and verticalised models?
Where does Arvind believe the most value will accrue in the model layer?
Is it possible for smaller companies or university research institutions to even play in the model space given the intense cash needed to fund model development?
3. Education, Healthcare and Misinformation: When AI Goes Wrong:
What are the single biggest dangers that AI poses to society today?
To what extent does Arvind believe misinformation through generative AI is going to be a massive problem in democracies and misinformation?
How does Arvind analyse AI impacting the future of education? What does he believe everyone gets wrong about AI and education?
Does Arvind agree that AI will be able to put a doctor in everyone's pocket? Where does he believe this theory is weak and falls down?
How seriously should governments take the threat of existential risk from AI, given the lack of consensus among researchers? On the one hand, existential risks (x-risks) are necessarily somewhat speculative: by the time there is concrete evidence, it may be too late. On the other hand, governments must prioritize — after all, they don’t worry too much about x-risk from alien invasions.
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at brave.com/api.
Sayash Kapoor is a computer science Ph.D. candidate at Princeton University's Center for Information Technology Policy. His research focuses on the societal impact of AI. Kapoor has previously worked on AI in both industry and academia, with experience at Facebook, Columbia University, and EPFL Switzerland. He is a recipient of a best paper award at ACM FAccT and an impact recognition award at ACM CSCW. Notably, Kapoor was included in TIME's inaugural list of the 100 most influential people in AI.
Sayash Kapoor
https://x.com/sayashk
https://www.cs.princeton.edu/~sayashk/
Arvind Narayanan (other half of the AI Snake Oil duo)
https://x.com/random_walker
AI existential risk probabilities are too unreliable to inform policy
https://www.aisnakeoil.com/p/ai-existential-risk-probabilities
Pre-order AI Snake Oil Book
https://amzn.to/4fq2HGb
AI Snake Oil blog
https://www.aisnakeoil.com/
AI Agents That Matter
https://arxiv.org/abs/2407.01502
Shortcut learning in deep neural networks
https://www.semanticscholar.org/paper/Shortcut-learning-in-deep-neural-networks-Geirhos-Jacobsen/1b04936c2599e59b120f743fbb30df2eed3fd782
77% Of Employees Report AI Has Increased Workloads And Hampered Productivity, Study Finds
https://www.forbes.com/sites/bryanrobinson/2024/07/23/employees-report-ai-increased-workload/
TOC:
00:00:00 Intro
00:01:57 How seriously should we take Xrisk threat?
00:02:55 Risk too unrealiable to inform policy
00:10:20 Overinflated risks
00:12:05 Perils of utility maximisation
00:13:55 Scaling vs airplane speeds
00:17:31 Shift to smaller models?
00:19:08 Commercial LLM ecosystem
00:22:10 Synthetic data
00:24:09 Is AI complexifying our jobs?
00:25:50 Does ChatGPT make us dumber or smarter?
00:26:55 Are AI Agents overhyped?
00:28:12 Simple vs complex baselines
00:30:00 Cost tradeoff in agent design
00:32:30 Model eval vs downastream perf
00:36:49 Shortcuts in metrics
00:40:09 Standardisation of agent evals
00:41:21 Humans in the loop
00:43:54 Levels of agent generality
00:47:25 ARC challenge
Today we're joined by Azarakhsh (Aza) Jalalvand, a research scholar at Princeton University, to discuss his work using deep reinforcement learning to control plasma instabilities in nuclear fusion reactors. Aza explains his team developed a model to detect and avoid a fatal plasma instability called ‘tearing mode’. Aza walks us through the process of collecting and pre-processing the complex diagnostic data from fusion experiments, training the models, and deploying the controller algorithm on the DIII-D fusion research reactor. He shares insights from developing the controller and discusses the future challenges and opportunities for AI in enabling stable and efficient fusion energy production.
The complete show notes for this episode can be found at twimlai.com/go/682.
Today we’re joined by Sayash Kapoor, a Ph.D. student in the Department of Computer Science at Princeton University. Sayash walks us through his paper: "On the Societal Impact of Open Foundation Models.” We dig into the controversy around AI safety, the risks and benefits of releasing open model weights, and how we can establish common ground for assessing the threats posed by AI. We discuss the application of the framework presented in the paper to specific risks, such as the biosecurity risk of open LLMs, as well as the growing problem of "Non Consensual Intimate Imagery" using open diffusion models.
The complete show notes for this episode can be found at twimlai.com/go/675.
Today, we continue our NeurIPS series with Dan Friedman, a PhD student in the Princeton NLP group. In our conversation, we explore his research on mechanistic interpretability for transformer models, specifically his paper, Learning Transformer Programs. The LTP paper proposes modifications to the transformer architecture which allow transformer models to be easily converted into human-readable programs, making them inherently interpretable. In our conversation, we compare the approach proposed by this research with prior approaches to understanding the models and their shortcomings. We also dig into the approach’s function and scale limitations and constraints.
The complete show notes for this episode can be found at twimlai.com/go/667.
Research has shown that humans possess strong inductive biases which enable them to quickly learn and generalize. In order to instill these same useful human inductive biases into machines, a paper was presented by Sreejan Kumar at the NeurIPS conference which won the Outstanding Paper of the Year award. The paper is called Using Natural Language and Program Abstractions to Instill Human Inductive Biases in Machines.
This paper focuses on using a controlled stimulus space of two-dimensional binary grids to define the space of abstract concepts that humans have and a feedback loop of collaboration between humans and machines to understand the differences in human and machine inductive biases.
It is important to make machines more human-like to collaborate with them and understand their behavior. Synthesised discrete programs running on a turing machine computational model instead of a neural network substrate offers promise for the future of artificial intelligence. Neural networks and program induction should both be explored to get a well-rounded view of intelligence which works in multiple domains, computational substrates and which can acquire a diverse set of capabilities.
Natural language understanding in models can also be improved by instilling human language biases and programs into AI models. Sreejan used an experimental framework consisting of two dual task distributions, one generated from human priors and one from machine priors, to understand the differences in human and machine inductive biases. Furthermore, he demonstrated that compressive abstractions can be used to capture the essential structure of the environment for more human-like behavior. This means that emergent language-based inductive priors can be distilled into artificial neural networks, and AI models can be aligned to the us, world and indeed, our values.
Humans possess strong inductive biases which enable them to quickly learn to perform various tasks. This is in contrast to neural networks, which lack the same inductive biases and struggle to learn them empirically from observational data, thus, they have difficulty generalizing to novel environments due to their lack of prior knowledge.
Sreejan's results showed that when guided with representations from language and programs, the meta-learning agent not only improved performance on task distributions humans are adept at, but also decreased performa on control task distributions where humans perform poorly. This indicates that the abstraction supported by these representations, in the substrate of language or indeed, a program, is key in the development of aligned artificial agents with human-like generalization, capabilities, aligned values and behaviour.
References
Using natural language and program abstractions to instill human inductive biases in machines [Kumar et al/NEURIPS]
https://openreview.net/pdf?id=buXZ7nIqiwE
Core Knowledge [Elizabeth S. Spelke / Harvard]
https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf
The Debate Over Understanding in AI's Large Language Models [Melanie Mitchell]
https://arxiv.org/abs/2210.13966
On the Measure of Intelligence [Francois Chollet]
https://arxiv.org/abs/1911.01547
ARC challenge [Chollet]
https://github.com/fchollet/ARC
Stephen Kotkin is a historian specializing in Stalin and Soviet history. Please support this podcast by checking out our sponsors:
– Lambda: https://lambdalabs.com/lex
– Scale: https://scale.com/lex
– Athletic Greens: https://athleticgreens.com/lex and use code LEX to get 1 month of fish oil
– ExpressVPN: https://expressvpn.com/lexpod and use code LexPod to get 3 months free
– ROKA: https://roka.com/ and use code LEX to get 20% off your first order
EPISODE LINKS:
Stephen’s Website: https://history.princeton.edu/people/stephen-kotkin
Stalin: 1878-1928 (Vol 1): https://amzn.to/3NvokpC
Stalin: 1929-1941 (Vol 2): https://amzn.to/3wIYqsT
PODCAST INFO:
Podcast website: https://lexfridman.com/podcast
Apple Podcasts: https://apple.co/2lwqZIr
Spotify: https://spoti.fi/2nEwCF8
RSS: https://lexfridman.com/feed/podcast/
YouTube Full Episodes: https://youtube.com/lexfridman
YouTube Clips: https://youtube.com/lexclips
SUPPORT & CONNECT:
– Check out the sponsors above, it’s the best way to support this podcast
– Support on Patreon: https://www.patreon.com/lexfridman
– Twitter: https://twitter.com/lexfridman
– Instagram: https://www.instagram.com/lexfridman
– LinkedIn: https://www.linkedin.com/in/lexfridman
– Facebook: https://www.facebook.com/lexfridman
– Medium: https://medium.com/@lexfridman
OUTLINE:
Here’s the timestamps for the episode. On some podcast players you should be able to click the timestamp to jump to that time.
(00:00) – Introduction
(10:17) – Putin and Stalin
(21:07) – Putin vs the West
(43:59) – Response to Oliver Stone
(55:05) – Russian invasion of Ukraine
(1:34:33) – Putin’s plan for the war
(1:42:32) – Henry Kissinger
(1:48:26) – Nuclear war
(1:59:00) – Parallels to World War II
(2:21:45) – China
(2:29:54) – World War III
(2:37:23) – Navalny
(2:41:40) – Meaning of life
Today we have a special treat. A conversation with Brian Kernighan! Brian’s been in the software game since the beginning of Unix. Yes, he was there at Bell Labs when it all began. And he is still at it today, writing books and teaching the next generation at Princeton.
This is an epic and wide ranging conversation. You’ll hear about the birth of Unix, Ken Thompson’s unique skillset, why Brian thinks C has stood the test of time, his thoughts on modern languages like Go and Rust, what’s changed in 50 years of software, what makes platforms like Unix and the web so powerful, his take as a professor on the trend of programmers skipping the university track, and so much more.
Seriously, this is a must-listen.
Join the discussion
Changelog++ members get a bonus 2 minutes at the end of this episode and zero ads. Join today!
Sponsors:
Square – Develop on the platform that sellers trust. There is a massive opportunity for developers to support Square sellers by building apps for today’s business needs. Learn more at changelog.com/square to dive into the docs, APIs, SDKs and to create your Square Developer account — tell them Changelog sent you.
InfluxData – The time series platform for building and operating time series applications — InfluxDB empowers developers to build IoT, analytics, and monitoring software. It’s purpose-built to handle massive volumes and countless sources of time-stamped data produced by sensors, applications, and infrastructure. Learn more at influxdata.com/changelog
FireHydrant – The reliability platform for every developer. Incidents impact everyone, not just SREs. FireHydrant gives teams the tools to maintain service catalogs, respond to incidents, communicate through status pages, and learn with retrospectives. Try FireHydrant free for 14 days at firehydrant.io
MongoDB – An integrated suite of cloud database and services — They have a FREE forever tier, so you can prove to yourself and to your team that they have everything you need. Check it out today at mongodb.com/changelog
Featuring:
Brian Kernighan – Website
Adam Stacoviak – Website, GitHub, LinkedIn, Mastodon, X
Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
Brian’s Wikipedia
PDP-7 on Wikipedia
Understanding the Digital World: What You Need to Know about Computers, the Internet, Privacy, and Security
The C Programming Language
UNIX: A History and a Memoir
The Go Programming Language
The Mythical Man-Month
Dick Hamming’s You and Your Research
Brian on Lex Fridman’s podcast
Something missing or broken? PRs welcome!
Daniel Kahneman is a Nobel prize-winning psychologist and economist and author of Thinking, Fast and Slow, a landmark book that decodes human decision-making. Yann LeCun is the chief AI scientist at Meta (Facebook) and a pioneer in the field of deep learning, which the cutting edge of AI is based on today. The two come together on Big Technology Podcast this week to discuss how machines and humans learn, whether there are parallels, and what each field can learn from each other.
Learn more about your ad choices. Visit megaphone.fm/adchoices
This week's guest is Princeton professor and co-founder of AI4All, Olga Russakovsky. Olga's research has focused on the societal biases in the data used to train machine learning AI. In our second episode, she talks about why we need more diversity in AI, problematic junk in, junk out data sets and what it was like being the only female researcher in early AI labs. Host: Pieter Abbeel Executive Producers: Ricardo Reyes & Henry Tobias Jones Audio Production: Kieron Matthew Banerji Hosted on Acast. See acast.com/privacy for more information.
Sam Harris speaks with Zeynep Tufekci about the problem of misinformation and group-think. They discuss the Covid-19 pandemic, the early failures of journalists and public health professionals to make sense of it, the sociology of mask wearing, the problem of correcting institutional errors, Covid as a dress rehearsal for something far worse, asymmetric information warfare, failures of messaging about vaccines, the paradox of scientific authority, the power of incentives, how to reform social media, and other topics.
If the Making Sense podcast logo in your player is BLACK, you can SUBSCRIBE to gain access to all full-length episodes at samharris.org/subscribe.
Social media-fueled protests are a force to be reckoned with in politics today. Movements like Occupy Wall Street, the Arab Spring, and Black Lives Matter have drawn millions into the streets in protest of central authorities. But can these movements be effective in the long term? Alex sits down with the field's leading writer and researcher, Zeynep Tufekci, to talk it through.
Learn more about your ad choices. Visit megaphone.fm/adchoices
Brian Kernighan is a professor of computer science at Princeton University. He co-authored the C Programming Language with Dennis Ritchie (creator of C) and has written a lot of books on programming, computers, and life including the Practice of Programming, the Go Programming Language, his latest UNIX: A History and a Memoir. He co-created AWK, the text processing language used by Linux folks like myself. He co-designed AMPL, an algebraic modeling language for large-scale optimization.
Support this podcast by supporting our sponsors:
– Eight Sleep: https://eightsleep.com/lex
– Raycon: http://buyraycon.com/lex
If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
OUTLINE:
00:00 – Introduction
04:24 – UNIX early days
22:09 – Unix philosophy
31:54 – Is programming art or science?
35:18 – AWK
42:03 – Programming setup
46:39 – History of programming languages
52:48 – C programming language
58:44 – Go language
1:01:57 – Learning new programming languages
1:04:57 – Javascript
1:08:16 – Variety of programming languages
1:10:30 – AMPL
1:18:01 – Graph theory
1:22:20 – AI in 1964
1:27:50 – Future of AI
1:29:47 – Moore’s law
1:32:54 – Computers in our world
1:40:37 – Life
John Hopfield is professor at Princeton, whose life’s work weaved beautifully through biology, chemistry, neuroscience, and physics. Most crucially, he saw the messy world of biology through the piercing eyes of a physicist. He is perhaps best known for his work on associate neural networks, now known as Hopfield networks that were one of the early ideas that catalyzed the development of the modern field of deep learning.
EPISODE LINKS:
Now What? article: http://bit.ly/3843LeU
John wikipedia: https://en.wikipedia.org/wiki/John_Hopfield
Books mentioned:
– Einstein’s Dreams: https://amzn.to/2PBa96X
– Mind is Flat: https://amzn.to/2I3YB84
This conversation is part of the Artificial Intelligence podcast. If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
This episode is presented by Cash App. Download it (App Store, Google Play), use code “LexPodcast”.
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
OUTLINE:
00:00 – Introduction
02:35 – Difference between biological and artificial neural networks
08:49 – Adaptation
13:45 – Physics view of the mind
23:03 – Hopfield networks and associative memory
35:22 – Boltzmann machines
37:29 – Learning
39:53 – Consciousness
48:45 – Attractor networks and dynamical systems
53:14 – How do we build intelligent systems?
57:11 – Deep thinking as the way to arrive at breakthroughs
59:12 – Brain-computer interfaces
1:06:10 – Mortality
1:08:12 – Meaning of life
Daniel Kahneman is winner of the Nobel Prize in economics for his integration of economic science with the psychology of human behavior, judgment and decision-making. He is the author of the popular book “Thinking, Fast and Slow” that summarizes in an accessible way his research of several decades, often in collaboration with Amos Tversky, on cognitive biases, prospect theory, and happiness. The central thesis of this work is a dichotomy between two modes of thought: “System 1” is fast, instinctive and emotional; “System 2” is slower, more deliberative, and more logical. The book delineates cognitive biases associated with each type of thinking.
This conversation is part of the Artificial Intelligence podcast. If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
This episode is presented by Cash App. Download it (App Store, Google Play), use code “LexPodcast”.
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
00:00 – Introduction
02:36 – Lessons about human behavior from WWII
08:19 – System 1 and system 2: thinking fast and slow
15:17 – Deep learning
30:01 – How hard is autonomous driving?
35:59 – Explainability in AI and humans
40:08 – Experiencing self and the remembering self
51:58 – Man’s Search for Meaning by Viktor Frankl
54:46 – How much of human behavior can we study in the lab?
57:57 – Collaboration
1:01:09 – Replication crisis in psychology
1:09:28 – Disagreements and controversies in psychology
1:13:01 – Test for AGI
1:16:17 – Meaning of life
Stephen Kotkin is a professor of history at Princeton university and one of the great historians of our time, specializing in Russian and Soviet history. He has written many books on Stalin and the Soviet Union including the first 2 of a 3 volume work on Stalin, and he is currently working on volume 3.
This conversation is part of the Artificial Intelligence podcast. If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
This episode is presented by Cash App. Download it (App Store, Google Play), use code “LexPodcast”.
Episode Links:
Stalin (book, vol 1): https://amzn.to/2FjdLF2
Stalin (book, vol 2): https://amzn.to/2tqyjc3
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
00:00 – Introduction
03:10 – Do all human beings crave power?
11:29 – Russian people and authoritarian power
15:06 – Putin and the Russian people
23:23 – Corruption in Russia
31:30 – Russia’s future
41:07 – Individuals and institutions
44:42 – Stalin’s rise to power
1:05:20 – What is the ideal political system?
1:21:10 – Questions for Putin
1:29:41 – Questions for Stalin
1:33:25 – Will there always be evil in the world?
This week on the podcast we’re featuring a series of conversations from the NIPs conference in Long Beach, California. I attended a bunch of talks and learned a ton, organized an impromptu roundtable on Building AI Products, and met a bunch of great people, including some former TWiML Talk guests. In this episode I speak with Yael Niv, professor of neuroscience and psychology at Princeton University. Yael joined me after her invited talk on “Learning State Representations.” In this interview Yael and I explore the relationship between neuroscience and machine learning. In particular, we discusses the importance of state representations in human learning, some of her experimental results in this area, and how a better understanding of representation learning can lead to insights into machine learning problems such as reinforcement and transfer learning. Did I mention this was a nerd alert show? I really enjoyed this interview and I know you will too. Be sure to send over any thoughts or feedback via the show notes page at twimlai.com/talk/92.