The Race to Production-Grade Diffusion LLMs with Stefano Ermon - #764
Today, we're joined by Stefano Ermon, associate professor at Stanford University and CEO of Inception Labs to discuss diffusion language models. We dig into how diffusion approaches—traditionally used for images—are being adapted for text and code generation, the technical challenges of applying continuous methods to discrete token spaces, and how diffusion models compare to traditional autoregressive LLMs. Stefano introduces Mercury 2, a commercial-scale diffusion LLM that can generate multiple tokens simultaneously and achieve inference speeds 5-10x faster than small frontier models, paving the way for latency-sensitive applications like voice interactions and fast agentic loops. We also cover the open research challenges in diffusion LLM training, serving infrastructure requirements, and post-training for diffusion-based systems. Finally, Stefano shares his perspective on whether diffusion models can rival or surpass autoregressive LLMs at scale, the advantages for highly controllable generation, and what the future of multimodal diffusion models might look like.
The complete show notes for this episode can be found at https://twimlai.com/go/764.
Inception Labs says its diffusion LLM is 10x faster than Claude, ChatGPT, Gemini
On a recent episode of the The New Stack Agents, Inception Labs CEO Stefano Ermon introduced Mercury 2, a large language model built on diffusion rather than the standard autoregressive approach. Traditional LLMs generate text token by token from left to right, which Ermon describes as “fancy autocomplete.” In contrast, diffusion models begin with a rough draft and refine it in parallel, similar to image systems like Stable Diffusion.
This parallel process allows Mercury 2 to produce over 1,000 tokens per second—five to ten times faster than optimized models from labs such as OpenAI, Anthropic, and Google, according to company tests. Ermon argues diffusion models better leverage GPUs, with support from investor Nvidia to optimize performance.
While Mercury 2 matches mid-tier models like Claude Haiku and Google Flash rather than top systems such as Claude Opus or GPT-4, Ermon believes diffusion’s speed and economic advantages will become increasingly compelling as AI applications scale.
Learn more from The New Stack about the latest developments around around large language model built on diffusion:
How Diffusion-Based LLM AI Speeds Up Reasoning
Get Ready for Faster Text Generation With Diffusion LLMs
Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Generating text with diffusion (and ROI with LLMs)
Two guests for the price of one! This episode has two interviews recorded at AWS re:Invent back in December. In part 1, Ryan chats with the co-founder and CEO of Inception, Stefano Ermon, about diffusion language models and how their multiple token generation compares to traditional LLMs (spoiler: they’re faster and more accurate). In the second half of the episode, Ryan and the chairman of Roomie, Aldo Luevano, dive into Roomie’s purpose built models for both physical and software AI, and how their ROI-first approach helps companies track the impact of their robotics and AI implementation.
Episode notes:
Inception researches and builds diffusion language models for faster and more efficient AI.
Roomie is a robotics and enterprise AI company with an ROI-first platform that tracks how well their AI solutions are actually working.
Connect with Stefano on LinkedIn.
Connect with Aldo on LinkedIn.
TRANSCRIPT
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
#310 Stefano Ermon: Why Diffusion Language Models Will Define the Next Generation of LLMs
This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents.
Visit https://agntcy.org/ and add your support. Most large language models today generate text one token at a time. That design choice creates a hard limit on speed, cost, and scalability. In this episode of Eye on AI, Stefano Ermon breaks down diffusion language models and why a parallel, inference-first approach could define the next generation of LLMs. We explore how diffusion models differ from autoregressive systems, why inference efficiency matters more than training scale, and what this shift means for real-time AI applications like code generation, agents, and voice systems. This conversation goes deep into AI architecture, model controllability, latency, cost trade-offs, and the future of generative intelligence as AI moves from demos to production-scale systems. Stay Updated: Craig Smith on X: https://x.com/craigssEye on A.I. on X: https://x.com/EyeOn_AI (00:00) Autoregressive vs Diffusion LLMs (02:12) Why Build Diffusion LLMs (05:51) Context Window Limits (08:39) How Diffusion Works (11:58) Global vs Token Prediction (17:19) Model Control and Safety (19:48) Training and RLHF (22:35) Evaluating Diffusion Models (24:18) Diffusion LLM Competition (30:09) Why Start With Code (32:04) Enterprise Fine-Tuning (33:16) Speed vs Accuracy Tradeoffs (35:34) Diffusion vs Autoregressive Future (38:18) Coding Workflows in Practice (43:07) Voice and Real-Time Agents (44:59) Reasoning Diffusion Models (46:39) Multimodal AI Direction (50:10) Handling Hallucinations
Beyond Uncanny Valley: Breaking Down Sora
In early 2024, the notion of high fidelity, believable AI-generated video seemed a distant future to many. Yet, a mere few weeks into the year, OpenAI unveiled Sora, its new state of the art text-to-video model producing videos of up to 60 seconds. The output shattered expectations – even for other builders and researchers within generative AI – sparking widespread speculation and awe.
How does Sora achieve such realism? And are explicit 3D modeling techniques or game engines at play?
In this episode of the a16z Podcast, a16z General Partner Anjney Midha connects with Stefano Ermon, Professor of Computer Science at Stanford and key figure at the lab behind the diffusion models now used in Sora, ChatGPT, and Midjourney. Together, they delve into the challenges of video generation, the cutting-edge mechanics of Sora, and what this all could mean for the road ahead.
Resources:
Find Stefano on Twitter: https://twitter.com/stefanoermon
Find Anjney on Twitter: https://twitter.com/anjneymidha
Learn more about Stefano’s Deep Generative Models course: :
https://deepgenerativemodels.github.io
Stay Updated:
Find a16z on Twitter: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Subscribe on your favorite podcast app: https://a16z.simplecast.com/
Follow our host: https://twitter.com/stephsmithio
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - TWiML Talk #15
My guest this week is Stefano Ermon, Assistant Professor of Computer Science at Stanford University, and Fellow at Stanford’s Woods Institute for the Environment. Stefano and I met at the Re-Work Deep Learning Summit earlier this year, where he gave a presentation on Machine Learning for Sustainability. Stefano and I spoke about a wide range of topics, including the relationship between fundamental and applied machine learning research, incorporating domain knowledge in machine learning models, dimensionality reduction, and his interest in applying ML & AI to addressing sustainability issues such as poverty, food security and the environment. The show notes can be found at twimlai.com/talk/15.