You’re probably underutilizing your GPUs
Ryan is joined by Jared Quincy Davis, CEO and co-founder of Mithril, to explore the importance of efficient resource allocation and GPU utilization in AI, the myth and misconceptions of the GPU shortage, and how the economics of GPU will change with new scheduling and utilization strategies.
Episode notes:
Mithril’s omnicloud platform aggregates and orchestrates multi-cloud GPUs, CPUs, and storage so you can access your infrastructure through a single platform.
Connect with Jared on Twitter and LinkedIn.
Shoutout to user Razzi Abuissa for winning a Populist badge on their answer to How to find last merge in git?.
TRANSCRIPT
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Infrastructure Scaling and Compound AI Systems with Jared Quincy Davis - #740
In this episode, Jared Quincy Davis, founder and CEO at Foundry, introduces the concept of "compound AI systems," which allows users to create powerful, efficient applications by composing multiple, often diverse, AI models and services. We discuss how these "networks of networks" can push the Pareto frontier, delivering results that are simultaneously faster, more accurate, and even cheaper than single-model approaches. Using examples like "laconic decoding," Jared explains the practical techniques for building these systems and the underlying principles of inference-time scaling. The conversation also delves into the critical role of co-design, where the evolution of AI algorithms and the underlying cloud infrastructure are deeply intertwined, shaping the future of agentic AI and the compute landscape.
The complete show notes for this episode can be found at https://twimlai.com/go/740.
The marketplace for AI compute with Jared Quincy Davis from Foundry
In this episode of No Priors, hosts Sarah and Elad are joined by Jared Quincy Davis, former DeepMind researcher and the Founder and CEO of Foundry, a new AI cloud computing service provider. They discuss the research problems that led him to starting Foundry, the current state of GPU cloud utilization, and Foundry's approach to improving cloud economics for AI workloads. Jared also touches on his predictions for the GPU market and the thinking behind his recent paper on designing compound AI systems.
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @jaredq_
Show Notes:
(00:00) Introduction
(02:42) Foundry background
(03:57) GPU utilization for large models
(07:29) Systems to run a large model
(09:54) Historical value proposition of the cloud
(14:45) Sharing cloud compute to increase efficiency
(19:17) Foundry’s new releases
(23:54) The current state of GPU capacity
(29:50) GPU market dynamics
(36:28) Compound systems design
(40:27) Improving open-ended tasks