Maximizing GPU Utilization: Heterogeneous Pipelines with Ray and Kubernetes
Summary
In this episode Robert Nishihara, co-founder of Anyscale and co-creator of Ray, talks about maximizing hardware utilization for AI and data-intensive workloads. He explores Ray’s evolution alongside Kubernetes and PyTorch, and why consolidation at these layers has enabled a new generation of complex, heterogeneous workloads. Robert explains how data preparation has shifted to GPU- and inference-heavy, multimodal pipelines; where Ray fits compared to Spark and workflow orchestrators; and why Ray excels at composing heterogeneous pools of compute, handling failures, and scaling complex systems like multi-node LLM inference and reinforcement learning. He digs into practical strategies for boosting GPU utilization across training and inference, elasticity and prioritization of workloads, topology-aware scheduling, and the importance of fast failure recovery as hardware scales from nodes to racks. If you’re wrestling with expensive GPUs, multimodal data curation, or cross-node LLM inference, this conversation offers concrete mental models and architectural guidance.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Your host is Tobias Macey and today I'm interviewing Robert Nishihara about the challenges of maximizing the utility of your available hardware for AI applications
Interview
Introduction
How did you get involved in the area of data management?
Can you start by giving an overview of the major contributors to wasted or idle compute?
Why does it matter if the available compute isn't being maximized?
What are some of the typical ad-hoc methods that teams might use to try to get the most out of their available hardware (especially GPUs)?
What are the most interesting, innovative, or unexpected ways that you have seen Ray used?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on Ray and distributed compute for data and AI?
When is Ray the wrong choice?
What do you have planned for the future of Ray?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
Links
AnyScale
Ray
Deep Learning
Computer Vision
Kubernetes
Cursor
Claude Code
Kube-Ray
PyTorch
Tensorflow
Theano
Caffe
vLLM
SGLang
Ray Tune
Neural Network
Learning Rates
Reinforcement Learning
AlphaGo
Cursor Composer 2
ImageNet
Transformer Architecture
Stochastic Gradient Descent
Airflow
Dagster
Flyte
Mixture of Experts
Prefill
Temporal
Actor Framework
RDMA == Remote Direct Memory Access
Neoclouds
AI Engineering Podcast Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
#547: Parallel Python at Anyscale with Ray
When OpenAI trained GPT-3, they didn't roll their own orchestration layer. They used Ray, an open source Python framework born out of the same Berkeley research lab lineage that gave us Apache Spark. And here's the twist: Ray was originally built for reinforcement learning research, then quietly faded as RL hit a wall. Until ChatGPT showed up. Suddenly reinforcement learning was back, as the post-training step that turns a raw language model into something genuinely useful.
Edward Oakes and Richard Liaw, two founding engineers behind Ray and Anyscale, join me on Talk Python to tell that story. We'll trace Ray from its RISE Lab origins at UC Berkeley to powering some of the largest training runs in the world. We'll talk about what Ray actually is, a distributed execution engine for AI workloads, and how a few lines of Python become work running across hundreds of GPUs. We'll cover Ray Data for multimodal pipelines, the dashboard, the VS Code remote debugger, KubRay for Kubernetes, and where Ray fits alongside Dask, multiprocessing, and asyncio.
If you've ever stared at a single-machine Python script and thought, "there has to be a better way to scale this", this one's for you
Episode sponsors
Sentry Error Monitoring, Code talkpython26
AgentField AI
Talk Python Courses
Links from the show
Guests
Richard Liaw: github.com
Edward Oakes: github.com
Ray: www.ray.io
Example code (we used for walk-through): docs.ray.io
Getting Started with Ray: docs.ray.io
Ray Libraries: docs.ray.io
kuberay: github.com
Watch this episode on YouTube: youtube.com
Episode #547 deep-dive: talkpython.fm/547
Episode transcripts: talkpython.fm
Theme Song: Developer Rap
🥁 Served in a Flask 🎸: talkpython.fm/flasksong
---== Don't be a stranger ==---
YouTube: youtube.com/@talkpython
Bluesky: @talkpython.fm
Mastodon: @talkpython@fosstodon.org
X.com: @talkpython
Michael on Bluesky: @mkennedy.codes
Michael on Mastodon: @mkennedy@fosstodon.org
Michael on X.com: @mkennedy
DeepSeek: separating fact from hype
Today, we’re bringing you a bonus episode zeroing in on DeepSeek, the Chinese AI lab that’s recently taken over the news and the app stores, beating out OpenAI's ChatGPT. Max Zeff is talking about it all with Ion Stoica, Professor of Computer Science Division at UC Berkeley and the cofounder and executive chairman of software startup Databricks.
Listen to the full episode to hear more about:
Why Stoica believes the future of AI lies in "doubling down on open source."
Microsoft's decision to host DeepSeek on Azure.
What the U.S. can do to foster accelerated innovation – with a look back at SB-1047 and a look ahead to 2025.
The controversy surrounding claims that DeepSeek used OpenAI’s models to train its own.
Equity will be back next week, so stay tuned!
Equity is TechCrunch’s flagship podcast, produced by Theresa Loconsolo, and posts every Wednesday and Friday.
Subscribe to us on Apple Podcasts, Overcast, Spotify and all the casts. You also can follow Equity on X and Threads, at @EquityPod. For the full episode transcript, for those who prefer reading over listening, check out our full archive of episodes here.
Credits: Equity is produced by Theresa Loconsolo with editing by Kell. We’d also like to thank TechCrunch’s audience development team. Thank you so much for listening, and we'll talk to you next time.
Learn more about your ad choices. Visit megaphone.fm/adchoices
Databricks Founder Ion Stoica: Turning Academic Open Source into Startup Success
Berkeley professor Ion Stoica, co-founder of Databricks and Anyscale, transformed the open source projects Spark and Ray into successful AI infrastructure companies. He talks about what mattered most for Databricks' success -- the focus on making Spark win and making Databricks the best place to run Spark. He highlights the importance of striking key partnerships -- the Microsoft partnership in particular that accelerated Databricks' growth and contributed to Spark's dominance among data scientists and AI engineers. He also shares his perspective on finding new problems to work on, which holds lessons for aspiring founders and builders: 1) building systems in new areas that, if widely adopted, put you in the best position to understand the new problem space, and 2) focusing on a problem that is more important tomorrow than today.
Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital
Mentioned in this episode:
Spark: The open source platform for data engineering that Databricks was originally based on.
Ray: Open source framework to manage, executes and optimizes compute needs across AI workloads, now productized through Anyscale
MosaicML: Generative AI startups founded by Naveen Rao that Databricks acquired in 2023.
Unity Catalog: Data and AI governance solution from Databricks.
CIB Berkeley: Multi-strategy hedge fund at UC Berkeley that commercializes research in the UC system.
Hadoop: A long-time leading platform for large scale distributed computing.
VLLM and Chatbot Arena: Two of Ion’s students’ projects that he wanted to highlight.
Ray & KubeRay, with Richard Liaw and Kai-Hsun Chen
In this episode, guest host and AI correspondent Mofi Rahman interviews Richard Liaw and Kai-Hsun Chen from Anyscale about Ray and KubeRay. Ray is an open-source unified compute framework that makes it easy to scale AI and Python workloads, while KubeRay integrates Ray's capabilities into Kubernetes clusters.
Do you have something cool to share? Some questions? Let us know:
- web: kubernetespodcast.com
- mail: kubernetespodcast@google.com
- twitter: @kubernetespod
News of the week CNCF Blog - LitmusChaos audit complete!
Kubernetes Podcast from Google episode 234 - LitmusChaos, with Karthik Satchitanand
Google Cloud Blog - Run your AI inference applications on Cloud Run with NVIDIA GPUs
Diginomica article - KubeCon China - at 33-and-a-third, Linux is a long player. So, why does Linus Torvalds hate AI?
CNCF-Hosted Co-Located Event Schedule for KubeCon NA 2024
Google Kubernetes Engine Release Notes - August 20, 2024 (1.31 available in Rapid Channel)
Kubernetes Podcast from Google - Kubernetes v1.31: "Elli", with Angelos Kolaitis
Red Hat Press Release - Red Hat OpenStack Services on OpenShift is Now Generally Available
Red Hat Enables OpenStack to Run Natively on OpenShift Platform
Broadcom Revamps Tanzu to Simplify Cloud-Native App Development and Deployment
Tanzu Platform 10 Offers Cloud Foundry Users Deep Visibility and Productivity Enhancements
VMware Explore Conference Website
CNCF Blog - Announcing 500 Kubestronauts
CNCF - Kubestronaut FAQ
Dapr Day 2024 Virtual Event Website
Links from the interview Kai-Hsun Chen on LinkedIn
Richard Liaw on LinkedIn
Ray from the RISE Lab at UC Berkeley
Ray: A Distributed System for AI by Robert Nishihara and Philipp Moritz - Jan 9, 2018
KubeRay Docs
KubeRay on GitHub
PyTorch
Apache Airflow
Apache Spark
Kubeflow
Apache Submarine (retired)
Jupyter Notebooks
VS Code
Examples of schedulers for Batch/AI workloads in Kubernetes
Kueue
Volcano
Apache Yunikorn
Examples of observability tools for Batch/AI workloads in Kubernetes
Prometheus
Grafana
Fluentbit
Examples of loadbalancers
Nginx
Istio
Ray Data: Scalable Datasets for ML
Dask Python - Parallel Python
Ray Serve: Scalable and Programmable Serving
HPA - Horizontal Pod Autoscaling in Kubernetes
Karpenter - "Just-in-time nodes for any Kubernetes cluster"
Lazy Computation Graphs with the Ray DAG API
Types of hardware accelerators
Google Cloud Tensor Processing Units (TPUs)
AMD Instinct
AMD Radeon
AWS Trainium
AWS Inferentia
Pandas
Numpy
KubeCon EU 2024 - Accelerators(FPGA/GPU) Chaining to Efficiently Handle Large AI/ML Workloads in K8s - Sampath Priyankara, Nippon Telegraph and Telephone Corporation & Masataka Sonoda, Fujitsu Limited
NVidia Megatron
Links from the post-interview chat DRA - Dynamic Resource Allocation in Kubernetes
Different ways of Running RayJob on Kubernetes
Ray framework diagram in the docs
SaaStr 534: Marketing Secrets to Hypergrowth from Building Elastic, Zuora, and Segment
Panelists:
Katrina Wong, VP Marketing and Demand Generation Segment
Asawari Samant, Head of Marketing, Anyscale
Jeffrey Yoshimura, CMO and Customer Experience Officer, Snyk
Three marketing leaders discuss the importance of community, data, and experimentation, as well as the do's and don'ts of using agencies.
This episode is an abbreviated version of the session. You can see the full session here: https://youtu.be/7wucuUQSQvQ
Blog post: https://www.saastr.com/marketing-secrets-to-hypergrowth-from-building-elastic-zuora-and-segment/
Want to join the SaaStr community? We're the 🌎largest community for B2B software.
Subscribe for weekly updates: https://www.saastr.com/subscribeform
Twitter: https://twitter.com/saastr
LinkedIn: https://www.linkedin.com/company/2724976
Quora Group: https://www.quora.com/q/cloud
Facebook: https://www.facebook.com/SaaStr/
Instagram: https://www.instagram.com/saastr/
Our North American Event: https://bit.ly/2OXeAYh
Our European Event: https://bit.ly/2OZTad8
Ion Stoica — Spark, Ray, and Enterprise Open Source
Ion Stoica is co-creator of the distributed computing frameworks Spark and Ray, and co-founder and Executive Chairman of Databricks and Anyscale. He is also a Professor of computer science at UC Berkeley and Principal Investigator of RISELab, a five-year research lab that develops technology for low-latency, intelligent decisions.
Ion and Lukas chat about the challenges of making a simple (but good!) distributed framework, the similarities and differences between developing Spark and Ray, and how Spark and Ray led to the formation of Databricks and Anyscale. Ion also reflects on the early startup days, from deciding to commercialize to picking co-founders, and shares advice on building a successful company.
The complete show notes (transcript and links) can be found here: http://wandb.me/gd-ion-stoica
---
Timestamps:
0:00 Intro
0:56 Ray, Anyscale, and making a distributed framework
11:39 How Spark informed the development of Ray
18:53 The story behind Spark and Databricks
33:00 Why TensorFlow and PyTorch haven't monetized
35:35 Picking co-founders and other startup advice
46:04 The early signs of sky computing
49:24 Breaking problems down and prioritizing
53:17 Outro
---
Subscribe and listen to our podcast today!
👉 Apple Podcasts: http://wandb.me/apple-podcasts
👉 Google Podcasts: http://wandb.me/google-podcasts
👉 Spotify: http://wandb.me/spotify
Robert Nishihara — The State of Distributed Computing in ML
The story of Ray and what lead Robert to go from reinforcement learning researcher to creating open-source tools for machine learning and beyond
Robert is currently working on Ray, a high-performance distributed execution framework for AI applications. He studied mathematics at Harvard. He’s broadly interested in applied math, machine learning, and optimization, and was a member of the Statistical AI Lab, the AMPLab/RISELab, and the Berkeley AI Research Lab at UC Berkeley.
robertnishihara.com
https://anyscale.com/
https://github.com/ray-project/ray
https://twitter.com/robertnishihara
https://www.linkedin.com/in/robert-nishihara-b6465444/
Topics covered:
0:00 sneak peak + intro
1:09 what is Ray?
3:07 Spark and Ray
5:48 reinforcement learning
8:15 non-ml use cases of ray
10:00 RL in the real world and and common uses of Ray
13:49 Ppython in ML
16:38 from grad school to ML tools company
20:40 pulling product requirements in surprising directions
23:25 how to manage a large open source community
27:05 Ray Tune
29:35 where do you see bottlenecks in production?
31:39 An underrated aspect of Machine Learning
Visit our podcasts homepage for transcripts and more episodes!
www.wandb.com/podcast
Get our podcast on Apple, Spotify, and Google!
Apple Podcasts: https://bit.ly/2WdrUvI
Spotify: https://bit.ly/2SqtadF
Google: http://tiny.cc/GD_Google
Subscribe to our YouTube channel for videos of these podcasts and more Machine learning-related videos:
https://www.youtube.com/c/WeightsBiases
We started Weights and Biases to build tools for Machine Learning practitioners because we care a lot about the impact that Machine Learning can have in the world and we love working in the trenches with the people building these models. One of the most fun things about these building tools has been the conversations with these ML practitioners and learning about the interesting things they’re working on. This process has been so fun that we wanted to open it up to the world in the form of our new podcast called Gradient Dissent. We hope you have as much fun listening to it as we had making it!
Join our bi-weekly virtual salon and listen to industry leaders and researchers in machine learning share their research:
http://tiny.cc/wb-salon
Join our community of ML practitioners where we host AMA's, share interesting projects and meet other people working in Deep Learning:
http://bit.ly/wb-slack
Our gallery features curated machine learning reports by researchers exploring deep learning techniques, Kagglers showcasing winning models, and industry leaders sharing best practices.
https://app.wandb.ai/gallery
a16z Podcast: On Data and Data Scientists in the Age of AI
Data, data, everywhere, nor any drop to drink. Or so would say Coleridge, if he were a big company CEO trying to use A.I. today -- because even when you have a ton of data, there's not always enough signal to get anything meaningful from AI.
Why? Because, "like they say, it's 'garbage in, garbage out' -- what matters is what you have in between," reminds Databricks co-founder (and director of the RISElab at U.C. Berkeley) Ion Stoica. And even then it's still not just about data operations, emphasizes SigOpt co-founder Scott Clark; your data scientists need to really understand "What's actually right for my business and what am I actually aiming for?" And then get there as efficiently as possible.
But beyond defining their goals, how do companies get over the "cold start" problem when it comes to doing more with AI in practice, asks a16z operating partner Frank Chen (who also released a microsite on getting started with AI earlier this year)? The guests on this short "a16z Bytes" episode of the a16z Podcast -- based on a conversation that took place at our recent annual Summit event -- share practical advice about this and more.
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
a16z Podcast: A New Lab Rises
with Ion Stoica, Peter Levine, and Sonal Chokshi
We’ve already talked quite a bit about the Algorithms, Machines, and People lab at U.C. Berkeley (AMPLab) — all about making sense of big data — so what happens when the entire world moves towards artificial intelligence — and the need to make intelligent decisions on that data? That’s where the new RISElab (Real-time Intelligence Secure Execution) comes in.
But what is a good “decision”, exactly? Beyond the existential question of that, what specific attributes make a “good” decision, both computationally and humanly? In this episode of the a16z Podcast (in conversation with general partner Peter Levine and Sonal Chokshi), computer science professor, entrepreneur (co-founder of Databricks), and RISElab director Ion Stoica answers that question. He also shares the “ingredients” of a working research lab model (one, dare we say, could also apply to many types of institutions?); the role of open source and building community; and the evolution of labs today given intense competition from industry and others… as well as what interesting projects — really, trends in decision making with AI — are coming next.
The views expressed here are those of the individual AH Capital Management, L.L.C. (“a16z”) personnel quoted and are not the views of a16z or its affiliates. Certain information contained in here has been obtained from third-party sources, including from portfolio companies of funds managed by a16z. While taken from sources believed to be reliable, a16z has not independently verified such information and makes no representations about the enduring accuracy of the information or its appropriateness for a given situation.
This content is provided for informational purposes only, and should not be relied upon as legal, business, investment, or tax advice. You should consult your own advisers as to those matters. References to any securities or digital assets are for illustrative purposes only, and do not constitute an investment recommendation or offer to provide investment advisory services. Furthermore, this content is not directed at nor intended for use by any investors or prospective investors, and may not under any circumstances be relied upon when making a decision to invest in any fund managed by a16z. (An offering to invest in an a16z fund will be made only by the private placement memorandum, subscription agreement, and other relevant documentation of any such fund and should be read in their entirety.) Any investments or portfolio companies mentioned, referred to, or described are not representative of all investments in vehicles managed by a16z, and there can be no assurance that the investments will be profitable or that other investments made in the future will have similar characteristics or results. A list of investments made by funds managed by Andreessen Horowitz (excluding investments and certain publicly traded cryptocurrencies/ digital assets for which the issuer has not provided permission for a16z to disclose publicly) is available at https://a16z.com/investments/.
Charts and graphs provided within are for informational purposes solely and should not be relied upon when making any investment decision. Past performance is not indicative of future results. The content speaks only as of the date indicated. Any projections, estimates, forecasts, targets, prospects, and/or opinions expressed in these materials are subject to change without notice and may differ or be contrary to opinions expressed by others. Please see https://a16z.com/disclosures for additional important information.
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Ray: A Distributed Computing Platform for Reinforcement Learning with Ion Stoica - TWiML Talk #55
The show you’re about to hear is part of a series of shows recorded in San Francisco at the Artificial Intelligence Conference. In this episode, I talk with Ion Stoica, professor of computer science & director of the RISE Lab at UC Berkeley. Ion joined us after he gave his talk “Building reinforcement learning applications with Ray.” We dive into Ray, a new distributed computing platform for RL, as well as RL generally, along with some of the other interesting projects RISE Lab is working on, like Clipper & Tegra. This was a pretty interesting talk. Enjoy! The notes for this show can be found at twimlai.com/talk/55