Open-source is for the people, by the people
Travis Oliphant, creator of NumPy and SciPy, joins Ryan to explore the development of Python as a data science tool, the evolution of these foundational libraries, and the importance of community and collaboration in open-source projects, including Travis’ current work to support sustainable open-source through the OpenTeams Incubator.
Episode notes:
NumPy and SciPy are the fundamental packages and algorithms for scientific computing with Python. NumPy 2.3.0 and SciPy 1.16.0 are out now.
The OpenTeams Incubator helps start, grow, and sustain open-source software communities.
Quansight is a data, science, and engineering firm rooted in the work of the Python Data, Science, and AI/ML open-source communities.
Connect with Travis on LinkedIn or email him at travis@OTincubator.com
Today we’re shouting out user RobinFrcd for answering pytest-asyncio has a closed event loop, but only when running all tests and winning a Populist badge.
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Python documentary companion pod (Interview)
Our friends at Cult.Repo launched their epic Python documentary on August 28th, 2025! To celebrate, we sat down with Travis Oliphant –creator of NumPy, SciPy, and more– to get his perspective on how Python took over the software world.
Stick around for the twist ending! We set aside Python and dissect Travis’ big idea to make open source projects financially sustainable through direct investment.
Join the discussion
Changelog++ members save 4 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
Depot – 10x faster builds? Yes please. Build faster. Waste less time. Accelerate Docker image builds, and GitHub Actions workflows. Easily integrate with your existing CI provider and dev workflows to save hours of build time.
Auth0 – The identity infrastructure for the age of AI. Built by developers, for developers—Auth0 helps you secure users, agents, and third-party access across modern AI workflows. Token vaulting, fine-grained authorization, and standards-based auth, all in one platform.
Start building at Auth0.com/ai
Fly.io – The home of Changelog.com — Deploy your apps close to your users — global Anycast load-balancing, zero-configuration private networking, hardware isolation, and instant WireGuard VPN connections. Push-button deployments that scale to thousands of instances. Check out the speedrun to get started in minutes.
Featuring:
Travis Oliphant – GitHub, LinkedIn, Mastodon, X
Adam Stacoviak – Website, GitHub, LinkedIn, Mastodon, X
Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
Python: The Documentary
NumPy
SciPy.org
Mojo 🔥: Powerful CPU+GPU Programming
FairOSS
tea.xyz
Cult.Repo
Something missing or broken? PRs welcome!
Travis Oliphant: SciPy, NumPy, and Fostering Scientific Python
<p>What went into developing the open-source Python tools data scientists use every day? This week on the show, we talk with Travis Oliphant about his work on SciPy, NumPy, Numba, and many other contributions to the Python scientific community.</p>
<p>Travis discusses his initial involvement in the open-source community and how he discovered Python while working in biomedical imaging. He was trying to find ways to manage large sets of numerical data, which led to his initial contributions and collaborations in building scientific libraries. </p>
<p>His appearance on the show coincides with the release of the Python documentary, in which he’s featured. We discuss the myriad organizations Travis founded, including Quansight, OpenTeams, and Anaconda. We dig into his underlying mission to continue fostering the growth of the open-source scientific computing community.</p>
<p>This episode is sponsored by InfluxData.</p>
<div class="alert alert-primary" role="alert">
<p><strong>Course Spotlight:</strong> <a href="https://realpython.com/courses/numpy-techniques-practical-examples/">NumPy Techniques and Practical Examples</a></p>
<p>In this video course, you’ll learn how to use NumPy by exploring several interesting examples. You’ll read data from a file into an array and analyze structured arrays to perform a reconciliation. You’ll also learn how to quickly chart an analysis and turn a custom function into a vectorized function.</p>
</div>
<p>Topics:</p>
<ul>
<li>00:00:00 – Introduction</li>
<li>00:02:41 – Python documentary</li>
<li>00:07:44 – Getting involved in open source</li>
<li>00:12:04 – Numeric Python</li>
<li>00:15:36 – SciPy and the SciPy community</li>
<li>00:17:35 – Starting to think about entrepreneurship </li>
<li>00:18:16 – NumPy evolving from the work of Numeric</li>
<li>00:22:01 – Sponsor: InfluxData</li>
<li>00:22:53 – Python as controlling code for lower-level libraries</li>
<li>00:23:37 – Numba open-source JIT compiler</li>
<li>00:30:09 – Starting to build in Python before learning it all</li>
<li>00:34:45 – Python as the language AI generates</li>
<li>00:36:31 – Guilds and sharing knowledge</li>
<li>00:40:15 – More NumPy backstory</li>
<li>00:46:36 – Contributing to Python</li>
<li>00:48:24 – Video Course Spotlight</li>
<li>00:49:41 – The investment of companies in Python</li>
<li>00:51:22 – Quansight and businesses in open source</li>
<li>00:53:09 – Open Teams and Quansight details</li>
<li>00:57:14 – NumFOCUS and Anaconda</li>
<li>00:58:51 – FairOSS</li>
<li>01:02:36 – Documenting these efforts</li>
<li>01:05:37 – What are you excited about in the world of Python?</li>
<li>01:07:12 – What do you want to learn next?</li>
<li>01:08:10 – How can people follow your work online?</li>
<li>01:10:03 – Thanks and goodbye</li>
</ul>
<p>Show Links:</p>
<ul>
<li><a href="https://www.youtube.com/watch?v=pqBqdNIPrbo">Python: The Documentary - OFFICIAL TRAILER - Coming August 28 - YouTube</a></li>
<li><a href="http://hugunin.net/papers/hugunin95numpy.html">The Python Matrix Object: Extending Python for Numerical Computation</a></li>
<li><a href="https://j1m.dev/">Jim Fulton</a></li>
<li><a href="http://hugunin.net/">Jim Hugunin - Home</a></li>
<li><a href="https://scipy.github.io/old-wiki/pages/History_of_SciPy">History of SciPy - SciPy wiki dump</a></li>
<li><a href="https://scipy.org/">SciPy</a></li>
<li><a href="https://numpy.org/">NumPy</a></li>
<li><a href="https://numba.pydata.org/">Numba: A High Performance Python Compiler</a></li>
<li><a href="https://lpython.org/">LPython - High performance typed Python compiler</a></li>
<li><a href="https://openteams.com/">OpenTeams: Open SaaS AI Solutions</a></li>
<li><a href="https://quansight.com/">Quansight Consulting</a></li>
<li><a href="https://otincubator.com/">OpenTeams Incubator</a></li>
<li><a href="https://numfocus.org/">NumFOCUS: A Nonprofit Supporting Open Code for Better Science</a></li>
<li><a href="https://anaconda.com/">Anaconda</a></li>
<li><a href="https://faiross.org/">FairOSS</a></li>
<li><a href="https://github.com/faster-cpython/">faster-cpython</a></li>
<li><a href="https://x.com/teoliphant">Travis Oliphant (@teoliphant) / X</a></li>
<li><a href="https://www.linkedin.com/in/teoliphant/">Travis Oliphant - LinkedIn</a></li>
</ul>
<p>Level up your Python skills with our expert-led courses:</p>
<ul>
<li><a href="https://realpython.com/courses/data-cleaning-with-pandas-and-numpy/">Data Cleaning With pandas and NumPy</a></li>
<li><a href="https://realpython.com/courses/stacks-queues-ideal-data-structure/">Stacks and Queues: Selecting the Ideal Data Structure</a></li>
<li><a href="https://realpython.com/courses/numpy-techniques-practical-examples/">NumPy Techniques and Practical Examples</a></li>
</ul> <p><a rel="payment" href="https://realpython.com/join">Support the podcast & join our community of Pythonistas</a></p>
AI at Anaconda with Greg Jennings
Anaconda is a software company that’s well-known for its solutions for managing packages, environments, and security in large-scale data workflows. The company has played a major role in making Python-based data science more accessible, efficient, and scalable. Anaconda has also invested heavily in AI tool development.
Greg Jennings is the VP of Engineering and AI at Anaconda. He joins the podcast with Kevin Ball to talk about the tooling ecosystem around AI app development, the Anaconda Toolbox, the rapidly evolving role of AI in engineering, and more.
Kevin Ball or KBall, is the vice president of engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript meetup, and organizes the AI inaction discussion group through Latent Space.
Please click here to see the transcript of this episode.
Sponsorship inquiries: sponsor@softwareengineeringdaily.com
The post AI at Anaconda with Greg Jennings appeared first on Software Engineering Daily.
#489: Anaconda Toolbox for Excel and more with Peter Wang
See the full show notes for this episode on the website at talkpython.fm/489
Open Source Monetization: 40M Users to 8-Figure ARR
Peter Wang gave away his product to 40 million users without requiring an email address. Then he built an 8-figure business on top of it through open source monetization. In this episode, you'll learn how Anaconda turned Python into the dominant language for data science and monetized the freemium SaaS model by selling to enterprise buyers instead of individual practitioners.
Peter reveals how three years of bootstrapping through consulting funded community-led growth via PyData conferences, why the first enterprise sale came from an inbound request by a law enforcement agency, and how internal teams struggled for years to align open source values with product-led growth revenue goals.
Anaconda now serves 40M+ users, generates 8-figure ARR, and employs 350+ people - proof that open source monetization works when you sell to the buyer, not the user.
🔑 Key Lessons
🚀 Open source monetization needs community investment first: Anaconda spent three years funding PyData conferences and advocacy before launching an enterprise product. Grassroots adoption of 40M users became the foundation for 8-figure ARR.
🎯 Sell to the buyer, not the user: Anaconda's free users are data scientists, but paying customers are IT managers and compliance officers who need governance and security - a completely different persona.
🛠️ Let inbound demand shape your first product: The first enterprise sale came when a law enforcement agency asked for a secure package repository behind their firewall. Peter built what the customer requested instead of guessing.
📉 Expect organizational confusion with open source monetization: New hires saw "a box of other people's parts" with no traditional upsell, creating years of tension between community advocates and revenue-focused teams.
💰 Compete on simplicity against incumbents: Rather than matching decades of specialized features, Peter bet Python would win because it "fit in people's heads" - domain experts chose ease of use over completeness.
Chapters
What Anaconda does and 8-figure ARR metrics
Bootstrapping with consulting for the first three years
Starting a nonprofit and a startup simultaneously
Overcoming enterprise skepticism about Python
How organic community growth reached 40 million users
The open source monetization model: no email required
First enterprise product from an inbound request
Internal confusion between open source and enterprise teams
How AI and ChatGPT affect Python and Anaconda
Lightning round and founder advice
Resources
Full show notes: https://saasclub.io/418
Join 5,000+ SaaS founders: https://saasclub.io/email
765: NumPy, SciPy and the Economics of Open-Source, with Dr. Travis Oliphant
Explore the origins of NumPy and SciPy with their creator, Dr. Travis Oliphant. Discover the journey from personal need to global impact, the challenges overcome, and the future of these essential Python libraries in scientific computing and data science.
This episode is brought to you by the DataConnect Conference, by Data Universe, the out-of-this-world data conference, and by CloudWolf, the Cloud Skills platform. Interested in sponsoring a SuperDataScience Podcast episode? Visit passionfroot.me/superdatascience for sponsorship information.
In this episode you will learn:
• Travis's journey to creating NumPy and SciPy [08:05]
• How Anaconda got started [42:24]
• How Numba, a high-performance Python compiler, was brought to market [54:48]
• Python's influence on the thought processes of scientists and engineers [1:04:21]
• The commercial projects that support Travis’s vast open-source efforts and communities [1:10:22]
• How to get involved in Travis's commercial projects and communities [1:22:34]
• The future of scientific computing and Python libraries [1:29:50]
Additional materials: www.superdatascience.com/765
#130 Mathew Lodge: The Future of Large Language Models in AI
Welcome to episode #130 of Eye on AI with Mathew Lodge. In this episode, we explore the world of reinforcement learning and code generation. Mathew Lodge, the CEO of Diffblue, shares insights into how reinforcement learning fuels generative AI. As we explore the intricacies of reinforcement learning, we uncover its potential in game playing and guiding us towards solutions. We shed light on the products that it powers, such as AlphaGo and AlphaDev. However, we also address the challenges of large language models and explain why they may not be the ultimate solution for code generation. In the last part of our conversation, we delve into the future of language models and intelligence. Mathew shares valuable insights on merging no-code and low-code solutions. We confront the skepticism of software developers towards AI for code products and the task of articulating program outcomes. Wrapping up, we reflect on the evolution of programming languages and the impact of abstraction on machine learning.
(00:00) Preview & sponsorship (01:51) Reinforcement Learning and Code Generation (04:39) Reinforcement Learning and Improving Algorithms (15:32) The Challenges of Large Language Models(23:58) Future of Language Models and Intelligence (35:50) Challenges and Potential of AI-generated Code (48:32) Programming Language Evolution and Higher-Level Languages
Craig Smith Twitter: https://twitter.com/craigss Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
#250 – Peter Wang: Python and the Source Code of Humans, Computers, and Reality
Peter Wang is the co-founder & CEO of Anaconda and one of the most impactful leaders and developers in the Python community. Also, he is a physicist and philosopher. Please support this podcast by checking out our sponsors:
– Quip: https://getquip.com/lex to get first refill free
– Magic Spoon: https://magicspoon.com/lex and use code LEX to get $5 off
– GiveWell: https://www.givewell.org/ and use code LEX to get donation matched up to $1k
– Four Sigmatic: https://foursigmatic.com/lex and use code LexPod to get up to 60% off
– BetterHelp: https://betterhelp.com/lex to get 10% off
EPISODE LINKS:
Peter’s Twitter: https://twitter.com/pwang
Anaconda’s Website: https://www.anaconda.com/
Books & resources mentioned:
Zen and the Art of Motorcycle Maintenance (book): https://amzn.to/3EnCELK
Lila (book): https://amzn.to/30VKIpE
PODCAST INFO:
Podcast website: https://lexfridman.com/podcast
Apple Podcasts: https://apple.co/2lwqZIr
Spotify: https://spoti.fi/2nEwCF8
RSS: https://lexfridman.com/feed/podcast/
YouTube Full Episodes: https://youtube.com/lexfridman
YouTube Clips: https://youtube.com/lexclips
SUPPORT & CONNECT:
– Check out the sponsors above, it’s the best way to support this podcast
– Support on Patreon: https://www.patreon.com/lexfridman
– Twitter: https://twitter.com/lexfridman
– Instagram: https://www.instagram.com/lexfridman
– LinkedIn: https://www.linkedin.com/in/lexfridman
– Facebook: https://www.facebook.com/lexfridman
– Medium: https://medium.com/@lexfridman
OUTLINE:
Here’s the timestamps for the episode. On some podcast players you should be able to click the timestamp to jump to that time.
(00:00) – Introduction
(06:49) – Python
(10:20) – Programming language design
(30:22) – Virtuality
(40:22) – Human layers
(47:21) – Life
(52:45) – Origin of ideas
(55:17) – Eric Weinstein
(1:00:16) – Human source code
(1:04:13) – Love
(1:18:32) – AI
(1:31:55) – Meaning crisis
(1:54:28) – Travis Oliphant
(2:00:53) – Python continued
(2:30:36) – Best setup
(2:37:54) – Advice for the youth
(2:46:28) – Meaning of Life
#224 – Travis Oliphant: NumPy, SciPy, Anaconda, Python & Scientific Programming
Travis Oliphant is a data scientist, entrepreneur, and creator of NumPy, SciPy, and Anaconda. Please support this podcast by checking out our sponsors:
– Novo: https://banknovo.com/lex
– Allform: https://allform.com/lex to get 20% off
– Onnit: https://lexfridman.com/onnit to get up to 10% off
– Athletic Greens: https://athleticgreens.com/lex and use code LEX to get 1 month of fish oil
– Blinkist: https://blinkist.com/lex and use code LEX to get 25% off premium
EPISODE LINKS:
Travis’s Twitter: https://twitter.com/teoliphant
Travis’s Wiki Page: https://en.wikipedia.org/wiki/Travis_Oliphant
NumPy: https://numpy.org/
SciPy: https://scipy.org/about.html
Anaconda: https://www.anaconda.com/products/individual
Quansight: https://www.quansight.com
PODCAST INFO:
Podcast website: https://lexfridman.com/podcast
Apple Podcasts: https://apple.co/2lwqZIr
Spotify: https://spoti.fi/2nEwCF8
RSS: https://lexfridman.com/feed/podcast/
YouTube Full Episodes: https://youtube.com/lexfridman
YouTube Clips: https://youtube.com/lexclips
SUPPORT & CONNECT:
– Check out the sponsors above, it’s the best way to support this podcast
– Support on Patreon: https://www.patreon.com/lexfridman
– Twitter: https://twitter.com/lexfridman
– Instagram: https://www.instagram.com/lexfridman
– LinkedIn: https://www.linkedin.com/in/lexfridman
– Facebook: https://www.facebook.com/lexfridman
– Medium: https://medium.com/@lexfridman
OUTLINE:
Here’s the timestamps for the episode. On some podcast players you should be able to click the timestamp to jump to that time.
(00:00) – Introduction
(07:06) – Early programming
(28:47) – SciPy
(45:41) – Open source
(57:23) – NumPy
(1:34:39) – Guido van Rossum
(1:46:57) – Efficiency
(1:55:49) – Objects
(2:02:47) – Numba
(2:11:53) – Anaconda
(2:16:20) – Conda
(2:31:56) – Quansight Labs
(2:35:32) – OpenTeams
(2:43:05) – GitHub
(2:48:35) – Marketing
(2:53:13) – Great programming
(3:04:03) – Hiring
(3:08:01) – Advice for young people
Anaconda + Pyston and more
In this episode, Peter Wang from Anaconda joins us again to go over their latest “State of Data Science” survey. The updated results include some insights related to data science work during COVID along with other topics including AutoML and model bias. Peter also tells us a bit about the exciting new partnership between Anaconda and Pyston (a fork of the standard CPython interpreter which has been extensively enhanced to improve the execution performance of most Python programs).
Sponsors:
SignalWire – Build what’s next in communications with video, voice, and messaging APIs powered by elastic cloud infrastructure. Try it today at signalwire.com and use code SHIPIT for $25 in developer credit.
The Brave Browser – Browse the web up to 8x faster than Chrome and Safari, block ads and trackers by default, and reward your favorite creators with the built-in Basic Attention Token. Download Brave for free and give tipping a try right here on changelog.com.
Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
Featuring:
Peter Wang – Website, X
Chris Benson – Website, GitHub, LinkedIn, X
Daniel Whitenack – Website, GitHub, X
Show Notes:
Anaconda’s State of Data Science
Pyston Team Joins Anaconda to Expand Open-Source Project Development
Upcoming Events:
Register for upcoming webinars here!
Peter Wang — Anaconda, Python, and Scientific Computing
Peter Wang talks about his journey of being the CEO of and co-founding Anaconda, his perspective on the Python programming language, and its use for scientific computing.
Peter Wang has been developing commercial scientific computing and visualization software for over 15 years. He has extensive experience in software design and development across a broad range of areas, including 3D graphics, geophysics, large data simulation and visualization, financial risk modeling, and medical imaging.
Peter’s interests in the fundamentals of vector computing and interactive visualization led him to co-found Anaconda (formerly Continuum Analytics). Peter leads the open source and community innovation group.
As a creator of the PyData community and conferences, he devotes time and energy to growing the Python data science community and advocating and teaching Python at conferences around the world. Peter holds a BA in Physics from Cornell University.
Follow peter on Twitter: https://twitter.com/pwang
https://www.anaconda.com/
Intake: https://www.anaconda.com/blog/intake-...
https://pydata.org/
Scientific Data Management in the Coming Decade paper: https://arxiv.org/pdf/cs/0502008.pdf
Topics covered:
0:00 (intro) Technology is not value neutral; Don't punt on ethics
1:30 What is Conda?
2:57 Peter's Story and Anaconda's beginning
6:45 Do you ever regret choosing Python?
9:39 On other programming languages
17:13 Scientific Data Management in the Coming Decade
21:48 Who are your customers?
26:24 The ML hierarchy of needs
30:02 The cybernetic era and Conway's Law
34:31 R vs python
42:19 Most underrated: Ethics - Don't Punt
46:50 biggest bottlenecks: open-source, python
Visit our podcasts homepage for transcripts and more episodes!
www.wandb.com/podcast
Get our podcast on these other platforms:
YouTube: http://wandb.me/youtube
Soundcloud: http://wandb.me/soundcloud
Apple Podcasts: http://wandb.me/apple-podcasts
Spotify: http://wandb.me/spotify
Google: http://wandb.me/google-podcasts
Join our bi-weekly virtual salon and listen to industry leaders and researchers in machine learning share their work:
http://wandb.me/salon
Join our community of ML practitioners where we host AMA's, share interesting projects and meet other people working in Deep Learning:
http://wandb.me/slack
Our gallery features curated machine learning reports by researchers exploring deep learning techniques, Kagglers showcasing winning models, and industry leaders sharing best practices.
https://wandb.ai/gallery
#285: Dask as a Platform Service with Coiled
See the full show notes for this episode on the website at talkpython.fm/285
Reining in Complexity: Data Science & Future of AI/ML Businesses
There is no spoon. Or rather, “There is no such thing as ‘data’, there’s just frozen models”, argues Peter Wang, the co-founder and CEO of Anaconda — who also created the PyData conferences and grew the early data science community there, while on the frontlines of trying to make Python useful for business analytics. He views both models and data as fluid, more like metaphysics than typical data management… Or perhaps it’s that when it comes to data, those with a physics background just better appreciate the mind-bending complexity and challenges of reining in the natural world, and therefore get the unique challenges of AI/ML development, observes a16z general partner Martin Casado — whose first job after college involved computational physics simulation and high-performance computing in Python at Lawrence Livermore National Laboratory. (Wang, meanwhile, graduated in physics.)
But this not just a philosophical question — the answer has real implications for the margins, organizational structures, and building of AI/ML businesses. Especially as we’re in a tricky time of transition, where customers don’t even know what they’re asking for, yet are looking for AI/ML help or know it’s the future. So what does this all mean for the software value chain; for open source collaboration and commodification; and for the future of software businesses? After all, it’s not written in stone that “All information systems must be deconstructed into hardware, and software, and data” and that “software must have these margins”… Will there be a new type of company?
image: Pawel Loj / Wikimedia Commons
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Building the world's most popular data science platform
Everyone working in data science and AI knows about Anaconda and has probably “conda” installed something. But how did Anaconda get started and what are they working on now? Peter Wang, CEO of Anaconda and creator of PyData and popular packages like Bokeh and DataShader, joins us to discuss that and much more. Peter gives some great insights on the Python AI ecosystem and very practical advice for scaling up your data science operation.
Sponsors:
DigitalOcean – DigitalOcean’s developer cloud makes it simple to launch in the cloud and scale up as you grow. They have an intuitive control panel, predictable pricing, team accounts, worldwide availability with a 99.99% uptime SLA, and 24/7/365 world-class support to back that up. Get your $100 credit at do.co/changelog.
Changelog++ – You love our content and you want to take it to the next level by showing your support. We’ll take you closer to the metal with no ads, extended episodes, outtakes, bonus content, a deep discount in our merch store (soon), and more to come. Let’s do this!
Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com.
Rollbar – We move fast and fix things because of Rollbar. Resolve errors in minutes. Deploy with confidence. Learn more at rollbar.com/changelog.
Featuring:
Peter Wang – Website, X
Chris Benson – Website, GitHub, LinkedIn, X
Daniel Whitenack – Website, GitHub, X
Show Notes:
Anaconda
PyData
NumFOCUS
Jupyter
Numba
JAX
Masakhane
Upcoming Events:
Register for upcoming webinars here!
#222: Interactive graphs with Bokeh and Python
See the full show notes for this episode on the website at talkpython.fm/222
#217: Notebooks vs data science-enabled scripts
See the full show notes for this episode on the website at talkpython.fm/217
#207: Parallelizing computation with Dask
See the full show notes for this episode on the website at talkpython.fm/207
#198: Catching up with the Anaconda distribution
See the full show notes for this episode on the website at talkpython.fm/198
Dask with Matthew Rocklin - Episode 2
Summary
There is a vast constellation of tools and platforms for processing and analyzing your data. In this episode Matthew Rocklin talks about how Dask fills the gap between a task oriented workflow tool and an in memory processing framework, and how it brings the power of Python to bear on the problem of big data.
Preamble
Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
You can help support the show by checking out the Patreon page which is linked from the site.
To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
Your host is Tobias Macey and today I’m interviewing Matthew Rocklin about Dask and the Blaze ecosystem.
Interview with Matthew Rocklin
Introduction
How did you get involved in the area of data engineering?
Dask began its life as part of the Blaze project. Can you start by describing what Dask is and how it originated?
There are a vast number of tools in the field of data analytics. What are some of the specific use cases that Dask was built for that weren’t able to be solved by the existing options?
One of the compelling features of Dask is the fact that it is a Python library that allows for distributed computation at a scale that has largely been the exclusive domain of tools in the Hadoop ecosystem. Why do you think that the JVM has been the reigning platform in the data analytics space for so long?
Do you consider Dask, along with the larger Blaze ecosystem, to be a competitor to the Hadoop ecosystem, either now or in the future?
Are you seeing many Hadoop or Spark solutions being migrated to Dask? If so, what are the common reasons?
There is a strong focus for using Dask as a tool for interactive exploration of data. How does it compare to something like Apache Drill?
For anyone looking to integrate Dask into an existing code base that is already using NumPy or Pandas, what does that process look like?
How do the task graph capabilities compare to something like Airflow or Luigi?
Looking through the documentation for the graph specification in Dask, it appears that there is the potential to introduce cycles or other bugs into a large or complex task chain. Is there any built-in tooling to check for that before submitting the graph for execution?
What are some of the most interesting or unexpected projects that you have seen Dask used for?
What do you perceive as being the most relevant aspects of Dask for data engineering/data infrastructure practitioners, as compared to the end users of the systems that they support?
What are some of the most significant problems that you have been faced with, and which still need to be overcome in the Dask project?
I know that the work on Dask is largely performed under the umbrella of PyData and sponsored by Continuum Analytics. What are your thoughts on the financial landscape for open source data analytics and distributed computation frameworks as compared to the broader world of open source projects?
Keep in touch
@mrocklin on Twitter
mrocklin on GitHub
Links
http://matthewrocklin.com/blog/work/2016/09/22/cluster-deployments?utm_source=rss&utm_medium=rss
https://opendatascience.com/blog/dask-for-institutions/?utm_source=rss&utm_medium=rss
Continuum Analytics
2sigma
X-Array
Tornado
Website
Podcast Interview
Airflow
Luigi
Mesos
Kubernetes
Spark
Dryad
Yarn
Read The Docs
XData
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
#34: Continuum: Scientific Python and The Business of Open Source
See the full show notes for this episode on the website at talkpython.fm/34