Astronomer's Role in the Airflow Ecosystem: A Deep Dive with Pete DeJoy
Summary
In this episode of the Data Engineering Podcast Pete DeJoy, co-founder and product lead at Astronomer, talks about building and managing Airflow pipelines on Astronomer and the upcoming improvements in Airflow 3. Pete shares his journey into data engineering, discusses Astronomer's contributions to the Airflow project, and highlights the critical role of Airflow in powering operational data products. He covers the evolution of Airflow, its position in the data ecosystem, and the challenges faced by data engineers, including infrastructure management and observability. The conversation also touches on the upcoming Airflow 3 release, which introduces data awareness, architectural improvements, and multi-language support, and Astronomer's observability suite, Astro Observe, which provides insights and proactive recommendations for Airflow users.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
Your host is Tobias Macey and today I'm interviewing Pete DeJoy about building and managing Airflow pipelines on Astronomer and the upcoming improvements in Airflow 3
Interview
Introduction
Can you describe what Astronomer is and the story behind it?
How would you characterize the relationship between Airflow and Astronomer?
Astronomer just released your State of Airflow 2025 Report yesterday and it is the largest data engineering survey ever with over 5,000 respondents. Can you talk a bit about top level findings in the report?
What about the overall growth of the Airflow project over time?
How have the focus and features of Astronomer changed since it was last featured on the show in 2017?
Astro Observe GA’d in early February, what does the addition of pipeline observability mean for your customers?
What are other capabilities similar in scope to observability that Astronomer is looking at adding to the platform?
Why is Airflow so critical in providing an elevated Observability–or cataloging, or something simlar - experience in a DataOps platform? What are the notable evolutions in the Airflow project and ecosystem in that time?
What are the core improvements that are planned for Airflow 3.0?
What are the most interesting, innovative, or unexpected ways that you have seen Astro used?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on Airflow and Astro?
What do you have planned for the future of Astro/Astronomer/Airflow?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
Links
Astronomer
Airflow
Maxime Beauchemin
MongoDB
Databricks
Confluent
Spark
Kafka
DagsterPodcast Episode
Prefect
Airflow 3
The Rise of the Data Engineer blog post
dbt
Jupyter Notebook
Zapier
cosmos library for dbt in Airflow
Ruff
Airflow Custom Operator
Snowflake
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Exploring Incident Management Strategies For Data Teams
Summary
Data assets and the pipelines that create them have become critical production infrastructure for companies. This adds a requirement for reliability and management of up-time similar to application infrastructure. In this episode Francisco Alberini and Mei Tao share their insights on what incident management looks like for data platforms and the teams that support them.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their managed Kubernetes platform it’s now even easier to deploy and scale your workflows, or try out the latest Helm charts from tools like Pulsar and Pachyderm. With simple pricing, fast networking, object storage, and worldwide data centers, you’ve got everything you need to run a bulletproof data platform. Go to dataengineeringpodcast.com/linode today and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
Atlan is a collaborative workspace for data-driven teams, like Github for engineering or Figma for design teams. By acting as a virtual hub for data assets ranging from tables and dashboards to SQL snippets & code, Atlan enables teams to create a single source of truth for all their data assets, and collaborate across the modern data stack through deep integrations with tools like Snowflake, Slack, Looker and more. Go to dataengineeringpodcast.com/atlan today and sign up for a free trial. If you’re a data engineering podcast listener, you get credits worth $3000 on an annual subscription
RudderStack helps you build a customer data platform on your warehouse or data lake. Instead of trapping data in a black box, they enable you to easily collect customer data from the entire stack and build an identity graph on your warehouse, giving you full visibility and control. Their SDKs make event streaming from any app or website easy, and their state-of-the-art reverse ETL pipelines enable you to send enriched data to any cloud tool. Sign up free… or just get the free t-shirt for being a listener of the Data Engineering Podcast at dataengineeringpodcast.com/rudder.
Are you looking for a structured and battle-tested approach for learning data engineering? Would you like to know how you can build proper data infrastructures that are built to last? Would you like to have a seasoned industry expert guide you and answer all your questions? Join Pipeline Academy, the worlds first data engineering bootcamp. Learn in small groups with likeminded professionals for 9 weeks part-time to level up in your career. The course covers the most relevant and essential data and software engineering topics that enable you to start your journey as a professional data engineer or analytics engineer. Plus we have AMAs with world-class guest speakers every week! The next cohort starts in April 2022. Visit dataengineeringpodcast.com/academy and apply now!
Your host is Tobias Macey and today I’m interviewing Francisco Alberini and Mei Tao about patterns and practices for incident management in data teams
Interview
Introduction
How did you get involved in the area of data management?
Can you start by describing some of the ways that an "incident" can manifest in a data system?
At a high level, what are the steps and participants required to bring an incident to resolution?
The principle of incident management is familiar to application/site reliability teams. What is the current state of the art/adoption for these practices among data teams?
What are the signals that teams should be monitoring to identify and alert on potential incidents?
Alerting is a subjective and nuanced practice, regardless of the context. What are some useful practices that you have seen and enacted to reduce alert fatigue and provide useful context in the alerts that do get sent?
Another aspect of this problem is the proper routing of alerts to ensure that the right person sees and acts on it. How have you seen teams deal with the challenge of delivering alerts to the right people?
When there is an active incident, what are the steps that you commonly see data teams take to understand the cause and scope of the issue?
How can teams augment their systems to make incidents faster to resolve?
What are the most interesting, innovative, or unexpected ways that you have seen teams approch incident response?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on incident management strategies?
What are the aspects of incident management for data teams that are still missing?
Contact Info
Mei
@tao_mei on Twitter
Email
Francisco
@falberini on Twitter
Email
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
To help other people find the show please leave a review on iTunes and tell your friends and co-workers
Links
Monte Carlo
Learn more about RCA best practices
Segment
Podcast Episode
Segment Protocols
Redshift
Airflow
dbt
Podcast Episode
The Goal by Eliahu Golratt
Data Mesh
Podcast Episode
Follow-Up Podcast Episode
PagerDuty
OpsGenie
Grafana
Prometheus
Sentry
Podcast.__init__ Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
Roger & DJ — The Rise of Big Data and CA's COVID-19 Response
Roger and DJ share some of the history behind data science as we know it today, and reflect on their experiences working on California's COVID-19 response.
---
Roger Magoulas is Senior Director of Data Strategy at Astronomer, where he works on data infrastructure, analytics, and community development. Previously, he was VP of Research at O'Reilly and co-chair of O'Reilly's Strata Data and AI Conference.
DJ Patil is a board member and former CTO of Devoted Health, a healthcare company for seniors. He was also Chief Data Scientist under the Obama administration and the Head of Data Science at LinkedIn.
Roger and DJ recently volunteered for the California COVID-19 response, and worked with data to understand case counts, bed capacities and the impact of intervention.
Connect with Roger and DJ:
📍 Roger's Twitter: https://twitter.com/rogerm
📍 DJ's Twitter: https://twitter.com/dpatil
---
🌟 Transcript: http://wandb.me/gd-roger-and-dj 🌟
⏳ Timestamps:
0:00 Sneak peek, intro
1:03 Coining the terms "big data" and "data scientist"
7:12 The rise of data science teams
15:28 Big Data, Hadoop, and Spark
23:10 The importance of using the right tools
29:20 BLUF: Bottom Line Up Front
34:44 California's COVID response
41:21 The human aspects of responding to COVID
48:33 Reflecting on the impact of COVID interventions
57:06 Advice on doing meaningful data science work
1:04:18 Outro
🍀 Links:
1. "MapReduce: Simplified Data Processing on Large Clusters" (Dean and Ghemawat, 2004): https://research.google/pubs/pub62/
2. "Big Data: Technologies and Techniques for Large-Scale Data" (Magoulas and Lorica, 2009): https://academics.uccs.edu/~ooluwada/courses/datamining/ExtraReading/BigData
3. The O'RLY book covers: https://www.businessinsider.com/these-hilarious-memes-perfectly-capture-what-its-like-to-work-in-tech-2016-4
4. "The Premonition" (Lewis, 2021): https://www.npr.org/2021/05/03/991570372/michael-lewis-the-premonition-is-a-sweeping-indictment-of-the-cdc
5. Why California's beaches are glowing with bioluminescence: https://www.youtube.com/watch?v=AVYSr19ReOs
6.
7. Sturgis Motorcyle Rally: https://en.wikipedia.org/wiki/Sturgis_Motorcycle_Rally
---
Get our podcast on these platforms:
👉 Apple Podcasts: http://wandb.me/apple-podcasts
👉 Spotify: http://wandb.me/spotify
👉 Google Podcasts: http://wandb.me/google-podcasts
👉 YouTube: http://wandb.me/youtube
👉 Soundcloud: http://wandb.me/soundcloud
Join our community of ML practitioners where we host AMAs, share interesting projects and meet other people working in Deep Learning:
http://wandb.me/slack
Check out Fully Connected, which features curated machine learning reports by researchers exploring deep learning techniques, Kagglers showcasing winning models, industry leaders sharing best practices, and more:
https://wandb.ai/fully-connected
Astronomer with Ry Walker - Episode 6
Summary
Building a data pipeline that is reliable and flexible is a difficult task, especially when you have a small team. Astronomer is a platform that lets you skip straight to processing your valuable business data. Ry Walker, the CEO of Astronomer, explains how the company got started, how the platform works, and their commitment to open source.
Preamble
Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at www.dataengineeringpodcast.com/linode?utm_source=rss&utm_medium=rss and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
You can help support the show by checking out the Patreon page which is linked from the site.
To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
This is your host Tobias Macey and today I’m interviewing Ry Walker, CEO of Astronomer, the platform for data engineering.
Interview
Introduction
How did you first get involved in the area of data management?
What is Astronomer and how did it get started?
Regulatory challenges of processing other people’s data
What does your data pipelining architecture look like?
What are the most challenging aspects of building a general purpose data management environment?
What are some of the most significant sources of technical debt in your platform?
Can you share some of the failures that you have encountered while architecting or building your platform and company and how you overcame them?
There are certain areas of the overall data engineering workflow that are well defined and have numerous tools to choose from. What are some of the unsolved problems in data management?
What are some of the most interesting or unexpected uses of your platform that you are aware of?
Contact Information
Email
@rywalker on Twitter
Links
Astronomer
Kiss Metrics
Segment
Marketing tools chart
Clickstream
HIPAA
FERPA
PCI
Mesos
Mesos DC/OS
Airflow
SSIS
Marathon
Prometheus
Grafana
Terraform
Kafka
Spark
ELK Stack
React
GraphQL
PostGreSQL
MongoDB
Ceph
Druid
Aries
Vault
Adapter Pattern
Docker
Kinesis
API Gateway
Kong
AWS Lambda
Flink
Redshift
NOAA
Informatica
SnapLogic
Meteor
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast