What even is the modern data stack (Interview)
Benn Stancil’s weekly Substack on data and technology provides a fascinating perspective on the modern data stack & the industry building it. On this episode, Benn joins Jerod to dissect a few of his essays, discuss opportunities he sees during this slowdown & explain why he thinks maybe we should disband the analytics team.
Join the discussion
Changelog++ members save 13 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
Cronitor – Cronitor helps you understand your cron jobs. Capture the status, metrics, and output from every cron job and background process. Name and organize each job, and ensure the right people are alerted when something goes wrong.
Retool – The low-code platform for developers to build internal tools — Some of the best teams out there trust Retool…Brex, Coinbase, Plaid, Doordash, LegalGenius, Amazon, Allbirds, Peloton, and so many more – the developers at these teams trust Retool as the platform to build their internal tools. Try it free at retool.com/changelog
Neon – Fleets of Postgres! Enterprises use Neon to operate hundreds of thousands of Postgres databases: Automated, instant provisioning of the world’s most popular database.
Featuring:
Benn Stancil – LinkedIn, X
Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
Benn’s Substack
It’s time to build
github.com/dbt-labs/dbt-core
Disband the analytics team
A gambler’s guide to giving talks
gilesbowkett/archaeopteryx
Something missing or broken? PRs welcome!
SaaS Branding: No Category Name Cost Years of Sales
Benn Stancil built a product customers loved but nobody could name. Mode's SaaS branding fell between BI dashboards and data science tools - a SaaS positioning problem that confused buyers for years.
Learn how Mode solved its SaaS branding challenge by narrowing to 'by analysts, for analysts,' using entertainment-first content for product positioning, and growing to 8-figure ARR.
Benn is the co-founder of Mode, a collaborative analytics platform. Their SaaS branding journey from category creation confusion to clarity serves clients like Lyft, DoorDash, and Bloomberg.
🔑 Key Lessons
🎯 SaaS branding needs a noun buyers can repeat: Mode fell between BI and data science with no category name - a SaaS positioning gap that made sales harder.
📉 Underinvesting in SaaS branding research costs years: Launching without product positioning research led to confused sales conversations for too long.
🛠️ Customers believe your SaaS branding labels: Mode called itself 'collaborative' and buyers repeated it as a purchase reason - despite zero collaboration features.
🚀 Entertainment-first content beats product marketing for SaaS branding: Mode's first blog post analyzed Miley Cyrus data, attracting analysts before they cared about the product.
🧠 Committed decisions beat perfect ones in SaaS branding: Three analytics founders learned that speed outperforms endless hedging on category creation decisions.
Chapters
Introduction
What Mode does and who it serves
Size of the business and key customers
How the idea started at Yammer
SaaS branding challenges after launch
The cost of underinvesting in product positioning
Why finding a category name took years
Handling conflicting customer requests
How SaaS branding language shapes expectations
Acquiring the first 10 customers
Why committed decisions beat perfect decisions
Content marketing and category creation
The startup marathon
Lightning round
Resources
Full show notes: https://saasclub.io/307
Join 5,000+ SaaS founders: https://saasclub.io/email
A Reflection On The Data Ecosystem For The Year 2021
Summary
This has been an active year for the data ecosystem, with a number of new product categories and substantial growth in existing areas. In an attempt to capture the zeitgeist Maura Church, David Wallace, Benn Stancil, and Gleb Mezhanskiy join the show to reflect on the past year and share their thought son the year to come.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their managed Kubernetes platform it’s now even easier to deploy and scale your workflows, or try out the latest Helm charts from tools like Pulsar and Pachyderm. With simple pricing, fast networking, object storage, and worldwide data centers, you’ve got everything you need to run a bulletproof data platform. Go to dataengineeringpodcast.com/linode today and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
Struggling with broken pipelines? Stale dashboards? Missing data? If this resonates with you, you’re not alone. Data engineers struggling with unreliable data need look no further than Monte Carlo, the world’s first end-to-end, fully automated Data Observability Platform! In the same way that application performance monitoring ensures reliable software and keeps application downtime at bay, Monte Carlo solves the costly problem of broken data pipelines. Monte Carlo monitors and alerts for data issues across your data warehouses, data lakes, ETL, and business intelligence, reducing time to detection and resolution from weeks or days to just minutes. Start trusting your data with Monte Carlo today! Visit dataengineeringpodcast.com/montecarlo to learn more. The first 10 people to request a personalized product tour will receive an exclusive Monte Carlo Swag box.
Are you bored with writing scripts to move data into SaaS tools like Salesforce, Marketo, or Facebook Ads? Hightouch is the easiest way to sync data into the platforms that your business teams rely on. The data you’re looking for is already in your data warehouse and BI tools. Connect your warehouse to Hightouch, paste a SQL query, and use their visual mapper to specify how data should appear in your SaaS systems. No more scripts, just SQL. Supercharge your business teams with customer data using Hightouch for Reverse ETL today. Get started for free at dataengineeringpodcast.com/hightouch.
Your host is Tobias Macey and today I’m interviewing Maura Church, David Wallace, Benn Stancil, and Gleb Mezhanskiy about the key themes of 2021 in the data ecosystem and what to expect for next year
Interview
Introduction
How did you get involved in the area of data management?
What were the main themes that you saw data practitioners and vendors focused on this year?
What is the major bottleneck for Data teams in 2021? Will it be the same in 2022?
One of the ways to reason about progress in any domain is to look at what was the primary bottleneck of further progress (data adoption for decision making) at different points in time. In the data domain, we have seen a number of bottlenecks, for example, scaling data platforms, the answer to which was Hadoop and on-prem columnar stores and then cloud data warehouses such as Snowflake & BigQuery. Then the problem was data integration and transformation which was solved by data integration vendors and frameworks such as Fivetran / Airbyte, modern orchestration frameworks such as Dagster & dbt and “reverse-ETL” Hightouch. What is the main challenge now?
Will SQL be challenged as a primary interface to analytical data?
In 2020 we’ve seen a few launches of post-SQL languages such as Malloy, Preql, metric layer query languages from Transform and Supergrain.
To what extent does speed matter?
Over the past couple of months, we’ve seen the resurgence of “benchmark wars” between major data warehousing platforms. To what extent do speed benchmarks inform decisions for modern data teams? How important is query speed in a modern data workflow? What needs to be true about your current DWH solution and potential alternatives to make a move?
How has the way data teams work been changing?
In 2020 remote seemed like a temporary emergency state. In 2021, it went mainstream. How has that affected the day-to-day of data teams, how they collaborate internally and with stakeholders?
What’s it like to be a data vendor in 2021?
Vertically integrated vs. modular data stack?
There are multiple forces in play. Will the stack continue to be fragmented? Will we see major consolidation? If so, in which parts of the stack?
Contact Info
Maura
LinkedIn
Website
@outoftheverse on Twitter
David
LinkedIn
@davidjwallace on Twitter
dwallace0723 on GitHub
Benn
LinkedIn
@bennstancil on Twitter
Gleb
LinkedIn
@glebmm on Twitter
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
To help other people find the show please leave a review on iTunes and tell your friends and co-workers
Links
Patreon
Dutchie
Mode Analytics
Datafold
Podcast Episode
Locally Optimistic
RJ Metrics
Stitch
Mozart Data
Podcast Episode
Dagster
Podcast Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
Building Tools And Platforms For Data Analytics
Summary
Data engineers are responsible for building tools and platforms to power the workflows of other members of the business. Each group of users has their own set of requirements for the way that they access and interact with those platforms depending on the insights they are trying to gather. Benn Stancil is the chief analyst at Mode Analytics and in this episode he explains the set of considerations and requirements that data analysts need in their tools and. He also explains useful patterns for collaboration between data engineers and data analysts, and what they can learn from each other.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. And for your machine learning workloads, they just announced dedicated CPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management.For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Dataversity, Corinium Global Intelligence, and Data Counsil. Upcoming events include the O’Reilly AI conference, the Strata Data conference, the combined events of the Data Architecture Summit and Graphorum, and Data Council in Barcelona. Go to dataengineeringpodcast.com/conferences to learn more about these and other events, and take advantage of our partner discounts to save money when you register today.
Your host is Tobias Macey and today I’m interviewing Benn Stancil, chief analyst at Mode Analytics, about what data engineers need to know when building tools for analysts
Interview
Introduction
How did you get involved in the area of data management?
Can you start by describing some of the main features that you are looking for in the tools that you use?
What are some of the common shortcomings that you have found in out-of-the-box tools that organizations use to build their data stack?
What should data engineers be considering as they design and implement the foundational data platforms that higher order systems are built on, which are ultimately used by analysts and data scientists?
In terms of mindset, what are the ways that data engineers and analysts can align and where are the points of conflict?
In terms of team and organizational structure, what have you found to be useful patterns for reducing friction in the product lifecycle for data tools (internal or external)?
What are some anti-patterns that data engineers can guard against as they are designing their pipelines?
In your experience as an analyst, what have been the characteristics of the most seamless projects that you have been involved with?
How much understanding of analytics are necessary for data engineers to be successful in their projects and careers?
Conversely, how much understanding of data management should analysts have?
What are the industry trends that you are most excited by as an analyst?
Contact Info
LinkedIn
@bennstancil on Twitter
Website
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
To help other people find the show please leave a review on iTunes and tell your friends and co-workers
Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
Links
Mode Analytics
Data Council Presentation
Yammer
StitchFix Blog Post
SnowflakeDB
Re:Dash
Superset
Marquez
Amundsen
Podcast Episode
Elementl
Dagster
Data Council Presentation
DBT
Podcast Episode
Great Expectations
Podcast.__init__ Episode
Delta Lake
Podcast Episode
Stitch
Fivetran
Podcast Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast