The AI Data Paradox: High Trust in Models, Low Trust in Data
Summary
In this episode of the Data Engineering Podcast Ariel Pohoryles, head of product marketing for Boomi's data management offerings, talks about a recent survey of 300 data leaders on how organizations are investing in data to scale AI. He shares a paradox uncovered in the research: while 77% of leaders trust the data feeding their AI systems, only 50% trust their organization's data overall. Ariel explains why truly productionizing AI demands broader, continuously refreshed data with stronger automation and governance, and highlights the challenges posed by unstructured data and vector stores. The conversation covers the need to shift from manual reviews to automated pipelines, the resurgence of metadata and master data management, and the importance of guardrails, traceability, and agent governance. Ariel also predicts a growing convergence between data teams and application integration teams and advises leaders to focus on high-value use cases, aggressive pipeline automation, and cataloging and governing the coming sprawl of AI agents, all while using AI to accelerate data engineering itself.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
Composable data infrastructure is great, until you spend all of your time gluing it together. Bruin is an open source framework, driven from the command line, that makes integration a breeze. Write Python and SQL to handle the business logic, and let Bruin handle the heavy lifting of data movement, lineage tracking, data quality monitoring, and governance enforcement. Bruin allows you to build end-to-end data workflows using AI, has connectors for hundreds of platforms, and helps data teams deliver faster. Teams that use Bruin need less engineering effort to process data and benefit from a fully integrated data platform. Go to dataengineeringpodcast.com/bruin today to get started. And for dbt Cloud customers, they'll give you $1,000 credit to migrate to Bruin Cloud.
Your host is Tobias Macey and today I'm interviewing Ariel Pohoryles about data management investments that organizations are making to enable them to scale AI implementations
Interview
Introduction
How did you get involved in the area of data management?
Can you start by describing the motivation and scope of your recent survey on data management investments for AI across your respondents?What are the key takeaways that were most significant to you?
The survey reveals a fascinating paradox: 77% of leaders trust the data used by their AI systems, yet only half trust their organization's overall data quality. For our data engineering audience, what does this suggest about how companies are currently sourcing data for AI? Does it imply they are using narrow, manually-curated "golden datasets," and what are the technical challenges and risks of that approach as they try to scale?
The report highlights a heavy reliance on manual data quality processes, with one expert noting companies feel it's "not reliable to fully automate validation" for external or customer data. At the same time, maturity in "Automated tools for data integration and cleansing" is low, at only 42%. What specific technical hurdles or organizational inertia are preventing teams from adopting more automation in their data quality and integration pipelines?
There was a significant point made that with generative AI, "biases can scale much faster," making automated governance essential. From a data engineering perspective, how does the data management strategy need to evolve to support generative AI versus traditional ML models? What new types of data quality checks, lineage tracking, or monitoring for feedback loops are required when the model itself is generating new content based on its own outputs?
The report champions a "centralized data management platform" as the "connective tissue" for reliable AI. How do you see the scale and data maturity impacting the realities of that effort?How do architectural patterns in the shape of cloud warehouses, lakehouses, data mesh, data products, etc. factor into that need for centralized/unified platforms?
A surprising finding was that a third of respondents have not fully grasped the risk of significant inaccuracies in their AI models if they fail to prioritize data management. In your experience, what are the biggest blind spots for data and analytics leaders?
Looking at the maturity charts, companies rate themselves highly on "Developing a data management strategy" (65%) but lag significantly in areas like "Automated tools for data integration and cleansing" (42%) and "Conducting bias-detection audits" (24%). If you were advising a data engineering team lead based on these findings, what would you tell them to prioritize in the next 6-12 months to bridge the gap between strategy and a truly scalable, trustworthy data foundation for AI?
The report states that 83% of companies expect to integrate more data sources for their AI in the next year. For a data engineer on the ground, what is the most important capability they need to build into their platform to handle this influx?
What are the most interesting, innovative, or unexpected ways that you have seen teams addressing the new and accelerated data needs for AI applications?
What are some of the noteworthy trends or predictions that you have for the near-term future of the impact that AI is having or will have on data teams and systems?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
Links
BoomiData Management
Integration & Automation Demo
Agentstudio
Data Connector Agent Webinar
Survey Results
Data Governance
Shadow ITPodcast Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
#271 Steve Lucas: Why AI Agents Will Automate 75% of Business Operations by 2026
This episode is sponsored by Oracle. OCI is the next-generation cloud designed for every workload – where you can run any application, including any AI projects, faster and more securely for less. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking. Join Modal, Skydance Animation, and today's innovative AI tech companies who upgraded to OCI…and saved.
Try OCI for free at http://oracle.com/eyeonai
What if the future of enterprise wasn't human-driven, but agent-driven?
In this groundbreaking episode, Steve Lucas, CEO of Boomi, unveils a radical vision for the next era of business: one where AI agents will power 75% of enterprise operations by 2026. From eliminating traditional user interfaces to transforming legacy systems with no-code automation, Steve walks us through how Boomi is building the infrastructure for a self-driving enterprise, and why businesses that fail to prepare will be left behind.
This episode will shift your perspective on where the enterprise is headed and who (or what) will be running it.
Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
Unpacking The Seven Principles Of Modern Data Pipelines
Summary
Data pipelines are the core of every data product, ML model, and business intelligence dashboard. If you're not careful you will end up spending all of your time on maintenance and fire-fighting. The folks at Rivery distilled the seven principles of modern data pipelines that will help you stay out of trouble and be productive with your data. In this episode Ariel Pohoryles explains what they are and how they work together to increase your chances of success.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
This episode is brought to you by Datafold – a testing automation platform for data engineers that finds data quality issues before the code and data are deployed to production. Datafold leverages data-diffing to compare production and development environments and column-level lineage to show you the exact impact of every code change on data, metrics, and BI tools, keeping your team productive and stakeholders happy. Datafold integrates with dbt, the modern data stack, and seamlessly plugs in your data CI for team-wide and automated testing. If you are migrating to a modern data stack, Datafold can also help you automate data and code validation to speed up the migration. Learn more about Datafold by visiting dataengineeringpodcast.com/datafold
Your host is Tobias Macey and today I'm interviewing Ariel Pohoryles about the seven principles of modern data pipelines
Interview
Introduction
How did you get involved in the area of data management?
Can you start by defining what you mean by a "modern" data pipeline?
At Rivery you published a white paper identifying seven principles of modern data pipelines:
Zero infrastructure management
ELT-first mindset
Speaks SQL and Python
Dynamic multi-storage layers
Reverse ETL & operational analytics
Full transparency
Faster time to value
What are the applications of data that you focused on while identifying these principles?
How do the application of these principles influence the ability of organizations and their data teams to encourage and keep pace with the use of data in the business?
What are the technical components of a pipeline infrastructure that are necessary to support a "modern" workflow?
How do the technologies involved impact the organizational involvement with how data is applied throughout the business?
When using managed services, what are the ways that the pricing model acts to encourage/discourage experimentation/exploration with data?
What are the most interesting, innovative, or unexpected ways that you have seen these seven principles implemented/applied?
What are the most interesting, unexpected, or challenging lessons that you have learned while working with customers to adapt to these principles?
What are the cases where some/all of these principles are undesirable/impractical to implement?
What are the opportunities for further advancement/sophistication in the ways that teams work with and gain value from data?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
Links
Rivery
7 Principles Of The Modern Data Pipeline
ELT
Reverse ETL
Martech Landscape
Data Lakehouse
Databricks
Snowflake
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Sponsored By:
Datafold: 
This episode is brought to you by Datafold – a testing automation platform for data engineers that finds data quality issues before the code and data are deployed to production. Datafold leverages data-diffing to compare production and development environments and column-level lineage to show you the exact impact of every code change on data, metrics, and BI tools, keeping your team productive and stakeholders happy. Datafold integrates with dbt, the modern data stack, and seamlessly plugs in your data CI for team-wide and automated testing. If you are migrating to a modern data stack, Datafold can also help you automate data and code validation to speed up the migration. Learn more about Datafold by visiting [dataengineeringpodcast.com/datafold](https://www.dataengineeringpodcast.com/datafold) today!
Rudderstack: 
Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)
Support Data Engineering Podcast
SaaStr 182: Marketo CEO, Steve Lucas on What Makes A Truly Great SaaS CEO Today, The Top Considerations You Must make Before Going To Enterprise & Why The Way We Sell Has To Fundamentally Change
Steve Lucas is the CEO @ Marketo, the world leader in marketing automation for companies of any size. Prior to their IPO and eventual sale to Vista Equity partners for $1.79Bn they raised over $100m in VC funding from the likes of Battery Ventures, IVP, Mayfield and Lead Edge Capital. As for Steve, prior to joining Marketo, he served in many leadership positions at SAP, Salesforce, Microsoft, BusinessObjects, and Crystal Decisions. If that wasn't enough Steve also sits on the board of Tivo, SendGrid and The American Diabetes Society.
In Today's Episode You Will Learn:
How did Steve make his way into the world of SaaS and come to be CEO @ Marketo?
Why does Steve describe his experience at Salesforce to be life-changing? What were the core takeaways for Steve? How has that impacted how he operates today with Marketo? What does Steve mean when he says Marc Benioff is a "master of relevance"?
Why does Steve believe the key to success as a CEO is accessibility? How can CEOs be both vulnerable and strong in today's SaaS world? What are the 2 different types of CEOs and how they engage with their CMOs? What do the best do? What do the worst do?
Why does Steve believe that the "CRM" term is incomplete? How does Steve fundamentally believe the way that customers want to be engaged with has changed? How can marketers enact this level of personalisation and engagement with such large customer bases? How does the role of artificial intelligence fit into this mass scale personalisation?
How does Steve view the broader martech landscape? Why does Steve strongly believe that we will be entering a period of consolidation in martech? How does Steve view the emergence of new categories such as ABM? How does this impact his overarching view on the next wave for martech?
Steve's 60 Second SaaStr
What does Steve know now that he wishes he had known when he started?
Management upgrade is the most important role of CEO, agree?
What keeps Steve up at night? How does that influence his running and operations of Marketo?
Read the full transcript on our blog.
If you would like to find out more about the show and the guests presented, you can follow us on Twitter here:
Jason Lemkin
Harry Stebbings
SaaStr
Steve Lucas