Why AI engineering needs old-school discipline
In this episode of The New Stack Makers, Nimisha Asthagiri of Thoughtworks explores why many AI initiatives stall between proof of concept and production. A key issue is that organizations focus on speed—asking how to move faster—rather than rethinking what new capabilities AI actually enables. Successful companies take a systems-thinking approach, investing in organizational literacy and aligning teams around meaningful use cases instead of retrofitting AI into existing workflows.
Asthagiri highlights that core engineering practices are ফিরে to prominence. As AI-generated code increases, so does the risk of “cognitive debt,” where developers lose understanding of their own systems. To counter this, teams are reviving fundamentals like test-driven development, mutation testing, observability, and zero-trust security, especially as autonomous agents contribute to production code.
She also introduces the concept of “dark code”—AI-generated code that may never be used—and argues for more intentional lifecycle management, including ephemeral code. Ultimately, the focus shifts from code itself to specifications, context management, and disciplined engineering practices.
Learn more from The New Stack around the latest about system-thinking approaches:
System Two AI: The Dawn of Reasoning Agents in Business
A practical systems engineering guide: Architecting AI-ready infrastructure for the agentic era
Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
How AI will change software engineering – with Martin Fowler
Brought to You By:
• Statsig — The unified platform for flags, analytics, experiments, and more. AI-accelerated development isn’t just about shipping faster: it’s about measuring whether, what you ship, actually delivers value. This is where modern experimentation with Statsig comes in. Check it out.
• Linear — The system for modern product development. I had a jaw-dropping experience when I dropped in for the weekly “Quality Wednesdays” meeting at Linear. Every week, every dev fixes at least one quality isse, large or small. Even if it’s one pixel misalignment, like this one. I’ve yet to see a team obsess this much about quality. Read more about how Linear does Quality Wednesdays – it’s fascinating!
—
Martin Fowler is one of the most influential people within software architecture, and the broader tech industry. He is the Chief Scientist at Thoughtworks and the author of Refactoring and Patterns of Enterprise Application Architecture, and several other books. He has spent decades shaping how engineers think about design, architecture, and process, and regularly publishes on his blog, MartinFowler.com.
In this episode, we discuss how AI is changing software development: the shift from deterministic to non-deterministic coding; where generative models help with legacy code; and the narrow but useful cases for vibe coding. Martin explains why LLM output must be tested rigorously, why refactoring is more important than ever, and how combining AI tools with deterministic techniques may be what engineering teams need.
We also revisit the origins of the Agile Manifesto and talk about why, despite rapid changes in tooling and workflows, the skills that make a great engineer remain largely unchanged.
—
Timestamps
(00:00) Intro
(01:50) How Martin got into software engineering
(07:48) Joining Thoughtworks
(10:07) The Thoughtworks Technology Radar
(16:45) From Assembly to high-level languages
(25:08) Non-determinism
(33:38) Vibe coding
(39:22) StackOverflow vs. coding with AI
(43:25) Importance of testing with LLMs
(50:45) LLMs for enterprise software
(56:38) Why Martin wrote Refactoring
(1:02:15) Why refactoring is so relevant today
(1:06:10) Using LLMs with deterministic tools
(1:07:36) Patterns of Enterprise Application Architecture
(1:18:26) The Agile Manifesto
(1:28:35) How Martin learns about AI
(1:34:58) Advice for junior engineers
(1:37:44) The state of the tech industry today
(1:42:40) Rapid fire round
—
The Pragmatic Engineer deepdives relevant for this episode:
• Vibe coding as a software engineer
• The AI Engineering stack
• AI Engineering in the real world
• What changed in 50 years of computing
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.
Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe
What will the dream car of the future be like? | Alex Koster
Fasten your seat belt as software engineer Alex Koster takes us on a journey in what he calls the "software dream car" of the future. He breaks down how massive technological shifts are transforming the automotive industry and paints a vivid picture of where cars are headed -- from AI drivers to interiors and exteriors shaped by augmented and virtual reality.
Hosted on Acast. See acast.com/privacy for more information.
The Realities of Working in Data with Emily Gorcenski
Emily Gorcenski, Data & AI Service Line Lead at Thoughtworks, joins Corey on Screaming in the Cloud to discuss how big data is changing our lives - both for the better, and the challenges that come with it. Emily explains how data is only important if you know what to do with it and have a plan to work with it, and why it’s crucial to understand the use-by date on your data. Corey and Emily also discuss how big data problems aren’t universal problems for the rest of the data community, how to address the ethics around AI, and the barriers to entry when pursuing a career in data.
About Emily
Emily Gorcenski is a principal data scientist and the Data & AI Service Line Lead of ThoughtWorks Germany. Her background in computational mathematics and control systems engineering has given her the opportunity to work on data analysis and signal processing problems from a variety of complex and data intensive industries. In addition, she is a renowned data activist and has contributed to award-winning journalism through her use of data to combat extremist violence and terrorism. The opinions expressed are solely her own.
Links Referenced:
ThoughtWorks: https://www.thoughtworks.com/
Personal website: https://emilygorcenski.com
Twitter: https://twitter.com/EmilyGorcenski
Mastodon: https://mastodon.green/@emilygorcenski@indieweb.social
609: Data Mesh
Jon Krohn speaks with Zhamak Dehghani, the empathetic technologist who coined the term “data mesh”. They explore what a data mesh is, and how its approach toward secure interconnectivity will help solve a roster of data-led business problems.
This episode is brought to you by Zencastr (zen.ai/sds), the easiest way to make high-quality podcasts. Interested in sponsoring a SuperDataScience Podcast episode? Visit JonKrohn.com/podcast for sponsorship information.
In this episode you will learn:
• The importance of data meshes [3:29]
• How standardizing database interfaces helps tech giants like Amazon [6:40]
• Current challenges with data meshes [9:33]
• How data meshes give users the freedom to work with data [17:09]
• The missing piece of the puzzle for data meshes [22:11]
• How data meshes connect with the metaverse and Web3 [33:18]
• The times when data meshes aren’t fit for purpose [42:24]
Additional materials: www.superdatascience.com/609
Revisiting The Technical And Social Benefits Of The Data Mesh
Summary
The data mesh is a thesis that was presented to address the technical and organizational challenges that businesses face in managing their analytical workflows at scale. Zhamak Dehghani introduced the concepts behind this architectural patterns in 2019, and since then it has been gaining popularity with many companies adopting some version of it in their systems. In this episode Zhamak re-joins the show to discuss the real world benefits that have been seen, the lessons that she has learned while working with her clients and the community, and her vision for the future of the data mesh.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their managed Kubernetes platform it’s now even easier to deploy and scale your workflows, or try out the latest Helm charts from tools like Pulsar and Pachyderm. With simple pricing, fast networking, object storage, and worldwide data centers, you’ve got everything you need to run a bulletproof data platform. Go to dataengineeringpodcast.com/linode today and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
Atlan is a collaborative workspace for data-driven teams, like Github for engineering or Figma for design teams. By acting as a virtual hub for data assets ranging from tables and dashboards to SQL snippets & code, Atlan enables teams to create a single source of truth for all their data assets, and collaborate across the modern data stack through deep integrations with tools like Snowflake, Slack, Looker and more. Go to dataengineeringpodcast.com/atlan today and sign up for a free trial. If you’re a data engineering podcast listener, you get credits worth $3000 on an annual subscription
Modern Data teams are dealing with a lot of complexity in their data pipelines and analytical code. Monitoring data quality, tracing incidents, and testing changes can be daunting and often takes hours to days. Datafold helps Data teams gain visibility and confidence in the quality of their analytical data through data profiling, column-level lineage and intelligent anomaly detection. Datafold also helps automate regression testing of ETL code with its Data Diff feature that instantly shows how a change in ETL or BI code affects the produced data, both on a statistical level and down to individual rows and values. Datafold integrates with all major data warehouses as well as frameworks such as Airflow & dbt and seamlessly plugs into CI workflows. Go to dataengineeringpodcast.com/datafold today to start a 30-day trial of Datafold.
Your host is Tobias Macey and today I’m welcoming back Zhamak Dehghani to talk about her work on the data mesh book and the lessons learned over the past 2 years
Interview
Introduction
How did you get involved in the area of data management?
Can you start by giving a brief recap of the principles of the data mesh and the story behind it?
How has your view of the principles of the data mesh changed since our conversation in July of 2019?
What are some of the ways that your work on the data mesh book influenced your thinking on the practical elements of implementing a data mesh?
What do you view as the as-yet-unknown elements of the technical and social design constructs that are needed for a sustainable data mesh implementation?
In the opening of your book you state that "Data Mesh is a new approach in sourcing, managing, and accessing data for analytical use cases at scale". As with everything, scale is subjective, but what are some of the heuristics that you rely on for determining when a data mesh is an appropriate solution?
What are some of the ways that data mesh concepts manifest at the boundaries of organizations?
While the idea of federated access to data product quanta reduces the amount of coordination necessary at the organizational level, it raises the spectre of more complex logic required for consumers of multiple quanta. How can data mesh implementations mitigate the impact of this problem?
What are some of the technical components that you have found to be best suited to the implementation of data elements within a mesh?
What are the technological components that are still missing for a mesh-native data platform?
How should an organization that wishes to implement a mesh style architecture think about the roles and skills that they will need on staff?
How can vendors factor into the solution?
What is the role of application developers in a data mesh ecosystem and how do they need to change their thinking around the interfaces that they provide in their products?
What are the most interesting, innovative, or unexpected ways that you have seen data mesh principles used?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on data mesh implementations?
When is a data mesh the wrong approach?
What do you think the future of the data mesh will look like?
Contact Info
LinkedIn
@zhamakd on Twitter
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
Data Engineering Podcast Data Mesh Interview
Data Mesh Book
Thoughtworks
Expert Systems
OpenLineage
Podcast Episode
Data Mesh Learning
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
Discussing Type Hints, Protocols, and Ducks in Python
<p>There seem to be three kinds of Python developers: those unaware of type hints or have no opinion, ones that embrace them, and others who have an allergic reaction at the mention of them. Python is famously a dynamically typed language, but there are advantages to adding type hints to your code. This week on the show, we have Luciano Ramalho to discuss his recent talk titled, “Type hints, protocols, and good sense.”</p>
<p>Luciano was not a fan of type hints. He’s only recently come around to their potential with the introduction of protocols in PEP 544. Python has adopted a gradual type system that is optional at all levels. We discuss the advantages, pitfalls, and recent developments around type hinting in Python.</p>
<p>We also talk about the second edition of Luciano’s book Fluent Python. He researched type hints in-depth for the book, which led to his recent conference talks on the subject. He also shares his experience with adding opinionated asides to the book in a fun and unique way.</p>
<div class="alert alert-primary" role="alert">
<p><strong>Course Spotlight:</strong> <a href="https://realpython.com/courses/python-type-checking/">Python Type Checking</a> </p>
<p>In this course, you’ll look at Python type checking. Traditionally, types have been handled by the Python interpreter in a flexible but implicit way. Recent versions of Python allow you to specify explicit type hints that can be used by different tools to help you develop your code more efficiently.</p>
</div>
<p>Topics:</p>
<ul>
<li>00:00:00 – Introduction</li>
<li>00:02:02 – Are you interested in creative uses for Python?</li>
<li>00:04:41 – Protocol: The keystone of type hints</li>
<li>00:08:14 – What is duck typing?</li>
<li>00:12:44 – Protocols declaring one method and emerging from a code base</li>
<li>00:17:04 – An example where type hint was too lax</li>
<li>00:21:20 – What if Python always had a strict type system?</li>
<li>00:33:23 – Sponsor: Cloudsmith</li>
<li>00:34:09 – Bias in companies using type hints, and projects that fail checking</li>
<li>00:40:27 – Background on personal use of type hints and added complexity</li>
<li>00:45:07 – Unsuitability of type hints for checking business rules</li>
<li>00:52:30 – Video Course Spotlight</li>
<li>00:53:46 – Fluent Python, 2nd edition</li>
<li>00:56:05 – Who is the intended developer for the book?</li>
<li>00:58:12 – Soapbox sections of the book</li>
<li>00:59:35 – What were things you were excited to update or add to the book?</li>
<li>01:05:46 – Metaprogramming portion of the book</li>
<li>01:08:17 – What are you excited about in the world of Python?</li>
<li>01:10:35 – What do you want to learn next?</li>
<li>01:18:41 – Shoutouts, plugs, and/or social connections</li>
<li>01:19:47 – Thanks and goodbye</li>
</ul>
<p>Show Links:</p>
<ul>
<li><a href="https://www.oreilly.com/library/view/fluent-python-2nd/9781492056348/">Fluent Python, 2nd Edition</a></li>
<li><a href="https://www.youtube.com/watch?v=kDDCKwP7QgQ">Protocol: The keystone of type hints - Luciano Ramalho | PyCon US 2021</a></li>
<li><a href="https://speakerdeck.com/ramalho/type-hints-protocols-and-good-sense">Type hints, protocols, and good sense: PyCon India 2021 - Speaker Deck</a></li>
<li><a href="https://www.youtube.com/watch?v=eKEjkB2bXK4&list=PL2Uw4_HvXqvYk1Y5P8kryoyd83L_0Uk5K&index=49&t=2s">Generate buzz with realtime FM audio synthesis - Łukasz Langa | PyCon US 2021</a></li>
<li><a href="https://garoa.net.br/wiki/P%C3%A1gina_principal">Garoa Hacker Clube</a></li>
<li><a href="https://py.processing.org/tutorials/">Processing.py - Tutorials</a></li>
<li><a href="https://www.python.org/dev/peps/pep-0544/">PEP 544 – Protocols: Structural subtyping (static duck typing) | Python.org</a></li>
<li><a href="https://github.com/python/typeshed/">typeshed: Collection of library stubs for Python, with static types</a></li>
<li><a href="https://realpython.com/python-type-checking/#duck-types-and-protocols">Python Type Checking (Guide) – Real Python</a></li>
<li><a href="https://mypy.readthedocs.io/en/latest/protocols.html">Protocols and structural subtyping — Mypy documentation</a></li>
<li><a href="https://en.wikipedia.org/wiki/Dependent_type">Dependent type - Wikipedia</a></li>
<li><a href="https://github.com/Microsoft/pyright">microsoft/pyright: Static type checker for Python</a></li>
<li><a href="https://mypy.readthedocs.io/en/stable/">Welcome to mypy documentation!</a></li>
<li><a href="https://www.python.org/dev/peps/pep-0487/">PEP 487 – Simpler customisation of class creation | Python.org</a></li>
<li><a href="https://www.python.org/dev/peps/pep-0636/">PEP 636 – Structural Pattern Matching: Tutorial | Python.org</a></li>
<li><a href="https://docs.python.org/3/whatsnew/3.10.html#better-error-messages">What’s New In Python 3.10 — Better error messages</a></li>
<li><a href="https://flutter.dev/">Flutter - Build apps for any screen</a></li>
<li><a href="https://ramalho.org/wiki/doku.php?id=start/">Ramalho.org/wiki</a></li>
<li><a href="https://twitter.com/ramalhoorg">Luciano Ramalho Twitter(@ramalhoorg)</a></li>
</ul>
<p>Level up your Python skills with our expert-led courses:</p>
<ul>
<li><a href="https://realpython.com/courses/asteroids-game-python-pygame/">Using Pygame to Build an Asteroids Game in Python</a></li>
<li><a href="https://realpython.com/courses/records-sets-ideal-data-structure/">Records and Sets: Selecting the Ideal Data Structure</a></li>
<li><a href="https://realpython.com/courses/python-type-checking/">Python Type Checking</a></li>
</ul> <p><a rel="payment" href="https://realpython.com/join">Support the podcast & join our community of Pythonistas</a></p>
Straining Your Data Lake Through A Data Mesh
Summary
The current trend in data management is to centralize the responsibilities of storing and curating the organization’s information to a data engineering team. This organizational pattern is reinforced by the architectural pattern of data lakes as a solution for managing storage and access. In this episode Zhamak Dehghani shares an alternative approach in the form of a data mesh. Rather than connecting all of your data flows to one destination, empower your individual business units to create data products that can be consumed by other teams. This was an interesting exploration of a different way to think about the relationship between how your data is produced, how it is used, and how to build a technical platform that supports the organizational needs of your business.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. And for your machine learning workloads, they just announced dedicated CPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
And to grow your professional network and find opportunities with the startups that are changing the world then Angel List is the place to go. Go to dataengineeringpodcast.com/angel to sign up today.
You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management.For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Dataversity, and the Open Data Science Conference. Upcoming events include the O’Reilly AI Conference, the Strata Data Conference, and the combined events of the Data Architecture Summit and Graphorum. Go to dataengineeringpodcast.com/conferences to learn more and take advantage of our partner discounts when you register.
Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
To help other people find the show please leave a review on iTunes and tell your friends and co-workers
Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
Your host is Tobias Macey and today I’m interviewing Zhamak Dehghani about building a distributed data mesh for a domain oriented approach to data management
Interview
Introduction
How did you get involved in the area of data management?
Can you start by providing your definition of a "data lake" and discussing some of the problems and challenges that they pose?
What are some of the organizational and industry trends that tend to lead to this solution?
You have written a detailed post outlining the concept of a "data mesh" as an alternative to data lakes. Can you give a summary of what you mean by that phrase?
In a domain oriented data model, what are some useful methods for determining appropriate boundaries for the various data products?
What are some of the challenges that arise in this data mesh approach and how do they compare to those of a data lake?
One of the primary complications of any data platform, whether distributed or monolithic, is that of discoverability. How do you approach that in a data mesh scenario?
A corollary to the issue of discovery is that of access and governance. What are some strategies to making that scalable and maintainable across different data products within an organization?
Who is responsible for implementing and enforcing compliance regimes?
One of the intended benefits of data lakes is the idea that data integration becomes easier by having everything in one place. What has been your experience in that regard?
How do you approach the challenge of data integration in a domain oriented approach, particularly as it applies to aspects such as data freshness, semantic consistency, and schema evolution?
Has latency of data retrieval proven to be an issue in your work?
When it comes to the actual implementation of a data mesh, can you describe the technical and organizational approach that you recommend?
How do team structures and dynamics shift in this scenario?
What are the necessary skills for each team?
Who is responsible for the overall lifecycle of the data in each domain, including modeling considerations and application design for how the source data is generated and captured?
Is there a general scale of organization or problem domain where this approach would generate too much overhead and maintenance burden?
For an organization that has an existing monolothic architecture, how do you suggest they approach decomposing their data into separately managed domains?
Are there any other architectural considerations that data professionals should be considering that aren’t yet widespread?
Contact Info
LinkedIn
@zhamakd on Twitter
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh
Thoughtworks
Technology Radar
Data Lake
Data Warehouse
James Dixon
Azure Data Lake
"Big Ball Of Mud" Anti-Pattern
ETL
ELT
Hadoop
Spark
Kafka
Event Sourcing
Airflow
Podcast.__init__ Episode
Data Engineering Episode
Data Catalog
Master Data Management
Podcast Episode
Polyseme
REST
CNCF (Cloud Native Computing Foundation)
Cloud Events Standard
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
#166: Continuous delivery with Python
See the full show notes for this episode on the website at talkpython.fm/166
Database Refactoring Patterns with Pramod Sadalage - Episode 22
Summary
As software lifecycles move faster, the database needs to be able to keep up. Practices such as version controlled migration scripts and iterative schema evolution provide the necessary mechanisms to ensure that your data layer is as agile as your application. Pramod Sadalage saw the need for these capabilities during the early days of the introduction of modern development practices and co-authored a book to codify a large number of patterns to aid practitioners, and in this episode he reflects on the current state of affairs and how things have changed over the past 12 years.
Preamble
Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
You can help support the show by checking out the Patreon page which is linked from the site.
To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
Your host is Tobias Macey and today I’m interviewing Pramod Sadalage about refactoring databases and integrating database design into an iterative development workflow
Interview
Introduction
How did you get involved in the area of data management?
You first co-authored Refactoring Databases in 2006. What was the state of software and database system development at the time and why did you find it necessary to write a book on this subject?
What are the characteristics of a database that make them more difficult to manage in an iterative context?
How does the practice of refactoring in the context of a database compare to that of software?
How has the prevalence of data abstractions such as ORMs or ODMs impacted the practice of schema design and evolution?
Is there a difference in strategy when refactoring the data layer of a system when using a non-relational storage system?
How has the DevOps movement and the increased focus on automation affected the state of the art in database versioning and evolution?
What have you found to be the most problematic aspects of databases when trying to evolve the functionality of a system?
Looking back over the past 12 years, what has changed in the areas of database design and evolution?
How has the landscape of tooling for managing and applying database versioning changed since you first wrote Refactoring Databases?
What do you see as the biggest challenges facing us over the next few years?
Contact Info
Website
pramodsadalage on GitHub
@pramodsadalage on Twitter
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
Database Refactoring
Website
Book
Thoughtworks
Martin Fowler
Agile Software Development
XP (Extreme Programming)
Continuous Integration
The Book
Wikipedia
Test First Development
DDL (Data Definition Language)
DML (Data Modification Language)
DevOps
Flyway
Liquibase
DBMaintain
Hibernate
SQLAlchemy
ORM (Object Relational Mapper)
ODM (Object Document Mapper)
NoSQL
Document Database
MongoDB
OrientDB
CouchBase
CassandraDB
Neo4j
ArangoDB
Unit Testing
Integration Testing
OLAP (On-Line Analytical Processing)
OLTP (On-Line Transaction Processing)
Data Warehouse
Docker
QA==Quality Assurance
HIPAA (Health Insurance Portability and Accountability Act)
PCI DSS (Payment Card Industry Data Security Standard)
Polyglot Persistence
Toplink Java ORM
Ruby on Rails
ActiveRecord Gem
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
#24: Fluent Python
See the full show notes for this episode on the website at talkpython.fm/24