The True Costs of Legacy Systems: Technical Debt, Risk, and Exit Strategies
Summary
In this episode Kate Shaw, Senior Product Manager for Data and SLIM at SnapLogic, talks about the hidden and compounding costs of maintaining legacy systems—and practical strategies for modernization. She unpacks how “legacy” is less about age and more about when a system becomes a risk: blocking innovation, consuming excess IT time, and creating opportunity costs. Kate explores technical debt, vendor lock-in, lost context from employee turnover, and the slippery notion of “if it ain’t broke,” especially when data correctness and lineage are unclear. Shee digs into governance, observability, and data quality as foundations for trustworthy analytics and AI, and why exit strategies for system retirement should be planned from day one. The discussion covers composable architectures to avoid monoliths and big-bang migrations, how to bridge valuable systems into AI initiatives without lock-in, and why clear success criteria matter for AI projects. Kate shares lessons from the field on discovery, documentation gaps, parallel run strategies, and using integration as the connective tissue to unlock data for modern, cloud-native and AI-enabled use cases. She closes with guidance on planning migrations, defining measurable outcomes, ensuring lineage and compliance, and building for swap-ability so teams can evolve systems incrementally instead of living with a “bowl of spaghetti.”
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
Your host is Tobias Macey and today I'm interviewing Kate Shaw about the true costs of maintaining legacy systems
Interview
Introduction
How did you get involved in the area of data management?
What are your crtieria for when a given system or service transitions to being "legacy"?
In order for any service to survive long enough to become "legacy" it must be serving its purpose and providing value. What are the common factors that prompt teams to deprecate or migrate systems?
What are the sources of monetary cost related to maintaining legacy systems while they remain operational?
Beyond monetary cost, economics also have a concept of "opportunity cost". What are some of the ways that manifests in data teams who are maintaining or migrating from legacy systems?How does that loss of productivity impact the broader organization?
How does the process of migration contribute to issues around data accuracy, reliability, etc. as well as contributing to potential compromises of security and compliance?
Once a system has been replaced, it needs to be retired. What are some of the costs associated with removing a system from service?
What are the most interesting, innovative, or unexpected ways that you have seen teams address the costs of legacy systems and their retirement?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on legacy systems migration?
When is deprecation/migration the wrong choice?
How have evolutionary architecture patterns helped to mitigate the costs of system retirement?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
Links
SnapLogic
SLIM == SnapLogic Intelligent Modernizer
Opportunity Cost
Sunk Cost Fallacy
Data Governance
Evolutionary Architecture
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
HPC Workload Scheduling, with Ricardo Rocha
Ricardo Rocha leads the Platform Infrastructure team at CERN with a strong focus on cloud native deployments and machine learning. He has led the internal effort to transition services and workloads to use cloud native technologies, as well as dissemination and training for several years. Ricardo got CERN to join the CNCF and is a member of the Technical Oversight Committee (TOC), currently chairs the End User Technical Advisory Board (TAB), as well as leading the Research User Group (RUG).
Do you have something cool to share? Some questions? Let us know:
- web: kubernetespodcast.com
- mail: kubernetespodcast@google.com
- twitter: @kubernetespod
- bluesky: @kubernetespodcast.com
News of the week Kubernetes Blog: Image Compatibility In Cloud Native Environments
Gemini CLI on GitHub
Cloud Native Glossary — The Vietnamese Version is Live!
CNCF Blog: Joining CNCF as Executive Director: Let's Build What's Next
OpenStack Foundation
OpenInfra Foundation
Links from the interview Ricardo Rocha on LinkedIn
CERN
Infiniband
Kubernetes Jobs
HTCondor
Slurm Workload Manager
Kueue
Volcano
Kube-batch (archived)
Kubefed (archived)
Yunikorn (Unicorn)
KubeAdmiral (formerly Kubefed v2)
CNCF End User Awards - CERN
Dynamic Resource Allocation (DRA)
CNCF TAG & WG Restructure (Reboot)
Interlink
Slinky (Slurm-Kubernetes integration)
XPK: a container-native platform for HPC
Gateway API
KubeRay
KubeCon+CloudNativeCon 2022 Rolls into Detroit
It's that time of the year again, when cloud native enthusiasts and professionals assemble to discuss all things Kubernetes. KubeCon+CloudNativeCon 2023 is being held later this month in Detroit, October 24-28.
In this latest edition of The New Stack Makers podcast, we spoke with Priyanka Sharma, general manager of the Cloud Native Computing Foundation — which organizes KubeCon —and CERN computer engineer and KubeCon co-chair Ricardo Rocha. For this show, we discussed what we can expect from the upcoming event.
This year, there will be a focus on Kubernetes in the enterprise, Sharma said. "We are reaching a point where Kubernetes is becoming the de facto standard when it comes to container orchestration. And there's a reason for it. It's not just about Kubernetes. Kubernetes spawned the cloud native ecosystem and the heart of the cloud native movement is building fast, resiliently observable software that meets customer needs. So ultimately, it's making you a better provider to your customers, no matter what kind of business you are."
Of this year's topics, security will be a big theme, Rocha said. Technologies such as Falco and Cilium will be discussed. Linux kernel add-on eBPF is popping up in a lot of topics, especially around networking. Observability and hybrid deployments also weigh heavily on the agenda. "The number of solutions [around Hybrid] are quite large, so it's interesting to see what people come up with," he said.
In addition to KubeCon itself, this year there are a number of co-located events, held during or before the conference itself. Some of them hosted by CNCF while others are hosted by other companies such as Canonical. They include the Network Application Day, BackstageCon, CloudNative eBPF Day, CloudNativeSecurityCon, CloudNative WASM Day, Data-on-Kubernetes Day, EnvoyCon, gRPCConf, KNativeCon, Spinnaker Summit, Open Observability Day, Cloud Native Telco Day, Operator Day, The Continuous Delivery Summit, among others.
What's amazing is not only the number of co-located events, but the high quality of talks being held there.
"Co-located events are a great way to know what's exciting to folks in the ecosystem right now," Sharma said. "Cloud native has really become the scaffolding of future progress. People want to build on cloud native, but have their own focus areas."
WebAssembly (WASM) is a great example of this. "In the beginning, you wouldn't have thought of WebAssembly as part of the cloud native narrative, but here we are," Sharma said. "The same thinking from professionals who conceptualized cloud native in the beginning are now taking it a step further."
"There's a lot of value in co-located events, because you get a group of people for a longer period in the same room, focusing on one topic," Rocha said.
Other topics discussed in the podcast include the choice of Detroit as a conference hub, the fun activities that CNCF have planned in between the technical sessions, surprises at the keynotes, and so much more! Give it a listen.
KubeCon EU 2022, with Ricardo Rocha
Live from Valencia, it's KubeCon EU! Craig talks to conference co-chair and CERN computer scientist Ricardo Rocha about the event, and what it's like to be in a room full of people again.
Do you have something cool to share? Some questions? Let us know:
web: kubernetespodcast.com
mail: kubernetespodcast@google.com
twitter: @kubernetespod
Chatter of the week 9am Karaoke
News of the week CNCF news from KubeCon EU: SlashData survey
800 members
Boeing
Coinbase
Prometheus Certified Associate
Google Cloud improves GitOps usability with Config Sync and Porch kpt
Other Google news from KubeCon
Tetragon from Isovalent
Envoy Gateway
Infra Ask HN with the creators
Cloud Foundry launches Korifi
SUSE NeuVector is open source
CloudNativePG from EnterpriseDB All the other options
Assured Open Source Software from Google Cloud
Recent Guest news: Akuity announces $20m Series A (episode 172)
Komodor raises $42 million Series B (episode 153)
Deepfence launches Deepfence Cloud (episode 173)
Lightning Round Armory announced public early access to their new Continuous Deployment-as-a-Service product
Aserto announces its "better together" approach to authorization by bringing together OPA, OCI, and Sigstore
Bunnyshell Introduces support for multi-repository Terraform with full-stack drift management and GitOps
Calyptia announces the General Availability of Calyptia for Fluent Bit,
CAST AI introduces advanced Autoscaler for AKS
Clastix launches Kamaji, a new open source tool for Managed Kubernetes Service
CloudCasa by Catalogic expands to support Microosft AKS
Codenotary combines Community Attestation Service with background vulnerability scanning
CodeZero Launches Surf, a new developer tool for observability in pre-production Kubernetes environments
CrateDB introduces Logical Replication
D2iQ Partners with GitLab
DataCore Bolt container-native storage software now GA; built on their acquisition of Mayadata
Datadog launches Application Security Monitoring and support for OpenTelemetry Protocol in the Datadog Agent,
Deepfactor partners with Synopsys to help developers resolve cloud native supply chain security risks
env0 enables full-stack IaC deployment and management with native Kubernetes support
Era Software introduces EraStreams
Fairwinds Insights unifies DevSecOps with additional shift-left enhancements
GitLab free tier adds pull-based Kubernetes deployments
Google announced a new low-cost, high-usage pricing tier for Google Cloud Managed Service for Prometheus
HCL Technologies launches Kubernetes migration platform
Kasten by Veeam launches K10 v5.0 released
Runecast adds CI/CD integration and image scanning
Lacework introduces new Kubernetes Audit Logs monitoring
Loft Labs announces a Cluster API provider for vcluster
NetFoundry embeds zero trust into Prometheus
New Relic introduces low-overhead Kubernetes monitoring and Pixie plug-in framework
Pure Storage's new Database as a Service platform is GA
Replicated introduces community licensing and pre-flight checks
SphereEx releases DB-Plus Suite
Snapt announces security package to run Kubernetes in public cloud
SPIRE now runs on Windows
Sysdig launches new Advisor and Sysdig Open Source leverages Falco plugins
SysEleven unveils MetaKube Operator
Timescale announces OpenTelemetry Tracing support for Promscale
Vultr Kubernetes Engine now Generally Available
Zesty Disk for Kubernetes introduced
Links from the interview Episode 62 Lukas Heinrich
Clemens Lange
CERN LHC Computing Grid
Large Hadron Collider
Kubeflow
Data on Kubernetes Community
CNCF Research User Group
CNCF TOC
Volcano moves to incubation
KubeCon EU 2022
Episode 165, with Jasmine James
Selection process report for KubeCon EU
KubeCon China 2021
Research track
Puppies at KubeCon NA 2019
Code, mountains and flying
Kubernetes on an F/16
Ricardo Rocha on Twitter and on the web
Large Hadron Kubernetes at CERN, with Ricardo Rocha, Lukas Heinrich, and Clemens Lange
Back in 2012, CERN announced one of its most important achievements; the discovery of the Higgs boson. This work led to the 2013 Nobel Prize in Physics. Ricardo Rocha, Lukas Heinrich and Clemens Lang of CERN redid the data analysis on top of Kubernetes this year, which Ricardo and Lukas demonstrated at a keynote at KubeCon EU. All three join Adam and Craig for a short physics lesson and a view into computing at the largest scale, for particles at the smallest.
Do you have something cool to share? Some questions? Let us know:
web: kubernetespodcast.com
mail: kubernetespodcast@google.com
twitter: @kubernetespod
Chatter of the week 50th anniversary of the launch of Apollo 11 by NASA's Astronomy Picture of the Day, and as reported by CBS News in real time
LEGO Saturn V - mid-completion
47th annual Seafair Milk Carton Derby Adam's pictures, including the Saturn V rocket
News of the week IBM announced it has closed its acquisition of Red Hat
Hashicorp Consul 1.6
Benchmarking best practices for Istio by Megan O'Keefe, Mandar Jog and John Howard
IPv6 enhancement proposal for Kubernetes Now passing tests!
Architecting with Google Kubernetes Engine specialization
Weave Ignite
Cloud Native CI/CD with OpenShift Pipelines
k3v
Avoid time-of-measurement bias with Prometheus Prometheus client tracer for Ruby
Links from the interview CERN LHC Computing Grid
ATLAS experiment
CMS experiment
Standard model of particle physics
Cosmos: A Spacetime Odyssey, with Neil deGrasse Tyson Dark Matter is a misnomer
Baryonic matter
Dark matter
History of computing at CERN
Where the web was born
Large Hadron Collider
Higgs boson Discovery of the Higgs boson
Servicing the first web server - Tim Berners-Lee's NeXT cube
CERN Program Library (FORTRAN)
KubeCon EU keynote: Reperforming a Nobel Prize Discovery on Kubernetes Slides
YouTube video
CERN openlab partnership
ROOT Data Analysis Framework
Particle physics is embarassingly parallel Kubeflow
Spark Operator on Kubernetes
Open Data Initiative Find a Higgs boson in LHC public data
Clemens' shirt
Our guests on Twitter: Ricardo Rocha
Lukas Heinrich
Clemens Lange
ATLAS and the LHC
What does the ATLAS detector do at the LHC? We explore the detector, the LHC, and hear from Kate Shaw and Steven Goldfarb who both work with ATLAS.
Learn more about your ad-choices at https://www.iheartpodcastnetwork.comSee omnystudio.com/listener for privacy information.
#29: Python at the Large Hadron Collider and CERN
See the full show notes for this episode on the website at talkpython.fm/29