Designing Data-intensive Applications with Martin Kleppmann
Brought to You By:
• Statsig — The unified platform for flags, analytics, experiments, and more.
• Sonar – The makers of SonarQube, the industry standard for automated code review
• WorkOS – Everything you need to make your app enterprise ready.
—
Martin Kleppmann is a researcher and the author of Designing Data-Intensive Applications, one of the most influential books on modern distributed systems. As of this month, the second, heavily updated edition of the book is out.
In this episode of Pragmatic Engineer, we discuss Martin’s career in tech building startups, how he ended up writing this iconic book, and what he’s focused on now after moving into academia.
We talk about the tradeoffs behind modern infrastructure, how the cloud has changed what it means to scale, and the thinking behind Designing Data-Intensive Applications, including what’s changing in the second edition.
Martin reflects on lessons from building startups like Rapportive, which he sold to LinkedIn, and shares how his experience in both academia and industry shaped his perspective.
We also explore what’s ahead: why formal verification may become more important in an AI-assisted world, the challenges of building local-first software, and his recent research into using cryptography to improve transparency in supply chains without exposing sensitive data.
—
Timestamps
(00:00) Early career
(05:46) Building Rapportive
(10:47) Working at LinkedIn
(14:09) Writing Designing Data-Intensive Applications
(23:00) Reliability, scalability, and repeatability
(26:24) DDIA: the second edition
(30:50) Tradeoffs of using cloud services
(39:02) How the cloud changed scaling
(42:53) The trouble with distributed systems
(49:02) Ethics for software engineers
(52:45) Formal verification
(1:00:12) Academia vs. industry
(1:03:50) Local-first software
(1:09:50) Computer science education
(1:18:32) Martin’s current research and advice
—
The Pragmatic Engineer deepdives relevant for this episode:
• Building Bluesky: a distributed social network
• Inside Uber’s move to the cloud
• The history of servers, the cloud, and what’s next
• The past and future of modern backend practices
• How Kubernetes is built
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.
Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe
The ethics of using AI to immortalize the dead
There's an emerging industry that uses artificial intelligence to create simulations of people who've died. These post mortem avatars are also called griefbots.
Some critics, including Tomasz Hollanek, a researcher at the University of Cambridge, say this practice raises a number of ethical issues. He walks us through the mechanics of how this technology works, and how it may or may not be used responsibly.
The Story: Embracing AI and All Life’s Uncertainties w/ David Spiegelhalter
Statistician David Spiegelhalter is no stranger to AI – he used it to help him research his recent book and, back in the late 70s, he helped develop foundational algorithms for the tech. So, he understands the pandora’s box that technology can represent, as well as the uncertainty embedded in its future development. Spiegelhalter sits down with Oz to unpack how we should interpret AI predictions, why better data matters and why we should consciously embrace uncertainty in our own lives.
See omnystudio.com/listener for privacy information.
Seeing beyond the scan in neuroimaging
In this episode, we explore the intersection of AI, machine learning, and healthcare through the lens of neuroimaging and epilepsy diagnosis. Dr. Gavin Winston shares insights from his work using MRI data and machine learning to uncover subtle abnormalities in brain function. We discuss the cultural and ethical barriers to AI adoption in medicine, how predictive data analysis could transform the diagnostic workflow, and what the future holds for medical imaging in a world increasingly shaped by intelligent systems.
Featuring:
Gavin Winston – LinkedIn, Website
Chris Benson – Website, GitHub, LinkedIn, X
Daniel Whitenack – Website, GitHub, X
Links:
Detection of Epileptogenic Focal Cortical Dysplasia Using Graph Neural Networks: A MELD Study
Machine Learning in Neuroimaging across Disciplines
Automated and Interpretable Detection of Hippocampal Sclerosis in Temporal Lobe Epilepsy: AID-HS
Literature review and protocol for a prospective multicentre cohort study on multimodal prediction of seizure recurrence after unprovoked first seizure
Deep learning in neuroimaging of epilepsy
Non-parametric combination of multimodal MRI for lesion detection in focal epilepsy
Detection of covert lesions in focal epilepsy using computational analysis of multimodal magnetic resonance imaging data
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought with Chengzu Li - #722
Today, we're joined by Chengzu Li, PhD student at the University of Cambridge to discuss his recent paper, “Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.” We explore the motivations behind MVoT, its connection to prior work like TopViewRS, and its relation to cognitive science principles such as dual coding theory. We dig into the MVoT framework along with its various task environments—maze, mini-behavior, and frozen lake. We explore token discrepancy loss, a technique designed to align language and visual embeddings, ensuring accurate and meaningful visual representations. Additionally, we cover the data collection and training process, reasoning over relative spatial relations between different entities, and dynamic spatial reasoning. Lastly, Chengzu shares insights from experiments with MVoT, focusing on the lessons learned and the potential for applying these models in real-world scenarios like robotics and architectural design.
The complete show notes for this episode can be found at https://twimlai.com/go/722.
#062 - Dr. Guy Emerson - Linguistics, Distributional Semantics
Dr. Guy Emerson is a computational linguist and obtained his Ph.D from Cambridge university where he is now a research fellow and lecturer. On panel we also have myself, Dr. Tim Scarfe, as well as Dr. Keith Duggar and the veritable Dr. Walid Saba. We dive into distributional semantics, probability theory, fuzzy logic, grounding, vagueness and the grammar/cognition connection.
The aim of distributional semantics is to design computational techniques that can automatically learn the meanings of words from a body of text. The twin challenges are: how do we represent meaning, and how do we learn these representations? We want to learn the meanings of words from a corpus by exploiting the fact that the context of a word tells us something about its meaning. This is known as the distributional hypothesis. In his Ph.D thesis, Dr. Guy Emerson presented a distributional model which can learn truth-conditional semantics which are grounded by objects in the real world.
Hope you enjoy the show!
https://www.cai.cam.ac.uk/people/dr-guy-emerson
https://www.repository.cam.ac.uk/handle/1810/284882?show=full
https://www.semanticscholar.org/paper/Computational-linguistics-and-grammar-engineering-Bender-Emerson/bbd6f3b92a0f1ea8212f383cc4719bfe86b3588c
Patreon: https://www.patreon.com/mlst
Applications of Variational Autoencoders and Bayesian Optimization with José Miguel Hernández Lobato - #510
Today we’re joined by José Miguel Hernández-Lobato, a university lecturer in machine learning at the University of Cambridge. In our conversation with Miguel, we explore his work at the intersection of Bayesian learning and deep learning. We discuss how he’s been applying this to the field of molecular design and discovery via two different methods, with one paper searching for possible chemical reactions, and the other doing the same, but in 3D and in 3D space. We also discuss the challenges of sample efficiency, creating objective functions, and how those manifest themselves in these experiments, and how he integrated the Bayesian approach to RL problems. We also talk through a handful of other papers that Miguel has presented at recent conferences, which are all linked at twimlai.com/go/510.
The Physics of Data with Alpha Lee - #377
Today we’re joined by Alpha Lee, Winton Advanced Fellow in the Department of Physics at the University of Cambridge. Our conversation centers around Alpha’s research which can be broken down into three main categories: data-driven drug discovery, material discovery, and physical analysis of machine learning. We discuss the similarities and differences between drug discovery and material science, his startup, PostEra which offers medicinal chemistry as a service powered by machine learning, and much more
#92 – Harry Cliff: Particle Physics and the Large Hadron Collider
Harry Cliff is a particle physicist at the University of Cambridge working on the Large Hadron Collider beauty experiment that specializes in searching for hints of new particles and forces by studying a type of particle called the “beauty quark”, or “b quark”. In this way, he is part of the group of physicists who are searching answers to some of the biggest questions in modern physics. He is also an exceptional communicator of science with some of the clearest and most captivating explanations of basic concepts in particle physics I’ve ever heard.
Support this podcast by signing up with these sponsors:
– ExpressVPN at https://www.expressvpn.com/lexpod
– Cash App – use code “LexPodcast” and download:
– Cash App (App Store): https://apple.co/2sPrUHe
– Cash App (Google Play): https://bit.ly/2MlvP5w
EPISODE LINKS:
Harry’s Website: https://www.harrycliff.co.uk/
Harry’s Twitter: https://twitter.com/harryvcliff
Beyond the Higgs Lecture: https://www.youtube.com/watch?v=edvdzh9Pggg
Harry’s stand-up: https://www.youtube.com/watch?v=dnediKM_Sts
This conversation is part of the Artificial Intelligence podcast. If you would like to get more information about this podcast go to https://lexfridman.com/ai or connect with @lexfridman on Twitter, LinkedIn, Facebook, Medium, or YouTube where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on Apple Podcasts, follow on Spotify, or support it on Patreon.
Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.
OUTLINE:
00:00 – Introduction
03:51 – LHC and particle physics
13:55 – History of particle physics
38:59 – Higgs particle
57:55 – Unknowns yet to be discovered
59:48 – Beauty quarks
1:07:38 – Matter and antimatter
1:10:22 – Human side of the Large Hadron Collider
1:17:27 – Future of large particle colliders
1:24:09 – Data science with particle physics
1:27:17 – Science communication
1:33:36 – Most beautiful idea in physics
Making Algorithms Trustworthy with David Spiegelhalter - TWiML Talk #212
Today we’re joined by David Spiegelhalter, Chair of Winton Center for Risk and Evidence Communication at Cambridge University and President of the Royal Statistical Society. David, an invited speaker at NeurIPS, presented on “Making Algorithms Trustworthy: What Can Statistical Science Contribute to Transparency, Explanation and Validation?”. In our conversation, we explore the nuanced difference between being trusted and being trustworthy, and its implications for those building AI systems.