How AWS S3 is built
Brought to You By:
• Statsig — The unified platform for flags, analytics, experiments, and more.
• Sonar – The makers of SonarQube, the industry standard for automated code review
• WorkOS – Everything you need to make your app enterprise ready.
—
Amazon S3 is one of the largest distributed systems ever built, storing and serving data for a significant portion of the internet. Behind its simple interfaces hides an enormous amount of engineering work, careful tradeoffs, and long-term thinking.
In this episode, I sit down with Mai-Lan Tomsen Bukovec, VP of Data and Analytics at AWS, who has been running Amazon S3 for more than a decade. Mai-Lan shares how S3 operates at extreme scale, what it takes to design for durability and availability across millions of servers, and why building for failure is a core principle.
We also go deep into how AWS approaches correctness using formal methods, how storage tiers and limits shape system design, and why simplicity remains one of the hardest and most important goals at S3’s scale.
—
Timestamps
(00:00) Intro
(01:03) S3’s scale
(03:58) How S3 started
(07:25) Parquet, Iceberg, and S3 tables
(09:46) S3 for developers
(13:37) Why AWS keeps S3 prices low
(17:10) AWS pricing tiers
(19:38) Availability and durability
(26:21) The cost of S3's consistency
(31:22) Automated reasoning and proof of correctness
(35:14) Durability at AWS scale
(39:58) Correlated failure and crash consistency
(43:22) Failure allowances
(46:04) Two opposing principles in S3 design
(49:09) S3’s evolution
(52:21) S3 Vectors
(1:01:16) The 50 TB limit on AWS
(1:07:54) The simplicity principle
(1:10:10) Types of engineers working on S3
(1:14:15) Closing recommendations
—
The Pragmatic Engineer deepdives relevant for this episode:
• Inside Amazon’s engineering culture
• How AWS deals with a major outage
• A Day in the Life of a Senior Manager at Amazon
• What is a Principal Engineer at Amazon? – with Steve Huynh
• Working at Amazon as a software engineer – with Dave Anderson
Amazon papers recommended by Mai-Lan:
• Using lightweight formal methods to validate a key-value storage node in Amazon S3
• Formally verified cloud-scale authorization
• Analyzing metastable failures
• Amazon’s engineering tenets
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.
Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe
Amazon S3: The Backbone of Modern Data Systems
Summary
In this episode of the Data Engineering Podcast Mai-Lan Tomsen Bukovec, Vice President of Technology at AWS, talks about the evolution of Amazon S3 and its profound impact on data architecture. From her work on compute systems to leading the development and operations of S3, Mylan shares insights on how S3 has become a foundational element in modern data systems, enabling scalable and cost-effective data lakes since its launch alongside Hadoop in 2006. She discusses the architectural patterns enabled by S3, the importance of metadata in data management, and how S3's evolution has been driven by customer needs, leading to innovations like strong consistency and S3 tables.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
This is a pharmaceutical Ad for Soda Data Quality. Do you suffer from chronic dashboard distrust? Are broken pipelines and silent schema changes wreaking havoc on your analytics? You may be experiencing symptoms of Undiagnosed Data Quality Syndrome — also known as UDQS. Ask your data team about Soda. With Soda Metrics Observability, you can track the health of your KPIs and metrics across the business — automatically detecting anomalies before your CEO does. It’s 70% more accurate than industry benchmarks, and the fastest in the category, analyzing 1.1 billion rows in just 64 seconds. And with Collaborative Data Contracts, engineers and business can finally agree on what “done” looks like — so you can stop fighting over column names, and start trusting your data again.Whether you’re a data engineer, analytics lead, or just someone who cries when a dashboard flatlines, Soda may be right for you. Side effects of implementing Soda may include: Increased trust in your metrics, reduced late-night Slack emergencies, spontaneous high-fives across departments, fewer meetings and less back-and-forth with business stakeholders, and in rare cases, a newfound love of data. Sign up today to get a chance to win a $1000+ custom mechanical keyboard. Visit dataengineeringpodcast.com/soda to sign up and follow Soda’s launch week. It starts June 9th.
Your host is Tobias Macey and today I'm interviewing Mai-Lan Tomsen Bukovec about the evolutions of S3 and how it has transformed data architecture
Interview
Introduction
How did you get involved in the area of data management?
Most everyone listening knows what S3 is, but can you start by giving a quick summary of what roles it plays in the data ecosystem?
What are the major generational epochs in S3, with a particular focus on analytical/ML data systems?The first major driver of analytical usage for S3 was the Hadoop ecosystem. What are the other elements of the data ecosystem that helped shape the product direction of S3?
Data storage and retrieval have been core primitives in computing since its inception. What are the characteristics of S3 and all of its copycats that led to such a difference in architectural patterns vs. other shared data technologies? (e.g. NFS, Gluster, Ceph, Samba, etc.)
How does the unified pool of storage that is exemplified by S3 help to blur the boundaries between application data, analytical data, and ML/AI data?
What are some of the default patterns for storage and retrieval across those three buckets that can lead to anti-patterns which add friction when trying to unify those use cases?
The age of AI is leading to a massive potential for unlocking unstructured data, for which S3 has been a massive dumping ground over the years. How is that changing the ways that your customers think about the value of the assets that they have been hoarding for so long?What new architectural patterns is that generating?
What are the most interesting, innovative, or unexpected ways that you have seen S3 used for analytical/ML/Ai applications?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on S3?
When is S3 the wrong choice?
What do you have planned for the future of S3?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
Links
AWS S3
Kinesis
Kafka
SQS
EMR
Drupal
Wordpress
Netflix Blog on S3 as a Source of Truth
Hadoop
MapReduce
Nasa JPL
FINRA == Financial Industry Regulatory Authority
S3 Object Versioning
S3 Cross Region
S3 Tables
Iceberg
Parquet
AWS KMS
Iceberg REST
DuckDB
NFS == Network File System
Samba
GlusterFS
Ceph
MinIO
S3 Metadata
Photoshop Generative Fill
Adobe Firefly
Turbotax AI Assistant
AWS Access Analyzer
Data Products
S3 Access Point
AWS Nova Models
LexisNexis Protege
S3 Intelligent Tiering
S3 Principal Engineering Tenets
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Being Present in the Moment Through Balcony-Hopping with Mai-Lan Tomsen Bukovec
Mai-Lan Tomsen Bukovec, Vice President of Foundational Data Services at AWS, joins Corey on Screaming in the Cloud to discuss her technique for spending time intentionally and prioritizing work-life balance called balcony-hopping. Mai-Lan explains how she created the concept of balcony-hopping and how it has helped her to be a better leader, mother, wife, and boxer. Corey and Mai-Lan discuss how in today’s age, attention is a form of currency and why it’s so important to be intentional with how and where you spend your attention. Mai-Lan also offers practical insights to anyone seeking to feel more productive, present, and balanced.
About Mai-Lan
Mai-Lan Tomsen Bukovec is Vice President, Foundational Data Services (FDS) at Amazon Web Services (AWS) and leads a number of high-scale AWS cloud services that provide storage and streaming of petabytes or exabytes of data and essential building blocks for modern application architecture like queuing and notifications, monitoring, alarming, logging and reliability validation. Mai-Lan’s teams include some of AWS’ first and largest-scale services like Amazon S3 and Simple Queue Service (SQS) to more recent and fast-growing services like managed open source streaming (Amazon Managed Streaming for Apache Kafka).
Prior to joining Amazon, Mai-Lan spent almost 15 years in engineering and product leadership roles at technology companies including Microsoft and early stage startups. She began her technology career after serving in the U.S. Peace Corps in the Mopti region of Africa as a Forestry volunteer after earning her degree from University of California, San Diego.
At Amazon, Mai-Lan is an advisor to Asians@Amazon, creator and sponsor of internal leadership development programs for Amazon employees, and is passionate about AWS initiatives and cloud services that maximize human potential everywhere.
Mai-Lan has three children and lives in Seattle with her family. When she is not working on Amazon cloud services and spending time with her husband and kids, Mai-Lan trains primarily in boxing with additional practice in the martial art Savate.
Links Referenced:
LinkedIn post “Live Your Best Life Through Balcony Hopping”: https://www.linkedin.com/pulse/live-your-best-life-through-balcony-hopping-mai-lan-tomsen-bukovec/
LinkedIn: https://www.linkedin.com/in/mailan/
Episode 46: Don't Be Afraid of the Bold Ask
If you’re looking for older services at AWS, there really aren’t any. For example, Simple Storage Service (S3) has been with us since the beginning. It was the first publicly launched service that was quickly followed by Simple Queue Service (SQS). Still today, when it comes to these services, simplicity is key!
Today, we’re talking to Mai-Lan Tomsen Bukovec, vice president of S3 at AWS. Many people use S3 the same way that they have for years, such as for backups in the Cloud. However, others have taken S3 and ran with it to find a myriad of different use cases.
Some of the highlights of the show include:
Data: Where do I put it? What do I do with it?
S3 Select and Cross-Region Replication (CRR) make it easier and cheaper to use and manage data
Customer feedback drives AWS S3 price options and tiers
Using Glacier and S3 together for archive data storage; decisions and constraints that affect people’s use and storage of data
Feature requests should meet customers where they are, rather than having to invest in time and training
Different design patterns and best practices to use when building applications
Batch operations make it easier for customers to manage objects stored in S3
AWS considers compliance and retention when building features
Mentorship: Don’t be afraid of the bold ask
Links:
re:Invent
AWS S3
Amazon SQS
AWS Glacier
Lambda
CHAOSSEARCH
.