Observability and human intuition in an AI world
In this two for one episode recorded at HumanX, Ryan is first joined by Christine Yen, CEO of Honeycomb, to discuss how AI compresses the software development lifecycle, making observability about capturing the right telemetry. Then, Spiros Xanthos, founder and CEO of Resolve AI, shares with us how AI coding increases code volume but decreases human intuition, making production operations harder than ever.
Episode notes:
Honeycomb is an observability platform that enables deep, high-dimensional exploration so you can debug unpredictable behavior with precision.
Resolve AI allows you to resolve incidents, optimize costs, and code with production context using AI that works across your code, infrastructure, and telemetry.
Connect with Christine on LinkedIn.
Connect with Spiros on LinkedIn.
TRANSCRIPT
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Building Systems That Work Even When Everything Breaks with Ben Hartshorne
When AWS has a major outage, what actually happens behind the scenes? Ben Hartshorne, a principal engineer at Honeycomb, joins Corey Quinn to discuss a recent AWS outage and how they kept customer data safe even when their systems couldn't fully work. Ben explains why building services that expect things to break is the only way to survive these outages. Ben also shares how Honeycomb used its own tools to cut their AWS Lambda costs in half by tracking five different things in a spreadsheet and making small changes to all of them.
About Ben Hartshorne:
Ben has spent much of his career setting up monitoring systems for startups and now is thrilled to help the industry see a better way. He is always eager to find the right graph to understand a service and will look for every excuse to include a whiteboard in the discussion.
Show highlights:
(02:41)Two Stories About Cost Optimization
(04:20) Cutting Lambda Costs by 50%
(08:01) Surviving the AWS Outage
(09:20) Preserving Customer Data During the Outage
(13:08) Should You Leave AWS After an Outage?
(15:09) Multi-Region Costs 10x More
(18:10) Vendor Dependencies
(22:06) How LaunchDarkly's SDK Handles Outages
(24:40) Rate Limiting Yourself
(29:00) How Much Instrumentation Is Too Much?
(34:28) Where to Find Ben
Links:
Linkedin: https://www.linkedin.com/in/benhartshorne/
GitHub: https://github.com/maplebed
Sponsored by:
duckbillhq.com
What’s Driving the Rising Cost of Observability?
Observability is expensive because traditional tools weren’t designed for the complexity and scale of modern cloud-native systems, explains Christine Yen, CEO of Honeycomb.io. Logging tools, while flexible, were optimized for manual, human-scale data reading. This approach struggles with the massive scale of today’s software, making logging slow and resource-intensive. Monitoring tools, with their dashboards and metrics, prioritized speed over flexibility, which doesn’t align with the dynamic nature of containerized microservices. Similarly, traditional APM tools relied on “magical” setups tailored for consistent application environments like Rails, but they falter in modern polyglot infrastructures with diverse frameworks.
Additionally, observability costs are rising due to evolving demands from DevOps, platform engineering, and site reliability engineering (SRE). Practices like service-level objectives (SLOs) emphasize end-user experience, pushing teams to track meaningful metrics. However, outdated observability tools often hinder this, forcing teams to cut back on crucial data. Yen highlights the potential of AI and innovations like OpenTelemetry to address these challenges.
Learn more from The New Stack about the latest trends in observability:
Honeycomb.io’s Austin Parker: OpenTelemetry In-Depth
Observability in 2025: OpenTelemetry and AI to Fill In Gaps
Observability and AI: New Connections at KubeCon
Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Observability: the present and future, with Charity Majors
Supported by Our Partners
• Sonar — Trust your developers – verify your AI-generated code.
• Vanta —Automate compliance and simplify security with Vanta.
—
In today's episode of The Pragmatic Engineer, I'm joined by Charity Majors, a well-known observability expert – as well as someone with strong and grounded opinions. Charity is the co-author of "Observability Engineering" and brings extensive experience as an operations and database engineer and an engineering manager. She is the cofounder and CTO of observability scaleup Honeycomb.
Our conversation explores the ever-changing world of observability, covering these topics:
• What is observability? Charity’s take
• What is “Observability 2.0?”
• Why Charity is a fan of platform teams
• Why DevOps is an overloaded term: and probably no longer relevant
• What is cardinality? And why does it impact the cost of observability so much?
• How OpenTelemetry solves for vendor lock-in
• Why Honeycomb wrote its own database
• Why having good observability should be a prerequisite to adding AI code or using AI agents
• And more!
—
Timestamps
(00:00) Intro
(04:20) Charity’s inspiration for writing Observability Engineering
(08:20) An overview of Scuba at Facebook
(09:16) A software engineer’s definition of observability
(13:15) Observability basics
(15:10) The three pillars model
(17:09) Observability 2.0 and the shift to unified storage
(22:50) Who owns observability and the advantage of platform teams
(25:05) Why DevOps is becoming unnecessary
(27:01) The difficulty of observability
(29:01) Why observability is so expensive
(30:49) An explanation of cardinality and its impact on cost
(34:26) How to manage cost with tools that use structured data
(38:35) The common worry of vendor lock-in
(40:01) An explanation of OpenTelemetry
(43:45) What developers get wrong about observability
(45:40) A case for using SLOs and how they help you avoid micromanagement
(48:25) Why Honeycomb had to write their database
(51:56) Companies who have thrived despite ignoring conventional wisdom
(53:35) Observability and AI
(59:20) Vendors vs. open source
(1:00:45) What metrics are good for
(1:02:31) RUM (Real User Monitoring)
(1:03:40) The challenges of mobile observability
(1:05:51) When to implement observability at your startup
(1:07:49) Rapid fire round
—
The Pragmatic Engineer deepdives relevant for this episode:
• How Uber Built its Observability Platform https://newsletter.pragmaticengineer.com/p/how-uber-built-its-observability-platform
• Building an Observability Startup https://newsletter.pragmaticengineer.com/p/chronosphere
• How to debug large distributed systems https://newsletter.pragmaticengineer.com/p/antithesis
• Shipping to production https://newsletter.pragmaticengineer.com/p/shipping-to-production
—
See the transcript and other references from the episode at https://newsletter.pragmaticengineer.com/podcast
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.
Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe
Observability & Engineering Management, with Charity Majors
Charity Majors is the co-founder and CTO of honeycomb.io. She pioneered the concept of modern Observability, drawing on her years of experience building and managing massive distributed systems at Parse (acquired by Facebook), then subsequently at Facebook, and at Linden Lab building Second Life. She is the co-author of Observability Engineering and Database Reliability Engineering (O'Reilly). She loves free speech, free software and single malt scotch.
Do you have something cool to share? Some questions? Let us know:
- web: kubernetespodcast.com
- mail: kubernetespodcast@google.com
- twitter: @kubernetespod
News of the week CNCF Blog: Vitess 20 is now Generally Available
Vitess Blog: Announcing Vitess 20
Anthropic Blog: Claude 3.5 Sonnet
KubeCon India 2024 CFP
Apps on Azure Blog: Announcing support of OCI v1.1 specification in Azure Container Registry
VMware Tanzu Blog: Announcing VMware Tanzu Greenplum 7.2: Powering Your Business with Enhanced Performance and Advanced Capabilities
VMware Tanzu Blog: Join the public beta for GenAI on Tanzu Platform today!
CNCF: Adobe End User Journey Report
Links from the interview Honeycomb.io
O'Reilly Book: Observability Engineering
O'Reilly Book: Database Reliability Engineering
Charity's blog site: charity.wtf
Charity Blog: Questionable Advice: "My boss says we don't need any engineering managers. Is he right?"
Daniel H. Pink book: "Drive: The Surprising Truth About What Motivates Us"
In which, "He examines the three elements of true motivation—autonomy, mastery, and purpose-and offers smart and surprising techniques for putting these into action in a unique book that will change how we think and transform how we live."
Charity blog on Stack Overflow: "Generative AI is not going to build your engineering team for you"
In which she talks about how the tech industry is an apprenticeship industry.
Charity Majors in the Google Cloud Next 2024 Developer Keynote
honeycomb.io blog: "How Time Series Databases Work—And Where They Don't" by Alex Vondrak
honeycomb.io blog: "Why Observability Requires a Distributed Column Store" by Alex Vondrak
Links from the post-interview chat CNCF Kubernetes Community Days (KCDs)
CNCF Kubernetes Community Days (KCDs) on GitHub
Julia Evans Blog
Wizard Zines by Julia Evans
"Help! I Have a Manager!" zine by Julia Evans
Aja Hammerly aka "thagomizer" blog
"The Toaster Parable"
"Manager Toolkit: Manage The Person In Front Of You"
"Manager Toolkit: Useful Manager Phrases for 1:1s"
"Manager Toolkit: You Talk, I Type"
Shifting from Observability 1.0 to 2.0 with Charity Majors
This week on Screaming in the Cloud, Corey is joined by good friend and colleague, Charity Majors. Charity is the CTO and Co-founder of Honeycomb.io, the widely popular observability platform. Corey and Charity discuss the ins and outs of observability 1.0 vs. 2.0, why you should never underestimate the power of software to get worse over time, and the hidden costs of observability that could be plaguing your monthly bill right now. The pair also shares secrets on why speeches get better the more you give them and the basic role they hope AI plays in the future of computing. Check it out!
Show Highlights:
(00:00 - Reuniting with Charity Majors: A Warm Welcome
(03:47) - Navigating the Observability Landscape: From 1.0 to 2.0
(04:19) - The Evolution of Observability and Its Impact
(05:46) - The Technical and Cultural Shift to Observability 2.0
(10:34) - The Log Dilemma: Balancing Cost and Utility
(15:21) - The Cost Crisis in Observability
(22:39) - The Future of Observability and AI's Role
(26:41) - The Challenge of Modern Observability Tools
(29:05) - Simplifying Observability for the Modern Developer
(30:42) - Final Thoughts and Where to Find More
About Charity
Charity is an ops engineer and accidental startup founder at honeycomb.io. Before this she worked at Parse, Facebook, and Linden Lab on infrastructure and developer tools, and always seemed to wind up running the databases. She is the co-author of O'Reilly's Database Reliability Engineering, and loves free speech, free software, and single malt scotch.
Links:
https://charity.wtf/
Honeycomb Blog: https://www.honeycomb.io/blog
Twitter: @mipsytipsy
Building a Strong Company Culture at Honeycomb with Mike Goldsmith
Mike Goldsmith, Staff Software Engineer at Honeycomb, joins Corey on Screaming in the Cloud to talk about Open Telemetry, company culture, and the pros and cons of Go vs. .NET. Corey and Mike discuss why OTel is such an important tool, while pointing out its double-edged sword of being fully open-source and community-driven. Opening up about Honeycomb’s company culture and how to find a work-life balance as a fully-remote employee, Mike points out how core-values and social interaction breathe life into a company like Honeycomb.
About Mike
Mike is an OpenSource focused software engineer that builds tools to help users create, shape and deliver system & application telemetry. Mike contributes to a number of OpenTelemetry initiatives including being a maintainer for Go Auto instrumentation agent, Go proto packages and an emeritus .NET SDK maintainer..
Links Referenced:
Honeycomb: https://www.honeycomb.io/
Twitter: https://twitter.com/Mike_Goldsmith
Honeycomb blog: https://www.honeycomb.io/blog
LinkedIn: https://www.linkedin.com/in/mikegoldsmith/
Honeycomb on Observability as Developer Self-Care with Brooke Sargent
Brooke Sargent, Software Engineer at Honeycomb, joins Corey on Screaming in the Cloud to discuss how she fell into the world of observability by adopting Honeycomb. Brooke explains how observability was new to her in her former role, but she quickly found it to enable faster learning and even a form of self care for herself as a developer. Corey and Brooke discuss the differences of working at a large company where observability is a new idea, versus an observability company like Honeycomb. Brooke also reveals the importance of helping people reach a personal understanding of what observability can do for them when trying to introduce it to a company for the first time.
About Brooke
Brooke Sargent is a Software Engineer at Honeycomb, working on APIs and integrations in the developer ecosystem. She previously worked on IoT devices at Procter and Gamble in both engineering and engineering management roles, which is where she discovered an interest in observability and the impact it can have on engineering teams.
Links Referenced:
Honeycomb: https://www.honeycomb.io/
Twitter: https://twitter.com/codegirlbrooke
The Need for Reliability with Lex Neva
Lex Neva, Staff Site Reliability Engineer at Honeycomb and Curator of SRE Weekly, joins Corey on Screaming in the Cloud to discuss reliability and the life of a newsletter curator. Lex shares some interesting insights on how he keeps his hobbies and side projects separate, as well as the intrusion that open-source projects can have on your time. Lex and Corey also discuss the phenomenon of newsletter curators being much more demanding of themselves than their audience typically is. Lex also shares his views on how far reliability has come, as well as how far we have to go, and the critical implications reliability has on our day-to-day lives.
About Lex
Lex Neva is interested in all things related to running large, massively multiuser online services. He has years of SRE, Systems Engineering, tinkering, and troubleshooting experience and perhaps loves incident response more than he ought to. He’s previously worked for Linden Lab, DeviantArt, Heroku, and Fastly, and currently works as an SRE at Honeycomb while also curating the SRE Weekly newsletter on the side.
Lex lives in Massachusetts with his family including 3 adorable children, 3 ridiculous cats, and assorted other awesome humans and animals. In his copious spare time he likes to garden, play tournament poker, tinker with machine embroidery, and mess around with Arduinos.
Links Referenced:
SRE Weekly: https://sreweekly.com/
Honeycomb: https://www.honeycomb.io/
Charity Majors: Taking an Outsider's Approach to a Startup
In the early 2000s, Charity Majors was a homeschooled kid who’d gotten a scholarship to study classical piano performance at the University of Idaho.
“I realized, over the course of that first year, that music majors tended to still be hanging around the music department in their 30s and 40s,” she said. “And nobody really had very much money, and they were all doing it for the love of the game. And I was just like, I don't want to be poor for the rest of my life.”
Fortunately, she said, it was pretty easy at that time to jump into the much more lucrative tech world. “It was buzzing, they were willing to take anyone who knew what Unix was,” she said of her first tech job, running computer systems for the university.
Eventually, she dropped out of college, she said, “made my way to Silicon Valley, and I’ve been here ever since.”
Majors, co-founder and chief technology officer of the six-year-old Honeycomb.io, an observability platform company, told her story for The New Stack’s podcast series, The Tech Founder Odyssey, which spotlights the personal journeys of some of the most interesting technical startup creators in the cloud native industry.
It’s been a busy year for her and the company she co-founded with Christine Yen, a colleague from Parse, a mobile application development company that was bought by Facebook. In May, O’Reilly published “Observability Engineering,” which Majors co-wrote with George Miranda and Liz Fong-Jones. In June, Gartner named Honeycomb.io as a Leader in the Magic Quadrant for Application Performance Monitoring and Observability.
Thus far Honeycomb.io, now employing about 200 people, has raised just under $97 million, including a $50 million Series C funding round it closed in October, led by Insight Partners (which owns The New Stack).
This Tech Founder Odyssey conversation was co-hosted by Colleen Coll and Heather Joslyn of TNS.
‘Rage-Driven Development’
Honeycomb.io grew from efforts at Parse to solve a stubborn observability problem: systems crashed frequently, and rarely for the same reasons each time. “We invested a lot in the last generation of monitoring technology, we had all these dashboards, we have all these graphs,” Majors said. “But in order to figure out what's going on, you kind of had to know in advance what was going to break.”
Once Parse was acquired by Facebook, Majors, Yen and their teams began piping data into a Facebook tool called Scuba, which ”was aggressively hostile to users,” she recalled.
But, “it did one thing really well, which is let you slice and dice in real time on dimensions that have very high cardinality,” meaning those that contain lots of unique terms. This set it apart from the then-current monitoring technologies, which were built around assessing low cardinality dimensions.
Scuba allowed Majors’ organization to gain more control over its reliability problem. And it got her and Yen thinking about how a platform tool that could analyze high cardinality data about system health in real time. “Everything is a high cardinality dimension now,” Majors said. “And [with] the old generation of tools, you hit a wall really fast and really hard.”
And so, Honeycomb.io was created to build that platform. “My entire career has been rage-driven development,” she said. “Like: sounds cool, I'm gonna go play with that. This isn't working — I'm gonna go fix it from anger.”
A Reluctant CEO
Yen now holds the CEO role at Honeycomb.io, but Majors wound up with the job for roughly the first half of the company’s life.
Did Majors like being the boss? “Hated it,” she said. “Constitutionally what you want in a CEO is someone who is reliable, predictable, dependable, someone who doesn't mind showing up every Tuesday at 10:30 to talk to the same people.
“I am not structured. I really chafe against that stuff.”
However, she acknowledged, she may have been the right leader in the startup’s beginning: “It was a state of chaos, like we didn't think we were going to survive. And that's where I thrive.”
Fortunately, in Honeycomb.io’s early days, raising money wasn’t a huge challenge, due to its founders’ background at Facebook. “There were people who were coming to us, like, do you want $2 million for a seed thing? Which is good, because I've seen the slides that we put together, and they are laughable. If I had seen those slides as an investor, I would have run the other way.”
The “pedigree” conferred on her by investors due to her association with Facebook didn’t sit comfortably with her. “I really hated it,” she said. “Because I did not learn to be a better engineer at Facebook. And part of me kind of wanted to just reject it. But I also felt this like responsibility on behalf of all dropouts, and queer women everywhere, to take the money and do something with it. So that worked out.”
Majors, a frequent speaker at tech conferences, has established herself as a thought leader in not only observability but also engineering management. For other women, people of color, or people in the tech field with an unconventional story, she advised “investing a little bit in your public speaking skills, and making yourself a bit of a profile. Being externally known for what you do is really helpful because it counterbalances the default assumptions that you're not technical or that you're not as good.”
She added, “if someone can Google your name plus a technology, and something comes up, you're assumed to be an expert. And I think that that really works to people's advantage.“
Majors had a lot more to say about how her outsider perspective has shaped the way she approaches hiring, leadership and scaling up her organization. Check out this latest episode of the Tech Founder Odyssey.
ONE MORE thing every dev should know (Interview)
The incomparable Jessica Kerr is back with another grab-bag of amazing topics. We talk about her journey to Honeycomb, devs getting satisfaction from the code they write, why step one for her is “get that new project into production” and step two is observe it, her angst for the context switching around pull requests, some awesome book recommendations, how game theory and design can translate to how we skill up and level up our teams, and so much more.
Join the discussion
Changelog++ members save 5 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
InfluxData – The time series platform for building and operating time series applications — InfluxDB empowers developers to build IoT, analytics, and monitoring software. It’s purpose-built to handle massive volumes and countless sources of time-stamped data produced by sensors, applications, and infrastructure. Learn more at influxdata.com/changelog
Sentry – Working code means happy customers. That’s exactly why teams choose Sentry. From error tracking to performance monitoring, Sentry helps teams see what actually matters, resolve problems quicker, and learn continuously about their applications - from the frontend to the backend. Use the code CHANGELOG and get the team plan free for three months.
Retool – The low-code platform for developers to build internal tools — Some of the best teams out there trust Retool…Brex, Coinbase, Plaid, Doordash, LegalGenius, Amazon, Allbirds, Peloton, and so many more – the developers at these teams trust Retool as the platform to build their internal tools. Try it free at retool.com/changelog
MongoDB – An integrated suite of cloud database and services — They have a FREE forever tier, so you can prove to yourself and to your team that they have everything you need. Check it out today at mongodb.com/changelog
Featuring:
Jessica Kerr – Website, GitHub, X
Adam Stacoviak – Website, GitHub, LinkedIn, Mastodon, X
Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
The ONE thing every dev should know (and other words of wisdom) with Jessica Kerr
Graceful.Dev - the successor to Ruby Tapas
Games: Agency As Art
Trespassing on Einstein’s Lawn
Shigeru Miyamoto on Wikipedia
Better coordination, or better software?
Something missing or broken? PRs welcome!
Best Practices Don’t Exist with Paul Osman
About Paul Osman
Paul Osman is a Software Engineer with 20 years of experience in the industry. He's the Lead Instrumentation Engineer at Honeycomb.io and is passionate about making production a less scary word. Having spent most of his career in the ill-defined space between software development and operations, Paul spends a lot of time thinking about making on-call experiences better, responding to and learning from incidents, and improving ways for software engineers to share knowledge. Before joining Honeycomb.io, Paul worked in Platform and SRE teams at Under Armour, PagerDuty, and SoundCloud.
Links Referenced:
Honeycomb.io
Follow Paul on Twitter
Paul’s Blog
Making Outages Boring with Danyel Fisher
About Danyel Fisher
Danyel Fisher is a Principal Design Researcher for Honeycomb.io. He focuses his passion for data visualization on helping SREs understand their complex systems quickly and clearly. Before he started at Honeycomb, he spent thirteen years at Microsoft Research, studying ways to help people gain insights faster from big data analytics.
Links Referenced:
Danyel’s section on Honeycomb’s website
Danyel’s Personal Site
Follow Danyel on Twitter
The Era of Virtual Events with Shelby Spees
About Shelby Spees
Shelby Spees has been developing software professionally since 2015 in a range of domains, which has made her appreciate the importance of learning how to learn and creating support systems for lifelong skill development. When she’s not helping teams level up their observability practice, you can find her at home playing on her Switch or singing karaoke with her rescue pitbull Nova.
Links Referenced
Follow Shelby on Twitter
Connect with Shelby on LinkedIn
Shelby’s Personal Site
Email Shelby directly at shelby@hey.com
DevOpsy Security with Jam Leomi
About Jam Leomi
Jam Leomi is a penmaker who just so happens to computer. When not found ranting on equality and equity in #infosec and beyond on twitter, they're found doing their day job as Lead Security Engineer at Honeycomb.
Links Referenced
Honeycomb
Jam's Personal Blog
Follow Jam on Twitter
Connect with Jam on LinkedIn
Managing Humans with Charity Majors
Links Referenced:
Honeycomb: https://www.honeycomb.io/
Personal Blog: https://charity.wtf/
Honeycomb Blog: https://www.honeycomb.io/blog
Building Ethical Tech Companies with Liz Fong-Jones
About Liz Fong-Jones
Liz is a developer advocate, labor and ethics organizer, and Site Reliability Engineer (SRE) with 16+ years of experience. She is an advocate at Honeycomb for the SRE and Observability communities, and previously was an SRE working on products ranging from the Google Cloud Load Balancer to Google Flights.
Links Referenced
Company Site: https://www.honeycomb.io/
The Duckbill Group: https://www.duckbillgroup.com/
Honeycomb Liz: https://www.honeycomb.io/liz
Personal site: https://www.lizthegrey.com/
Twitter: https://twitter.com/lizthegrey
The ONE thing every dev should know (Interview)
The incomparable Jessica Kerr drops by with a grab-bag of amazing topics. Understanding software systems, transferring knowledge between devs, building relationships, using VS Code & Docker to code together, observability as a logical extension of TDD, and a whole lot more.
Join the discussion
Changelog++ members support our work, get closer to the metal, and make the ads disappear. Join today!
Sponsors:
Linode – Our cloud of choice and the home of Changelog.com. Deploy a fast, efficient, native SSD cloud server for only $5/month. Get 4 months free using the code changelog2019 OR changelog2020. To learn more and get started head to linode.com/changelog.
Tidelift – If you use open source to develop applications as part of your day job…our friends at Tidelift would like you to share your thoughts in their annual open source survey. Take the survey — it takes around 10 minutes on average.
Algolia – Our search partner. As the world is spending even more time online, search will be a critical lever for engaging your own users and customers. Algolia made their Pro Plan free to any developer or team working on a COVID-19-related, not-for-profit website or app.
Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com.
Featuring:
Jessica Kerr – Website, GitHub, X
Adam Stacoviak – Website, GitHub, LinkedIn, Mastodon, X
Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
Get an email as soon as we ship new shows ~> subscribe via email here
Symmathecist (n) - A quick definition, without the narrative
newsletter.jessitron.com
Listen to Jessica (and others) elsewhere on Arrested DevOps and Greater Than Code
Something missing or broken? PRs welcome!
Episode 195: Ad-hoc promotion and quitting a huge company with Charity Majors
<p>We’re excited to have special guest <a href="https://twitter.com/mipsytipsy">Charity Majors</a> on the show! Charity is the CTO and former CEO of <a href="https://www.honeycomb.io">Honeycomb</a>. She has worked at Second Life, Parse, Facebook, and more. She blogs at <a href="https://charity.wtf/">charity.wtf</a>.</p>
<p>Dave, Jamison, and Charity answer these questions:</p>
<ol>
<li>
<p>I’ve had the role of tech lead informally for the past two years at a fast-growing tech startup. We were a team of 6 developers, and now we are 16. Recently, we had a department meeting in which the Software Development VP communicated that we have 3 teams and I was the tech lead of two of them. I was surprised. He hasn’t mentioned his decision of splitting the teams nor that I’ve been officially promoted to tech lead. I was expecting a one-on-one where he would “pop the question”: Will you be my tech lead?</p>
<p>I asked him privately if that meant I would be officially promoted and would have my title changed. He said that he was going to have this conversation with the HR Manager and would get back to me, but potentially.</p>
<p>He doesn’t spend time on one-on-ones, nor is he very good at managing people although he’s good technically. How weird is this situation? A manager tells his team that they now have a tech lead along some org changes. I haven’t been informed, haven’t had my title changed yet, and haven’t been offered a raise yet.</p>
</li>
<li>
<p>Hi! I love your show and have been listening to it almost since day one. I was an engineer for about 10 years, and I’ve been a manager for about 1 year, and I love my team. They’re high performers, we have a high level of trust. I also like my boss! But the larger org has some issues, and in time-honored Soft Skills Engineering tradition, I plan to quit. I would like to stay in management. So I have these questions:</p>
<p>1) My employer is a very large public company. How much should I care about negative headlines and Wall Street’s opinion?</p>
<p>2) How long should I stay in my role as a manager before looking for a new job?</p>
<p>3) How do I message this to my team when I leave?</p>
</li>
</ol>
Observability is for your unknown unknowns (Interview)
Christine Yen (co-founder and CEO of Honeycomb) joined the show to talk about her upcoming talk at Strange Loop titled “Observability: Superpowers for Developers.” We talk practically about observability and how it delivers on these superpowers. We also cover the biggest hurdles to observability, the cultural shifts needed in teams to implement observability, and even the gains the entire organization can enjoy when you deliver high-quality code and you’re able to respond to system failure with resilience.
Join the discussion
Changelog++ members support our work, get closer to the metal, and make the ads disappear. Join today!
Sponsors:
DigitalOcean – The simplest cloud platform for developers and teams Whether you’re running one virtual machine or ten thousand, makes managing your infrastructure too easy. Get started for free with a $50 credit. Learn more at do.co/changelog.
GoCD + Kubernetes – With GoCD running on Kubernetes, you define your build workflow and let GoCD provision and scale build infrastructure on the fly. GoCD installs as a Kubernetes native application. Scale your build infrastructure elastically. Learn more at gocd.org/kubernetes
CrossBrowserTesting – The ONLY all-in-one testing platform that can run automated, visual, and manual UI tests – on thousands of real desktops and mobile browsers.
Strange Loop – A conference for software developers in St. Louis, MO. covering programming languages, databases, distributed systems, security, machine learning, creativity, and more! Sep 12-14, 2019 / Oct 1-3, 2020 / Sep 30-Oct 2, 2021
Featuring:
Christine Yen – Website, GitHub, X
Adam Stacoviak – Website, GitHub, LinkedIn, Mastodon, X
Jerod Santo – Website, GitHub, LinkedIn, Mastodon, X
Show Notes:
“testing is for known knowns, monitoring is for known unknowns, observability is for unknown unknowns” – Jez Humble
Check out Strange Loops’ impressive lineup of speakers this year
Observability: Superpowers for Developers
Framework for an observability maturity model
Something missing or broken? PRs welcome!
Episode 19: I want to build a world spanning search engine on top of GCP
Some companies that offer services expect you to do things their way or take the highway. However, Google expects people to simply adapt the tech company’s suggestions and best practices for their specific context. This is how things are done at Google, but this may not work in your environment.
Today, we’re talking to Liz Fong-Jones, a Senior Staff Site Reliability Engineer (SRE) at Google. Liz works on the Google Cloud Customer Reliability Engineering (CRE) team and enjoys helping people adapt reliability practices in a way that makes sense for their companies.
Some of the highlights of the show include:
Liz figures out an appropriate level of reliability for a service and how a service is engineered to meet that target
Staff SRE involves implementation, and then identifying and solving problems
Google’s CRE team makes sure Google Cloud customers can build seamless services on the Google Cloud Platform (GCP)
Service Level Objectives (SLOs) include error budgets, service level indicators, and key metrics to resolve issues when technology fails
Learn from failures through instant reports and shared post-mortems; be transparent with customers and yourself
GCP: Is it part of Google or not? It’s not a division between old and new.
Perceptions and misunderstandings of how Google does things and how it’s a different environment
Google’s efforts toward customer service and responsiveness to needs
Migrating between different Cloud providers vs. higher level services
How to use Cloud machine learning-based products
GCP needs to focus on usability to maintain a phase of growth
Offer sensible APIs; tear up, turn down, and update in a programmatic fashion
Promotion vs. Different Job: When you’ve learned as much as you can, look for another team to teach something new
What is Cloud and what isn’t? Cloud deployments require SRE to be successful but SREs can work on systems that do not necessarily run in the Cloud.
Links:
Cloud Spanner
Kubernetes
Cloud Bigtable
Google Cloud Platform blog - CRE Life Lessons
Google SRE on YouTube
.
Honeycomb Data Infrastructure with Sam Stokes - Episode 20
Summary
One of the sources of data that often gets overlooked is the systems that we use to run our businesses. This data is not used to directly provide value to customers or understand the functioning of the business, but it is still a critical component of a successful system. Sam Stokes is an engineer at Honeycomb where he helps to build a platform that is able to capture all of the events and context that occur in our production environments and use them to answer all of your questions about what is happening in your system right now. In this episode he discusses the challenges inherent in capturing and analyzing event data, the tools that his team is using to make it possible, and how this type of knowledge can be used to improve your critical infrastructure.
Preamble
Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
You can help support the show by checking out the Patreon page which is linked from the site.
To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
A few announcements:
There is still time to register for the O’Reilly Strata Conference in San Jose, CA March 5th-8th. Use the link dataengineeringpodcast.com/strata-san-jose to register and save 20%
The O’Reilly AI Conference is also coming up. Happening April 29th to the 30th in New York it will give you a solid understanding of the latest breakthroughs and best practices in AI for business. Go to dataengineeringpodcast.com/aicon-new-york to register and save 20%
If you work with data or want to learn more about how the projects you have heard about on the show get used in the real world then join me at the Open Data Science Conference in Boston from May 1st through the 4th. It has become one of the largest events for data scientists, data engineers, and data driven businesses to get together and learn how to be more effective. To save 60% off your tickets go to dataengineeringpodcast.com/odsc-east-2018 and register.
Your host is Tobias Macey and today I’m interviewing Sam Stokes about his work at Honeycomb, a modern platform for observability of software systems
Interview
Introduction
How did you get involved in the area of data management?
What is Honeycomb and how did you get started at the company?
Can you start by giving an overview of your data infrastructure and the path that an event takes from ingest to graph?
What are the characteristics of the event data that you are dealing with and what challenges does it pose in terms of processing it at scale?
In addition to the complexities of ingesting and storing data with a high degree of cardinality, being able to quickly analyze it for customer reporting poses a number of difficulties. Can you explain how you have built your systems to facilitate highly interactive usage patterns?
A high degree of visibility into a running system is desirable for developers and systems adminstrators, but they are not always willing or able to invest the effort to fully instrument the code or servers that they want to track. What have you found to be the most difficult aspects of data collection, and do you have any tooling to simplify the implementation for user?
How does Honeycomb compare to other systems that are available off the shelf or as a service, and when is it not the right tool?
What have been some of the most challenging aspects of building, scaling, and marketing Honeycomb?
Contact Info
@samstokes on Twitter
Blog
samstokes on GitHub
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
Honeycomb
Retriever
Monitoring and Observability
Kafka
Column Oriented Storage
Elasticsearch
Elastic Stack
Django
Ruby on Rails
Heroku
Kubernetes
Launch Darkly
Splunk
Datadog
Cynefin Framework
Go-Lang
Terraform
AWS
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast
Honeycomb, Complex Systems, Saving Sanity
Charity Majors joined the show to talk about debugging complex systems, using go to save one’s sanity, hiring smart people who can learn, and collectively working to make “on-call” life not miserable.
Join the discussion
Changelog++ members support our work, get closer to the metal, and make the ads disappear. Join today!
Sponsors:
Linode – Our cloud server of choice. Get one of the fastest, most efficient SSD cloud servers for only $5/mo. Use the code changelog2017 to get 4 months free!
Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform.
Toptal – Scale your team and hire from the top 3% of developers and designers with Toptal. Email adam@changelog.com for a personal introduction.
Compose – Production ready, cloud hosted databases. Pick your flavor - MongoDB, Elasticsearch, RethinkDB, Redis, Postgres, etcd, or RabbitMQ. When you’re ready to sign up use our special URL compose.com/changelog to get 60-days free on Compose
Featuring:
Charity Majors – Website, GitHub, X
Erik St. Martin – GitHub, X
Carlisia Thompson – GitHub, LinkedIn, X
Brian Ketelsen – GitHub, X
Show Notes:
Honeycomb :: Powerful, Exploratory Learning with Richer Data
go package libhoney (it’s the APM of the future!)
How We Moved Our API From Ruby to Go and Saved Our Sanity
CHARITY.WTF
Database Reliability Engineering book
Interesting Go Projects and News
Charity wants to give big shout outs (shouts out?) to Naitik Shah and Matt Silverlock!
Go 1.8 is released
Implementing a Debugger: The Fundamentals
Building a Go Debugger
Gobot - 1.2 Released
Pixterm - Draw images in your ANSI terminal with true color
1.8 Release Parties Everywhere
Change to Go CoC
Free Software Friday!
Each week on the show we give a shout out to an open source project or community that’s made an impact in our day to day developer lives.
Brian - Eclipse Che
Erik - Kube-Lego
Carlisia - Visual Studio Code
Something missing or broken? PRs welcome!