Cloud Native Observability: Fighting Rising Costs, Incidents
Observability in multi-cloud environments is becoming increasingly complex, as highlighted by Martin Mao, CEO and co-founder of Chronosphere. This challenge has two main components: a rise in customer-facing incidents, which demand significant engineering time for debugging, and the ineffectiveness and high cost of existing tools. These issues are creating a problematic return on investment for the industry.
Mao discussed these observability challenges on The New Stack Makers podcast with host Heather Joslyn, emphasizing the need to help teams prioritize alerts and encouraging a shift left approach for security responsibility among developers. With the adoption of distributed cloud architectures, organizations are not only dealing with a surge in data but also facing a cultural shift towards DevOps, where developers are expected to be more accountable for their software in production.
Historically, operations teams handled software in production, but in the cloud-native world, developers must take on these responsibilities themselves. Many current observability tools were designed for centralized operations teams, which creates a gap in addressing developer needs.
Mao suggests that cloud-native observability tools should empower developers to run and maintain their software in production, providing insights into the complex environments they work in. Moreover, observability tools can assist developers in understanding the intricacies of their software, such as its dependencies and operational aspects.
To streamline the data obtained from observability efforts and manage costs, Chronosphere introduced the "Observability Data Optimization Cycle." This framework starts with establishing centralized governance to set budgets for teams generating data. The goal is to optimize data usage to extract value without incurring unnecessary costs. This approach applies financial operations (FinOps) concepts to the observability space, helping organizations tackle the challenges of cloud-native observability.
Learn more from The New Stack about Observability and Chronosphere:
Observability Overview, News and Trends
4 Key Observability Best Practices
Top Ways to Reduce Your Observability Costs
Top 4 Factors for Cloud Native Observability Tool Selection
Unpacking the Costs and Value of Observability with Martin Mao
Martin Mao, CEO & Cofounder at Chronosphere, joins Corey on Screaming in the Cloud to discuss the trends he sees in the observability industry. Martin explains why he feels measuring observability costs isn’t nearly as important as understanding the velocity of observability costs increasing, and why he feels efficiency is something that has to be built into processes as companies scale new functionality. Corey and Martin also explore how observability can now be used by business executives to provide top line visibility and value, as opposed to just seeing observability as a necessary cost.
About Martin
Martin is a technologist with a history of solving problems at the largest scale in the world and is passionate about helping enterprises use cloud native observability and open source technologies to succeed on their cloud native journey. He's now the Co-Founder & CEO of Chronosphere, a Series C startup with $255M in funding, backed by Greylock, Lux Capital, General Atlantic, Addition, and Founders Fund. He was previously at Uber, where he led the development and SRE teams that created and operated M3. Previously, he worked at AWS, Microsoft, and Google. He and his family are based in the Seattle area, and he enjoys playing soccer and eating meat pies in his spare time.
Links Referenced:
Chronosphere: https://chronosphere.io/
LinkedIn: https://www.linkedin.com/in/martinmao/
Open Core, Real-Time Observability Born in the Cloud with Martin Mao
About Martin
Martin Mao is the co-founder and CEO of Chronosphere. He was previously at Uber, where he led the development and SRE teams that created and operated M3. Prior to that, he was a technical lead on the EC2 team at AWS and has also worked for Microsoft and Google. He and his family are based in our Seattle hub and he enjoys playing soccer and eating meat pies in his spare time.
Links:
Chronosphere: https://chronosphere.io/
Email: contact@chronosphere.io
Monitoring, Metrics and M3, with Martin Mao and Rob Skillington
Martin Mao and Rob Skillington are co-founders of Chronosphere; CEO and CTO respectively. They both worked on the monitoring team at Uber, where they created M3: a metrics platform with an open source time-series database built for scale. They join Craig and Adam to talk about monitoring, metrics and M3 on the last episode of 2019.
Do you have something cool to share? Some questions? Let us know:
web: kubernetespodcast.com
mail: kubernetespodcast@google.com
twitter: @kubernetespod
Chatter of the week Test message from Delta Airlines
News of the week CSI migration and CSI volume snapshots
AKS Private Clusters in preview
GKE maintenance Windows and exclusions is GA
Google Cloud E2 VMs: introduction and understanding dynamic resource management
New features in Cloud Run for Anthos
Best practices for performing forensics on containers
Infrastructure at Cliqz, and introducing Hydra
Envoy CVEs Istio security bulletin
The Top 3 Service Mesh Developments in 2019 by Zack Jory
Istio Service Mesh Explained in 5 Minutes by Ram Vennam
Ambassador Edge Stack
Solo.io WebAssembly Hub Episode 55, with Idit Levine
Kafka Envoy Protocol Filter
Talos 0.3 beta
AutoTiKV tuning
OpenPolicyAgent's KubeCon recap Episode 42, with John Murray
A first look at Antrea from Alex Brand
TODO: read this article by Patrick DeVivo
Does Testing Kubernetes Conformance Leave You in the Dark? Get Progress Updates as Tests Run by John Schnake
Demystifying Kubernetes as a Service – How Alibaba Cloud Manages 10,000s of Kubernetes Clusters
How Jaeger Helped Grafana Labs Improve Query Performance and Root Out Tough Bugs
Adopting Kubernetes at Quora by Taylor Barrella,
CNCF announces schedule for Bengaluru/Delhi Forums
Links from the interview M3 website
M3: Uber's Open Source, Large-scale Metrics Platform for Prometheus
Before: Graphite and its Whisper database
Prometheus Why pull rather than push?
AlertManager
PromQL
RRDtool
M3 on GitHub: open source from the start
Chronosphere
Rob's 2019 KubeCon's talks: EU: M3 and Prometheus, Monitoring at Planet Scale for Everyone
NA: Deep Linking Metrics and Traces with OpenTelemetry, OpenMetrics and M3
Twitter: Rob Skillington
Martin Mao
M3
Chronosphere