Replay - Chaos Engineering for Gremlins with Jason Yee
On this Replay, we’re revisiting our conversation with Jason Yee, Staff Technical Advocate at Datadog. At the time of this recording, he was the Director of Advocacy at Gremlin, an enterprise-grade chaos engineering platform. Join Corey and Jason as they talk about what Gremlin is and what a director of advocacy does, making chaos engineering more accessible for the masses, how it’s hard to calculate ROI for developer advocates, how developer advocacy and DevRel changes from one company to the next, why developer advocates need to focus on meaningful connections, why you should start chaos engineering as a mental game, qualities to look for in good developer advocates, the Break Things On Purpose podcast, and more.
Show Highlights
(0:00) Intro
(0:31) Blackblaze sponsor read
(0:58) The role of a Director of Advocacy
(3:34) DevRel and twisting job definitions
(5:50) How DevRel confusion manifests into marketing
(11:37) Being able to measure and define a team’s success
(13:42) Building respect and a community in tech
(15:22) Effectively courting a community
(18:02) The challenges of Jason’s job
(21:06) Planning for failure modes
(22:30) Determining your value in tech
(25:41) The growth of Gremlin
(30:16) Where you can find more from Jason
About Jason Yee
Jason Yee is Staff Technical Avdocate at Datadog, where he works to inspire developers and ops engineers with the power of metrics and monitoring. Previously, he was the community manager for DevOps & Performance at O’Reilly Media and a software engineer at MongoDB.
Links
Break Things On Purpose podcast: https://www.gremlin.com/podcast/
Twitter: https://twitter.com/gitbisect
Original episode
https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/chaos-engineering-for-gremlins-with-jason-yee/
Sponsor
Backblaze: https://www.backblaze.com/
Chaos Engineering for Gremlins with Jason Yee
About Jason
Jason Yee is Director of Advocacy at Gremlin where he helps companies build more resilient systems by learning from how they fail. He also leads the internal Chaos Engineering practices to make Gremlin more reliable. Previously, he worked at Datadog, O’Reilly Media, and MongoDB. His pandemic-coping activities include drinking whiskey, cooking everything in a waffle iron, and making craft chocolate.
Links:
Break Things On Purpose podcast: https://www.gremlin.com/podcast/
Twitter: https://twitter.com/gitbisect
Chaos Engineering, with Ana Margarita Medina
Chaos Engineering is the discipline of experimenting in identifying potential areas of failure before they express themselves in outages. Ana Margarita Medina is a Chaos Engineer and Developer Advocate at Gremlin, a chaos-as-a-service vendor that recently added Kubernetes support. She talks to Adam and Craig about the discipline, and her journey to it.
Do you have something cool to share? Some questions? Let us know:
web: kubernetespodcast.com
mail: kubernetespodcast@google.com
twitter: @kubernetespod
Chatter of the week Shopify's Black Friday
Craig's Black Friday
News of the week AWS announcements: Managed node groups
EventBridge support in ECR
Sagemaker operators for Kubernetes
Eirini 1.0 is here
Security considerations for GKE by Maya Kaczorowski Episode 8. with Maya Kaczorowski
Managing a multi-site Cassandra cluster on multiple Kubernetes with CassKop / MultiCassKop by Seb Allamand
Run Ansible Tower or AWX in Kubernetes or OpenShift with the Tower Operator by Jeff Geerling
Everything I know about Kubernetes I learned from a cluster of Raspberry Pis by Jeff Geerling
Prometheus OpenMetrics Integration
Develop a Kubernetes controller in Java by Min Kim and Tony Ado
Running Kubernetes locally on Linux with Microk8s by Ihor Dvoretskyi and Carmine Rimi Episode 21, with Ihor Dvoretski
Episode 60, with Mark Shuttleworth
Linux Foundation Cyber Monday sale
Barrons says Kubernetes is the future of computing by Tae Kim
Links from the interview Chaos Engineering Chaos Engineering: the history, principles, and practice
Chaos Monkey
Netflix Simian Army
Fuzzing
Site reliability engineering
Google DiRT testing Video: 10 years of crashing Google by Kripa Krishnan
Ana's re:Invent talk
Reggaetón
#hugops
Chaos Engineering Slack
Gremlin Gremlin Free
What is a Gremlin? The Gremlins (Roald Dahl book)
Gremlins (1984 film)
Ana Margarita Medina on Twitter
SaaStr 282: The Ultimate Guide To Remote Work; All vs Part Remote Teams, How To Maintain Culture Across Teams, Should Compensation Be Location Adjusted, How To Structure Internal Processes with Remote Teams, How Remote Teams Impact Hiring, Sales and Fundr
Michael Pryor, Co-Founder & CEO @ Trello, now Head of Trello Product with Atlassian following their recent acquisition.
Kolton Andrus is the Founder & CEO @ Gremlin, the failure as a service startup finds weaknesses in your system before they cause problems.
Dylan Serota, Co-Founder and Chief Strategy Officer @ Terminal, the startup that helps you create world-class technical teams through remote operations as a service.
Rachel Carlson, Co-Founder and CEO @ Guild Education, the leader in education benefits offering the single most scalable solution for preparing the workforce of today for the jobs of tomorrow.
Sid Sijbrandi, Founder & CEO @ Gitlab, a single application for the entire software development lifecycle.
Jeppe Rindom is the Founder & CEO @ Pleo, the simple spending solution for your company automating expense reports and simplifying company expenses.
In Today's Episode We Discuss:
How should founders think about the debate between all remote vs part remote teams? How does life and operations change with each? What are the pros and cons? Is it possible to move between the two overtime?
What can one do to maintain culture with remote teams? What processes need to be in place to ensure a cohesive and streamlined communication process? What technical architecture needs to be in place? Where are the breakpoints when it comes to communication? How often does one need to do in person off-sites?
How does being remote or part remote impact fundraising? How do VCs think about this new structure of operations? What is the right way to present it? How does being outside a core tech hub impact one's ability to raise? How should one run a fundraising process if outside a core hub?
How important is it for your team to be near your customers? How does this change according to sector and customer base? How important is it for your team to be near your investors? Does having an exec and sales team in one place and the rest of the team elsewhere work?
Jason Lemkin
Harry Stebbings
SaaStr
Read the full transcript on our blog.
Episode 22: The Chaos Engineering experiment that is us-east-1
Trying to convince a company to embrace the theory and idea of Chaos Engineering is an uphill battle. When a site keeps breaking, Gremlin’s plan involves breaking things intentionally. How do you introduce chaos as a step toward making things better?
Today, we’re talking to Ho Ming Li, lead solutions architect at Gremlin. He takes a strategic approach to deliver holistic solutions, often diving into the intersection of people, process, business, and technology. His goal is to enable everyone to build more resilient software by means of Chaos Engineering practices.
Some of the highlights of the show include:
Ho Ming Li previously worked as a technical account manager (TAM) at Amazon Web Services (AWS) to offer guidance on architectural/operational best practices
Difference between and transition to solutions architect and TAM at AWS
Role of TAM as the voice and face of AWS for customers
Ultimate goal is to bring services back up and make sure customers are happy
Amazon Leadership Principles: Mutually beneficial to have the customer get what they want, be happy with the service, and achieve success with the customer
Chaos Engineering isn’t about breaking things to prove a point
Chaos Engineering takes a scientific approach
Other than during carefully staged DR exercises, DR plans usually don’t work
Availability Theater: A passive data center is not enough; exercise DR plan
Chaos Engineering is bringing it down to a level where you exercise it regularly to build resiliency
Start small when dealing with availability
Chaos Engineering is a journey of verifying, validating, and catching surprises in a safe environment
Get started with Chaos Engineering by asking: What could go wrong?
Embrace failure and prepare for it; business process resilience
Gremlin’s GameDay and Chaos Conf allows people to share experiences
Links:
Ho Ming Li on Twitter
Gremlin
Gremlin on Twitter
Gremlin on Facebook
Gremlin on Instagram
Gremlin: It’s GameDay
Chaos Engineering Slack
Chaos Conf
Amazon Leadership Principles
Adrian Cockcroft and Availability Theater
Digital Ocean
.
SaaStr 169: The Secret To Ensure High Conversion On Trials, How To Start and Scale A Remote Team Successfully & Lessons From Amazon and Netflix on A Culture of Achievement and Goal-setting with Kolton Andrus, Founder & CEO @ Gremlin
Kolton Andrus is the Founder & CEO @ Gremlin, the failure as a service startup finds weaknesses in your system before they cause problems. To date, they have raised over $8m in VC funding from some of the best in the business including the likes of Mike Volpi @ Index Ventures and Mike Dauber @ Amplify Partners. Prior to Gremlin, Kolton was a Chaos Engineer at Netflix improving streaming reliability and operating the Edge services. Fun fact, Kolton also designed and built Netflix's failure injection service. Before that he improved the performance and reliability of the Amazon Retail website. At both companies he has served as a 'Call Leader', managing the resolution of company-wide incidents.
In Today's Episode You Will Learn:
How a conversation in the hallway of a conference with a VC gave Kolton the confidence that he could leave the corporate world of Netflix and Amazon and start a startup?
What were Kolton's biggest takeaways from seeing the first hand scaling of behemoths like Amazon and Netflix? How did they fundamentally alter how he views goal setting today? How does Kolton look to achieve the balance of ambitious goal setting without the team losing motivation if they do not hit the goals?
Why does Kolton believe that a decentralised workforce is merely an evolution in how we do business? What are the core fundamentals to achieving success in creating and scaling a remote workforce? What have been some of the biggest challenges in structuring the team this way?
What is the single biggest tip Kolton has for other founders in ensuring high conversion rates from trials? Where do most founders go wrong with this? Today, is engineering buy in the only necessity to succeed in a bottoms up sales world?
60 Second SaaStr?
What does Kolton know now that he wishes he had known at the beginning?
If an investor can provide one thing, what is most important for Kolton?
What are Kolton's favourite SaaS reading materials?
When is a stretch a stretch too far for a team member?
If you would like to find out more about the show and the guests presented, you can follow us on Twitter here:
Jason Lemkin
Harry Stebbings
SaaStr
Kolton Andrus