Introducing Streamhouse: the open data architecture for AI | Learn More
Running Kafka reliably can be challenging. Network instability, message corruption, slow disks, single-broker ""gray failure"", an unresponsive controller - these are all things that can happen to your clusters (and have happened to ours!).
We'll take a deep dive into the Kafka incidents Stripe has experienced, highlighting some of the most challenging and interesting ones. You'll get an inside look at the problems we've seen running 50 clusters with millions of messages per second for several years.
Most importantly, this talk will cover the remediations we've implemented. From configuration changes to alerting improvements to fully automated control plane actions, you'll learn how to run your own Kafka clusters more reliably.