Introducing Streamhouse: the open data architecture for AI | Learn More

Presentation

97 Incidents: Lessons From Running Kafka for 5 Years

« Kafka Summit Bangalore 2024

Running Kafka reliably can be challenging. Network instability, message corruption, slow disks, single-broker ""gray failure"", an unresponsive controller - these are all things that can happen to your clusters (and have happened to ours!).

We'll take a deep dive into the Kafka incidents Stripe has experienced, highlighting some of the most challenging and interesting ones. You'll get an inside look at the problems we've seen running 50 clusters with millions of messages per second for several years.

Most importantly, this talk will cover the remediations we've implemented. From configuration changes to alerting improvements to fully automated control plane actions, you'll learn how to run your own Kafka clusters more reliably.

Related Links

How Confluent Completes Apache Kafka eBook

Leverage a cloud-native service 10x better than Apache Kafka

Confluent Developer Center

Spend less on Kafka with Confluent, come see how