New in Confluent Cloud: Making Data & Pipelines Accessible for AI-Ready Streaming | Learn More

Feb 1, 2016Read Time: 3 min

Log Compaction | Highlights in the Kafka and Stream Processing Community | February 2016

Written By

Gwen ShapiraEngineering Manager, Confluent

Feb 1, 2016Read Time: 3 min

Welcome to the February 2016 edition of Log Compaction, a monthly digest of highlights in the Apache Kafka and stream processing community. Got a newsworthy item? Let us know.

We’ve been discussing many improvement proposals this month

- KIP-41 – a proposal to limit the number of records returned by KafkaConsumer.poll method has been accepted.
- KIP-42 – a proposal to add interceptors to producers and consumers has been accepted. This improvement creates interesting new monitoring options and once this is implemented, it will be interesting to hear how to community is using the new APIs.
- KIP-43 and KIP-44 propose improvements and extensions to Kafka’s authentication protocols. These are still under active discussion, and if you are interested in security in Kafka, I suggest reading the wiki and the discussion to see where we are heading.
- KIP-45 – a proposal to standardize the various collections that the KafkaConsumer API expects is still under discussion, with the benefits of more standardized approach being weighed against the desire to maintain backward compatibility for this new API.

Many of us are just learning the ins and outs of the new consumer. This recently published blog post, with a complete end-to-end example proves very useful.

This guy compared Kafka with Kinesis and shared some performance numbers.

Eric Sammer published a slide deck showing how Rocana uses Kafka and HDFS for large-scale search system.

A passionate developer wrote very detailed blog posts on Kafka integration with Spark Streaming. This includes the little-discussed question of how to write the results of the stream processing job back into Kafka.

LinkedIn wrote about new features in Samza. The blog post also includes sexy throughput numbers, description of their use-case and description of how Samza is used in their data products. Really cool stuff.

Google contributed their Dataflow API (but not implementation) to the Apache Software Foundation and are inviting other stream processing projects to implement their SDK. We are watching to see where this will take the active stream processing community.

There’s still time to register for Kafka Summit (April 26 in SF) – a gathering of the Apache Kafka community. Make the most of your week and sign up for a Kafka training classes or tutorial. If you need to stay in SF that week, there is a Kafka Summit room block at the Hilton with a special rate, be sure to book you room by March 30. Learn more at www.kafka-summit.org – check out the awesome program committtee while you’re there. Spots are filling up quickly so hurry and register now!

Gwen Shapira is a Software Enginner at Confluent. She has 15 years of experience working with code and customers to build scalable data architectures, integrating relational and big data technologies. She currently specialises in building real-time reliable data processing pipelines using Apache Kafka. Gwen is an Oracle Ace Director, an author of books including “Kafka, the Definitive Guide”, and a frequent presenter at data related conferences. Gwen is also a committer on the Apache Kafka and Apache Sqoop projects.