News, guides, and engineering deep dives.
Practical guidance on Kafka, Flink, Iceberg, and real-time data.
Kafka cluster monitoring
What to monitor at the Kafka cluster level: key JMX metrics, multi-broker collection, alerting thresholds, capacity signals, and a health check script.
Kafka consumer monitoring and performance tuning
Learn which Kafka consumer metrics matter most, how to interpret them, and which configuration changes will improve performance and reduce lag.
Kafka monitoring: a guide for platform engineers
A practical guide to Kafka monitoring for platform engineers: the metrics that matter, alert thresholds, JVM tuning, consumer lag, and KRaft changes.
A detailed guide to Kafka producer monitoring
A practical guide to Kafka producer metrics, JMX collection, alerting thresholds, and diagnostic scripts for Java-based Kafka producers.
How Adidas uses Apache Kafka in production
A deep-dive into Adidas's Kafka architecture — covering observability at 100 billion messages per day, self-service topic provisioning, and custom GoLang tooling.
How Apple uses Apache Kafka in production
A deep-dive into Apple's Kafka architecture — covering their managed internal platform, Strimzi on EKS, tiered storage, zero-data-movement balancing, and mTLS migration.
How Barclays uses Apache Kafka in production
A deep-dive into Barclays' Kafka architecture — covering dual-environment deployment on AWS and IBM Z-Linux, operating practices, and the broader streaming stack.
How Datadog uses Apache Kafka in production
A deep-dive into Datadog's Kafka architecture — covering use cases, scale, engineering decisions, and key contributors across hundreds of clusters.
How Netflix uses Apache Kafka in production
A deep-dive into Netflix's Kafka architecture, covering the Keystone pipeline, Data Mesh platform, scale figures from 700 billion to 2 trillion events per day, and the engineering decisions behind it.