Chad Harris
- Using Apache Kafka since 0.8.0 (2012)
- 18 years across engineering, architecture and engineering leadership
- 6+ years at Square / Block, Engineering Manager to Engineering Leader
- PCI-compliant, high-volume transactional platforms
In short
Chad Harris is Head of Product at Factor House. He has worked with Apache Kafka since 2012 and previously led engineering on PCI-compliant payments systems at Square (Block).
Chad Harris is Head of Product at Factor House, bringing 18 years of experience across software engineering, application architecture, and engineering leadership. He has deep hands-on expertise with Apache Kafka, high-volume transactional systems, and PCI-compliant architectures, most recently as an Engineering Leader at Block (formerly Square). At Factor House, Chad works directly with global enterprise customers to help them improve how they manage, govern, and observe their real-time data.
Expertise
Chad's expertise spans distributed systems, real-time data infrastructure, and enterprise application architecture. He has over six years of production experience with Apache Kafka and extensive background in NoSQL data stores including DynamoDB and Cassandra. His specialisation includes building highly scalable, tokenised security systems and high-volume transactional platforms built to PCI compliance standards. Beyond the technical, Chad is a seasoned engineering leader with a track record of mentoring teams and applying agile and lean methodologies in ways that deliver practical business outcomes.
Experience
Chad is currently Head of Product at Factor House. Before that, he spent over six years at Square (now Block), progressing from Engineering Manager to Engineering Leader. Prior to Square, he was Head of Engineering at Verrency, a payments technology company, where he led engineering across a high-security, high-availability fintech platform. Earlier roles include Technical Lead and Senior Software Engineer positions at Reece Australia, IOOF Holdings, National Australia Bank, and AIA, as well as consulting engagements across financial services and government.
Education
- University of Newcastle
Talks, appearances and mentions
Work published, hosted or co-presented by someone other than Factor House.
- Talk June 10, 2026 · Factor House and Aiven
Things that go bump in the night: Kafka operational issues
co-presented with Hugh Evans, Aiven
Real-world Kafka operational failures, from subtle misconfigurations to full-scale incidents, and the debugging workflows that help. Co-presented with Aiven, whose audience the session was run for.
- Video August 14, 2025 · Factor House
Chad Harris on Factor House's software
recorded while at Block, as a customer
Recorded while Chad was at Block, before he joined Factor House: a practitioner's account of running the tooling as a customer.
Writing and talks for Factor House
- Video September 16, 2026 · Factor House
Apache Kafka RBAC & multi-tenancy: Kpow demo
Scoping a virtual cluster view down to a single team's topics and metrics in Kpow, governed by SSO and fine-grained role-based access control.
- Video September 16, 2026 · Factor House
Apache Kafka broker monitoring & configuration: Kpow demo
Monitoring broker disk, throughput and replication health, filtering and exporting broker/topic tables, editing broker configuration under RBAC, and managing KRaft controller info and ACLs in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka topic management: Kpow demo
Truncating data by offset, electing leaders, increasing partitions, break-glass configuration changes like retention, and topic-level ACLs and reassignments in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka consumer group monitoring & lag: Kpow demo
Tracking consumer group stability over time, breaking lag down to the partition level, safely resetting offsets on a running group, and using group topology to trace lag back to a host or topic in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka Connect monitoring & task management: Kpow demo
Monitoring connector and task state, historical health charts, deploying new connector instances from the UI, and filtering to bulk-restart a subset of Kafka Connect tasks in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka schema registry management: Kpow demo
Viewing and editing schemas, creating new revisions, updating compatibility settings, and creating or deleting subjects across Confluent, Karapace, MSK, and other schema registries in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka cluster health monitoring: Kpow demo
An overall cluster health score tracked over time, a real example of partition data skew from a poor partitioning key, cost-saving cleanup signals, and what's coming next for Signals in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka data inspection & search: Kpow demo
Filtering topic data with kJQ, running high-volume streaming searches with no scan limit, narrowing scans by partition or key, and downloading, cloning, or producing result sets to other topics in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka data masking & PII protection: Kpow demo
Last-four, full, and email-domain masking rules applied during data inspection, and the data policy playground for testing redaction rules like show-first and hashing in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka tenant configuration: Kpow demo
Admin versus team-member tenant access, scoping a team to a single tenant by default, and configuring tenancy through a simple YAML file in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka audit logging: Kpow demo
How the audit log lives on a Kafka topic you can inspect directly, tracing an entry back to the exact query and data it exposed, and forwarding audit events to Slack, Teams, or a SIEM via webhook in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka access policies & SSO: Kpow demo
How access policies map out what a role can do, and how SSO integration with OAuth, SAML, Entra ID, and other providers drives permission assignment from existing identity roles and groups in Kpow.
- Video September 16, 2026 · Factor House
Apache Kafka temporary access & just-in-time permissions: Kpow demo
Configuring time-boxed, API-driven permissions that expire automatically, and wiring temporary policies to a service desk like Jira or ServiceNow for just-in-time approval workflows in Kpow.
- Video September 11, 2026 · Factor House
Kpow CLI, terminal UI, and agentic skills
A first look at Kpow's new CLI, terminal UI, and the agentic skills that let an AI assistant query, diagnose, and operate Kafka through Kpow under your existing SSO and RBAC controls.
- Article August 19, 2026 · Factor House
Dead letter queues in Kafka: a consumer-side, Kafka-only approach
The DLQ pattern for plain Kafka consumers: Kafka-only topics scoped per consumer group, replay through a retry topic rather than the main topic, and why using retries to ride out a downstream outage is the wrong move.
- Video August 7, 2026 · Factor House
Kpow Signals: automated operational insights
Introducing Signals, a new Kpow feature that continuously monitors your Kafka clusters for misconfigurations and early warning signs across brokers, topics, and consumer groups.
- Article June 27, 2026 · Factor House
Kafka UI: The Ultimate Guide
A 27-minute guide to what a Kafka UI is for, and what operators actually need visual control over: topics, consumers, brokers and connectors.
- Talk June 25, 2026 · Factor House webinar
Kafka operational issues and how to survive them
Four real production incidents, walked through end to end: a consumer group with tens of thousands of idle members overwhelming its coordinator, a config change that silently failed to roll back and filled a broker's disk six months later, a partition increase that left messages unread, and a 20-minute poll interval that turned one stuck batch into a 2am page.
The recording of the 25 June run. The same talk also ran on 29 April, and on 10 June with Aiven (both listed above).
- Video May 29, 2026 · Factor House
Data governance for Apache Kafka: lineage support in Factor Platform
A walkthrough of the OpenLineage metadata support for Apache Kafka in Factor Platform, presented to camera.
- Article May 29, 2026 · Factor House
Data governance for Kafka: introducing lineage support in Factor Platform
How OpenLineage metadata makes data ownership, PII classification and lineage visible by default in a Kafka environment.
- Talk April 29, 2026 · Factor House webinar
Kafka production failures, live tech talk
The first run of the Kafka production failures talk, delivered for a North American audience.
Registration page for a session that has ended.
- Post April 8, 2026 · LinkedIn
On StreamNative forking Kafka: why the Ursa engine changes the math
50 reactions, 2 comments, 3 reposts
"I always liked the idea of Pulsar and the Ursa engine. But there was just one problem with it for me: it wasn't Kafka." On why competing with the Kafka ecosystem is a Sisyphean task, and what changes once a challenger implements the Kafka protocol instead.
- Post March 27, 2026 · LinkedIn
Four questions to ask your platform team after the IBM–Confluent deal
80 reactions, 8 comments, 7 reposts
"You don't need to migrate anything today. But you should know your exposure." Four questions on Confluent-specific features versus standard Kafka APIs, what would break on a move to MSK, Redpanda or Aiven, and how much CI/CD and monitoring is coupled to Confluent tooling.
- Article March 26, 2026 · Factor House
What the IBM Confluent acquisition means for Kafka users
Assessing lock-in risk across Schema Registry, managed connectors and operational tooling after IBM's $11B acquisition.
Chad also posts on LinkedIn.
Talks
Kafka New at Factor House: September 2026
Chad Harris and Derek Troy-West demo the new Factor House CLI, terminal UI and agent skills for Kpow, then cover CVE hygiene, new hardened Docker images, Signals, and the redesigned Kpow UI coming in October.
September 23, 2026
Kafka Migrating to open source Kafka: cutting TCO without the operational burden
Chad Harris (Factor House) and Justin George (NetApp Instaclustr) work a real total-cost-of-ownership comparison across Confluent, Amazon MSK, and managed open source Kafka.
September 9, 2026
Kafka Reduce Kafka spend and operational risk: practical techniques for cluster consolidation
In this recorded webinar, Karel Sague and Chad Harris from Factor House share a practical framework for cutting Kafka costs and operational risk through cluster consolidation and migration.
August 27, 2026
Kafka Kafka operational issues: how to survive them
In this recorded session, Chad Harris, Solutions Architect at Factor House, walks through real Kafka operational failures and the debugging workflows that actually help in production.
June 25, 2026Latest articles
Best tools to monitor Spark Structured Streaming jobs that read Kafka
Five ways to monitor a Spark Structured Streaming job that reads Kafka, scored on rates, offsets behind latest, history, alerting and portability. Factor House has no Spark product and is not scored.
PySpark Kafka error: Failed to find data source: kafka
PySpark cannot read Kafka and fails with Failed to find data source: kafka. The cause is the missing Kafka connector package, and the fix is adding it with --packages at submit time.
Spark Structured Streaming Kafka offsets and lag: where Spark keeps its position
Spark Structured Streaming does not commit offsets to Kafka, so a consumer group view shows no lag for it. Where Spark keeps its position, how to measure lag from Spark, and how to restart safely.
Spark Structured Streaming from Kafka to Iceberg: how to write and keep tables healthy
How to write a Kafka topic into an Apache Iceberg table with Spark Structured Streaming, which settings matter for partitioned tables and commit rate, and what maintenance a streaming table needs.
Confluent alternatives after the IBM acquisition: how to choose
The realistic Confluent alternatives after IBM, compared on Kafka API, operator, schema registry and Connect, access control, cost model and migration path, plus what to check before leaving.
Best tools to monitor and operate Flink jobs
Flex ranks first of five tools for monitoring and operating Apache Flink jobs, scored on seeing a job's state, acting on it, alerting, controlling access and staying out of the data path.
Best Flink UI for submitting and managing jobs day to day
Flex and Ververica Platform tie on 81 of 100 among four Flink UIs for daily job work, scored on submitting, stopping and restoring jobs, inspecting them, and reading exceptions and logs.
Best tool to manage Kafka and Flink together: six options scored
Kpow and Flex rank first of six ways to manage Apache Kafka and Apache Flink together, scored on the data path, one access model, approvals, audit, Flink jobs and sign-in.
Best management tools for teams running Kafka to Iceberg pipelines
Kpow ranks first of seven Kafka and Flink tools for running a Kafka to Iceberg pipeline, scored on the data path, Connect sink tasks, sink lag, Flink checkpoints, access and audit.
How to find the bottleneck behind Flink backpressure
A Flink job is backpressured, sources slow down, latency grows and checkpoints time out. How to find the operator that causes it, tell a slow sink, skewed keys, a slow external call and too little parallelism apart, and fix the right one.
How to fix Flink checkpoints that fail or time out
A Flink job keeps restarting or falling behind because checkpoints fail, expire or take longer every hour. How to find the stage at fault, the metric that proves it, and the one config change to try first.
How to fix a Flink job stuck in a restart loop
A Flink job keeps cycling between RUNNING and RESTARTING. How to find the root exception, tell a transient failure from a permanent one, and set a restart strategy that stops the loop.
How to recover from a Flink JobManager failure with HA
The Flink JobManager pod was evicted, its node was lost or it lost leadership, and jobs restarted, looped or disappeared. What JobManager high availability protects, how to confirm it is configured and working, and how to fix a failover that loops or a cluster that did not recover.
How to fix Kafka consumer lag that looks wrong for a Flink job
Kafka consumer group lag for a Flink job keeps growing, stays flat, jumps after a restart or shows no committed offsets. Why the Flink Kafka source commits offsets only on checkpoints, how to read the Flink-side lag metrics, and what to change.
Best Flink management tool for platform teams serving many teams
Flex ranks first of six Flink management tools for platform teams running Apache Flink for many teams, scored on tenant boundaries, who may act on which job, one view across clusters and self-service.
Best Flink management tool for regulated teams: access control, audit and SSO
Flex ranks first of five Flink management tools for regulated teams, scored on the data path, production access on request, a per-person audit trail and SSO.
How to rescale a Flink job that cannot keep up
A Flink job falls behind after traffic grows. How to decide whether more parallelism will help, the limits set by max parallelism, Kafka partitions and task slots, and how to rescale safely with a savepoint or the Kubernetes operator autoscaler.
How to upgrade a Flink job from a savepoint without losing state
A Flink job needs new code, a new Flink version or a new parallelism, and the restore from a savepoint fails or comes up with empty state. The safe stop, verify, deploy and restore procedure, and how to diagnose a restore that fails.
Flink SQL to Kafka: fix duplicate or missing rows
A Flink SQL job writing to Kafka fails on update changes, or consumers see duplicate or missing rows. How to read the changelog mode, choose kafka or upsert-kafka, and set delivery.
How to find and fix Flink state that keeps growing
Flink checkpoints get bigger every day, disks fill and TaskManagers run short of memory. How to find the operator and state that hold the growth, tell a leak from real growth, and bring it down.
Best Flink tools for Flink on Kubernetes
Flex ranks first of five tools for teams running Apache Flink on Kubernetes, scored on governance per person and job, one view across clusters and staying out of the data path.
Best Flink tools for self-managed Apache Flink
Flex ranks first of five tools for self-managed Apache Flink, scored on governance per person and job, one view across clusters and staying out of the data path.
How to fix Flink windows that never fire or drop late events
An event-time Flink job emits no window output or silently drops late events. How to find the cause from the watermark metrics and fix idleness, timestamps, bounds and lateness.
Iceberg writer and reader see different tables
A Kafka to Iceberg writer commits, but a reader sees no data or fails with table not found. Find which catalog, warehouse or namespace each side resolves, then fix it without losing rows.
Iceberg commits failing or conflicting: find and fix the cause
A streaming writer stops committing to an Iceberg table with CommitFailedException or exhausted retries. How optimistic concurrency works, how to find the conflicting job, and how to fix it.
Iceberg storage keeps growing: expire snapshots and orphan files
A streaming-fed Iceberg table keeps growing because every commit adds a snapshot and old files stay referenced. How to measure it, then expire snapshots and remove orphan files safely.
Kafka to Iceberg pipeline breaks after a schema change
A producer changes a field and the Kafka to Iceberg sink fails or drops data. What Iceberg schema evolution allows, what the sink does with an unknown field, and how to fix it.
Kafka to Iceberg: too many small files, and how to compact them
A Kafka to Iceberg pipeline writes many tiny data files on every commit, so query planning slows and metadata grows. Why it happens, how to measure it, and how to fix it with compaction and sink settings.
Best Kafka management tools for transport and logistics companies
Kpow ranks first of six Kafka management tools for transport and logistics companies such as Schneider, Knight-Swift, US Foods and 90POE, scored on the data path, search, clusters and access.
Best Kafka management tools for sports betting and iGaming operators
Kpow, used by FanDuel, ranks first of six Kafka management tools for sports betting and iGaming operators, scored on the data path, production access, audit, masking and search.
Best Kafka management tools for technology and SaaS companies
Kpow ranks first of six Kafka tools for technology and SaaS companies, scored on the data path, multi-cluster reach and monitoring. HPE, ServiceNow, Garmin, TripleLift and JobNimbus run Kpow.
Best Kafka management tools for telecommunications companies
Kpow ranks first of six Kafka management tools for telecommunications companies such as Belong, scored on subscriber data masking, topic inspection and staying out of the data path.
Best Kafka tools for Aiven Kafka Connect
Kpow ranks first of six Kafka tools for Aiven's standalone Kafka Connect service, scored on the data path, production access, per-person audit, connector operations, sign-in and shared clusters.
Best Kafka tools for Amazon MSK Connect
Kpow ranks first of seven Kafka tools for Amazon MSK Connect, scored on the data path, MSK Connect API support, production access, per-person audit, sign-in and shared clusters.
Best Kafka tools for Apache Kafka Connect
Kpow ranks first of seven Kafka tools for Apache Kafka Connect, scored on the data path, production access, per-person audit, sign-in, shared clusters and connector operations, with TD Bank's use.
Best Kafka tools for teams using AWS IAM Identity Center (AWS SSO)
Kpow ranks first of four Kafka tools for teams that sign in with AWS IAM Identity Center, formerly AWS SSO, scored on the data path, documented Identity Center setup, group-to-role mapping, enforcement, production access and per-person audit.
Best Kafka tools for Confluent Cloud ksqlDB
Six Kafka tools for teams using Confluent Cloud ksqlDB, scored on the data path, access, per-person audit and ksqlDB operations. Kpow ranks first; NORD/LB moved its ksqlDB work into Kpow.
Best Kafka tools for Confluent Cloud Managed Connect
Kpow ranks first of six Kafka tools for Confluent Cloud managed connectors, scored on the data path, production access, per-person audit, managed connector operations, sign-in and shared clusters.
Best Kafka tools for teams using Debezium
Kpow ranks first of seven Kafka tools for teams running Debezium on Kafka Connect, scored on the data path, production access, per-person audit, sign-in, shared clusters and CDC operations.
Best Kafka tools for Instaclustr managed Kafka Connect
Kpow ranks first of six Kafka tools for NetApp Instaclustr managed Kafka Connect, scored on the data path, connector operations, production access, per-person audit, Instaclustr setup and sign-in.
Best Kafka tools for Kafka Streams
Kpow ranks first of six Kafka tools for teams running Kafka Streams applications, scored on the data path, production access, per-person audit, Kafka Streams visibility, sign-in and shared clusters.
Best Kafka tools for teams using Microsoft Entra ID (Azure AD)
Kpow ranks first of four Kafka tools for teams signing in with Microsoft Entra ID (Azure AD), scored on the data path, Entra ID setup, group-to-role mapping, enforcement, production access and audit.
Best Kafka tools for teams using Okta
Kpow ranks first of four Kafka tools for teams that sign in with Okta, scored on the data path, documented Okta setup, group-to-role mapping, enforcement behind the login, production access and per-person audit.
Best Kafka tools for self-managed ksqlDB
Kpow ranks first of six Kafka tools for teams running ksqlDB themselves, scored on the data path, production access, per-person audit, ksqlDB operations, sign-in and shared clusters, with NORD/LB's use.
Best Kafka tools for Spring Kafka and Spring Cloud Stream
Kpow ranks first of five Kafka tools for teams running Spring Kafka and Spring Cloud Stream applications, scored on the data path, Spring application visibility, production access, per-person audit, sign-in and shared clusters.
Best Kafka UI tools for Apicurio Registry
Kpow ranks first of seven Kafka UI tools for Apicurio Registry, scored on data path, access, audit, sign-in and work through Apicurio's Confluent-compatible API.
Best Kafka UI tools for Azure Event Hubs
Kpow ranks first of seven Kafka UI tools for the Azure Event Hubs Kafka endpoint, scored on the data path, production access, per-person audit, sign-in, shared namespaces and Event Hubs fit.
Best Kafka UI tools for Buf Schema Registry (BSR)
Kpow ranks first of six Kafka UI tools for teams on the Buf Schema Registry, scored on the data path, production access, per-person audit, sign-in, registries and Protobuf.
Best Kafka UI tools for Confluent Schema Registry
Kpow ranks first of seven Kafka UI tools for Confluent Schema Registry, with TD and NORD/LB among its users, on data path, access, audit and registry depth.
Best Kafka UI tools for Google Managed Kafka Schema Registry
Kpow ranks first of seven Kafka UI tools for Google Managed Kafka Schema Registry, scored on governance, the data path, Google token sign-in and schema management.
Best Kafka UI tools for IBM Event Streams
Kpow ranks first of seven Kafka UI tools for IBM Event Streams, scored on the data path, production access, audit, sign-in, shared clusters and Event Streams fit.
Best Kafka UI tools for Karapace Schema Registry
Kpow ranks first of six Kafka UI tools for teams using Karapace, on data path, access, audit, several registries and documented Karapace setup on Aiven and Instaclustr.
Best Kafka UI tools for Redpanda Schema Registry
Kpow ranks first of six Kafka UI tools for teams using Redpanda Schema Registry, on data path, access, audit, several registries and documented Redpanda setup.
Best Kafka UI tools for StreamNative Schema Registry
Kpow ranks first of six Kafka UI tools for StreamNative Schema Registry, scored on the data path, access, audit and how each tool handles the registry's API differences.
Best Kafka management tools for defence and national security organisations
Kpow ranks first of six Kafka tools for defence and national security, scored on the data path, air-gapped operation, time-boxed access, per-person audit and need-to-know scoping.
Best Kafka UI tools for Bufstream
Kpow ranks first of six Kafka UI tools for Bufstream, scored on the data path, production access, per-person audit, sign-in, shared clusters and Protobuf schemas.
Best Kafka UI tools for OCI Streaming with Apache Kafka
Kpow ranks first of seven Kafka UI tools for OCI Streaming with Apache Kafka on Oracle Cloud, scored on governance, the data path, self-run Connect and registries, and sign-in.
Best Kafka UI tools for Red Hat AMQ Streams (Streams for Apache Kafka on OpenShift)
Kpow ranks first of eight Kafka UI tools for Red Hat AMQ Streams on OpenShift, scored on mTLS, SCRAM and OAuth sign-in, per-person audit and the data path.
Best Kafka UI tools for WarpStream
Kpow ranks first of seven Kafka UI tools for WarpStream, scored on the data path, production access, per-person audit, sign-in, mixed fleets and WarpStream fit.
Best Flink tools for Ververica Platform
Flex and the Ververica Platform web UI tie for first of five Flink tools, scored on per-job governance, one view across installations and Deployment lifecycle.
Best Kafka management tools for airlines and travel companies
Kpow ranks first of six Kafka management tools for airlines and travel companies, scored on the data path, shared clusters, audit, hybrid clusters, masking and search.
Best Kafka management tools for asset and wealth managers
Kpow ranks first of six Kafka management tools for asset managers, wealth managers and superannuation funds, scored on data path, access, audit and masking.
Best Kafka management tools for automotive companies
Kpow ranks first of six Kafka management tools for carmakers, suppliers and connected-vehicle firms, scored on the data path, plant and cloud, and search.
Best Kafka management tools for Confluent Platform
Kpow ranks first of six Kafka management tools for Confluent Platform, Control Center included, on data path, production access, per-person audit and ksqlDB.
Best Kafka management tools for crypto exchanges
Kpow ranks first of eight Kafka management tools for crypto exchanges and CASPs, on staying out of the trading path, time-boxed access and per-person audit.
Best Kafka management tools for ecommerce companies and marketplaces
Kpow ranks first of six Kafka management tools for ecommerce and marketplaces, scored on the data path, footprint, finding orders, masking, audit and SSO.
Best Kafka management tools for energy trading firms
Kpow ranks first of six Kafka management tools for energy and commodity traders, scored on the data path, audit under REMIT, production access and hybrid reach.
Best Kafka management tools for energy and utility companies
Kpow ranks first of six Kafka management tools for grid operators and utilities, scored on the data path, audit, production access, Kubernetes and sign-in.
Best Kafka management tools for government agencies
Kpow ranks first of six Kafka tools for government and the public sector, scored on data path, time-boxed access, audit, directory sign-in and a published VPAT.
Best Kafka management tools for health technology and medical device companies
Kpow ranks first of six Kafka tools for health technology and medical device companies, scored on staying out of the data path, team scoping, audit and MSK.
Best Kafka management tools for healthcare organisations
Kpow ranks first of six Kafka management tools for healthcare organisations, scored on the data path, the audit trail, production access to patient data and teams.
Best Kafka management tools for insurers
Kpow ranks first of six Kafka management tools for insurers and reinsurers, scored on the data path, production access, audit, message tracing, SSO and teams.
Best Kafka management tools for lenders and consumer fintechs
Kpow ranks first of six Kafka tools for lenders and consumer fintechs, scored on staying out of the data path, time-boxed access, per-person audit and MSK.
Best Kafka management tools for manufacturers
Kpow ranks first of six Kafka management tools for manufacturers, scored on the data path, plant and cloud clusters, per-person audit, access, SSO and teams.
Best Kafka management tools for payments companies
Kpow ranks first of seven Kafka management tools for payments companies, scored on PCI DSS scope, card masking, time-boxed access and per-person audit.
Best Kafka management tools for pharma and life sciences companies
Kpow ranks first of six Kafka management tools for pharma and life sciences companies, scored on the data path, audit trail, controlled access and hybrid reach.
Best Kafka management tools for retailers
Kpow ranks first of six Kafka management tools for retailers, scored on the data path, team scoping, SSO, production access, audit and data inspect.
Best Kafka management tools for supermarkets and food distributors
Kpow ranks first of six Kafka management tools for supermarket chains and food distributors, scored on the data path, mixed clusters, inspection, teams and SSO.
Best Kafka management tools for trading firms
Kpow ranks first of six Kafka management tools for trading firms and capital markets, scored on the trading path, audit, production access, search and SSO.
Best Kafka tools for Aiven for Apache Kafka
Kpow ranks first of six Kafka UI and management tools for Aiven for Apache Kafka, scored on governance beyond ACLs, the data path, Aiven sign-in and Karapace.
Best Kafka tools for NetApp Instaclustr managed Kafka
Kpow ranks first of seven Kafka UI and management tools for NetApp Instaclustr, scored on governance beyond ACLs, Karapace, managed Connect and the data path.
Best Kafka tools for self-managed Apache Kafka
Kpow ranks first of seven Kafka tools for self-managed Apache Kafka, scored on the data path, production access, audit, sign-in, shared clusters and brokers.
Best Kafka UI tools for AWS Glue Schema Registry
Kpow ranks first of seven Kafka UI tools for AWS Glue Schema Registry on governance, Glue decoding, schema management, cross-account roles and the data path.
Best Kafka UI tools for Confluent Cloud
Kpow ranks first of seven Kafka UI tools for Confluent Cloud, scored on per-person governance, the data path, Managed Connect, Schema Registry and sign-in.
Best Kafka UI tools for Google Cloud Managed Service for Apache Kafka
Kpow ranks first of seven Kafka UI tools for Google Cloud Managed Service for Apache Kafka, scored on governance, the data path, IAM sign-in and the registry.
Best Kafka UI tools for Redpanda
Kpow ranks first of six Kafka UI tools for Redpanda, scored on the data path, production access, per-person audit, sign-in, shared clusters and Redpanda fit.
Best Kafka UI tools for StreamNative Cloud
Kpow ranks first of six Kafka UI tools for StreamNative Cloud, scored on governance beyond StreamNative RBAC, the data path, sign-in and the Schema Registry.
Best Kafka UI tools for Strimzi (Kafka on Kubernetes)
Kpow ranks first of eight Kafka UI tools for Strimzi, including StreamsHub Console, on mTLS, SCRAM and OAuth sign-in, per-person audit and Kubernetes fit.
Best Kafka management tools for banks
Kafka management tools for banks, scored on time-boxed production access, a per-person audit trail, directory sign-in, shared clusters, on-prem plus cloud, and staying out of the data path.
Best Kafka UI tools for Amazon MSK
Kpow ranks first of seven Kafka UI tools for Amazon MSK, scored on IAM and SCRAM sign-in, MSK Serverless, MSK Connect, AWS Glue, per-person governance and VPC fit.
How to let an AI agent operate Kafka safely
Give an AI agent Kafka access without the keys: its own identity and ACLs, read-only first, human approval for destructive changes, audit of every call, and no sensitive payloads in the model.
How to inspect Kafka topics, messages and consumer groups from the terminal
Inspect Kafka over SSH with no web UI: brokers, under-replicated partitions, safe message peeks, Avro, consumer lag and Connect status, checked against Kafka 4.3 and kcat.
Best tools to manage Kafka ACLs
kafka-acls.sh, Terraform, kafka-gitops, Julie Ops, Strimzi, Klaw, Apache Ranger, OPA and the Kafka UIs, scored on auditability, pattern handling, automation, drift and visibility.
Kafka agent skills compared
Kafka agent skills for Claude Code and other agents, scored on one rubric. Covers guardrails, SSO and RBAC reuse, audit, cost, and Kpow as the self-hosted option for regulated teams.
How to troubleshoot Kafka with an AI agent
Troubleshoot Kafka with an AI agent: the read-only access to give it, the questions to ask, how to tell a connector fault from a database or catalog fault, and how to check its answer before you act.
Best tools for Kafka audit logging
Kafka audit logging happens at three layers: broker authorization logs, the audit trail of the tools people use, and data lineage. Every option for each layer, scored on the same rubric.
Best tools to manage Kafka broker configs
kafka-configs.sh, Strimzi, Terraform, JulieOps, AKHQ, Kafbat UI, Confluent Control Center and Kpow, scored on seeing the running config, catching drift, changing it safely and knowing who changed it.
Best tools to monitor Kafka broker health
Prometheus with the JMX Exporter, kafka_exporter, KMinion, Cruise Control, Datadog, New Relic, Confluent Control Center and Kpow, scored on the broker signals they actually see.
Kafka CLI commands for day-2 operations
The Kafka CLI commands operators run in production, checked against Kafka 4.3: topics, produce and consume, lag and offset resets, configs, cluster health and ACLs, with KRaft-era flag changes.
Best Kafka CLI tools
Kafka CLIs scored on coverage, contexts, auth, agent safety, maintenance and annual cost: the bundled scripts, kcat, kafkactl, kaf, kcl, rpk, Confluent CLI, Zoe, kt and the Kpow CLI.
Best tools to monitor Kafka Connect connectors
Kafka Connect monitoring tools, from the REST API and JMX to Strimzi, Kpow and Lenses, scored on task state, per-task metrics, restarts, alerting and access.
How to diagnose and fix a failed Kafka Connect connector
A connector reports RUNNING while its tasks have FAILED. The diagnostic sequence, and how to tell a connector fault from a limit in the external system it writes to.
Best tools to monitor Kafka consumer lag
Kafka consumer lag tools compared on one rubric: client-side against cluster-side measurement, offset lag against time lag, per-partition detail and alerting. The CLI, JMX, Burrow, Kpow and more.
How to find and mask PII in Kafka topics
An audit finds card numbers or emails readable in a Kafka topic. How to find every field that carries PII, choose where to mask it, and prove the masking holds, with the commands for each step.
Best tools for Kafka data masking
Kafka data masking tools hide sensitive fields at ingestion, in the client, in a proxy or at view time. Eleven options scored on what each placement protects against.
How to diagnose a Kafka deserialization error
A consumer throwing SerializationException and looping on one offset has one of five causes. How to tell which, and how to unblock the partition without losing data.
Best tools to control destructive Kafka operations
Kpow suits regulated teams that must control Kafka topic deletes and offset resets with RBAC, audit logging and staged approvals. Every control, from ACLs to GitOps, scored on one rubric.
Best tools to manage a dead letter queue in Kafka
Kafka DLQ tools compared on one rubric: error context, triage, repair, targeted replay and governance. Kpow suits regulated teams needing RBAC, audit logging and data masking.
Best Kafka governance tools for financial services
Kafka governance tools for banks, payments and insurance, scored against what DORA, PCI DSS, GDPR and SOX ask for: limited access, strong authentication, controlled change, audit and masked data.
Best Kafka MCP servers
Every Kafka MCP server and AI-agent access route, scored on governance: read vs write scope, permission model, audit, approval for destructive ops, message exposure to the model, and deployment.
Best tools to search messages across Kafka topics
Kafka is a log, not an index, so every search is a scan. Eleven ways to search messages across Kafka topics, scored on scan scope, filter language, multi-topic reach and masking.
Best tools to manage multiple Kafka clusters from one place
Twelve tools to manage multiple Kafka clusters from one place, scored on mixed distributions, deployment shape, per-cluster RBAC, cross-cluster views and cost at N clusters.
Best tools to reset Kafka consumer group offsets
How to reset Kafka consumer group offsets with the native CLI, dry run first, then tools scored on preview, scope, strategies, the inactive-group rule and audit: Kpow, AKHQ, Kafbat UI and Conduktor.
Best tools to reassign Kafka partitions
kafka-reassign-partitions.sh, Cruise Control, Strimzi, topicctl, topicmappr, Kpow and the managed-service options, scored on planning, throttling, batching, progress, cancellation and audit.
How to find and skip a poison pill in Kafka
A consumer stuck on one offset, lag climbing on one partition, the same deserialization error looping. How to confirm it is a poison pill, inspect the record, and skip it safely with a dry run first.
Best tools to manage poison pills in Kafka
Kafka poison pill tools compared on one rubric: isolation in the consumer, finding the bad record, skipping it safely, keeping a copy, and who may do it. Spring Kafka, the CLI, Kpow, AKHQ and more.
Best tools for Kafka role-based access control (RBAC)
Apache Kafka authorizes with ACLs, not roles. Eleven ways to add RBAC, scored on where each one enforces, who it binds, what it covers and how it audits change.
Best tools for Kafka schema registry management
Confluent Schema Registry, Apicurio, Karapace, AWS Glue, Redpanda, Kpow, Kafbat UI and AKHQ, scored on formats, compatibility checks, diffs, multiple registries and RBAC.
Best tools for Kafka SSO integration
Kafka SSO is two jobs: people signing in to Kafka UIs through Okta, Entra ID or Keycloak, and services authenticating to brokers with OAuth tokens. The tools for each half, scored on one rubric.
Best tools to manage Kafka topics at scale
Strimzi, Terraform, topicctl, JulieOps, Conduktor, AKHQ, Kafbat UI and Kpow, scored on topics as code, creation guardrails, scoped self-service, audit, multi-cluster reach and bulk operations.
Best Kafka terminal UIs (TUI)
Kafka terminal UIs scored on what you can see and change, multi-cluster, auth, keyboard UX and annual cost: ktea, Kaskade, kafui, Yozefu, kaftui, karat, the Kpow TUI and plain scripts.
How to diagnose and fix an unbalanced Kafka cluster
One broker hot, disk alerts on another, leader counts uneven. How to tell leader imbalance from data imbalance, the commands that fix each, how to throttle the move, and how to stop it coming back.
MirrorMaker 2 on Connect: migrating off connect-mirror-maker.sh
Moving MirrorMaker 2 off connect-mirror-maker.sh onto Kafka Connect: replication progress doesn't migrate, offset translation is off by default, and when not to bother.
Apache Flink: the complete guide
Apache Flink is a distributed stream processing framework for stateful computation over unbounded and bounded data. This hub indexes what we have written about running it in production.
Flink use cases
Companies running Apache Flink in production, the architectures behind their deployments, and what to read next. Indexed as we publish new use-case research.
Kpow vs AKHQ
AKHQ is free and costs operator time. Kpow is licensed per cluster with a published price. How the two differ on masking, audit, and support.
Kpow vs CMAK
CMAK connects through ZooKeeper and stops at Kafka 4.0. Kpow bills per cluster and talks to brokers. What each one costs, and which one fits your team.
Kpow vs Confluent Control Center
Control Center needs a Confluent JAR in the broker classpath. Kpow runs against any distribution, per cluster, at a published price. Which fits which team.
Kpow vs Kadeck
Kadeck bills per user, Kpow bills per cluster. What each one costs, what each needs in order to start, and which one fits your team.
Kpow vs Kafbat UI
Kafbat UI is free and carries real RBAC, masking and an audit log. Kpow is licensed per cluster with support behind it. Which fits which team.
Kpow vs Kafdrop
Kafdrop is free with no tier above it. Kpow is licensed per cluster. What each one does well, where each runs out, and which fits your cluster.
Kpow vs Lenses.io
Lenses.io bills by capability with a user cap. Kpow bills per cluster. How the pricing, the control plane and the exit differ, and which fits which team.
Kpow vs Offset Explorer
Offset Explorer is licensed per named user and installed on a laptop. Kpow is licensed per cluster and shared. What each costs, and which fits which team.
Kpow vs Redpanda Console
Redpanda Console is free, and its governance is licensed to the broker vendor. Kpow prices per cluster and publishes the price. Which one fits which team.
Apache Kafka vs Confluent Kafka
The broker is the same. What differs is licensing, bundled components, deployment models and support. Where the Apache line sits, what Confluent adds, and the dependency questions to ask.
Kafka-compatible cloud brokers
Redpanda and the Kafka-compatible brokers, evaluated from the operator's seat: compatibility proof, operational claims, TCO against tiered-storage Kafka, and benchmarks.
Confluent Kafka in Docker
Running Confluent's Kafka images and the official apache/kafka image in Docker: a working compose file, the advertised-listeners fix, log levels and custom Connect images.
Data governance policy examples for streaming platforms
Worked data governance policy examples for Kafka: schema compatibility rules, topic lifecycle, PII classification tiers, ACL and RBAC templates, and lineage requirements.
Debezium vs Kafka Connect
Debezium and Kafka Connect are not alternatives: Debezium is a CDC connector family that runs on Connect. The real decisions are log-based CDC vs JDBC polling, and Connect cluster vs Debezium Server.
Kafka deployment automation
Deployment automation for Kafka: zero-downtime rolling changes, declarative topics and ACLs, Kubernetes operators, Terraform and GitOps for cluster state.
The difference between Kafka and RabbitMQ
The difference between Kafka and RabbitMQ is the data model: a replayable log against a delete-on-acknowledge queue. Storage, routing, scaling and when to pick which.
How RabbitMQ works
How RabbitMQ works, mapped to Kafka concepts: exchanges and bindings, per-message acknowledgement, competing consumers, quorum queues and dead lettering.
Kafka ACL
A Kafka ACL allows or denies a principal an operation on a resource from a host. Syntax, copy-paste commands, production best practices and GitOps automation.
Kafka authentication
Kafka authentication verifies every client and broker connection with SASL or mutual TLS. Listener and JAAS blueprints, mechanism choice, rotation and troubleshooting.
Kafka consumer groups
A Kafka consumer group shares the work of consuming a topic, one partition per member. Troubleshooting lag and rebalances, the CLI cheat sheet, coordinator architecture and offset commits.
Kafka Connect
Kafka Connect moves data between Kafka and external systems through source and sink connectors. What a deployment is made of, the operational tasks, and where the sub-pages go deeper.
Kafka Connect MongoDB example
A production Kafka Connect MongoDB example in both directions: source connector via change streams, sink connector with idempotent writes, DLQ settings and secrets handled properly.
Kafka Connect pricing
Kafka Connect is free open source software. The cost is in running it: managed per-task and per-GB billing, network surcharges, or the engineering hours of self-hosting. The TCO math, honestly.
Kafka Tool download
Kafka Tool is the former name of Offset Explorer. What each download installs, and why regulated teams that need RBAC, audit logging, data masking and SSO on shared Kafka access should look at Kpow.
Kafka UI and console comparisons
Every major Kafka UI and console compared head to head: Kpow against each, and each against the others. Pricing, plan limits and where each one runs out.
Kafka vs other brokers
How Kafka compares with RabbitMQ and other message brokers: retention models, routing, use cases and operational trade-offs, from the operator's seat.
Kafka vs RabbitMQ performance
Kafka and RabbitMQ perform differently because their storage models differ. Throughput and latency behaviour, durability trade-offs, benchmark methodology and the workload facts that decide the fit.
Managed vs unmanaged databases
Managed vs unmanaged databases for teams running Kafka: operational overhead, SLAs, cost architecture, connector and CDC integration, control and security.
Multi-tenant architecture
A multi-tenant Kafka architecture shares one cluster across teams with quotas, ACLs and naming conventions. Isolation, namespaces, chargeback and topology.
RBAC roles
RBAC roles bundle permissions into named sets like viewer, operator and admin. How role definitions, resource patterns and operation mappings work across the Kafka ecosystem.
Kafka stream governance
Stream governance applies data governance to data in motion: schemas enforced at produce time, lineage across topics and jobs, catalogs, and access and quality rules on live streams.
What is a data governance policy?
A data governance policy is an enforceable rule for how data is structured, accessed, retained and traced. On Kafka it is implemented as configuration and code, not documents.
What is envelope encryption?
Envelope encryption encrypts data with a local data key, then wraps that key with a KMS-held key. On Kafka it is the pattern that makes per-field encryption work at full throughput.
What is Kafka Connect?
Kafka Connect streams data between Kafka and other systems as managed, fault-tolerant tasks. The internal architecture, production scaling, error handling and custom development.
What is Kafka rebalancing?
Kafka rebalancing redistributes a consumer group's partitions when membership changes. The triggers, the three timeouts, cooperative rebalancing, KIP-848 and the metrics that explain incidents.
What is RabbitMQ?
RabbitMQ explained for Kafka operators: the AMQP queue model, broker-side routing, push delivery, streams, and where it fits beside a Kafka deployment.
What is Redpanda?
Redpanda reimplements the Kafka wire protocol in a single C++ binary. The architecture, the drop-in compatibility boundaries, and how to evaluate it against tiered-storage Kafka.
Kafka: The Complete Guide
Apache Kafka is a distributed event streaming platform that stores ordered, replayable records in partitioned topics. This hub covers Kafka fundamentals, operations, governance and tooling.
Apache Kafka architecture: a complete guide
A complete guide to Apache Kafka architecture: internals, components, KRaft, replication, consumers, Connect, Streams, and deployment options.
Kafka brokers in production: configuration, troubleshooting and metrics
A Kafka broker stores partition logs and serves client requests. Broker configuration, bootstrap timeouts, crash loops and the JMX metrics to watch.
Kafka consumers in production
A Kafka consumer reads records from topic partitions, tracking its own offset. The configuration that decides message loss, rebalance troubleshooting, and poll-loop patterns that survive production.
Kafka in Docker
Running Kafka in Docker: the official images, reliable Docker Compose topologies, the advertised.listeners trap that breaks local connections, and why a single-node container is not a deployment.
Kafka offsets
A Kafka offset is a record's position in its partition, and a committed offset is a consumer group's bookmark. Reading lag from CURRENT-OFFSET and LOG-END-OFFSET, plus the reset strategies.
Kafka producers in production
A Kafka producer appends records to topic partitions. Configuration and tuning with real numbers, idempotence and delivery guarantees, and the client-library decision that quietly matters most.
Kafka Streams
Kafka Streams is a Java library for stream processing that runs inside your application, with state in local RocksDB stores and changelog topics. The operational realities, and where Flink wins.
Kafka Streams documentation, mapped
The Kafka Streams documentation divides into four layers: API reference, configuration surface, state store internals, and the upgrade guide. Which layer answers which production question.
Kafka topics
A Kafka topic is a named, append-only log, and in production it is the unit you operate on. The CLI verbs, partition and replication mechanics, retention, compaction and the troubleshooting moves.
A Kafka topic example, fully specified
A production Kafka topic example: the exact creation command with durability properties, the same topic as Terraform and Strimzi code, the naming convention that scales, and the schema contract.
Kafka topic vs partition
A topic is the logical name; a partition is the physical log. How replication, ordering, parallelism and key hashing follow the physical unit, and why keyed topics lock their count.
A production Kafka tutorial
A Kafka tutorial for people who run it in production: zero-downtime upgrades, broker tuning, layered security, troubleshooting signals, and the client settings that decide delivery guarantees.
Kafka use cases
The production Kafka use cases with the numbers behind them: real-time analytics, event-driven microservices, change data capture, event sourcing and log aggregation, plus the business case for each.
Kafka KRaft
KRaft replaces ZooKeeper with a Raft quorum built into Kafka. The migration path and deadlines, controller sizing, the real scale limits, and day-2 operations for KRaft clusters.
ksqlDB on Kafka
ksqlDB is the streaming SQL layer for Kafka: streams and tables in SQL, persistent queries with state in internal topics. The push-vs-pull trap, and where it stands as Confluent shifts to Flink.
Kafka with Spring Boot
Spring Boot integrates with Kafka through spring-kafka: KafkaTemplate, @KafkaListener and auto-configuration. Production setup, dead letter topics, non-blocking retries, and listener tuning.
Kafka DLQs per consumer group: a consumer-side, Kafka-only design
Why each Kafka consumer group should own its DLQ and retry topic, why replay never goes back to the main topic, and why retrying transient downstream failures is an anti-pattern.
What is Apache Kafka?
Apache Kafka is an open-source distributed event streaming platform that stores records in ordered, partitioned, replayable logs. This is what it is, how the pieces fit, and when to choose it.
Best free Kafka UI tools in 2026
Compare the best free Kafka UI and management tools in 2026: Kpow Community Edition, Conduktor Console Community, Lenses Community Edition, AKHQ, and Kafbat UI.
Kafdrop: pricing and alternatives
Kafdrop review for 2026: strengths, limitations, pricing, and the best alternatives for platform and data engineers running production Kafka clusters.
Kafka dashboard: features that matter
A Kafka dashboard gives you real-time visibility into consumer lag, broker health, and partition state. Here's what to look for and how Kpow delivers it in production.
Kafka management console: what to look for
A Kafka management console gives your team full control of topics, consumers, schemas, and connectors from one UI. See what to look for and how Kpow delivers it.
Kafka message key best practices
A technical guide to Kafka message key best practices covering partitioning, ordering guarantees, hot keys, log compaction, and serialization for production systems.
Kafka security architecture for production
Kafka ships insecure by default. Learn how to build a production-ready Kafka security architecture covering TLS encryption, SASL authentication, ACLs, audit logging, and network isolation.
Kafka UI: The Ultimate Guide
A Kafka UI is a web interface for managing Apache Kafka, giving operators visual control over topics, consumers, brokers, and connectors without the CLI.
Dead letter queues in Kafka: patterns and pitfalls
How to implement a dead letter queue in Apache Kafka, with Spring Kafka, Connect, and Streams examples, and the production failure modes to avoid.
Best Kafka management tools for 2026
Compare the 10 best Kafka management tools for 2026, including Kpow, AKHQ, Conduktor, and Confluent Control Center. Covers pricing, RBAC, and deployment requirements.
Best Kafka monitoring tools for 2026
Compare 12 Kafka monitoring tools for 2026, from enterprise-grade Kpow to open-source AKHQ and Prometheus. Covers deployment, pricing, and key trade-offs.
Kafka broker monitoring
How to monitor Kafka brokers and the cluster they form: key JMX metrics, alert rules, capacity signals, health check scripts and step-by-step diagnosis.
Kafka consumer performance tuning: the config changes that reduce lag
Learn which Kafka consumer metrics matter most, how to interpret them, and which configuration changes will improve performance and reduce lag.
Kafka performance monitoring: metrics that matter
The Kafka metrics that matter when something breaks: the ten critical broker, replication, JVM and storage metrics, alert thresholds, and KRaft changes.
A detailed guide to Kafka producer monitoring
A practical guide to Kafka producer metrics, JMX collection, alerting thresholds, and diagnostic scripts for Java-based Kafka producers.
How Apple uses Apache Kafka in production
A deep-dive into Apple's Kafka architecture, covering their managed internal platform, Strimzi on EKS, tiered storage, zero-data-movement balancing, and mTLS migration.
How Barclays uses Apache Kafka in production
A deep-dive into Barclays' Kafka architecture, covering dual-environment deployment on AWS and IBM Z-Linux, operating practices, and the broader streaming stack.
How Datadog uses Apache Kafka in production
A deep-dive into Datadog's Kafka architecture, covering use cases, scale, engineering decisions, and key contributors across hundreds of clusters.
How Netflix uses Apache Kafka in production
A deep-dive into Netflix's Kafka architecture, covering the Keystone pipeline, Data Mesh platform, scale figures from 700 billion to 2 trillion events per day, and the engineering decisions behind it.
How New Relic uses Apache Kafka in production
A deep-dive into New Relic's Kafka architecture, covering use cases, scale, engineering decisions and key contributors.
How Notion uses Apache Kafka in production
A deep-dive into Notion's Kafka architecture, covering use cases, scale, engineering decisions, and key contributors across their data lake and AI pipelines.
How PagerDuty uses Apache Kafka in production
A deep-dive into PagerDuty's Kafka architecture, covering event ingestion, notification scheduling, task execution, and the engineering decisions behind each.
How Pinterest uses Apache Kafka in production
A deep-dive into Pinterest's Kafka architecture, covering use cases, scale, engineering decisions, and key contributors. From 15 million to 40 million messages per second across 3,000 brokers.
How Salesforce uses Apache Kafka in production
A deep-dive into Salesforce's Kafka architecture, covering use cases, scale, engineering decisions and key contributors across a fleet of 100+ clusters processing 3+ trillion events per day.
How Shopify uses Apache Kafka in production
A deep-dive into Shopify's Kafka architecture, covering CDC at 100,000 records/sec, Kubernetes deployment, the Sarama Go client library, and BFCM scale engineering.
How Tencent uses Apache Kafka in production
A deep-dive into Tencent's Kafka architecture, covering their federated cluster design, 20 trillion messages per day, KIP contributions, and tiered storage at Tencent Cloud.
How Wix uses Apache Kafka in production
A deep-dive into Wix's Kafka architecture: 66 billion daily messages, 2,200+ microservices, the Greyhound SDK, Confluent Cloud migration, and operating 500,000+ partitions across 4 regions.
Conduktor: pricing and alternatives
Conduktor review for 2026: pricing, strengths, deployment trade-offs, and how it compares to alternatives for enterprise Kafka governance teams.
How Airbnb uses Apache Kafka in production
A deep-dive into Airbnb's Kafka architecture, covering six production systems, 35+ billion daily events, SpinalTap CDC, Flink-based personalisation, and Kafka as a write-ahead log.
How Bytedance uses Apache Kafka in production
ByteDance ran Kafka at tens of TB/s before replacing it with ByteMQ, a Kafka-compatible platform separating storage from compute. How it works, and why migrating cut resource cost by roughly 70%.
How Cloudflare uses Apache Kafka in production
A deep-dive into Cloudflare's Kafka architecture: use cases at trillion-message scale, 14 clusters, internal tooling decisions, and the engineering lessons behind a decade of Kafka operations.
How DoorDash uses Apache Kafka in production
A deep dive into DoorDash's Kafka architecture, covering the Iguazu event platform, Flink-based ML feature pipelines, self-serve topic governance, and hundreds of billions of daily events.
How Goldman Sachs uses Apache Kafka in production
A deep-dive into Goldman Sachs's Kafka architecture, covering use cases across three divisions, migration to Amazon MSK, resilience design, and key engineering decisions.
How Grab uses Apache Kafka in production
A deep-dive into Grab's Kafka architecture, how the Coban team built a terabyte-per-hour streaming platform serving 300 billion events a week across GrabFood, GrabPay, mobility, and more.
How JPMorgan uses Apache Kafka in production
A deep dive into JPMorgan Chase's Kafka architecture, covering multi-tenant cluster design, managed Kafka Connect, the Photon Framework, and decisions behind a large-scale deployment.
How LinkedIn uses Apache Kafka in production
A deep-dive into LinkedIn's Kafka architecture, covering use cases, scale, engineering decisions, and key contributors.
How PayPal uses Apache Kafka in production
A deep-dive into PayPal's Kafka architecture, covering use cases, scale, engineering decisions, and key contributors across a fleet handling 1.3 trillion messages per day.
How Reddit uses Apache Kafka in production
A deep-dive into Reddit's Kafka architecture, covering use cases, scale, engineering decisions and key contributors.
How Robinhood uses Apache Kafka in production
A deep-dive into Robinhood's Kafka architecture: use cases, scale, and engineering decisions. Robinhood processes 2.2 million messages per second across equities, crypto, and fraud detection.
How Spotify used Apache Kafka in production
A deep-dive into Spotify's Kafka architecture, covering their event delivery system, 700K events/second scale, engineering decisions, and why they ultimately migrated to Google Cloud Pub/Sub.
How The New York Times uses Kafka
A deep-dive into The New York Times' Kafka publishing pipeline, covering the Monolog architecture, single-partition design, Kafka Streams usage, and treating Kafka as a permanent content store.
How Uber uses Apache Kafka in production
A deep-dive into Uber's Kafka architecture - covering use cases, scale, engineering decisions, and key contributors. From one region to trillions of messages a day.
How Walmart uses Apache Kafka in production
A deep-dive into Walmart's Kafka architecture, covering real-time inventory, fraud detection, the Customer Data Platform, and the Messaging Proxy Service handling trillions of messages per day.
Data lineage support in Factor Platform
Learn how Factor Platform brings OpenLineage metadata into your Kafka environment, making data ownership, PII classification, and lineage visible by default.
Kafka scaling best practices: An in-depth primer
A practical guide to scaling Apache Kafka in production, covering partitioning strategy, consumer group design, broker sizing, KRaft migration, and more.
AKHQ: pricing and alternatives
AKHQ review for 2026: features, known limitations, pricing, and the best alternatives for teams that need more than open-source tooling.
CMAK: pricing and alternatives
CMAK is a free, open-source Kafka admin tool from Yahoo. This review covers features, KRaft limitations, security gaps, and the best alternatives for 2026.
Confluent Control Center: pricing and alternatives
An honest technical review of Confluent Control Center in 2026, covering features, deployment, pricing, and the best alternatives for Kafka teams.
Kadeck: pricing and alternatives
Kadeck review for 2026: features, deployment, pricing, and how it compares to AKHQ, Kafbat, Conduktor, and Kpow for Kafka management teams.
Kafbat UI: pricing and alternatives
A practical review of Kafbat, the open-source kafka-ui fork: features, deployment, security, pricing, and the best alternatives in 2026.
Lenses.io review: pricing and alternatives
Lenses.io review for 2026: honest assessment of SQL Studio, deployment complexity, pricing, and when to consider alternatives like Conduktor or Kpow.
Redpanda Console: pricing and alternatives
Redpanda Console reviewed for 2026: features, pricing, limitations, and the best alternatives for engineering teams running Apache Kafka or Redpanda.
Top Kafka UI tools in 2026: a practical comparison
Honest comparison of Kafka UI tools for enterprise teams, evaluating AKHQ, Kafbat, Redpanda Console, Conduktor, Confluent Control Center, and Kpow.
Best practices for Kafka data observability
12 best practices for Kafka data observability covering consumer lag monitoring, schema enforcement, end-to-end auditing, DLQs, and lineage, with an implementation roadmap.
Kafka message size best practice
Kafka's default max message size is about 1 MB (message.max.bytes is 1,048,588 bytes). Which configs to raise, cloud ceilings, and when to use claim-check.
Kafka partition key best practices
How Kafka partition keys work, what makes a good key, and practical guidance on cardinality, hot partitions, compaction, cross-language hashing, and safe key migration.
Kafka cluster management: a practical guide
A practical guide to Kafka cluster management: architecture sizing, day-to-day operations, performance tuning, KRaft migration, and monitoring for production clusters.
Kafka topic partition best practices
Size Kafka topic partitions correctly from day one. Covers the throughput formula, the keyed topic asymmetry, KRaft-era limits, and operational best practices.
The complete guide to Kafka change data capture
Learn how to implement change data capture with Kafka using Debezium. Includes working PostgreSQL CDC examples, architecture patterns, and monitoring.
What the IBM-Confluent deal means for Kafka users
IBM's $11B Confluent acquisition raises questions for Kafka users. Assess your lock-in risk across Schema Registry, managed connectors, and operational tooling.
How to prove Kafka compliance in an audit
What Apache Kafka records on its own, the evidence an auditor asks for, and where the EU Data Act fits, with each regulatory point cited to the Regulation itself.