Skip to content

Best tools to monitor Spark Structured Streaming jobs that read Kafka

Comparisons
Chad Harris·October 6, 2026·9 min read

To monitor a Spark Structured Streaming job that reads Kafka, send the query’s progress reports to a metrics store with a StreamingQueryListener. It is the one option here that carries offsets behind latest for each Kafka source on any Spark deployment, and it ranks first on this page’s five-criterion rubric. Factor House makes Kpow for Kafka and Flex for Flink, has no Spark product, and does not appear in the scoring.

Tools compared

Five ways to monitor a Spark Structured Streaming job that reads Kafka, scored on the five criteria below (read 6 October 2026). Total is the unweighted score out of 50, and the criteria are explained under how these tools were scored. Factor House has no Spark product, so none of its products is scored here.
Rank Tool Total (out of 50) Query rates and batch timing Offsets behind latest per Kafka source History after a job stops Alerting and push Any Spark deployment
1 StreamingQueryListener to your own metrics store 43 Input and process rate per batch Min, max and average per source As long as your store keeps it Callback per report, rules are yours Yes, needs code in the job
2 Databricks Streaming tab and listener 31 Streaming tab and progress fields Min, max, average and estimated bytes Not described on the page Listener push to external services No, Databricks only
3 Spark metrics to Prometheus 26 Input rate, processing rate, latency Not in the metric list Whatever Prometheus retains Metrics only, rules are yours Yes, sink marked experimental
4 Spark History Server 25 UI rebuilt from event logs Not shown Completed and incomplete applications None Yes, needs a server and event logs
5 Spark web UI Structured Streaming tab 22 Rates, batch and operation duration Not shown Only while the application runs None Yes, built in

The tools, ranked for monitoring Spark and Kafka

Rank 1

StreamingQueryListener to your own metrics store

spark.apache.org

43 out of 50 Total

What it is
A Spark API that calls your code when a query starts, stops or reports progress
Cost
No license, plus the code and the store you run
Query rates and batch timing
9 out of 10
Offsets behind latest
9 out of 10
History after a job stops
8 out of 10
Alerting and push
8 out of 10
Runs on any Spark deployment
9 out of 10
Why these scores for StreamingQueryListener to your own metrics store
Query rates and batch timing 9 out of 10
The programming guide’s progress report carries numInputRows, inputRowsPerSecond and processedRowsPerSecond, and the listener receives that report after each batch.
Offsets behind latest 9 out of 10
Spark’s Kafka source adds minOffsetsBehindLatest, maxOffsetsBehindLatest and avgOffsetsBehindLatest to the progress report, per its KafkaMicroBatchStream source, and Databricks documents the same fields.
History after a job stops 8 out of 10
Spark keeps nothing itself here. History lasts as long as the store you send the reports to keeps them, which is why this is not a 10.
Alerting and push 8 out of 10
The listener fires a callback on every progress report, so a rule on the maximum offsets behind latest is a few lines of your own code. Spark ships no alert rules.
Runs on any Spark deployment 9 out of 10
It is part of the Spark API, available on any deployment that runs Structured Streaming. It needs code in the job, which is the cost of this option.

What it shows on your pipeline. The programming guide describes attaching a listener with sparkSession.streams.addListener() and receiving callbacks when a query starts, stops and makes progress. Each progress report names the Kafka source and its end offsets, and the Kafka source adds the offsets-behind-latest metrics. Alert on the maximum, because one stalled partition hides inside the average.

Where it falls short. Someone has to write, test and own the listener and the store behind it. Databricks warns that processing latency in a listener can slow the query, and advises writing to a fast system such as Kafka.

Rank 2

Databricks Streaming tab and listener

databricks.com

31 out of 50 Total

What it is
The Spark UI Streaming tab and StreamingQueryListener on the Databricks platform
Cost
Part of a Databricks subscription
Query rates and batch timing
9 out of 10
Offsets behind latest
9 out of 10
History after a job stops
5 out of 10
Alerting and push
8 out of 10
Runs on any Spark deployment
0 out of 10
Why these scores for Databricks Streaming tab and listener
Query rates and batch timing 9 out of 10
Databricks documents built-in monitoring through the Spark UI Streaming tab, and its progress fields include batchDuration, inputRowsPerSecond and processedRowsPerSecond.
Offsets behind latest 9 out of 10
Its documentation lists sources.metrics.avgOffsetsBehindLatest, maxOffsetsBehindLatest, minOffsetsBehindLatest and estimatedTotalBytesBehindLatest for the Kafka source.
History after a job stops 5 out of 10
The monitoring page describes live metrics and pushing them out. It does not say how long the Streaming tab keeps history, so this scores a middle 5.
Alerting and push 8 out of 10
Databricks says streaming metrics can be pushed to external services for alerting or dashboarding with StreamingQueryListener, available in Databricks Runtime 11.3 LTS and above.
Runs on any Spark deployment 0 out of 10
It runs only on Databricks. A team on its own Spark cannot use it.

What it shows on your pipeline. The Databricks monitoring page shows the Kafka source’s topic, start and end offsets, offsets behind latest and an estimate of the bytes not yet consumed, next to the batch timing of the query.

Where it falls short. It is available only to teams that run on Databricks, and Unity Catalog compute modes add runtime version requirements for the listener.

Rank 3

Spark metrics to Prometheus

spark.apache.org

26 out of 50 Total

What it is
Spark's PrometheusServlet metrics sink, read by Prometheus and charted in Grafana
Cost
No license, plus the Prometheus stack you run
Query rates and batch timing
7 out of 10
Offsets behind latest
0 out of 10
History after a job stops
6 out of 10
Alerting and push
5 out of 10
Runs on any Spark deployment
8 out of 10
Why these scores for Spark metrics to Prometheus
Query rates and batch timing 7 out of 10
With spark.sql.streaming.metricsEnabled set to true, the Spark monitoring guide lists inputRate-total, processingRate-total and latency for Structured Streaming.
Offsets behind latest 0 out of 10
The streaming metrics the guide lists include rates, latency, watermark and state size. Offsets behind latest is not among them.
History after a job stops 6 out of 10
Prometheus keeps time series, but the guide describes only the endpoint, so retention is whatever you configure.
Alerting and push 5 out of 10
Spark exposes the metrics and nothing more. Alert rules are yours to write in your own Prometheus setup.
Runs on any Spark deployment 8 out of 10
Any deployment can serve the endpoint, but the monitoring guide marks the PrometheusServlet experimental and the streaming metrics are off by default.

What it shows on your pipeline. The monitoring guide says the PrometheusServlet serves metrics in Prometheus format from the existing Spark UI. The streaming namespace holds the input rate, processing rate, latency, watermark and state metrics, and applies to Structured Streaming only.

Where it falls short. It gives no per-Kafka-source lag. To chart offsets behind latest you still need the listener, or a comparison against the topic’s own end offsets.

Rank 4

Spark History Server

spark.apache.org

25 out of 50 Total

What it is
A web interface that rebuilds the Spark UI from saved event logs
Cost
No license, plus a server and a log directory
Query rates and batch timing
7 out of 10
Offsets behind latest
1 out of 10
History after a job stops
9 out of 10
Alerting and push
0 out of 10
Runs on any Spark deployment
8 out of 10
Why these scores for Spark History Server
Query rates and batch timing 7 out of 10
The monitoring guide says it constructs the UI of an application from its event logs. It does not list which tabs return, so this scores below the live UI.
Offsets behind latest 1 out of 10
It replays what the UI showed, and the Structured Streaming tab’s metrics do not include offsets behind latest.
History after a job stops 9 out of 10
It is the history option in Spark, because it lists incomplete and completed applications and attempts, provided spark.eventLog.enabled was true.
Alerting and push 0 out of 10
It is a viewer and sends nothing anywhere.
Runs on any Spark deployment 8 out of 10
It starts with sbin/start-history-server.sh on any deployment, but needs the server and a shared event log directory.

What it shows on your pipeline. The monitoring guide says to set spark.eventLog.enabled to true before starting the application, then start the history server, which serves on port 18080 by default.

Where it falls short. A streaming job writes one very long event log. The guide recommends rolling event logs with compaction for long-running applications, and the history server stays a viewer you open after something has gone wrong.

Rank 5

Spark web UI Structured Streaming tab

spark.apache.org

22 out of 50 Total

What it is
The web UI built into every running Spark application, on port 4040 by default
Cost
Included with Spark
Query rates and batch timing
9 out of 10
Offsets behind latest
1 out of 10
History after a job stops
2 out of 10
Alerting and push
0 out of 10
Runs on any Spark deployment
10 out of 10
Why these scores for Spark web UI Structured Streaming tab
Query rates and batch timing 9 out of 10
The web UI documentation lists Input Rate, Process Rate, Input Rows, Batch Duration and Operation Duration for each streaming query.
Offsets behind latest 1 out of 10
None of the listed tab metrics is offsets behind latest, so lag has to be inferred from the input and process rates.
History after a job stops 2 out of 10
The monitoring guide says the information is available only for the duration of the application by default.
Alerting and push 0 out of 10
It is a page to look at, with no alerting.
Runs on any Spark deployment 10 out of 10
It is built into every Spark application.

What it shows on your pipeline. The web UI documentation describes the Structured Streaming tab: a table of active and completed queries, and a statistics page per run id with rates, batch duration and the time taken by operations such as addBatch and latestOffset.

Where it falls short. It disappears with the application unless event logging is on, and it cannot tell you how far behind a Kafka partition the query is.

What a Spark job reading Kafka needs from monitoring

A Spark job that reads Kafka fails in three ways: it falls behind the topic, it stops making progress, or it loses data because retention deleted records before the query read them. The offsets and lag page explains why none of these shows up as consumer group lag, because the Kafka source commits no offsets. Monitoring has to read the position from Spark itself.

See how far behind each Kafka source is

The offsets behind latest metrics are the closest thing Spark has to consumer lag. Alert on the maximum across partitions, not the average.

Keep a record after the job stops

A job that restarts or fails loses its live UI. History needs event logs, or progress reports kept in a store, or both.

Push reports somewhere that can page someone

Spark exposes metrics. Deciding what counts as too far behind, and who is told, is work for the team running the job.

What the Kafka side adds

Spark’s progress report shows how far the query has read. It does not show retention, the topic’s growth, or what other consumers do with the same topic, so a team running Spark on Kafka usually watches the topic from the Kafka side as well. The consumer lag monitoring tools comparison ranks the tools that do that for consumers that commit offsets.

This page is published by Factor House, which makes Kpow, a Kafka management product. Kpow does not read Spark checkpoints or Spark progress reports, which is why it is not scored above. Factor House’s open source Factor House Local environment includes a Spark Structured Streaming lab that runs beside Kpow, and that lab is the only Spark material Factor House publishes.

FAQ

What is the best tool to monitor Spark Structured Streaming?

On this page’s rubric, a StreamingQueryListener that sends progress reports to a metrics store you already run ranks first with 43 out of 50. It is the only option here that gives offsets behind latest per Kafka source on any Spark deployment, and the alert rules are yours to write.

How do I see Kafka lag for a Spark Structured Streaming job?

Read maxOffsetsBehindLatest from the Kafka source’s metrics in the query progress, or subtract the end offsets in the progress report from the topic’s latest offsets. A Kafka consumer group view will not show it, because the Spark Kafka source does not commit offsets.

Does the Spark web UI show Kafka lag?

No. The web UI documentation lists input rate, process rate, input rows, batch duration and operation duration for the Structured Streaming tab, and offsets behind latest is not among them.

Does Factor House offer a Spark monitoring tool?

No. Kpow manages Apache Kafka and Flex manages Apache Flink. Factor House publishes a Spark lab in its open source Factor House Local environment and nothing else for Spark.

How these tools were scored

The five criteria come from the three ways a Spark job reading Kafka fails and what a team needs to act on each one. They carry equal weight, and each is scored 0 to 10: 10 where a tool is the best here and needs nothing extra, 8 or 9 for a documented capability that needs some work from you, 5 or 6 for a partial one or one the documentation does not describe fully, 1 to 4 for a weak or indirect form, and 0 where it is absent. Each score on a card links its reason to the documentation it came from.

1. Query rates and batch timing. Whether the tool shows input rate, processing rate and batch duration for a streaming query.

2. Offsets behind latest per Kafka source. Whether the tool shows how far the query is behind the latest offset of its Kafka source. This is the measure closest to consumer lag.

3. History after a job stops or restarts. Whether the information is still there after the application ends.

4. Alerting and push to other systems. Whether the tool can send its numbers to something that can notify a person. A tool that only exposes a page scores 0.

5. Runs on any Spark deployment. Whether a team running its own Spark, on its own machines or any cloud, can use the tool.

Databricks scores 0 on the fifth criterion because it cannot be used off its own platform. Costs are not scored, because the license price of the open source options is nil and the real cost is the work to run them, which varies too much by team to model from public pages. No Factor House product is scored, because none of them reads a Spark job.

Every option is scored from 0 to 10 on each criterion, from the evidence and sources this page cites, and the reason for each score is on its card. Each criterion counts once, for a total out of 50. The options are listed by total.

Related reading