Best tools to monitor Spark Structured Streaming jobs that read Kafka
ComparisonsTo monitor a Spark Structured Streaming job that reads Kafka, send the query’s progress reports to a metrics store with a StreamingQueryListener. It is the one option here that carries offsets behind latest for each Kafka source on any Spark deployment, and it ranks first on this page’s five-criterion rubric. Factor House makes Kpow for Kafka and Flex for Flink, has no Spark product, and does not appear in the scoring.
Tools compared
| Rank | Tool | Total (out of 50) | Query rates and batch timing | Offsets behind latest per Kafka source | History after a job stops | Alerting and push | Any Spark deployment |
|---|---|---|---|---|---|---|---|
| 1 | StreamingQueryListener to your own metrics store | 43 | Input and process rate per batch | Min, max and average per source | As long as your store keeps it | Callback per report, rules are yours | Yes, needs code in the job |
| 2 | Databricks Streaming tab and listener | 31 | Streaming tab and progress fields | Min, max, average and estimated bytes | Not described on the page | Listener push to external services | No, Databricks only |
| 3 | Spark metrics to Prometheus | 26 | Input rate, processing rate, latency | Not in the metric list | Whatever Prometheus retains | Metrics only, rules are yours | Yes, sink marked experimental |
| 4 | Spark History Server | 25 | UI rebuilt from event logs | Not shown | Completed and incomplete applications | None | Yes, needs a server and event logs |
| 5 | Spark web UI Structured Streaming tab | 22 | Rates, batch and operation duration | Not shown | Only while the application runs | None | Yes, built in |
The tools, ranked for monitoring Spark and Kafka
StreamingQueryListener to your own metrics store
43 out of 50 Total
- What it is
- A Spark API that calls your code when a query starts, stops or reports progress
- Cost
- No license, plus the code and the store you run
- Query rates and batch timing
- 9 out of 10
- Offsets behind latest
- 9 out of 10
- History after a job stops
- 8 out of 10
- Alerting and push
- 8 out of 10
- Runs on any Spark deployment
- 9 out of 10
Why these scores for StreamingQueryListener to your own metrics store
- Query rates and batch timing 9 out of 10
- The programming guide’s progress report carries numInputRows, inputRowsPerSecond and processedRowsPerSecond, and the listener receives that report after each batch.
- Offsets behind latest 9 out of 10
- Spark’s Kafka source adds minOffsetsBehindLatest, maxOffsetsBehindLatest and avgOffsetsBehindLatest to the progress report, per its KafkaMicroBatchStream source, and Databricks documents the same fields.
- History after a job stops 8 out of 10
- Spark keeps nothing itself here. History lasts as long as the store you send the reports to keeps them, which is why this is not a 10.
- Alerting and push 8 out of 10
- The listener fires a callback on every progress report, so a rule on the maximum offsets behind latest is a few lines of your own code. Spark ships no alert rules.
- Runs on any Spark deployment 9 out of 10
- It is part of the Spark API, available on any deployment that runs Structured Streaming. It needs code in the job, which is the cost of this option.
What it shows on your pipeline. The programming guide describes attaching a listener with sparkSession.streams.addListener() and receiving callbacks when a query starts, stops and makes progress. Each progress report names the Kafka source and its end offsets, and the Kafka source adds the offsets-behind-latest metrics. Alert on the maximum, because one stalled partition hides inside the average.
Where it falls short. Someone has to write, test and own the listener and the store behind it. Databricks warns that processing latency in a listener can slow the query, and advises writing to a fast system such as Kafka.
Databricks Streaming tab and listener
31 out of 50 Total
- What it is
- The Spark UI Streaming tab and StreamingQueryListener on the Databricks platform
- Cost
- Part of a Databricks subscription
- Query rates and batch timing
- 9 out of 10
- Offsets behind latest
- 9 out of 10
- History after a job stops
- 5 out of 10
- Alerting and push
- 8 out of 10
- Runs on any Spark deployment
- 0 out of 10
Why these scores for Databricks Streaming tab and listener
- Query rates and batch timing 9 out of 10
- Databricks documents built-in monitoring through the Spark UI Streaming tab, and its progress fields include batchDuration, inputRowsPerSecond and processedRowsPerSecond.
- Offsets behind latest 9 out of 10
- Its documentation lists sources.metrics.avgOffsetsBehindLatest, maxOffsetsBehindLatest, minOffsetsBehindLatest and estimatedTotalBytesBehindLatest for the Kafka source.
- History after a job stops 5 out of 10
- The monitoring page describes live metrics and pushing them out. It does not say how long the Streaming tab keeps history, so this scores a middle 5.
- Alerting and push 8 out of 10
- Databricks says streaming metrics can be pushed to external services for alerting or dashboarding with StreamingQueryListener, available in Databricks Runtime 11.3 LTS and above.
- Runs on any Spark deployment 0 out of 10
- It runs only on Databricks. A team on its own Spark cannot use it.
What it shows on your pipeline. The Databricks monitoring page shows the Kafka source’s topic, start and end offsets, offsets behind latest and an estimate of the bytes not yet consumed, next to the batch timing of the query.
Where it falls short. It is available only to teams that run on Databricks, and Unity Catalog compute modes add runtime version requirements for the listener.
Spark metrics to Prometheus
26 out of 50 Total
- What it is
- Spark's PrometheusServlet metrics sink, read by Prometheus and charted in Grafana
- Cost
- No license, plus the Prometheus stack you run
- Query rates and batch timing
- 7 out of 10
- Offsets behind latest
- 0 out of 10
- History after a job stops
- 6 out of 10
- Alerting and push
- 5 out of 10
- Runs on any Spark deployment
- 8 out of 10
Why these scores for Spark metrics to Prometheus
- Query rates and batch timing 7 out of 10
- With spark.sql.streaming.metricsEnabled set to true, the Spark monitoring guide lists inputRate-total, processingRate-total and latency for Structured Streaming.
- Offsets behind latest 0 out of 10
- The streaming metrics the guide lists include rates, latency, watermark and state size. Offsets behind latest is not among them.
- History after a job stops 6 out of 10
- Prometheus keeps time series, but the guide describes only the endpoint, so retention is whatever you configure.
- Alerting and push 5 out of 10
- Spark exposes the metrics and nothing more. Alert rules are yours to write in your own Prometheus setup.
- Runs on any Spark deployment 8 out of 10
- Any deployment can serve the endpoint, but the monitoring guide marks the PrometheusServlet experimental and the streaming metrics are off by default.
What it shows on your pipeline. The monitoring guide says the PrometheusServlet serves metrics in Prometheus format from the existing Spark UI. The streaming namespace holds the input rate, processing rate, latency, watermark and state metrics, and applies to Structured Streaming only.
Where it falls short. It gives no per-Kafka-source lag. To chart offsets behind latest you still need the listener, or a comparison against the topic’s own end offsets.
Spark History Server
25 out of 50 Total
- What it is
- A web interface that rebuilds the Spark UI from saved event logs
- Cost
- No license, plus a server and a log directory
- Query rates and batch timing
- 7 out of 10
- Offsets behind latest
- 1 out of 10
- History after a job stops
- 9 out of 10
- Alerting and push
- 0 out of 10
- Runs on any Spark deployment
- 8 out of 10
Why these scores for Spark History Server
- Query rates and batch timing 7 out of 10
- The monitoring guide says it constructs the UI of an application from its event logs. It does not list which tabs return, so this scores below the live UI.
- Offsets behind latest 1 out of 10
- It replays what the UI showed, and the Structured Streaming tab’s metrics do not include offsets behind latest.
- History after a job stops 9 out of 10
- It is the history option in Spark, because it lists incomplete and completed applications and attempts, provided spark.eventLog.enabled was true.
- Alerting and push 0 out of 10
- It is a viewer and sends nothing anywhere.
- Runs on any Spark deployment 8 out of 10
- It starts with sbin/start-history-server.sh on any deployment, but needs the server and a shared event log directory.
What it shows on your pipeline. The monitoring guide says to set spark.eventLog.enabled to true before starting the application, then start the history server, which serves on port 18080 by default.
Where it falls short. A streaming job writes one very long event log. The guide recommends rolling event logs with compaction for long-running applications, and the history server stays a viewer you open after something has gone wrong.
Spark web UI Structured Streaming tab
22 out of 50 Total
- What it is
- The web UI built into every running Spark application, on port 4040 by default
- Cost
- Included with Spark
- Query rates and batch timing
- 9 out of 10
- Offsets behind latest
- 1 out of 10
- History after a job stops
- 2 out of 10
- Alerting and push
- 0 out of 10
- Runs on any Spark deployment
- 10 out of 10
Why these scores for Spark web UI Structured Streaming tab
- Query rates and batch timing 9 out of 10
- The web UI documentation lists Input Rate, Process Rate, Input Rows, Batch Duration and Operation Duration for each streaming query.
- Offsets behind latest 1 out of 10
- None of the listed tab metrics is offsets behind latest, so lag has to be inferred from the input and process rates.
- History after a job stops 2 out of 10
- The monitoring guide says the information is available only for the duration of the application by default.
- Alerting and push 0 out of 10
- It is a page to look at, with no alerting.
- Runs on any Spark deployment 10 out of 10
- It is built into every Spark application.
What it shows on your pipeline. The web UI documentation describes the Structured Streaming tab: a table of active and completed queries, and a statistics page per run id with rates, batch duration and the time taken by operations such as addBatch and latestOffset.
Where it falls short. It disappears with the application unless event logging is on, and it cannot tell you how far behind a Kafka partition the query is.
What a Spark job reading Kafka needs from monitoring
A Spark job that reads Kafka fails in three ways: it falls behind the topic, it stops making progress, or it loses data because retention deleted records before the query read them. The offsets and lag page explains why none of these shows up as consumer group lag, because the Kafka source commits no offsets. Monitoring has to read the position from Spark itself.
See how far behind each Kafka source is
The offsets behind latest metrics are the closest thing Spark has to consumer lag. Alert on the maximum across partitions, not the average.
Keep a record after the job stops
A job that restarts or fails loses its live UI. History needs event logs, or progress reports kept in a store, or both.
Push reports somewhere that can page someone
Spark exposes metrics. Deciding what counts as too far behind, and who is told, is work for the team running the job.
What the Kafka side adds
Spark’s progress report shows how far the query has read. It does not show retention, the topic’s growth, or what other consumers do with the same topic, so a team running Spark on Kafka usually watches the topic from the Kafka side as well. The consumer lag monitoring tools comparison ranks the tools that do that for consumers that commit offsets.
This page is published by Factor House, which makes Kpow, a Kafka management product. Kpow does not read Spark checkpoints or Spark progress reports, which is why it is not scored above. Factor House’s open source Factor House Local environment includes a Spark Structured Streaming lab that runs beside Kpow, and that lab is the only Spark material Factor House publishes.
FAQ
What is the best tool to monitor Spark Structured Streaming?
On this page’s rubric, a StreamingQueryListener that sends progress reports to a metrics store you already run ranks first with 43 out of 50. It is the only option here that gives offsets behind latest per Kafka source on any Spark deployment, and the alert rules are yours to write.
How do I see Kafka lag for a Spark Structured Streaming job?
Read maxOffsetsBehindLatest from the Kafka source’s metrics in the query progress, or subtract the end offsets in the progress report from the topic’s latest offsets. A Kafka consumer group view will not show it, because the Spark Kafka source does not commit offsets.
Does the Spark web UI show Kafka lag?
No. The web UI documentation lists input rate, process rate, input rows, batch duration and operation duration for the Structured Streaming tab, and offsets behind latest is not among them.
Does Factor House offer a Spark monitoring tool?
No. Kpow manages Apache Kafka and Flex manages Apache Flink. Factor House publishes a Spark lab in its open source Factor House Local environment and nothing else for Spark.
How these tools were scored
The five criteria come from the three ways a Spark job reading Kafka fails and what a team needs to act on each one. They carry equal weight, and each is scored 0 to 10: 10 where a tool is the best here and needs nothing extra, 8 or 9 for a documented capability that needs some work from you, 5 or 6 for a partial one or one the documentation does not describe fully, 1 to 4 for a weak or indirect form, and 0 where it is absent. Each score on a card links its reason to the documentation it came from.
1. Query rates and batch timing. Whether the tool shows input rate, processing rate and batch duration for a streaming query.
2. Offsets behind latest per Kafka source. Whether the tool shows how far the query is behind the latest offset of its Kafka source. This is the measure closest to consumer lag.
3. History after a job stops or restarts. Whether the information is still there after the application ends.
4. Alerting and push to other systems. Whether the tool can send its numbers to something that can notify a person. A tool that only exposes a page scores 0.
5. Runs on any Spark deployment. Whether a team running its own Spark, on its own machines or any cloud, can use the tool.
Databricks scores 0 on the fifth criterion because it cannot be used off its own platform. Costs are not scored, because the license price of the open source options is nil and the real cost is the work to run them, which varies too much by team to model from public pages. No Factor House product is scored, because none of them reads a Spark job.
Every option is scored from 0 to 10 on each criterion, from the evidence and sources this page cites, and the reason for each score is on its card. Each criterion counts once, for a total out of 50. The options are listed by total.