Skip to content

Spark Structured Streaming vs Flink: how to choose for Kafka pipelines

Comparisons
Karel Sague·October 6, 2026·8 min read

Spark Structured Streaming and Apache Flink are both distributed engines that read Kafka topics, keep state and write results to a sink. Spark runs a streaming query as a series of micro-batches by default, with an experimental continuous mode, while Flink is built around stateful computation over unbounded data streams. A team already running Spark for batch and lakehouse work usually gets further with Structured Streaming, and a team whose workload is stateful, event-time logic with tight latency usually gets further with Flink. Factor House makes Flex, a Flink management product, and has no Spark product, so weigh the comparison with that in mind.

The two engines at a glance

The Apache Spark documentation says the default micro-batch engine “can achieve exactly-once guarantees but achieve latencies of ~100ms at best”. It describes continuous processing as an experimental mode, introduced in Spark 2.3, that offers about 1 ms end-to-end latency with at-least-once guarantees. The Structured Streaming guide lists end-to-end exactly-once semantics as one of the key goals behind the design.

The Apache Flink documentation describes Flink as “a framework and distributed processing engine for stateful computations over unbounded and bounded data streams”. Its concepts overview says the lowest-level abstraction, the process function, lets users process events from one or more streams freely, provides consistent fault-tolerant state, and lets them register event-time and processing-time callbacks.

F1 Spark Structured Streaming and Apache Flink side by side
Spark Structured Streaming Apache Flink
Default execution Micro-batches, with latencies of about 100 ms at best. Stateful computation over unbounded and bounded data streams.
Low-latency mode Continuous processing, experimental, about 1 ms with at-least-once guarantees. Chosen through application and cluster settings.
Kafka read position Kept in the query checkpoint. The Kafka source commits no offsets. Kept in Flink's state snapshots.
Recovery guarantee End-to-end exactly-once is a stated design goal of the micro-batch engine. At-most-once, at-least-once or exactly-once depending on your choices.
State Managed by the engine for aggregations and windows. Keyed and operator state, with event-time and processing-time callbacks.
Sources: the Apache Spark and Apache Flink documentation, linked in the text, and the Flink fault tolerance guide

How each one reads Kafka

Spark’s Kafka source commits no offsets. The Kafka integration guide says Structured Streaming manages which offsets are consumed internally, and the position lives in the query’s checkpoint. That is why a consumer group view shows no lag for a Spark job, covered in where Spark keeps its position.

Flink also keeps the read position in its own state snapshots. The Flink fault tolerance guide says that depending on the choices you make for your application and the cluster you run it on, a failure can mean lost results, duplicated results, or neither. The Flink side of the same question is in Flink Kafka source offsets and lag.

So on both engines, lag has to be read from the engine and not from Kafka’s committed offsets alone.

When Spark is the better fit

  • The team already runs Spark. Batch jobs, SQL and the lakehouse tooling sit on the same engine, and a streaming query is the same DataFrame API.
  • Micro-batch latency is acceptable. If about 100 ms at best meets the requirement, the default engine gives exactly-once behavior without choosing a mode.
  • The sink is a table format. Spark is well supported by Iceberg. Factor House’s own Lab 8 defines its Iceberg sink table with Spark SQL, because of Flink’s partitioning limitations, while Lab 10 writes Kafka orders to Iceberg from a Spark Structured Streaming job. See the Apache Iceberg guide.
  • The workload is stateful and event-time driven. Flink’s process function and its timers are the core of its API.
  • Latency requirements are tighter than micro-batches allow. Spark’s own documentation puts its default engine at about 100 ms at best, and its roughly 1 ms mode is experimental with at-least-once guarantees.
  • You want to choose the delivery guarantee per application. Flink’s documentation describes the three outcomes and the choices that lead to each.

Operating the two

Both need monitoring of state, checkpoints and lag. For Spark, the monitoring tools ranking scores five options. For Flink, Factor House’s Flex shows a job’s state, checkpoints and events, and the comparison of tools for self-managed Apache Flink ranks the options. Flex does not manage Spark, and neither does Kpow, which manages Kafka. Many teams run both engines on the same Kafka cluster, and Kafka is the common layer. The complete Kafka guide covers it.

FAQ

Is Flink faster than Spark Structured Streaming?

The documentation does not compare them. Spark states about 100 ms at best for its default micro-batch engine and about 1 ms for its experimental continuous mode with at-least-once guarantees. Flink’s documentation describes stateful stream processing without a latency figure, so test your own workload.

Does Spark Structured Streaming support exactly-once with Kafka?

Structured Streaming lists end-to-end exactly-once semantics as a design goal, and the micro-batch engine is described as achieving exactly-once guarantees. The experimental continuous mode offers at-least-once.

Does Factor House support Spark?

No. Factor House’s products manage Kafka (Kpow) and Flink (Flex). Spark appears only in its open source Factor House Local labs.

Which should a Kafka to Iceberg pipeline use?

Either can write Iceberg. Use the engine your team already operates. Factor House’s labs show a Spark job and a Flink job writing Iceberg from Kafka.

Related reading