News, guides, and engineering deep dives.
Practical guidance on Kafka, Flink, Iceberg, and real-time data.
How Airbnb uses Apache Iceberg in production
How Airbnb moved its Kafka-fed warehouse ingestion pipeline off Tez and Hive onto Spark 3 and Apache Iceberg, consolidating tables via partition-spec evolution and cutting compute by over 50%.
How Apple uses Apache Iceberg in production
How Apple built GDPR- and DMA-compliant row-level deletes on Iceberg tables holding tens of petabytes of data, sourced from a peer-reviewed paper by eight Apple engineers and named conference talks.
How LinkedIn uses Apache Iceberg in production
How LinkedIn's OpenHouse control plane governs 300,000+ Apache Iceberg tables holding over an exabyte of data, with automated compaction and cross-region replication.
How Netflix uses Apache Iceberg in production
How Netflix moved its S3 data warehouse off Hive onto an Iceberg-only architecture, and how Maestro, Psyberg, and its Cassandra engine lean on Iceberg's own snapshot and partition metadata.
How Pinterest uses Apache Iceberg in production
How Pinterest rebuilt CDC ingestion and ML feature backfills on Apache Iceberg, cutting pipeline latency from 24+ hours to minutes and backfill time from 140 days to 26.
How Airbus uses Apache Flink in production
How Airbus's AirSense unit processes more than 2 billion aircraft-position events a day with Apache Flink, fusing multi-provider ADS-B feeds into real-time global flight tracking.
How Alibaba uses Apache Flink in production
How Alibaba built Blink, its internal Apache Flink fork, to handle 4 billion records a second during Double 11, then merged its isolation and checkpointing innovations into Apache Flink 1.9 and 1.10.
How Booking.com uses Apache Flink in production
How Booking.com's Security Platform Services team runs Apache Flink as the engine behind an internal security-as-a-service platform, scaling to more than 250 jobs under Ververica Platform.
Apache Flink: the complete guide
Apache Flink is a distributed stream processing framework for stateful computation over unbounded and bounded data. This hub indexes what we have written about running it in production.