News, guides, and engineering deep dives.
Practical guidance on Kafka, Flink, Iceberg, and real-time data.
Industry Apache Kafka 4.3.0: A guide for platform engineers
Kafka 4.3.0 covers broker cordoning, partition size metrics, share group tuning, and tiered storage fixes. Here's what platform engineers need to act on.
Product Data lineage support in Factor Platform
Learn how Factor Platform brings OpenLineage metadata into your Kafka environment, making data ownership, PII classification, and lineage visible by default.
Industry What the IBM-Confluent deal means for Kafka users
IBM's $11B Confluent acquisition raises questions for Kafka users. Assess your lock-in risk across Schema Registry, managed connectors, and operational tooling.
Apache Iceberg: the complete guide
Apache Iceberg is an open table format bringing ACID transactions, schema evolution, and time travel to huge analytic tables in S3. This hub indexes what we've written about running it in production.
Iceberg use cases
Companies running Apache Iceberg in production, the architectures behind their deployments, and what to read next. Indexed as we publish new use-case research.
How Airbnb uses Apache Iceberg in production
How Airbnb moved its Kafka-fed warehouse ingestion pipeline off Tez and Hive onto Spark 3 and Apache Iceberg, consolidating tables via partition-spec evolution and cutting compute by over 50%.
How Apple uses Apache Iceberg in production
How Apple built GDPR- and DMA-compliant row-level deletes on Iceberg tables holding tens of petabytes of data, sourced from a peer-reviewed paper by eight Apple engineers and named conference talks.
How LinkedIn uses Apache Iceberg in production
How LinkedIn's OpenHouse control plane governs 300,000+ Apache Iceberg tables holding over an exabyte of data, with automated compaction and cross-region replication.
How Netflix uses Apache Iceberg in production
How Netflix moved its S3 data warehouse off Hive onto an Iceberg-only architecture, and how Maestro, Psyberg, and its Cassandra engine lean on Iceberg's own snapshot and partition metadata.
How Pinterest uses Apache Iceberg in production
How Pinterest rebuilt CDC ingestion and ML feature backfills on Apache Iceberg, cutting pipeline latency from 24+ hours to minutes and backfill time from 140 days to 26.
How Airbus uses Apache Flink in production
How Airbus's AirSense unit processes more than 2 billion aircraft-position events a day with Apache Flink, fusing multi-provider ADS-B feeds into real-time global flight tracking.
How Alibaba uses Apache Flink in production
How Alibaba built Blink, its internal Apache Flink fork, to handle 4 billion records a second during Double 11, then merged its isolation and checkpointing innovations into Apache Flink 1.9 and 1.10.