Siphon : Modern Data Stack with SF-CH & Iceberg

📺 Click to view the youtube embed player and accept their cookies.
Session Abstract

Tired of waiting for batch jobs? See how we transformed our data pipeline using Apache Iceberg to stream quality data into Snowflake and Clickhouse simultaneously. Learn about our battle-tested architecture, performance gains, and how we maintain data consistency across dual analytics engines

Session Description

Ever wondered how to stream data reliably to multiple warehouses without compromising data quality? We’ll show you how Siphon uses Apache Iceberg’s time travel and ACID properties to ensure data consistency across Snowflake and Clickhouse. Dive into our journey from batch to streaming – covering architecture evolution, data quality frameworks, and performance optimizations. We’ll share our battle-tested patterns for handling schema evolution, managing data contracts, and implementing quality gates. Learn how we achieved sub-minute latency while preventing bad data from corrupting our warehouses. Perfect for data engineers and architects looking to modernize their data infrastructure with real-world proven solutions.

Maschinenhaus
17.Jun 2025
16:00pm - 16:20pm
Short Talk