Posts
All the articles I've posted.
- 30 MIN READ•Aug 24, 2026
Iceberg Row Lineage: The Feature AI and CDC Workloads Will Eventually Depend On
Iceberg row lineage gives rows a durable identity across rewrites. Why CDC pipelines and AI workloads will eventually depend on it.
Apache Icebergrow lineageCDC - 31 MIN READ•Aug 24, 2026
Iceberg v4's Adaptive Metadata Tree, Explained From First Principles
Iceberg v4's adaptive metadata tree, explained from first principles: why commits rewrite too much today and how the tree makes change cheaper.
Apache IcebergIceberg v4metadata - 30 MIN READ•Aug 24, 2026
Why Iceberg v4 Is Really About Making the Cost of Change Proportional to the Change
Iceberg v4 is really about making the cost of a change proportional to the change. The principle, the current tax, and what the redesign pays down.
Apache IcebergIceberg v4metadata - 31 MIN READ•Aug 24, 2026
The Catalog Can Now Plan Your Iceberg Query: Inside REST Scan Planning
A mechanics walkthrough of Iceberg REST scan planning: client-side planning, remote endpoints, pagination, and where engine support stands in 2026.
Apache IcebergREST catalogscan planning - 31 MIN READ•Aug 24, 2026
Can Seven Different Iceberg REST Catalogs Really Run the Same DuckDB Code?
Can seven Iceberg REST catalogs run the same DuckDB script? What the protocol makes portable, what still differs, and a test matrix you can rerun.
Apache IcebergREST catalogDuckDB - 31 MIN READ•Aug 24, 2026
Stop Flattening Your JSON: How Iceberg Variant Changes Semi-Structured Analytics
Iceberg Variant stores JSON as navigable binary with shredding for columnar filters. Why flattening wide tables is no longer the only performance path.
Apache IcebergVariantJSON - 31 MIN READ•Aug 24, 2026
What Actually Happens When Two Engines Write the Same Iceberg Table at Once?
What happens when two engines write the same Iceberg table at once: snapshot isolation, optimistic commits, conflict detection, and when retries fail.
Apache Icebergconcurrencycommits - 31 MIN READ•Aug 24, 2026
Variant Shredding Explained: How Iceberg Gets Columnar Performance From Messy JSON
Variant shredding turns messy JSON into Parquet columns with statistics. How the layout works, how readers reassemble values, and why some queries prune.
Apache IcebergVariantParquet - 30 MIN READ•Aug 24, 2026
Who Actually Owns an Iceberg Table? Managed, External, and the New Vocabulary of Lakehouse Control
Managed and external Iceberg tables mean different things on every platform. Five ownership dimensions and a translation method for vendor vocabulary.
Apache Icebergcatalogsgovernance - 31 MIN READ•Aug 19, 2026
Mastering Apache Iceberg v3 Deletion Vectors for High-Throughput Streaming Ingest
Apache Iceberg v3 deletion vectors for high-throughput streaming ingest: how bitmaps and Puffin files fix CDC write amplification and read decay.
Apache Icebergv3deletion vectors - 31 MIN READ•Aug 19, 2026
The Decoupled Data Lakehouse: Multi-Engine Freedom with Open REST Catalogs
The decoupled data lakehouse: multi-engine freedom with open REST catalogs, credential vending, and an estate that outlives its tools.
REST catalogdecoupled lakehousemulti-engine - 31 MIN READ•Aug 19, 2026
The Five Layers of an Agentic Lakehouse
The five layers of an agentic lakehouse: Storage, Catalog, Semantic, Gateway, and Agent Surface, and how one question travels through all of them.
agentic lakehousearchitectureMCP