Tag: Apache Iceberg
All the articles with the tag "Apache Iceberg".
- 14 MIN READ•Jun 8, 2026
Apache Iceberg v3 Deletion Vectors on Snowflake
Deletion vectors matter because row-level changes should not require a full rewrite of every affected data file.
Apache Icebergopen table formatlakehouse - 15 MIN READ•Jun 8, 2026
CDC Without Complexity Using Iceberg v3 Row Lineage
Row lineage gives Iceberg a native way to tell incremental consumers which rows changed and when they changed.
Apache Icebergopen table formatlakehouse - 15 MIN READ•Jun 8, 2026
The 2026 Guide to Iceberg View Federation
Portable views are the missing logic layer between open tables and multi-engine analytics.
Apache Icebergopen table formatlakehouse - 15 MIN READ•Jun 8, 2026
Modern Python Tooling for Apache Iceberg
Python has become a practical Iceberg control plane for metadata work, catalog automation, and smaller operational workflows.
Apache Icebergopen table formatlakehouse - 15 MIN READ•Jun 8, 2026
Bidirectional Iceberg Writes with Horizon Catalog
Bidirectional Iceberg interoperability changes managed Iceberg from a read surface into a shared write contract.
Apache Icebergopen table formatlakehouse - 24 MIN READ•May 23, 2026
Single-Node Data Engineering: DuckDB, DataFusion, Polars, and LakeSail
Optimize single-node data engineering with DuckDB, DataFusion, Polars, and LakeSail. Compare architectures and learn when to transition to Dremio MPP.
DuckDBApache ArrowDataFusion - 20 MIN READ•May 23, 2026
An In-Depth Overview of the Apache Iceberg 1.11.0 Release
Apache Iceberg 1.11.0 delivers manifest list encryption, the new pluggable File Format API, credential lifecycle refreshes, and Spark/Flink improvements.
Apache IcebergData LakehouseOpen Table Format - 22 MIN READ•May 22, 2026
Open Table Format Benchmarks: Why They Require Critical Evaluation
An in-depth analysis of open table format benchmarks comparing Apache Iceberg, Delta Lake, and Apache Hudi, detailing the pitfalls of standard benchmarks and how to choose a format.
open table formatsapache icebergdelta lake - 27 MIN READ•May 22, 2026
Apache Iceberg SCD Type 2 and CDC Patterns: Building Historical Lakehouse Tables
A deep dive into implementing Slowly Changing Dimension Type 2 (SCD Type 2) patterns and Change Data Capture (CDC) pipelines on Apache Iceberg, using PySpark and Dremio.
apache icebergcdcscd type 2 - 25 MIN READ•May 22, 2026
Setting Up an AWS-Native Open Lakehouse: Querying Apache Iceberg with AWS Athena and AWS Glue Catalog
A comprehensive guide to building an open, high-performance lakehouse on AWS using Apache Iceberg, AWS Glue Catalog, Amazon S3, and S3 Tables, with query acceleration via the Dremio engine.
Apache IcebergAWS AthenaAWS Glue Catalog - 24 MIN READ•May 22, 2026
Apache Iceberg Catalogs Explained: REST, Glue, Hive Metastore, Polaris, Nessie, and Snowflake
A deep dive into Apache Iceberg catalog architecture, comparing REST catalogs, AWS Glue, Project Nessie, Polaris, and Snowflake. Learn catalog role, credential vending, and cross-engine configurations.
apache icebergcatalogsNessie - 24 MIN READ•May 22, 2026
Maintaining Apache Iceberg Tables: Compaction, Snapshot Expiration, and Orphan File Cleanup
An in-depth guide to orchestrating maintenance operations on Apache Iceberg tables, covering bin-packing, sort-based, Z-Order compaction, snapshot expiration, and orphan file removal, with query acceleration details for the Dremio engine.
Apache IcebergCompactionData Engineering