Welcome!
RSS FeedIceberg Lakehouse is the technical encyclopedia for Apache Iceberg, lakehouse catalogs, the Agentic Lakehouse, and modern data architecture. Whether you are learning what table formats are, how to deploy Apache Polaris, or how to connect engines to Iceberg tables, you will find the definitive reference material here.
This blog is not affiliated with the Apache Foundation or the Apache Iceberg project whose official page is iceberg.apache.org.
Join the Data Lakehouse Hub Slack Community: Join Now!
Subscribe to our calendar of Data Lakehouse events: Subscribe!
Recent Posts
- 21 MIN READ•Aug 4, 2026
Budgeting for Agentic Analytics When Every Question Costs Something Different
Budgeting for agentic analytics when every question costs something different: token economics, query economics, instrumentation, and the cost controls that actually return.
AI AgentsTCOCost Management - 21 MIN READ•Aug 4, 2026
The Five Layers of an Agentic Lakehouse and Where the MCP Server Sits
The five layers of an agentic lakehouse and where the MCP server sits: storage, catalog, semantic layer, MCP gateway, and agent surface, plus identity, session isolation, and budgets.
AI AgentsMCPAgentic Lakehouse - 21 MIN READ•Aug 4, 2026
Autonomous Table Optimization When Your Query Workload Stops Being Predictable
Autonomous table optimization when query workloads stop being predictable: observing file layout and query patterns, scoring compaction work, adaptive sort order, and cost discipline.
Apache IcebergTable OptimizationCompaction - 21 MIN READ•Aug 4, 2026
Building Apache Iceberg Lakehouses That Run Without an Internet Connection
How to build an Apache Iceberg lakehouse that runs fully offline: storage, catalog, compute, cross-zone transfer, compliance, and the failure modes that bite.
Apache IcebergAir-GappedOn-Premises
Must Reads on Iceberg, Agentic AI and Lakehouse from Around the Web
-
The Definitive Guide to the Semantic Layer
Understand what a semantic layer is, why it matters for modern data architectures, and how it creates a consistent, governed layer between raw data and business consumers.
Read Article -
Apache Polaris: The Catalog Standard for Lakehouses and AI
A deep dive into Apache Polaris, the open-source catalog that is emerging as the standard for managing Iceberg tables across multi-engine Lakehouses and AI workloads.
Read Article -
What Are Table Formats and Why Were They Needed?
Explore the history and motivations behind open table formats like Apache Iceberg, Delta Lake, and Apache Hudi, and why they solved critical problems in big data engineering.
Read Article -
What is Dremio?
A comprehensive overview of Dremio's Lakehouse platform — how it unifies data access, accelerates queries, and powers self-service analytics across cloud and on-premise sources.
Read Article -
What Apache Iceberg Native Actually Means
Not all Iceberg integrations are equal. This article breaks down what it truly means for a platform to be 'Apache Iceberg native' and why the distinction matters for your architecture.
Read Article -
Open Source and the Data Lakehouse
A survey of the open source ecosystem powering modern Data Lakehouses — from Apache Iceberg and Nessie to Apache Arrow and Spark — and how they work together.
Read Article -
What is Agentic Analytics?
Discover how AI agents are transforming analytics pipelines — autonomously querying data, generating insights, and taking actions — and what it means for the future of the Lakehouse.
Read Article