Posts
All the articles I've posted.
- 28 MIN READ•Jul 6, 2026
Federation and the Lakehouse: Two Roads to Unified Data Access, and How to Know Which One to Take
Every data strategy document written this decade contains some version of the same sentence: we need a single place to access all our data. The sen...
data federationdata lakehouseunified access - 27 MIN READ•Jul 6, 2026
A Deep Dive Into File Compression: How Data Gets Smaller, Why Codecs Differ, and What to Actually Use in the Lakehouse
Somewhere in your data platform right now, a single configuration property is quietly deciding a meaningful percentage of your storage bill, your q...
file compressioncodecsparquet - 27 MIN READ•Jul 6, 2026
The File Format Renaissance: Parquet, Lance, Vortex, Nimble, BtrBlocks, and the New Physics of Columnar Storage
For a decade, the file format layer was the most settled real estate in data. Apache Parquet held the analytical world, ORC held the Hive legacy es...
parquetlancevortex - 15 MIN READ•Jul 6, 2026
Enforcing Fine-Grained Security at Machine Speed: Dynamic Access Control for High-Frequency AI Agents
AI agents change the security model for analytics. A human user may run a handful of queries, pause, interpret the answer, and ask a follow-up. An...
fine-grained securityai agentsaccess control - 16 MIN READ•Jul 6, 2026
Implementing Positional Deletes in Iceberg v3: Streamlining Merge-on-Read for Fast-Inbound Event Lakes
Event data has a way of humbling neat architecture diagrams. It arrives late. It arrives twice. It arrives with incorrect attributes. It needs priv...
iceberg v3positional deletesevent lakes - 16 MIN READ•Jul 6, 2026
Mapping the Variant Type in Iceberg v3: Standardizing Semi-Structured AI JSON Payloads
AI applications are messy data producers. They create prompts, completions, tool calls, retrieval traces, ranking signals, evaluation scores, safet...
iceberg v3variant typejson - 15 MIN READ•Jul 6, 2026
Designing Idempotent Pipelines in the Agentic Lakehouse: Eliminating Double-Write Anomalies
Agents retry. Networks fail. Jobs time out after doing some work. APIs return ambiguous responses. Schedulers run the same workflow twice. A human...
idempotent pipelinesagentic lakehousedouble-write - 27 MIN READ•Jul 6, 2026
File Encryption for the Lakehouse: The Terminology, the Machinery, and the Hard Problem of Interoperable Encrypted Tables
For years, the open lakehouse had an honest gap that practitioners whispered about and slide decks skipped: encryption. Not the checkbox kind, ever...
encryptiondata lakehousesecurity - 15 MIN READ•Jul 6, 2026
The 2026-07-28 Model Context Protocol Release Candidate: What the Stateless Spec Means for Data Platforms
The date in this topic matters. Today is July 6, 2026. A release candidate dated July 28, 2026 is still in the future. That means this article shou...
model context protocolmcpdata platforms - 15 MIN READ•Jul 6, 2026
The Metric Contract Mandate: Standardizing Semantic Layers Before AI Agent Access
AI agents are very good at moving quickly. That is the opportunity and the risk. If an agent can inspect metadata, generate queries, compare result...
semantic layersmetric contractsai agents - 15 MIN READ•Jul 6, 2026
Multi-Engine Catalog Federation with Apache Polaris: Syncing Google Cloud, AWS, and Azure Metadata
Open table formats changed the data lakehouse conversation, but they did not finish it. A table can be stored in an open format and still be hard t...
apache polariscatalog federationmulticloud - 28 MIN READ•Jul 6, 2026
Open Source Foundations, Explained: What Apache, Linux, Eclipse, and Their Peers Actually Do, and Why Governance Differences Matter
Writing about open data and AI means repeating the same phrases over and over: donated to the Apache Software Foundation, incubating at the Linux F...
open sourcefoundationsapache