7:50 Β· Music video
The Iceberg Open Lakehouse
A visual tour of the open data architecture beneath modern analytics and AI.
Open videoA working reference for data engineers. 200 terms, each with a definition you can use in a meeting and an explanation you can build from.
Pipelines, storage, modeling, orchestration, governance, and the lakehouse and AI architectures they are all converging on. Written and maintained byAlex Merced, co-author of Apache Iceberg: The Definitive Guide.
Featured films
Music-driven explainers for the open lakehouse and the open agentic stack.
7:50 Β· Music video
A visual tour of the open data architecture beneath modern analytics and AI.
Open video7:37 Β· Music video
A manifesto for composable agents, open standards, model choice, and portable execution.
Open videoThe terms most people look up first. Each one links out to the concepts around it, so you can follow a thread as far as you need to.
A data engineer builds and operates the systems that move data from the places it is created into the places it gets analyzed. That covers ingestion, storage, transformation and serving, plus the orchestration, testing and monitoring that keep those pipelines producing correct results on every run, not just the first one.
Read the full entryETL transforms data on a separate tier before loading it, so the destination only ever stores modeled tables. ELT loads raw data into the warehouse or lakehouse first and transforms it there using the platform's own compute. ELT keeps a replayable raw copy; ETL discards it at transform time.
Read the full entryA data lakehouse is open files on object storage with a table format layered over them. The table format, usually Apache Iceberg, tracks every data file in metadata rather than inferring a table from a directory, which restores ACID transactions, schema evolution and time travel to data that any engine can still read.
Read the full entryBatch is enough whenever the business can act on data that is minutes or hours old, and it is cheaper and simpler to operate. Streaming earns its complexity when a decision has to happen within seconds of the event. Micro-batching sits between the two and covers a large share of real requirements.
Read the full entryStart with how analytical data is shaped and where it lives: data modeling, the difference between a warehouse, a lake and a lakehouse, and how a pipeline is scheduled. Those three ideas explain most architecture decisions you will meet, and every other term connects back to them.
Read the full entryWhy a shared metric layer sits above the tables rather than inside each BI tool.
Read ArticleHow an open catalog keeps the lakehouse readable by more than one engine.
Read ArticleWhat a table format does, and the problems it was invented to fix.
Read ArticleSeparating real Iceberg support from a connector with an Iceberg label on it.
Read ArticleWhat changes when an agent, not a person, is the one asking the questions.
Read ArticleCommunity Slack for data lakehouse practitioners and enthusiasts.
Slack community for organizing and coordinating data events.
Join the Dremio developer community for support and collaboration.
Direct Slack workspace for Dremio developers.
Official Apache Iceberg open source community.
Official Apache Polaris open source community.
Official Apache Arrow open source community.
Register for upcoming Dremio workshops, webinars, and hands-on sessions.
Luma calendar for worldwide lakehouse events and meetups.
Global meetup group for open data lakehouse enthusiasts.
Connect with lakehouse practitioners across North America.
Community-organized Apache Iceberg meetups across North America.
Lakehouse linkup events calendar on Luma.
Data lakehouse events in the New York City area.
Data community events in the Orlando, Florida area.
Events in SF, Seattle, Denver, and more.
Events in NYC, Boston, Atlanta, Austin, and more.