A data lakehouse is a way to query files in object storage with the reliability you expect from a database. The name merges two older ideas: the data lake, where large datasets live as cheap files, and the data warehouse, where SQL runs fast and consistently. This article explains one concrete open-source stack: Apache Iceberg as the table format, MinIO as S3-compatible object storage, and DuckDB as the query engine. It is written for analysts, developers, and small data teams who want warehouse-style analytics without standing up a warehouse cluster or paying per query.
What is a data lakehouse?
On its own, a folder of files lacks the things that make a database trustworthy: a defined schema, consistent snapshots, and safe changes. Apache Iceberg adds them as metadata layers on top of plain files, usually Parquet. It tracks every version of a table, how it is partitioned, and which files belong to each snapshot, so readers always see a consistent state even while data changes underneath them. MinIO stores those files behind an Amazon S3-compatible API on hardware or a VM you control. DuckDB, an in-process SQL analytics engine, reads Iceberg tables directly over that API and answers queries in seconds. Together the three deliver warehouse behavior on a plain object store, on a single server, with no JVM cluster to feed.
What is a data lakehouse used for?
- SQL on object storage. Query Parquet and CSV files in MinIO with standard SQL instead of writing scripts that walk folders and guess at schemas.
- One dataset, many tools. An Iceberg table can be read by DuckDB today and by Spark or Trino later, without copying or converting the data.
- Time travel and audit. Iceberg snapshots let you query a table exactly as it looked yesterday or last month, for audits and reproducible reports.
- Schema evolution. Add, rename, or reorder columns as source systems change, without rewriting the dataset or breaking older queries.
- Long-term history you can query. Keep years of records in cheap object storage and still reach any month in seconds when questions come up.
- Local analytics on curated data. Analysts query the same tables from a laptop with DuckDB, with no warehouse account and no queue behind a shared cluster.
Key features
- An open table format. Iceberg is engine-neutral, so Spark, Trino, Flink, DuckDB, and most cloud warehouses can read the same tables.
- Atomic snapshots. Commits are all-or-nothing, so readers never see a half-written table, even during load jobs.
- Hidden partitioning. Partitioning lives in metadata rather than in folder paths, which keeps queries fast and lets you change the layout later without rewriting data.
- S3-compatible storage. MinIO provides Amazon S3-style APIs on your own server, with erasure coding for durability.
- Fast in-process SQL. DuckDB is a columnar analytics engine that runs inside a container or a laptop, with no separate cluster to operate.
- Time travel. Historical snapshots stay available for auditing, reproducibility, and easy rollbacks of bad writes.
Data Lakehouse on OpenSysLab
Every OpenSysLab server is a private virtual machine in the region you chose at purchase. The app installs automatically in about ten minutes, and the Vibe-code agent inside Open WebUI can read, configure, and manage it for you — you just connect your own AI API key. Servers start at $6/month. See the Data Lakehouse product page for plans and regions.
The bottom line
A lakehouse does not assemble or maintain itself, and it will not replace a managed warehouse at every scale. But if you want open formats, your own storage, and fast SQL without a data platform team, Iceberg with MinIO and DuckDB is a pragmatic way to start. Begin with one bucket, one catalog, and a few tables, then grow from there.