If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
eddietejeda 2 days ago [-]
Thanks for the shoutout.
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
prpl 2 days ago [-]
What latencies were you targeting?
eddietejeda 2 days ago [-]
Extremely high concurrency at sub-second response times.
I always thought the catalogue was a duckdb file. E.g, data lives in partitioned parquet files, but which parquet files are current or soft deleted, etc, etc, is managed in a duckdb data file.
However, looking at https://ducklake.select/, it seems the catalogue lives in PostgresSQL - so it is not really a ducklake, but a postgresslake.
The more you know.
pdet 1 days ago [-]
The fundamental requirements of DBMS so it can be a DuckLake catalog are the following:
1. It must support primary keys
2. It must support basic data types (INTEGER,VARCHAR,TIMESTAMP)
3. ACID
There are ofc some quirks from the SQL supported on each DBMS. Hence, the DuckLake extension from DuckDB currently supports DuckDB, SQLite, Postgres, DuckDB + Quack and MySQL. With MySQL being in a rather experimental state.
Disclaimer: I'm the lead developer in DuckLake.
paragraft 2 days ago [-]
That's a tabbed interface on the site that just defaults to postgres. SQLite and duckdb are supported too.
celias 3 days ago [-]
Motherduck is offering a free copy of O'reilly's "DuckLake: The Definitive Guide" book on their DuckLake web page
It's alright, it's pretty alpha software. On v1.5.4, catalog filtered counts are broken, afaik. I went to main/v2 to fix it, and then the SQL parser in duckdb v2 is 10x slower, which was another wrench in the gears. It's been a bit of a pain tbh
jauco 2 days ago [-]
Yep, they made the spec 1.0 but it isn’t 1.0 software. Browse the bugs before use.
When it works well it’s really nice. And it beats handrolling a multi level parquet store.
engineeringwoke 2 days ago [-]
Absolutely. I love it, but you need a fork for now.
tomwphillips 2 days ago [-]
I think this is an elegant design that's superior to the competition, but I think lakehouses are not as generally useful as vendors would like us to believe. The access controls are limited to what's possible on the underlying bucket.
For example I think a lakehouse is a bad choice for standard enterprise BI type analytics - you've got no column or row access controls, and no column masking. I don't see how this could ever be bolted on to the bucket and catalog.
I think you can model this as locked down buckets, wider engine access (spark/presto/etc), and apply the controls at the engine level. (it's not inherently different from making sure your DB files are locked down, if you squint at it). This does obviously block any non-engine access which has downsides. I agree that it is generally much less mature with lakehouses than eneterprise DBs.
tomwphillips 1 days ago [-]
You could, but the whole idea of lakehouse is that you can use whatever engine suits your needs. Now you've got to make sure that every engine enforces your access control policies properly. It just seems like a lot of work.
snapetom 2 days ago [-]
Is this basically a table format like Delta/Iceberg but with an SQL engine built in via DuckDB?
Lucasoato 2 days ago [-]
Nope, from my understanding the delta log (the files that say which of your data files are actually valid or not) isn’t saved in json/parquet but directly in a database.
Much faster, but adds a dependency... that you would have added anyway with database based catalogs (that are not the only kind of catalogs)
snapetom 2 days ago [-]
Ah, I see. Thanks. Having the metadata in a DB sounds a lot more robust.
efromvt 1 days ago [-]
I unironically love that we've come back around to the hive metastore (there are pros to the decentralized and centralized catalog, it's good to have options)
vira28 2 days ago [-]
[flagged]
bnuttall 3 days ago [-]
[flagged]
loufe 2 days ago [-]
Why program greenfield in C++? I know "made with rust" is a meme but seriously, why not a memory safe language in 2026?
pdet 1 days ago [-]
The main reason was due to a much easier tight-integration with the DuckDB internals, so we can take the most of the DuckDB engine (and existing code) to use. As the extension must interact with table scanners, casting functions and whatnot.
However, DuckLake doesn't need to be implemented in C++ at all, I believe that whatever fits the engine that will be performing the reads/writes in Parquet and the connectors for the DBMSs should be a good fit.
DuckDB has been around since 2018. Naturally its offspring use C++.
OutOfHere 2 days ago [-]
Rust has been out since 2010. It already was rated "most-loved" by 2016. Systems programmers had adopted it by 2018.
2 days ago [-]
meredithbloom 2 days ago [-]
And the compiler was horrendously slow for big projects on x86_84 until 2023.
Tanjreeve 1 days ago [-]
There is nothing stopping creation of rust implementations of tools or things that support these specs and protocols. But in DBMS world the centre of gravity is C/C++ mostly with some exceptions.
OutOfHere 19 hours ago [-]
If I was writing a DBMS today, I'd consider Zig, never C/C++, as the latter continue to remain fraught with memory-safety footguns. So no, I don't concur with what's your center of gravity.
Tanjreeve 15 hours ago [-]
Fair enough. You wouldn't be the first (Tigerbeetle is famously implemented in Zig).
As someone who's personal spare time tinkering project is writing DBMS + storage engine in Rust I will say the language choice or it's memory safety is not that consequential.
If you're doing anything dealing with an actual storage engine you will be writing a bunch of unsafe code and using raw pointers all over the place. My hypothesis is once the guts are in place then I can go a lot faster with Rust but it's very much swimming against the tide and it is not magically safe out of the box once you are dealing with raw storage and paging.
But nonetheless if you go around the industry and do a count of commercial and Open source DBMS systems that have any kind of significant adoption it'll be about 95% C/C++ even if you filter to recent systems and there are legitimate reasons for that being a common choice.
TL:DR The hard part of a DBMS implementation is not picking the language
There's a cool alternate rust/datafusion ecosystem initiative going on at https://github.com/datafusion-contrib/datafusion-ducklake, and think the Quack protocol opens up a lot of cool possibilities too.
If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
Here is our write up on the Ducklake blog: https://ducklake.select/2026/07/29/bringing-ducklake-to-data...
I always thought the catalogue was a duckdb file. E.g, data lives in partitioned parquet files, but which parquet files are current or soft deleted, etc, etc, is managed in a duckdb data file.
However, looking at https://ducklake.select/, it seems the catalogue lives in PostgresSQL - so it is not really a ducklake, but a postgresslake.
The more you know.
There are ofc some quirks from the SQL supported on each DBMS. Hence, the DuckLake extension from DuckDB currently supports DuckDB, SQLite, Postgres, DuckDB + Quack and MySQL. With MySQL being in a rather experimental state.
Disclaimer: I'm the lead developer in DuckLake.
https://motherduck.com/product/ducklake/
https://duckdb.org/2025/05/19/the-lost-decade-of-small-data....
When it works well it’s really nice. And it beats handrolling a multi level parquet store.
For example I think a lakehouse is a bad choice for standard enterprise BI type analytics - you've got no column or row access controls, and no column masking. I don't see how this could ever be bolted on to the bucket and catalog.
https://www.tomwphillips.co.uk/2026/08/the-benefits-of-data-...
Much faster, but adds a dependency... that you would have added anyway with database based catalogs (that are not the only kind of catalogs)
However, DuckLake doesn't need to be implemented in C++ at all, I believe that whatever fits the engine that will be performing the reads/writes in Parquet and the connectors for the DBMSs should be a good fit.
For example https://github.com/borchero/ducklake-sdk - Rust https://github.com/datafusion-contrib/datafusion-ducklake - Rust https://github.com/motherduckdb/ducklake-spark - Scala https://github.com/brikk/trino-ducklake - Kotlin
Disclaimer: I'm the lead developer of DuckLake.
As someone who's personal spare time tinkering project is writing DBMS + storage engine in Rust I will say the language choice or it's memory safety is not that consequential.
If you're doing anything dealing with an actual storage engine you will be writing a bunch of unsafe code and using raw pointers all over the place. My hypothesis is once the guts are in place then I can go a lot faster with Rust but it's very much swimming against the tide and it is not magically safe out of the box once you are dealing with raw storage and paging.
But nonetheless if you go around the industry and do a count of commercial and Open source DBMS systems that have any kind of significant adoption it'll be about 95% C/C++ even if you filter to recent systems and there are legitimate reasons for that being a common choice.
TL:DR The hard part of a DBMS implementation is not picking the language