-
-
Save RobinL/daff2396e32346791c7f08a1758b2de7 to your computer and use it in GitHub Desktop.
Transcript of Introducing DuckLake
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Good morning. | |
| Good morning, everyone. | |
| Hello, world. | |
| Hello, world. | |
| This is amazing. | |
| This is something new that we're trying. | |
| We are in a studio. | |
| Yes. | |
| In Amsterdam. | |
| And we are talking about our new and exciting project. | |
| Yes. | |
| Duck Lake. | |
| Yes, Duck Lake. | |
| Duck Lake. | |
| Duck Lake. | |
| Well, nobody has heard of it yet. | |
| Nobody knows what it is yet. | |
| And this is recorded on Friday. | |
| And you're going to learn about this on Tuesday, most likely. | |
| So we're talking to you from the past. | |
| And this is supposed to be just a chat between the two of us. | |
| Hopefully not too ranty. | |
| Yes, yes. | |
| It will be. | |
| It will be. | |
| Our PR people are a little bit concerned, but there's not a lot we can do about that. | |
| Let's do a quick introduction. | |
| So, Mark, who are you? | |
| Well, I'm Mark. | |
| I'm at this point the CTO of DuckDB Labs, one of the co-founders of DuckDB, the database | |
| system that we all know and love, I hope. | |
| And that's who I am right now. | |
| And yeah. | |
| Yeah. | |
| Hannes. | |
| Yeah. | |
| My name is Hannes. | |
| I'm the other half of the founding team of DuckDB. | |
| And I'm currently the CEO of DuckDB Labs. | |
| And that, you know, that is what I do on a daily basis. | |
| How do you know each other? | |
| Ah, it's been, I think, 10 years now, 10 years? | |
| More than 10 years. | |
| Yeah. | |
| Yeah. | |
| So, well, we've known each other since I did my master's, actually, when you came to my | |
| university presenting about a not very widely known at this point system called Munedibi. | |
| Yeah. | |
| The database course. | |
| And I was so impressed that I thought, oh, that looks like fun to work on. | |
| So I reached out to you about doing my master's thesis project with you at the CWI, the Center | |
| of Math and Computer Science, Center for Viskun and Informatica. | |
| Yeah. | |
| And did my master's thesis there. | |
| Then figured database were kind of cool. | |
| Stuck around, did my PhD, did my postdoc, started a company together. | |
| The whole nine yards. | |
| Yeah. | |
| And I mean, obviously, you impressed me back then. | |
| And we've been working together for all of these 10 years, which has been a great privilege. | |
| So, yeah. | |
| And obviously, you're no longer a student. | |
| No, no, no. | |
| I think some people call you a master even at this point. | |
| Yes, yes, yes. | |
| That is one of the titles. | |
| It says it on the piece of paper. | |
| Yeah, yeah. | |
| Wonderful. | |
| So, yeah. | |
| Obviously, DuckDB. | |
| How is DuckDB coming along? | |
| I think it's been a quite wild ride over the last, what did you say, seven years? | |
| Seven years. | |
| Yeah, yeah, yeah. | |
| So, interesting. | |
| We've known each other for 10 years. | |
| Seven of those years have been DuckDB somehow. | |
| But I don't know. | |
| Somehow it still feels like it was only like two years ago that we started this. | |
| It's been. | |
| Yeah. | |
| I mean, there was a little pandemic in the middle that may have done some time dilution there. | |
| But, yeah, it has been a while. | |
| And I think when we started this, we never expected it to go to how far it has now. | |
| Like, we made this because we thought it would be useful. | |
| Yes. | |
| But we made this mostly because we thought it would be like a nice R package that would be used by a few thousand people. | |
| We never thought this would start the next data revolution. | |
| No. | |
| Which I think it's fair to say that it has in some capacity done that. | |
| Like. | |
| Yeah. | |
| Yeah. | |
| It's kind of wild. | |
| Indeed. | |
| Like the original approach to make a database that can be used by data scientists has accidentally spilled over into the rest of database land. | |
| And I think what surprised me the most was that we were fully expecting to be ridiculed for the next two decades or so for not making a distributed system. | |
| Yes. | |
| And we went with this sort of unapologetic, we'll go for single nodes, single computers are big enough, which is actually an idea that I think Martin Kersen, the creator of MonoDB, first popularized. | |
| Absolutely. | |
| But we ran with it. | |
| And I think we've conclusively shown the world that this is a viable way of looking at your data. | |
| In case you've been following our blog posts, we've last week, when you hear this, we've released this blog post about running a huge analytical benchmark on a 10-year-old computer. | |
| 13-year-old. | |
| 13-year-old. | |
| I'm sorry. | |
| Yes. | |
| My first MacBook Pro, actually. | |
| And it worked. | |
| That's the wild thing that it worked. | |
| And I think we can kind of claim that we have sort of won the single player campaign. | |
| I mean, you play computer games. | |
| Yes, yes, yes. | |
| The single player campaign of DuckDB, how would you characterize that? | |
| Well, I think for the single player campaign, you're on your MacBook, you have some data you want to analyze, you have some data stored somewhere, right? | |
| You grab some CSV files from the internet, you have some parquet files, maybe you have a database from your organization you connect to, you load it onto your MacBook and you run it. | |
| Because it's kind of like some people refer to it as the last mile analytics, which I think is a bit odd because, of course, that for most organizations encompasses literally all analytics that they will ever need, right? | |
| But for some, that is the last mile analytics. | |
| And I think that's kind of like you are by yourself on your MacBook using DuckDB to analyze data. | |
| I think that's the single player scenario for DuckDB. | |
| Yeah, PR people want me to note that DuckDB does not only run on MacBooks. | |
| No, that's a joke. | |
| We don't have PR people in the room. | |
| But it runs anywhere, really. | |
| Like it doesn't have to be a MacBook, can be Windows, can be Linux even. | |
| And I think it's also kind of crazy to see how many people are doing that, right? | |
| I mean, we have more than 20 million downloads per month at this point, per month. | |
| And we accidentally made one of the bigger websites of this country with millions of people visiting the website every month, which is also something that I still fail to comprehend somehow in my brain. | |
| I don't know what a million is. | |
| But we talked about the single player campaign, the last mile. | |
| I think this is an excellent continuation here. | |
| DuckDB has, of course, always, not always, but for a long time, been able to run queries in parallel, right? | |
| Yeah. | |
| But there have been some restrictions for like the multiplayer experience in that sense, right? | |
| Because in DuckDB, the database runs inside a process, which means that while within that process, you can have multiple individual CPU threads that do things in parallel on the same database. | |
| A sort of a longstanding complaint, criticism, what do I want to say? | |
| The suggestion for improvement of the community has always been like, yeah, I mean, if you're doing multiplayer, this doesn't work so well because you're locking the database file. | |
| And we have to. | |
| We have to lock the database file in order to do the right thing, TM. | |
| Yeah. | |
| Because DuckDB is fully transactional. | |
| We go to great lengths to make sure that your data will be persistent, even when you pull the plug on your computer, things like that. | |
| But we have always had this restriction. | |
| So the restriction being if you have multiple computers, I think that's the most logical way of looking at it, trying to modify the same database at the same time. | |
| That wasn't a great solution. | |
| Yeah. | |
| Yeah. | |
| And people have worked around this because, as you mentioned, DuckDB does support multiple connections within the same process. | |
| And this is not an uncommon restriction in database systems, right? | |
| Like most database systems operate like this. | |
| They just get around it by having a client-server protocol. | |
| So you have one process, namely your database server, that has internally multiple connections that are then connected to different processes using this client-server protocol. | |
| But we have never implemented a client-server protocol. | |
| Why? | |
| We do have some experience with client-server protocols. | |
| There's some work we did on the MineteB and Postgres client-server protocols in 2016, I want to say. | |
| In the long, in the gray past. | |
| Where we actually looked at all the client-server protocols of all database and concluded they are all... | |
| There's room for improvement in all of them to keep it politically... | |
| Oh, man. | |
| Yeah, there's some recent work, however, that has improved on this. | |
| So I think partly inspired by this paper, the Aero people have built the Aero flight protocol. | |
| Yes, yes. | |
| And I should mention, I think as we're recording this yesterday or today, the airport extension for DuckDB launched, which does give you a client-server sort of setup for DuckDB. | |
| Yeah, yeah. | |
| There are solutions for this, but... | |
| We have never done anything. | |
| We have never done anything. | |
| Yeah. | |
| And I think it's also, once you start going down that hole, it's kind of a deep rabbit hole, right? | |
| Because, okay, you want a client-server. | |
| Sure, it's more than just implementing a protocol, right? | |
| You need to have, like, fallover. | |
| People are going to say, oh, now I want replication because I need to be able to handle... | |
| Failover. | |
| Failures if things go down. | |
| I need to have backups, right? | |
| Yeah. | |
| Like, it's not just a client-server protocol. | |
| It's a long tail of stuff that is necessary to make a robust service based on a database. | |
| Yeah. | |
| And that has been done in other systems, like Postgres has all this stuff, but we have not started on that journey, right? | |
| No, we have not. | |
| And I want to give a shout-out to the Tiger Beetle people, who, in my opinion, are doing the most exciting work on replication and sort of transactional safety in a multi-computer setup at the moment. | |
| I find this really impressive. | |
| Yeah. | |
| So, here we are. | |
| And in this sort of moment where we have self-declared won the single-player campaign and there are some challenges for multiplayer DuckDB. | |
| Yes. | |
| If you want, right? | |
| That's where we're at. | |
| That's where we're at. | |
| That's where you are at as of this moment. | |
| And so, we want to talk about what we have and what the, when we have already given you the name, the whole Duck Lake concept that we're going to explain to you how we are looking at this going forward. | |
| Okay. | |
| So, let's look at lakes. | |
| I like lakes. | |
| I like boats. | |
| Yes. | |
| You like boats. | |
| I love boats. | |
| I love lakes. | |
| I love water. | |
| Yeah. | |
| Water is great. | |
| Yeah. | |
| So, fundamentally what we have seen in the last couple of years is this disconnect of storage and compute, right? | |
| Yes. | |
| Yeah. | |
| Which is this whole idea that you move on from this idea that from this shared nothing architecture where data is partitioned over the nodes, like these were early distributed systems. | |
| Yeah. | |
| So, for data management, we move to this other approach where we have the, you know, the data lake where we disconnect the computers that store data and the computers that do computation on the data. | |
| Yes. | |
| And we do that so we can kind of scale them independently. | |
| And there have been two systems that I want to point out that have been doing this very well, which is Snowflake and BigQuery. | |
| I think those are where the two OG. | |
| The pioneers. | |
| Yeah. | |
| The pioneers here. | |
| And I think it's really impressive, especially, I think, looking at Snowflake because we know the origin story better, I think, is how they realized very early on that basically Hadoop was doomed. | |
| Yeah. | |
| And that a data system for the next sort of 20 years would have to be designed around the cloud computing primitives like blob storage and so on. | |
| I found this extremely interesting. | |
| But, yeah, let's talk a bit more about data lakes. | |
| Yes. | |
| So, this is interesting because we have been looking a bit into the history, into sort of how this all came down. | |
| And we are interested as, you know, former recovering academics. | |
| We are interested in how we got to where we're at, right? | |
| Because very often you can learn a lot about why things are the way they are by understanding where they came from. | |
| So, we had Hadoop, which was generally regarded as a bad move. | |
| Is that a quote from – what's the quote from – it's made in the beginning – oh, I think it's Hitchhiker's Guide. | |
| Yeah, Hitchhiker's Guide. | |
| Yeah. | |
| It's like, in the beginning, the world was created. | |
| This was generally regarded as a bad move and made everyone very angry. | |
| Yes. | |
| So, in the beginning, there was Hadoop, which was generally regarded as a bad move and made everyone very angry. | |
| I think that's fair to say. | |
| Yes. | |
| And there has been a longstanding feud between database people and Hadoop people. | |
| Yes. | |
| Of course, Hadoop people having eventually evolved into Spark people. | |
| Yeah. | |
| I mean, there has been a long list of software built on top of Hadoop. | |
| Yeah. | |
| Somehow never actually removing the original Hadoop code, which is a mystery to me. | |
| Yeah. | |
| But, okay. | |
| Okay. | |
| So, in this longstanding feud, I think the famously Mike Stonebraker, the Turing Award winner and database person, has written about this very early on, how this was a giant step backwards. | |
| And then the Hadoop people were yelling at the old people. | |
| Yes. | |
| Old and crusty and pointless database people that were getting it wrong and so on and so forth. | |
| It does not scale. | |
| It does not scale. | |
| Databases do not scale. | |
| This is one of my favorite sentences. | |
| It does not scale. | |
| I get so many comments on the internet with that that I'm very aggressive about this at this point. | |
| But then, you know, things improved. | |
| We went from Hadoop HDFS. | |
| We went, as I mentioned, we went to this disconnect. | |
| We also improved on the horrible Hadoop sequence files to data. | |
| Ah, I remember that one. | |
| Yes, yes, yes. | |
| Also improved to things like Parquet. | |
| Yes. | |
| We like Parquet. | |
| But, yeah, Parquet is great. | |
| Why do we like Parquet? | |
| Parquet, it's a columnar format. | |
| Yay. | |
| It's compressed. | |
| Yay. | |
| It has a binary representation of the data. | |
| Hey. | |
| It's mostly self-describing, right? | |
| Yep. | |
| Like it has data types. | |
| Crazy, I know. | |
| I know. | |
| Don't tell Dr. Hip. | |
| It has data types. | |
| It has compression. | |
| So the files are small. | |
| It's a clever format, right? | |
| It's clever. | |
| And the way it works by writing like top down as well allows you to just dump this stuff | |
| in a streaming way right as you go. | |
| Yeah. | |
| And then only write the header at the end. | |
| Like it's clever for the use case it was intended for, right? | |
| Yeah, which is. | |
| Writing data to blob stores. | |
| A table. | |
| A table. | |
| A single table. | |
| A single table. | |
| This is which is important to note here. | |
| A single table. | |
| It's meant for a single table. | |
| Although if you really wanted to, we could make some unholy nested table struct. | |
| Yes, yes. | |
| We should try this at some point. | |
| That sounds like a fun challenge. | |
| Okay. | |
| So we have Parquet. | |
| We like Parquet. | |
| I actually got a chance to talk to Julian Jodem two weeks ago at Data Council, the original | |
| creator of Parquet. | |
| And he asked me, this is your chance. | |
| What do you hate about Parquet? | |
| Apparently the guy has been hated on by so many people. | |
| And I just told him, no, I love the format. | |
| The format is great. | |
| I mean, we were stuck in CSV land for so long. | |
| In some ways, we're still stuck in CSV land. | |
| I mean, Parquet is not perfect. | |
| But compared to CSV or JSON, like it's miles ahead. | |
| I also find it ridiculous that people are now jumping to try to replace it because it | |
| took us 10 years to standardize on it. | |
| And now for some reason, people are, I don't know. | |
| I don't get it. | |
| Let's just move on and innovate on something else. | |
| I mean, I get it. | |
| Like there's a lot of glory to be had with replacing Parquet at this point. | |
| But the fact is like Parquet, when it was first created, was useless. | |
| Yeah. | |
| Because the whole point of these data formats is interchange. | |
| So in order to make a format to replace Parquet, you need to have the same sort of interchange | |
| capacity as Parquet. | |
| Right. | |
| Which is going to take a long time, if ever, to build out. | |
| So you better have a really good story on why your format should be used instead of Parquet. | |
| And from the formats I've seen so far, they all, they have some stories like, okay, we | |
| use better lightweight compression algorithms. | |
| It's nice, you know, like, okay, you can jump into the file. | |
| All right. | |
| Great. | |
| Is this a 10x? | |
| No. | |
| Is it a 10% boost? | |
| Yeah, sure. | |
| It's a 10% boost. | |
| Yeah. | |
| And maybe for some specific use cases, it can be a 10x on jumping around. | |
| But that's usually not how Parquet files are processed anyway, right? | |
| No, the value is in the convention. | |
| Yeah. | |
| People don't get that. | |
| I don't know why. | |
| I mean, some people get it. | |
| But I think many people do get it. | |
| And I mean, I applaud them for trying something new and not against it. | |
| But I think Parquet, as it stands in today's world, it is the interchange format for high | |
| performance analytics. | |
| Not Org. | |
| Not Org. | |
| No. | |
| No, not Org. | |
| Just checking. | |
| Yes, yes. | |
| Not Avro either. | |
| No, Avro, absolutely. | |
| Absolutely not. | |
| Okay. | |
| But now, okay. | |
| So now we have Disconnect Storage Compute. | |
| We have a data lake. | |
| We have blob storage. | |
| We have all these wonderful things. | |
| Yes. | |
| But now we have the one, we have the, we ran. | |
| So people try to basically do that. | |
| People try to say, this is our new stack of new data stack. | |
| Yeah. | |
| We're going to run on top of that. | |
| Parquet files on S3 is like the, I would say the canonical. | |
| That's what I tell the students. | |
| Just use freaking Parquet files on S3. | |
| Just get a blob store, push your Parquet files in there. | |
| Data is solved. | |
| And done, right? | |
| This is how it goes. | |
| Yes, yes. | |
| And then this very pesky problem appears. | |
| Yes. | |
| The pesky problem of changing data. | |
| Oh, no. | |
| Oh, no. | |
| Change, change, change. | |
| Change. | |
| There's this wonderful quote that I use in my, in my, from Heraclit that is, everything | |
| flows, Pantare, which is like, yeah, they knew things changed, but somehow we pretend that | |
| things do not change, right? | |
| Yeah. | |
| With the whole Parquet on S3 thing. | |
| I mean, yeah, it's also in, I would say in database research, that's a thing, right? | |
| Like where all the benchmarks are run on static data sets. | |
| No. | |
| Not many. | |
| I was joking. | |
| So are you criticizing research people? | |
| I mean, no, no, no. | |
| Never, never. | |
| They're great. | |
| No, no, no. | |
| But the update problem remains. | |
| Yes. | |
| So, so if you have a Parquet file, Parquet is not a format design. | |
| Parquet as a format cannot be updated, right? | |
| Yeah. | |
| I mean, you can hack around it if you wanted to, but in principle, these files are once | |
| written, they are, they're static. | |
| So now you can rewrite them, which is great if they're a terabyte large. | |
| And there are some things that, that, that are kind of, that quickly, quickly, you know, | |
| turned out to be obvious. | |
| Like if you want to append to a table. | |
| Yeah. | |
| You just add another Parquet file and your reads become a glob. | |
| Spark is like defaulting to this, even that all the reads are globs. | |
| Yeah. | |
| Where they basically match any file with the same prefix. | |
| Ah. | |
| Yeah. | |
| Yeah. | |
| And then you have the data dash zero dash guide dot Parquet files who that contains all | |
| the parts. | |
| The order of which is like undefined, but different problem. | |
| Yeah. | |
| But it gets harder when you do things like deletes. | |
| Yes. | |
| There, you know, what do we do? | |
| I don't know. | |
| Like we'll store somewhere else. | |
| Again, we can rewrite the whole damn thing. | |
| Or we can come up with some format of storing which rows are deleted. | |
| Interesting. | |
| It gets even harder when we do things like schema changes. | |
| Schema evolution. | |
| Schema evolution is absolute nightmare. | |
| Right. | |
| How do we even know? | |
| And then also things like time travel. | |
| Like how do we know what the, or any transaction. | |
| Just transaction. | |
| Just transactionality and updates. | |
| Just having a snap, like a valid snapshot of your data at a certain point in time. | |
| Right. | |
| Right. | |
| That is not like full of half written data. | |
| Like consistency. | |
| Yeah. | |
| Consistency. | |
| Great, great idea. | |
| Great idea. | |
| Great. | |
| Atomicity. | |
| Yeah. | |
| Great ideas. | |
| Yeah. | |
| Yeah. | |
| Yeah. | |
| And of course the, you know, the old, you know, crusty database people had all this | |
| really sorted out. | |
| There were data warehouse systems. | |
| Yeah. | |
| There are still data warehouse systems like, you know, Teradata and so on that solve this | |
| problem. | |
| Snowflake solves this problem. | |
| Yeah. | |
| But then we have this, we had this Hadoop world which decided that databases are stupid. | |
| Yes. | |
| And they decided to try to solve the problems of updating tables with sort of the tools | |
| they had. | |
| Right. | |
| Yeah. | |
| Yeah. | |
| Like the tools, what tools did they have? | |
| They had files. | |
| Yes. | |
| Files on a blob store. | |
| Yes. | |
| Right? | |
| Yeah. | |
| So this is essentially what Iceberg and Delta are trying to do or try to do in the beginning, | |
| right? | |
| Yeah. | |
| Yeah. | |
| Yeah. | |
| In the very beginning, you know, there was, you know, there was Parquet and S3. | |
| And shortly after, on day two, people invented these metadata formats on top. | |
| Yeah. | |
| Yeah. | |
| That were trying to get, they were always yelling ACID, which got me a bit angry, but | |
| were trying to invent ACID using just files. | |
| Right? | |
| I think maybe it's jumping ahead a bit, of course, because we, in the past, we also had | |
| the Hive Metastore, the Hive partition stuff, right? | |
| Like it took a long time. | |
| It took a long time. | |
| To get to the sort of current state of the art in lake houses, right? | |
| Yeah. | |
| True. | |
| And we haven't even mentioned the term lake house yet. | |
| Yeah. | |
| What is a lake house? | |
| It's the house of a lake. | |
| I had the pleasure of staying in a house on a lake a couple of weeks ago. | |
| Now, the idea, yes, it's a great question. | |
| So the data lake is this whole collection of Parquet files. | |
| Yeah. | |
| The house is the, as in data warehouse, where you have, you know, some sanity about updates. | |
| Yeah. | |
| I think that's fair to say. | |
| Yeah. | |
| Exactly. | |
| And then lake house together means that we combine the two, right? | |
| Yes. | |
| We have a sanity in updates in a data lake. | |
| Yeah. | |
| And that is what Iceberg and Delta are doing. | |
| And that is what Iceberg and Delta are doing. | |
| Exactly. | |
| Yeah. | |
| So they try to build this, basically bring back the database to the zoo of files on the | |
| blob. | |
| That's right. | |
| Yeah. | |
| Exactly. | |
| In the Hadoop way, like you mentioned. | |
| In the Hadoop way. | |
| Yes, exactly. | |
| Like they tried so hard not to have a database in this. | |
| Yes. | |
| They tried so hard because databases are evil. | |
| Yes. | |
| Absolutely. | |
| They don't scale. | |
| They don't scale. | |
| They don't scale. | |
| They're not web scale. | |
| They're not web scale. | |
| No. | |
| They're not web scale. | |
| They're not web three scale either. | |
| No. | |
| Nothing web related can ever run on a database. | |
| No, no. | |
| Databases are stupid. | |
| But it turns out. | |
| Hard to say that with a straight face. | |
| Yes. | |
| It turns out that there was some interesting development. | |
| So we come back to Acid, not the Acid that people in Amsterdam like. | |
| We come back to the sense of sanity of transactional guarantees. | |
| Yes. | |
| And this is actually quite hard on data lakes. | |
| Yeah. | |
| And this is hard. | |
| There's one reason why this is hard. | |
| It has to do with the consistency guarantees offered by blob stores like S3. | |
| And yes, I know they're constantly saying how this is now fixed. | |
| I don't think it's fixed yet. | |
| But we might create a file and that might not necessarily be visible in the directory listing | |
| from all the nodes that are looking at this. | |
| We are maybe changing a file. | |
| But again, the update might not be immediately visible. | |
| Yeah. | |
| But they've tried really hard not to have a database. | |
| So these formats like Iceberg and Delta, they did a lot of interesting design decisions | |
| to kind of try to make that work. | |
| Like, for example, you cannot really change files. | |
| Yes. | |
| Which just means you have to rely on directory listings to get the latest. | |
| Yeah. | |
| You need to be very careful about the size of the files that you write. | |
| They can't be too big. | |
| They can't be too small. | |
| Which is interesting because your metadata files, they probably like to be small, but | |
| they can't be because if you have too many small files, again, these blob stores don't | |
| really like that because you have too many round trips. | |
| So that happened. | |
| These formats like Iceberg and Delta only ever really concerned themselves with a single | |
| table too. | |
| Yes. | |
| Yeah, absolutely. | |
| Single table. | |
| Single table. | |
| Yeah, yeah. | |
| Which is, in a sense, a natural evolution of Parquet, right? | |
| Parquet stores a single table. | |
| So Iceberg and Delta built on top of that single table. | |
| But in practice, when you give a person a single table, very soon they will want a second | |
| one. | |
| Entitled people. | |
| I can tell you. | |
| Like, they want a second table? | |
| A second table. | |
| A second table. | |
| Yes. | |
| Wow. | |
| That's amazing. | |
| And so the solution was, well, this was the one problem that we only have one table. | |
| And the other problem, as I said, was this, was the, I think the biggest problem was the | |
| atomic commit. | |
| Yeah, exactly. | |
| Right? | |
| The atomic commit where basically you're making some change to a table and you want to have | |
| an atomic as in like, it either happened or it didn't happen. | |
| Yes. | |
| And in its entirety, a moment in time where you switched from version one to version two | |
| of that table. | |
| Yeah. | |
| A guarantee that you're seeing a consistent state of the data at the time, that you can | |
| resolve conflicts, that data doesn't get lost, right? | |
| Acid. | |
| Acid. | |
| Yes. | |
| The good stuff. | |
| The thing that the databases had for a million years. | |
| For, yeah, 60 years, maybe. | |
| 50, I guess. | |
| We're both younger than that. | |
| Yeah. | |
| For us, it's eternity. | |
| Yes. | |
| Absolutely. | |
| And so the solution was interesting because like many problems in computer science, any | |
| problem can be solved with another layer of abstraction, right? | |
| Yes. | |
| Yes. | |
| And any sort of technology can always be solved by slapping another technology on top of | |
| it. | |
| Yes. | |
| Any technology problem can be solved. | |
| So we got these catalog servers. | |
| And I remember vividly the discussion I had with Ryan Blue from Tabula at the time, pre-Databricks, | |
| when he tried to explain to me why they needed this catalog server on top and they cannot have | |
| the file-only based thing anymore. | |
| And what happened is that people said, in order to solve the problem of multi-table management | |
| and atomic commits, we'll invent a new system. | |
| And this is no longer going to be file-based. | |
| Yes. | |
| A massive departure. | |
| Massive departure. | |
| Yeah. | |
| Yeah. | |
| It's not going to be file-based anymore. | |
| It's going to be a service. | |
| Yes. | |
| A service. | |
| Hmm. | |
| And the service that speaks REST. | |
| Yes. | |
| You know, like web server, that that stores the versions and which tables exist. | |
| And how does it store that? | |
| Well, it uses a, it's a bit of a, like a exotic thing. | |
| It uses a system that is specifically designed to manage data. | |
| A database management system, if you will. | |
| That is wild. | |
| It's crazy, right? | |
| How is that? | |
| How is that even? | |
| Who had thought? | |
| Yeah. | |
| Who would have thought? | |
| So, and what does this mean in practice? | |
| Postgres. | |
| It means Postgres. | |
| Yes. | |
| Yes. | |
| So Postgres, of course, again, also the brainchild of Mike Stonebreaker, who he mentions already. | |
| Yes. | |
| Yes. | |
| Uses Postgres to store which tables exist. | |
| Yeah. | |
| That's a list. | |
| Yeah. | |
| And for each table, it stores a version number. | |
| Yeah. | |
| That's it. | |
| The root pointer. | |
| The smallest, smallest table ever, because it has only a single name, a single number in it. | |
| Yes. | |
| So, and then it is, it's actually looked at the implementation. | |
| These are like unholy masses of Java and Postgres and Kubernetes and web servers and all of that. | |
| Yeah. | |
| Interesting. | |
| And of course, the fun part of this is that, like what we talked about before, right? | |
| Like what, it's more than just slapping a client protocol on something. | |
| If you want to make a server like that, you need to build the whole thing. | |
| You need to have a client protocol. | |
| You need to have replication. | |
| You need to have a backup strategy. | |
| You need to have all these components that have existed in Postgres for a long time. | |
| Yeah. | |
| But don't, they don't materialize out of thin air, right? | |
| Like, so once you build a new server, you need to build all those things again. | |
| It's not just about the client protocol. | |
| No. | |
| And this is, I think it's really interesting because it's like, you know, it's like taking | |
| the fruit from the forbidden tree from their perspective. | |
| Yes. | |
| Right. | |
| They went, they went, this is a major, major departure from the, it does not scale yelling | |
| department. | |
| Yes. | |
| To invent a service. | |
| Yes. | |
| That needs, you know, servers and state and not files. | |
| Yeah. | |
| And a protocol and authentication and all that stuff to manage their format that cannot really | |
| work without it. | |
| Yes. | |
| Right. | |
| It cannot really do the things they promised without this service. | |
| No, no. | |
| And I think the thing that, I mean, we are, of course, technology aesthetics people. | |
| Yes. | |
| We like, we like technology. | |
| We like data. | |
| We actually like databases. | |
| We love databases. | |
| We love, actually, we love databases. | |
| And we really care about the aesthetics of databases. | |
| Yes. | |
| Yeah. | |
| And this, this, this aesthetic of having something that was based on this idea of file only interaction | |
| with a, with a file system or blob store, um, then slapping a database on top of it, that, | |
| that has interest, questionable aesthetics. | |
| Yes. | |
| But I think the interesting, the most interesting thing is that the, the iceberg and Delta and | |
| there's others, I don't want to, you know, but these are the two ones that are, that | |
| are kind of out there right now. | |
| They never went back on their design decisions. | |
| Yeah. | |
| Right. | |
| They made all these design decisions that had to happen in order for these things for, | |
| for somehow fitting an asset like table onto a blob store with parquet files and, you | |
| know, other restrictions like immutable files and file sizes being problematic, but they | |
| never went back on these design decisions. | |
| They just said they just slapped this layer on top and called it a day. | |
| Yes. | |
| Which still does not compute for me. | |
| I don't know. | |
| I mean, I think it makes sense given that they already had, uh, committed a lot to building | |
| these file formats, right? | |
| There are a lot of cool ideas went into these file formats and there's a lot of cool problems | |
| they solve, right? | |
| Yeah. | |
| Like, but to never go, it will be hard for them to go back and say, actually all this | |
| file-based stuff was a mistake. | |
| Let's just take all this data and maybe not put it in a bunch of Avro files. | |
| We'll get to that. | |
| We'll get to that. | |
| Yeah. | |
| And we actually, I mean, I think there's, even with the current setup with the catalogs, | |
| there's still a bunch of unsolved problems. | |
| Like still, still, I think they're still figuring out how to do cross table commits. | |
| Yes. | |
| It's still completely up in the air. | |
| Yeah. | |
| Small commits are a complete, like a lot of small commits are a complete nightmare. | |
| Yeah. | |
| Um, nobody really knows. | |
| And small commits, by the way, that's not like small commits. | |
| It's like anything less than a few hundred megabytes. | |
| That's right. | |
| Like it's actually, it's not like a one row commit is a problem. | |
| It's like a 100,000 row commit is already starting to be a problem. | |
| It's going to be a problem. | |
| And there's entire startups. | |
| Yes. | |
| You know, funded by VCs. | |
| By the way, VCs, I'm sorry if you ruin your plans right now. | |
| That's, that's, I apologize for that. | |
| Not. | |
| Um, uh, we have actually been, yeah. | |
| So it's a problem with a small, with a small ish. | |
| I mean, it commits, it's actually, I think iceberg at this point, it's gonna, it's gonna | |
| work if you have like one change per, per minute at the very best. | |
| I think per minute, not, but like, I think per few seconds. | |
| Realistically. | |
| I mean, realistically. | |
| Okay, fine. | |
| But, but like, that's still far away from what, you know, the dinosaurs of data warehouses | |
| could do in the nineties. | |
| Yeah, yeah, exactly. | |
| Like where we're talking about, like less than a transaction per second versus hundreds | |
| of thousands or tens of thousands of transactions per second that any respectable database can | |
| do, right? | |
| We've been here before. | |
| We had the same argument against blockchain, remember? | |
| Yeah, yeah, blockchain. | |
| Also a terrible idea. | |
| And somehow blockchain has a higher transactions per second than these, than these formats, | |
| which is kind of interesting to think about. | |
| Yeah. | |
| Because people were like, blockchain can only do five transactions per second. | |
| It's ridiculous. | |
| And it was ridiculous. | |
| And now we're on like 0.5 transactions per second. | |
| Yeah, yeah. | |
| Great. | |
| Well, we went down another 10x. | |
| Another 10x down. | |
| Yay progress. | |
| Yes. | |
| Yay progress. | |
| So as I said, Lakehouse formats, in our opinion, have bad, existing Lakehouse formats have | |
| bad aesthetics. | |
| Yeah. | |
| And we like databases. | |
| So their logical conclusion can only be, we are moving on from the current formats. | |
| So we as database people, we have a higher power, which is not God, it's COD. | |
| COD, Ed COD, the father of the relational model. | |
| Yes. | |
| So we had an epiphany from COD directly. | |
| Let's rethink a data lake format using more database. | |
| Yes. | |
| Is that fair? | |
| Yeah, absolutely. | |
| And I think it was inspired by looking at the current technology stack of data lakes | |
| and spotting this Postgres server up there. | |
| Yeah, yeah, yeah. | |
| I think it's the famous iceberg architecture diagram. | |
| Yeah, the diagram. | |
| Yeah, everyone has a diagram. | |
| Yeah. | |
| It's a beautiful diagram, right? | |
| You have first a layer of Parquet files, then like several layers of files, like one layer | |
| of Avro files, another layer of Avro files, then JSON for some reason, because of course | |
| that layer could not have also been Avro. | |
| And then up there at the very roots, five layers deep sits a database. | |
| A database. | |
| A database. | |
| So the epiphany from COD is really, really not that crazy if you think about it afterwards. | |
| But somehow it's like, yeah, I don't know. | |
| I don't know. | |
| Somehow I wonder. | |
| But yeah, so the epiphany is we'll use the database for more, I think. | |
| And because it's there already anyway. | |
| Yeah, you have to. | |
| The world has shown that you need it anyway. | |
| Yeah. | |
| And what can you, and then you put, the idea is that, the basic idea is we put a lot of | |
| this metadata that's currently sitting in these JSON files, in these Avro files into the | |
| database. | |
| Also, I'm personally offended by the Avro file format. | |
| I wrote a reader for it. | |
| What can I say? | |
| The result of all this is Duck Lake. | |
| Yes, Duck Lake. | |
| Say the thing. | |
| Duck Lake. | |
| Duck Lake. | |
| Duck Lake. | |
| And what is Duck Lake? | |
| So Duck Lake, it's really essentially what you described right now, right? | |
| It's we take all the metadata that's required to make a lake house, right? | |
| Like an actual lake house, not a single table, but a catalog, right? | |
| A database. | |
| And we take all that metadata and we put it into another database, right? | |
| Using this standard that's this little known standard that has been around for a while called | |
| SQL or SQL, depending on your branch, your school of thought. | |
| And we have a bunch of SQL queries and they represent all the manipulations on the data. | |
| We have a bunch of tables. | |
| They represent the metadata that is stored. | |
| All of it. | |
| All of it. | |
| And we just put it in the database. | |
| Yeah. | |
| And not the data files, right? | |
| The data files, they can still be parquet files on S3. | |
| They can still be arbitrarily large. | |
| They can scale to whatever size you desire. | |
| But only the metadata. | |
| Only the metadata. | |
| Just the metadata. | |
| Just the metadata. | |
| And so we have a bunch of tables and we have a bunch of queries. | |
| Yes. | |
| And then we have systems that know very well how to deal with transactions in a bunch of | |
| tables. | |
| Absolutely. | |
| Yeah. | |
| So we can actually do very sane data transformations across multiple tables because all these changes | |
| to the various metadata structures are just going to be updates to tables that are wrapped | |
| in transactions as executed by, for example, Postgres. | |
| Exactly. | |
| Or by DuckDB. | |
| Or by DuckDB. | |
| Yes. | |
| By any database system that has assets and primary keys, we'll be able to run DuckLake, | |
| right? | |
| Like it's not a DuckDB specific format. | |
| No. | |
| It is a bunch of SQL statements that are like standard SQL statements with simple data types | |
| designed to be very portable across database servers. | |
| And all that's required is they can use a sane database server that follows the basic | |
| principles of assets and has primary key constraints. | |
| Wow. | |
| So we have some principles. | |
| Yes. | |
| We formulated some principles. | |
| We have some principles. | |
| But I want to stress before we go into that, I want to stress again, it's not DuckLake is | |
| not a DuckDB specific thing, as you said. | |
| Yeah. | |
| It's a standard. | |
| It's a standard. | |
| It's a convention of how do we manage large tables on blob stores that are stored in formats | |
| like Parquet in a sane way using a database. | |
| Using a database. | |
| Using a database. | |
| Which again, you need for these other formats. | |
| Which you need anyway. | |
| Right? | |
| Which already was in your tech stack. | |
| Yeah. | |
| I mean, any company besides us for some reason. | |
| We don't have databases. | |
| We don't. | |
| I mean, we run databases. | |
| We have monday.com. | |
| Okay. | |
| Yes. | |
| We do have a database. | |
| They have a database. | |
| But most, any company that is managing any sort of serious amount of data will have a | |
| database server somewhere. | |
| And that can be Postgres. | |
| That can be Oracle. | |
| That can be SQL Server. | |
| That can be MySQL. | |
| There's like a lot of these, but any serious company will already know how to host a database | |
| server. | |
| And they will know how to back it up. | |
| They will know how to back it up. | |
| They will know how to replicate it. | |
| How to talk to it. | |
| How to do authentication. | |
| Yeah. | |
| Right? | |
| Like authentication. | |
| All this stuff. | |
| Like all that stuff is done. | |
| Yeah. | |
| Like they know how to do it. | |
| They have experts to do it. | |
| Yeah. | |
| It's in some sense, it's a commodity. | |
| Right? | |
| Like. | |
| Absolutely. | |
| Like storage. | |
| Like storage. | |
| Like a blob store is a commodity. | |
| Right? | |
| Like you can switch from S3 to Azure blob store to Google cloud store. | |
| And there are some minor differences. | |
| But unless you're doing very crazy things, these things are kind of, it's kind of the same | |
| thing. | |
| Yeah. | |
| Right? | |
| And you can switch, which means there's no lock-in. | |
| You can switch. | |
| Okay. | |
| Databases are also a commodity. | |
| Like not between them. | |
| Right? | |
| Like there's differences between Postgres and MySQL. | |
| But host, like running a Postgres instance is a commodity. | |
| You can right now talk to 15 different companies. | |
| That's probably even a low number. | |
| To 100 different companies. | |
| Ask them to run a Postgres server for you and they will run it for you. | |
| Apparently, one of them costs a billion dollars. | |
| One of them costs a billion dollars. | |
| That's kind of crazy. | |
| That is. | |
| Yes. | |
| Yes. | |
| Apparently. | |
| It's a big number. | |
| Yeah. | |
| Okay. | |
| Let's talk about principles of Duck Lake. | |
| The first one. | |
| We have three principles. | |
| Simplicity, scalability and speed. | |
| Yes. | |
| Like Jeremy Clarkson. | |
| We start with simplicity though. | |
| Because we are, we here at DuckDB. | |
| Yes. | |
| Recommend DuckDB. | |
| Of course. | |
| But we are. | |
| Sorry, Mark. | |
| We recommend DuckDB. | |
| But we also are very much about simplicity. | |
| Simplicity. | |
| This is one of our sort of fundamental principles. | |
| Because again, we are aesthetics about, you know, we care about the aesthetics of technology. | |
| Yeah. | |
| Yeah. | |
| So things need to be simple. | |
| How does DuckDB, Duck Lake look like? | |
| It is. | |
| Essentially, you have two things. | |
| You have, as Mark said, we have a database to store metadata. | |
| And we have some storage that can be S3, can be anything, can be Google Cloud, can be Azure, | |
| can be a network attached file system. | |
| It can be your local disk. | |
| I don't care. | |
| A file, a thing to store files in. | |
| Right? | |
| So this is simplicity as in like, we only need these two things. | |
| We need a database to store metadata. | |
| And as you said, can be anything really. | |
| And we need a storage place to store files, arcade files specifically. | |
| Yes. | |
| Right? | |
| So it's extremely simple in that and also flexible in the sense that we really don't care what you use. | |
| No. | |
| I mean, it's, you can use Postgres for the metadata. | |
| You can use DuckDB for the metadata. | |
| You can use, as I said, you can use S3 for your files. | |
| You can use, you know, an FTP server. | |
| Why not? | |
| Yes. | |
| Anything DuckDB can talk to. | |
| I'm not sure about FTP at the moment. | |
| Yeah. | |
| But yeah, anything that DuckDB can talk to. | |
| I mean, actually not. | |
| The standard doesn't care. | |
| The standard doesn't care. | |
| The standard doesn't care. | |
| Right? | |
| The implementation, we'll talk about that in a second. | |
| Or a second. | |
| Maybe in 10 minutes. | |
| We'll talk about that later. | |
| But the standard, the definition of the queries and the storage format doesn't care. | |
| You need a path to a file. | |
| Yeah. | |
| Okay. | |
| To a directory. | |
| To a directory. | |
| I'm sorry. | |
| So that was simplicity. | |
| It really is that simple. | |
| And beyond that, you mentioned this already. | |
| We have tables with simple data types. | |
| We intentionally kept the data types of these tables simple. | |
| Yeah. | |
| No nest types. | |
| Nothing crazy. | |
| It's just varchar, big int. | |
| Varchar and big int. | |
| Yes. | |
| Yeah. | |
| Oh, there's a timestamp. | |
| Oh, no. | |
| Yeah. | |
| Yeah. | |
| Yeah. | |
| And then we also kept the query simple as well. | |
| Yeah. | |
| Because we want this to work on multiple different things. | |
| Exactly. | |
| Yeah. | |
| Okay. | |
| That was simplicity. | |
| Scalability. | |
| It doesn't, it does it scale, Mark? | |
| Mark, but does it scale? | |
| Does it scale? | |
| Does it scale? | |
| I think this is one of the sort of fundamental questions that people, like the lake house people | |
| will have about this, right? | |
| Because the biggest sort of, the only sort of criticism of this format is, does it scale? | |
| No. | |
| Because, of course, the existing stack, like Iceberg and Delta, they store files on S3. | |
| And as we know, blob stores, they scale infinitely. | |
| Like infinitely. | |
| There's no limits. | |
| No limits. | |
| Zero limits. | |
| No limits. | |
| Zero limits to the scalability of these things. | |
| And we are now taking some of those files and putting them in the database. | |
| So that's the question. | |
| Does it scale? | |
| Does it scale? | |
| Well, first, it's important to realize that we're not taking the data out to these database | |
| servers, right? | |
| We're taking only the metadata. | |
| And what do you think of in terms of metadata? | |
| Well, it's, I would say, around a factor of one to 100,000, more or less. | |
| So for every 100 kilobytes in your data, you have around a byte in your metadata. | |
| So there's, like, it grows, right? | |
| Like as your data grows, your metadata grows as well. | |
| But it's really like, it's like five orders of magnitude less data. | |
| Like if you have a petabyte of data, you're going to have maybe 10 gigabytes of metadata. | |
| Oh, no. | |
| Which is not zero, but it's not a lot, right? | |
| And then who actually has a petabyte of data? | |
| Very few people. | |
| It's actually true. | |
| People have scaling anxiety a lot. | |
| But our point is that this does scale because of the great reduction in sort of in volume | |
| that caused by the fact that we only store the metadata in the database and not the rest. | |
| And I think the does it scale? | |
| Yelling has also subsided a little bit with tools like, you know, BigQuery arriving, right? | |
| Like nothing keeps you from putting your metadata in BigQuery or Snowflake. | |
| And again, that's also one of the things that the server you put your data in, that is, | |
| you can put it in different database servers, right? | |
| You're not limited to putting it in Postgres. | |
| So maybe Postgres can't handle the metadata for 10 petabytes of data. | |
| Okay. | |
| Then you can put it in some other system like Spanner or CockroachDB, right? | |
| Like you can always use a different database system. | |
| If the time comes that this is actually necessary. | |
| But most likely you will not need to. | |
| No. | |
| I think the bigger point about scalability is that we can scale the storage, which is our | |
| blob storage. | |
| Yeah. | |
| Compute, which is the thing that actually runs queries. | |
| And the metadata managed independently. | |
| You have three dimensions of scalability and you can kind of choose your tools and your | |
| infrastructure accordingly. | |
| It disconnects. | |
| It actually adds another level of disconnect to this. | |
| Yeah. | |
| I mean, of course, the existing Iceberg catalog also added a disconnect to it. | |
| Yeah. | |
| But they never thought about scaling that one. | |
| I think, so what I think is interesting is that the direction that Iceberg and Delta are | |
| going or seem to be going is that they are going to do more and more stuff with these | |
| catalogs. | |
| Right. | |
| And that's kind of natural because when you talk to an Iceberg server, you don't want to | |
| for every single query do all these hops along all these different files, read all the JSON | |
| files, read all the Avro files. | |
| And there's already a database server. | |
| So you cache it there. | |
| I think that's kind of like it seems to be that they are slowly moving in the same direction | |
| of putting more and more stuff into these Calix servers as well. | |
| Yeah. | |
| It's just not part of the format necessarily. | |
| It's so it's not open. | |
| It's like it's an opaque thing that only certain metadata servers will do that is actually kind | |
| of critical to making this stuff performance. | |
| Right. | |
| Yeah. | |
| Yeah. | |
| It is super interesting to see those developments. | |
| And yes, indeed, there's also a bunch of companies around springing up that are indeed | |
| are trying to custom things to make the existing formats faster. | |
| Yeah. | |
| We are taking a radical departure here. | |
| Exactly. | |
| We're saying no more metadata files. | |
| Yeah. | |
| This is it is it does not make any sense in any world anymore if you have the database | |
| anyway. | |
| But we talked about scalability. | |
| One more thing I want to say about scalability is that, by the way, the design of Duck Lake is | |
| the exact same design that Snowflake and BigQuery have been using for a long time already successfully. | |
| Yeah. | |
| With the one difference that on the bottom we have the open format parquet. | |
| Yes. | |
| That is the one difference between the existing design of these hyper scalability databases | |
| and Duck Lake is that we have visibility into the database itself. | |
| You can look at it, which you can't do for them. | |
| And you have you can look at the files, which you also can't do for them. | |
| Right. | |
| Exactly. | |
| It's all open formats. | |
| It's all open formats. | |
| It's all transparent. | |
| The Duck Lake, the metadata is open. | |
| The parquet files, the data is open. | |
| But it's the same exact design. | |
| Same exact design. | |
| And I think it's fair to say that Snowflake and BigQuery, they have big data sets. | |
| They do. | |
| I think those are like one of the few companies that actually have big data sets. | |
| And they use this architecture. | |
| They use this architecture. | |
| It's been proven. | |
| It's been proven to work. | |
| Yeah. | |
| And Snowflake uses FoundationDB and BigQuery uses Spanner for this. | |
| Exactly. | |
| Yeah. | |
| And you can use either of those as well. | |
| Yeah. | |
| If you'd like. | |
| Right. | |
| Like you can use Spanner. | |
| You can rent Spanner. | |
| You can use Spanner for Duck Lake. | |
| From Google. | |
| Google will rent it to you. | |
| Yeah. | |
| You can use FoundationDB. | |
| It's open source. | |
| I haven't ever tried that one. | |
| I have not tried either. | |
| But you could. | |
| You could. | |
| Yeah. | |
| The third principle I want to talk about of Duck Lake is speed. | |
| Speed. | |
| Speed. | |
| Performance. | |
| Performance. | |
| We like performance. | |
| We like to go fast. | |
| We like things go fast. | |
| Yeah. | |
| I don't know. | |
| This is just... | |
| It really does not exist in the rest of my life. | |
| I'm not into motorcycles. | |
| I'm not into speed boats. | |
| I like slow boats. | |
| But for databases, we do care about performance. | |
| Absolutely. | |
| Yeah. | |
| And Duck Lake is actually quite fast. | |
| Not just because of the way it's designed. | |
| So you need like a single query to plan a complex query over many tables using this format. | |
| Like you can do this all in a single request to the database. | |
| You're just... | |
| Single round trip. | |
| Single round trip. | |
| Yeah. | |
| You're not reading a ton of Abro files from storage. | |
| And yes, some of them can be read in parallel. | |
| But there are like four steps that are dependent on each other. | |
| So you have at least four dependent round trips to storage for these formats. | |
| Exactly. | |
| In Duck Lake, you don't have that, right? | |
| You can run a query and this will be completing in milliseconds. | |
| And then you can run your query. | |
| Exactly. | |
| Yeah. | |
| I mean, if we go back to the iceberg architecture diagram, you can already see... | |
| You can see the lines. | |
| See the lines. | |
| You can see the lines. | |
| Like the first thing you do is you need to figure out the version. | |
| So you talk to the metadata server and you get the version. | |
| A service, by the way. | |
| The service. | |
| Sorry. | |
| The service. | |
| No, no, no. | |
| Yeah. | |
| This is... | |
| You're right. | |
| But it is... | |
| The service. | |
| Which then talks to a database again. | |
| So like... | |
| So in Duck Lake, all you do is do that step, right? | |
| Yeah. | |
| And then in Iceberg, what follows is first you need to actually read the JSON files. | |
| Then you go and read the Avro files. | |
| Like the first layer, because there's two layers of Avro files. | |
| Then you read the next layer of Avro files. | |
| And the reason they have all these indirections, by the way, is to work around the immutability of files on blob stores. | |
| Because metadata is a bit different from data. | |
| It does like... | |
| It's very small incremental changes usually. | |
| Yeah. | |
| But that does mean like imagine you do a million commits, right? | |
| Every single commit writes one Parquet file. | |
| If you have one Avro file for each of these commits, then you need to read a million Avro files just to start planning your query. | |
| So that's terrible. | |
| So in Iceberg, they do this thing where they layer all these Avro files. | |
| They kind of progressively get larger with more and more of the data. | |
| And that's to avoid these metadata queries taking so long. | |
| Yeah. | |
| But that also... | |
| It's a lot of data to read still, right? | |
| A lot of metadata needs to be read in a lot of these round trips. | |
| That's why these indirections exist. | |
| Yeah. | |
| So it's why DuckLake can be faster. | |
| Not because of the implementation, just because the way it's designed. | |
| Yes. | |
| DuckLake is also much better on dealing with smallish changes, right? | |
| Because we don't have to work around, again, work around the limitation of blob stores. | |
| Exactly. | |
| That we cannot have too many small files. | |
| But Postgres has no problem adding a single row to a table. | |
| In fact, it's very good at that. | |
| Absolutely. | |
| So adding a transaction to a table, it costs almost nothing. | |
| Exactly. | |
| Probably you already have that block. | |
| You're just writing to it. | |
| So you have no problem dealing with small changes. | |
| And again, also the cost of reading from that small change is negligible. | |
| I mean, it doesn't really matter in the grand scheme of things. | |
| So we can really deal with small changes and frequent changes. | |
| And I mean, running 1,000 transactions per second on a Postgres server is... | |
| No problem. | |
| No problem. | |
| Right? | |
| You can push that thing to 100,000 if you wanted to. | |
| But like 1,000 transactions per second, not a problem. | |
| So that's what we're talking about here. | |
| Yeah. | |
| I want to expand a bit on the small changes. | |
| So if you make a change to a lake house format like Iceberg, first, what you need to do is | |
| write a data file. | |
| So you write a parquet file. | |
| Then you write the first layer of Avro files. | |
| First layer. | |
| First layer. | |
| So that's already two files. | |
| Then the second layer, three files. | |
| Then the final layer, the version layer. | |
| So that's four files you need to write for every single commit you do. | |
| And then you update the database. | |
| So in Duck Lake, three of those layers are gone. | |
| Right? | |
| You only write the parquet file and the change to the database. | |
| So that's already like you're cutting the amount of round trips, the amount of files | |
| you're writing to the Blob Store by four. | |
| However, there's more. | |
| But wait, there's more. | |
| So one of the cool things that because everything is already in the database, one of the cool | |
| optimization we can do is store data in that database. | |
| I know this is radical. | |
| You can store data in the database. | |
| No. | |
| Does it scale? | |
| It scales enough for the concept of small changes, right? | |
| So if you're writing a low amount of rows, how about instead of writing that parquet file | |
| at all, you can just write those rows to the database. | |
| And this is a concept that we call data inlining. | |
| You inline the data into the database. | |
| And that's cool because now if you want to write one row to your format, all you do is | |
| write that row to the database. | |
| And because that database encapsulates all of the transactionality, right, of the whole | |
| system, right, like a transaction in Duck Lake is essentially equivalent to a transaction | |
| within that database. | |
| That means that data that you have just written, it behaves exactly the same as data that is | |
| written to a parquet file. | |
| Like you can query it directly. | |
| It's not like it's different from a cache, right? | |
| Because people have built these solutions on top of Iceberg where you have a cache, right? | |
| Like a Kafka stream. | |
| It buffers a bunch of data. | |
| And then every few seconds you write it out. | |
| However, during those few seconds, that data is not visible. | |
| This is not that, right? | |
| This is not the same because the data exists in the database immediately after it's written, | |
| it's visible, right? | |
| You can query it. | |
| It's like immediately visible. | |
| This is amazing. | |
| This is, and I think it's really solving the, one of the huge biggest problems I think that | |
| we, that we see with Iceberg and Delta deployments, which is the small changes. | |
| And again, that Postgres doesn't care about a few hundred extra rows. | |
| No. | |
| And it solves a huge problem. | |
| By using the database for more things, we are also spending a lot less time on the critical | |
| path of transaction commits, right? | |
| Yeah. | |
| So we can actually have far fewer conflicts to deal with far fewer, like back off if | |
| two transactions are trying to change the same data. | |
| It's going to be a much more efficient or less likely for these conflicts to happen because | |
| our time on the critical path of a transaction is much smaller with Duck Lake. | |
| Exactly. | |
| Because in Iceberg, if there is a conflict, you need to rewrite your version file, right? | |
| So what you do is you write your whole stack, right? | |
| Like we've talked about the stack a lot. | |
| You write a Parquet file, you write some ABR files, and then you write a JSON file, and | |
| then you write to the catalog server. | |
| If there is a conflict in this stage, and that's not a logical conflict, but like two | |
| tables, two connections that are inserting data, they also conflict in this way. | |
| You need to go back to your blob store and write a file again, right? | |
| And that doesn't need to happen in Duck Lake ever. | |
| If there is a conflict, you have written a Parquet file, two transactions are inserting | |
| data, there's a conflict. | |
| All you do is rerun the transaction in your database server. | |
| Like it's one more round trip. | |
| No more connection to the blob store. | |
| You reference the same Parquet file. | |
| So in the common case where there are no logical conflicts, right? | |
| Like there's no schema evolution. | |
| There's no complex operations happening. | |
| And there's just multiple connections inserting data. | |
| You never need to go to the blob store more than that one time to actually write the data. | |
| Yeah. | |
| This is great. | |
| This is really, I think, one of the biggest advantages of Duck Lake. | |
| Yeah. | |
| So these were the principles, simplicity, scalability, and speed. | |
| And again, we're talking about the format, not the implementation. | |
| We're going to talk about the implementation later. | |
| Yes. | |
| But before we do, we know that people do like making these matrices of features that like | |
| has this, does not have this. | |
| So we want to for, you know, in the interest of being explicit, which is sometimes that | |
| we're not always very good at a DuckDB is because we're not always great at explicitly | |
| stating what we are good at. | |
| So in the interest of not doing that this time, we will explicitly state the things that | |
| Duck Lake, the format, can do really well. | |
| Yes. | |
| So let's go over those. | |
| I don't, we don't have to talk about them at infinity length, but I'm just going to | |
| mention some, what we think is important. | |
| So Duck Lake, in the first instance, it's like you can run whatever query you want. | |
| There is no restriction on what kind of analytical queries you can run on your data, right? | |
| Like you can run arbitrary selects, you can run arbitrary updates, insertions, deletions, | |
| appends, it does really not matter. | |
| There's no, there's no technical restriction of, oh, you need to be doing this or you need | |
| to be doing that or you can't use the window functions or something like that. | |
| This is, keeps this intact. | |
| There's no change here. | |
| One of the bigger departures, and this is by the way, arbitrary queries, you can also | |
| run on iceberg. | |
| There's no, there's no difference there, but we can do that too. | |
| I think one of the bigger things that are different from existing formats is that Duck | |
| Lake is multi-schema, multi-table. | |
| A Duck Lake instance can manage thousands of tables in thousands of schemas. | |
| No problem, right? | |
| Yeah. | |
| So Duck Lake is a catalog format. | |
| Right. | |
| Iceberg is, Iceberg and Delta, they are table formats. | |
| Right. | |
| So Duck Lake is not necessarily a direct replacement for Iceberg alone. | |
| It's a replacement for the whole Lake House stack. | |
| Yeah. | |
| It's a replacement for Iceberg and Polaris or Delta plus Unity, right? | |
| Like it replaces the whole stack and it can do everything that that stack can do. | |
| Yeah. | |
| That's, I think this is important to note. | |
| And because we have, we are a catalog format, Duck Lake is a catalog format. | |
| It can also do transactions across multiple tables. | |
| No problem. | |
| Because we manage all the metadata for all the tables in one schema in a single place, | |
| a database. | |
| And making changes to those is something that is trivially sort of doable in Duck Lake and | |
| extremely hard in the other formats. | |
| We have multi-table transactions built in. | |
| You can do arbitrary complex changes to arbitrary complex table structures and arbitrary amount | |
| of tables, really. | |
| Yes. | |
| With no significant additional cost compared to doing single table changes. | |
| There's not really a difference. | |
| And that's also on the read path, right? | |
| Like whenever you read from Duck Lake, you get a consistent view of your entire catalog. | |
| You don't get a snapshot of one table and another snapshot of another table. | |
| Nope. | |
| You get a consistent view. | |
| So if you have a database that has, for example, a foreign key relation that you're going to join on, | |
| and you make sure that those keys are always inserted at the same time, | |
| or there's always a corresponding entry in the one table to the other table, | |
| you will also see that in your read queries, | |
| because you will always see a consistent snapshot of the entire catalog, right? | |
| That's amazing. | |
| Yeah. | |
| We have, in that sense, we also have, you mentioned that the schema level time travel, | |
| basically we can look at the snapshots on a schema level, not on a table level. | |
| Exactly, yeah. | |
| We also have a transactional multi-table schema evolution where we can change the structure | |
| of multiple, we can create tables, multiple tables, we can add columns to multiple tables, | |
| we can change column types of many tables in one single transaction, right? | |
| So there is no inconsistency that arises from that we have to make a change to multiple tables | |
| in two different sort of high-level transactions. | |
| We can make these changes in a consistent way across hundreds of tables. | |
| We can create a hundred table in a single transaction, not a problem, | |
| and we can drop them again in another transaction. | |
| Yeah, and it's important to note as well that the DDL, right, it's all transactional, right? | |
| So if you can create a table, no one will see it yet until you commit. | |
| You can alter a table, change the schema, no one will see it until you commit, right? | |
| You can roll that all back as well, like rollbacks. | |
| It's all transactional. | |
| Yeah, all happens. | |
| And that works with arbitrary complex types, by the way. | |
| So if you have, you can have lists, you can have structs, you can have maps, the whole nine yards of types. | |
| And of course, all the normal scalar types are supported as well. | |
| There's really, I would say, anything you can stuff in a parquet file is going to work. | |
| Exactly. | |
| I think that's generally the design here. | |
| And then, of course, we have other nice features in DuckLake. | |
| We have things like views. | |
| We can do SQL-level views in DuckLake because those are just stored in the database. | |
| Yeah, exactly. | |
| It's really amazing. | |
| We just, we can have, and again, transactional. | |
| So we can have a transaction that changes a table and changes a view, and it's not a problem. | |
| So there's really no restrictions on data types, and sort of you can have features that have so far been not available, let's say. | |
| We've also thought about scalability in terms of, in the features, we have full partitioning information available in the metadata, just like Iceberg, right? | |
| Exactly. | |
| So we have all this information about partitioning, so we can do pruning on many layers, actually. | |
| We can do pruning of scans on the table level, we can do it on the file level, we can do it on the partition level. | |
| We can easily select only the columns that we're interested in. | |
| There's all these sort of basic, I would say, table stakes at this point. | |
| Yeah. | |
| Standard sort of partition and pruning features available. | |
| Yeah. | |
| So we have all the statistics for all the different files in the metadata. | |
| Yeah. | |
| You can use that. | |
| Essentially, we run a SQL query to get the list of files, right? | |
| Like it's all SQL queries. | |
| And that uses the partitioning information and the statistics of the files to select only the files you're actually interested in. | |
| Only the columns in the files as well. | |
| Yeah. | |
| Right? | |
| Not just the rows on rows or row groups and columns. | |
| We even have information about the parquet footer size in the files. | |
| Yes, the footer size. | |
| It's brilliant. | |
| So you don't have to read the file in order to find out how big the footer is, which is great. | |
| There's also what we've already mentioned is the transactional data definition language. | |
| Again, you can create a table, populate it in a single transaction. | |
| There's all or nothing semantics on this. | |
| If the metadata transaction doesn't go through, nothing will have happened. | |
| And that's really amazing. | |
| Interesting. | |
| You mentioned already the inlining. | |
| Yeah. | |
| Where we can, for small changes, can optionally. | |
| This is optional. | |
| It's optional. | |
| It's optional. | |
| Can optionally be buffered in the database itself. | |
| And that is something. | |
| And yeah, it's different from the existing buffering approaches because it's transactional. | |
| Yeah. | |
| And these are already part of the table, even if you commit only a single row. | |
| Exactly. | |
| And at any point, you can flush those out again to a parquet file. | |
| And something that's quite cool that we do is that we allow snapshots to refer to part | |
| of a parquet file. | |
| And that sounds kind of technical, maybe. | |
| Really? | |
| Mark is being technical. | |
| But essentially, what that allows you to do is it allows you to have in one parquet file | |
| changes that have been committed in different snapshots, which means you never really need | |
| to worry about expiring or deleting snapshots. | |
| So because if you look at Eidsberg and Delta, their snapshots are all at the file level. | |
| So every single snapshot needs to have a file associated with it. | |
| That's not the case in Duck Lake. | |
| You can have snapshots that basically reference only a few rows in a file, which means you can | |
| have many more snapshots than you can have files, which means you don't need to worry about deleting | |
| snapshots again. | |
| It's very lightweight snapshots, essentially. | |
| Yeah, this is amazing. | |
| I think this is something that you came up with. | |
| I think that I'm still impressed by is this partial, is having snapshots that refer to parts | |
| of parquet files. | |
| That's really cool. | |
| We also have a cool feature, which we are Europeans. | |
| We care about privacy and that sort of thing. | |
| We realize that the Americans will stop listening at this point, but you have to bear with us | |
| for five seconds. | |
| We have encryption support in Duck Lake. | |
| Encryption, yes, encryption. | |
| We can encrypt the data. | |
| Finally. | |
| Finally. | |
| Finally, we can encrypt a data lake format. | |
| We have an encrypted data lake. | |
| And this takes a second to realize what this means. | |
| So basically in Duck Lake, you can say, hey, I want the data to be encrypted. | |
| Here's my encryption key. | |
| No, actually, you don't have to say that. | |
| No, you just say encrypted. | |
| You're right. | |
| You just say it's encrypted. | |
| And the encryption keys for the files are stored in the metadata server. | |
| Exactly. | |
| And this is interesting because the metadata server is in a different trust zone than your | |
| file storage, potentially. | |
| Yeah. | |
| And that means that you can have a zero trust data lake for your data lake. | |
| You can store your sensitive, potentially, tables on a file storage backend that is open | |
| to the world. | |
| And there's nothing that can happen here because Duck Lake will encrypt all these Parquet files | |
| using standard Parquet encryption. | |
| Exactly. | |
| Store the keys in the metadata. | |
| So anyone that gets access to your data lake has absolutely nothing. | |
| Yeah. | |
| Like you cannot do anything with these files. | |
| You might as well give them to people because it uses standard IES, industry standard. | |
| Yeah, it just uses industry standard. | |
| It uses whatever Parquet is doing. | |
| But this is like from a design perspective, it's really cool because nothing keeps you | |
| then from having a private table in Duck Lake exposed, let's say, using Cloudflare, using | |
| a caching service for better availability and all that stuff. | |
| And the only thing that you need to keep, that you need to keep, you should keep private, | |
| is the access, probably, is the access to your metadata storage. | |
| Yeah. | |
| And I think that's a pretty wild feature, actually. | |
| Yeah, and I think it's something that is very new because, of course, when we talk about | |
| the history of data lakes, as we did before, Hive cannot do this because Hive needs partitioning | |
| information in the directories, right? | |
| So if you write a Hive partitioned dataset to S3, you have all these directories that contain | |
| information, right? | |
| Like it's data, like partition key, one is A. It's all in the directories. | |
| It's all in the names. | |
| You cannot have that be encrypted because the whole system relies on it being encrypted. | |
| Iceberg can't do this because it builds on Avro and JSON files and those have no encryption | |
| support. | |
| Parquet has encryption support, but Iceberg needs to write Avro and JSON files, which don't | |
| have encryption support, to your data store. | |
| Yeah. | |
| Duck Lake can do this because the only thing it's writing to your blob store is Parquet | |
| files that can be encrypted. | |
| And you mentioned before, the keys for all the files, they're stored in the metadata store | |
| in your database and they're generated automatically, right? | |
| So every time you write to the blob store, the system will generate a new key, right? | |
| So every Parquet file is using a new key, right? | |
| And that key is then stored in the metadata store so that you can read it back. | |
| And it's like a single setting. | |
| It's literally just encrypted. | |
| You just pass a Boolean to the server and then it will just do that transparently. | |
| You don't have to worry about it at all. | |
| As long as you can connect to the Duck Lake instance, it will automatically fetch the correct | |
| keys when reading. | |
| It will generate the keys when writing. | |
| And it's like you don't have to think about it at all. | |
| It's like, yeah, zero thought required. | |
| Yeah. | |
| And zero trust storage, which I think is really something that I find exciting is that you | |
| don't need to trust whoever stores your files. | |
| Exactly. | |
| Yeah. | |
| Like they can't do anything with these files. | |
| And I think this is really, this is, I think, pretty groundbreaking. | |
| Because, yeah, you can have your local Postgres server in your organization that you trust | |
| to do the metadata management. | |
| Self-hosted, for example, self-host Postgres server. | |
| Yeah. | |
| Yeah. | |
| But still store your data on S3. | |
| Yeah. | |
| And then everybody in your organization can connect to that private Postgres server and fetch | |
| the encrypted files from S3 and it will work fine. | |
| Yeah, it's really cool. | |
| So it's something that we accidentally sort of came up with. | |
| Yes. | |
| Yeah. | |
| We accidentally bumped into that. | |
| Yeah. | |
| And then finally on the feature list, I want to say compatibility. | |
| So DuckLake uses the exact same file formats like Iceberg, right? | |
| Like it's just Parquet files. | |
| Yeah. | |
| I think we use the same deletion format as well. | |
| Yeah. | |
| So maybe to expand a bit on that, we use the exact same field ID sets that Iceberg does. | |
| So all the data files that DuckLake writes, they are exactly compatible with Iceberg. | |
| We use the same delete format as Iceberg V2 deletions or something. | |
| Because there's many different Iceberg deletion formats. | |
| We use the V2 deletions, which is, again, just a Parquet file. | |
| So DuckLake data files, they are compatible with Iceberg. | |
| And it won't be there at launch, but we're planning to make a importer for Iceberg. | |
| So you can do a metadata only import from Iceberg and an exporter to Iceberg as well. | |
| So you can export from DuckLake back to Iceberg database files. | |
| And that all works because the data files themselves, they use the same format. | |
| So they don't need to change. | |
| I think we also need to, again, differentiate between the standard and the implementation. | |
| Yeah. | |
| Which the standard says we use the Iceberg file format with the V2 deletes, as you mentioned, | |
| because they can't agree on what the deletes should look like for some reason. | |
| Which brings us neatly to the next and final topic of this conversation, which is, so we talked about DuckLake as a convention, as a standard, | |
| as a set of tables and queries and conventions around how to deal with storage and so on and so forth. | |
| Exactly, yeah, yeah. | |
| But we know that it's very easy to standardize, right? | |
| Like entire departments and companies of software architects that standardize into thin air and then some poor person in another country has to implement that. | |
| We are not like that. | |
| No, no. | |
| We are not like that. | |
| We think that you cannot invent things without having, you know, done the deed, if you want, right? | |
| Exactly. | |
| You cannot invent something without actually having tried it. | |
| It's actually for me personally, and I think for you as well, the way I understand the problem is to try to implement a solution for it. | |
| Absolutely. | |
| And this is just the way how we, I think, conceptualize the world of data as well. | |
| It's like, can we implement this? | |
| No. | |
| Okay, then we haven't understood enough, I suppose. | |
| Exactly. | |
| And I think one, I think actually part of our sort of, let's say, insights into the iceberg format specifically came from me trying to make a reader for it in the first place, right? | |
| Like I sat down years ago, two years ago, one year ago, no, maybe two years ago, and tried to implement an iceberg reader in DuckDB, and it was interesting, right? | |
| There's still a recording of a ranty presentation somewhere. | |
| But let's talk about the implementation. | |
| Yes. | |
| And the implementation is also called DuckLake. | |
| It's also called DuckLake. | |
| That's very confusing. | |
| Yes. | |
| Well, it's the DuckLake extension, right? | |
| For DuckDB. | |
| For DuckDB. | |
| Right. | |
| So in case you have not heard, we make a system called DuckDB. | |
| Yes. | |
| And DuckDB can have extensions for doing various things. | |
| We have an iceberg extension. | |
| We have an iceberg extension. | |
| We have a Delta extension. | |
| We have a Delta extension. | |
| We have an Avro extension. | |
| We have an Avro extension, which is used by the iceberg. | |
| For some reason, we have an Avro extension. | |
| Yes. | |
| Someone may have forced us to read Avro files. | |
| Yeah. | |
| I don't know. | |
| We have tons of extensions. | |
| We also have other extensions like Spatial or, you know, tons of community extensions. | |
| People contribute other stuff like the airport extension. | |
| Yeah, exactly. | |
| And so we made an implementation of DuckLake called DuckLake for DuckDB, which is going | |
| to be available from DuckDB 1.3.0, which was released yesterday in our world and last | |
| week in your world. | |
| And it's very easy to install because we do care about simplicity, as you might remember. | |
| You just type install space DuckLake semicolon enter. | |
| I mean, I'm going to be explicit. | |
| Yeah, fair enough. | |
| And what that will do is we'll fetch the binary for your platform, like with all the other | |
| extensions and then, you know, load it into your DuckDB instance and you will have DuckLake. | |
| That's it. | |
| There is no other dependencies. | |
| Yes. | |
| You don't have to start a Docker container. | |
| We really don't like Docker containers, as I mentioned. | |
| There's no like additional services, nothing. | |
| And then you can instantiate a DuckLake. | |
| Yeah, a DuckLake. | |
| A DuckLake. | |
| A DuckLake. | |
| A DuckLake. | |
| Why not? | |
| And there you have choices. | |
| Yes. | |
| So our DuckLake extension that you wrote, that can use various different metadata storage | |
| servers. | |
| For example, you can use DuckDB itself. | |
| You can just use a file. | |
| Exactly. | |
| Local. | |
| Yeah. | |
| DuckDB is a perfectly capable database. | |
| It has ACID. | |
| It has primary keys. | |
| It can manage your metadata for you as well. | |
| Yeah. | |
| You can use DuckDB for the metadata server. | |
| You can also use SQLite. | |
| You can use SQLite. | |
| Yes. | |
| Even though it has funky types. | |
| You can use MySQL. | |
| You can use Postgres and the various compatible friends of Postgres. | |
| Yes. | |
| You can use MotherDuck. | |
| Yes. | |
| If you want to use a hosted server, you can use MotherDuck as a backend for your metadata. | |
| Exactly. | |
| And that's one dimension of choice that you have in the DuckLake DuckDB extension. | |
| Yeah. | |
| So basically anything that DuckDB can talk to as a database that is competent enough to | |
| have primary keys. | |
| Yes. | |
| Can be used as a metadata storage. | |
| And to expand on that, this builds on top of the other extensions that we have built | |
| for DuckDB that allow DuckDB to connect to all these systems. | |
| Right. | |
| So we have a SQLite extension that allows DuckDB to connect to a SQLite database. | |
| We have a Postgres extension that allows it to connect to a Postgres database. | |
| We have a MySQL extension that allows it to connect to a MySQL database. | |
| This builds completely on top of those extensions. | |
| Right. | |
| So it will just use the functionality of those extensions to do the connection, to do the | |
| table creation, running the SQL statements. | |
| The transactions. | |
| Everything. | |
| The transactions. | |
| Yeah. | |
| Yeah. | |
| Yeah. | |
| Yeah. | |
| The transactions. | |
| So anything that DuckDB has an extension for, that will work. | |
| Yeah. | |
| That's one dimension. | |
| You need to have a metadata storage. | |
| Yes. | |
| And as I said, you can use a fully local setup with DuckDB file locally. | |
| That's perfectly fine and actually quite cool. | |
| Yeah. | |
| You can have a DuckLake that's fully local by just having a DuckDB file and a directory. | |
| We're getting to that. | |
| Oh, yes. | |
| We need a second dimension of configuration, which is where on earth you want your Parquet | |
| files to be stored. | |
| Exactly. | |
| Where on earth is quite literal in this, actually. | |
| Yes, yes. | |
| This is literally anywhere on earth. | |
| Yes. | |
| Which data center location. | |
| Which data center, all that stuff. | |
| So basically, what I've mentioned is already in the design of DuckLake earlier, but the DuckDB | |
| DuckLake extension, the implementation of it, the first implementation, there might be more. | |
| I think there might be more implementations. | |
| There might be more. | |
| I'm actually expecting that there might be more implementations. | |
| But our implementation can talk to local files. | |
| It can talk to network attached storage. | |
| It can talk to S3, Google Cloud Storage, Azure Blob, Cloudflare R2, you name it. | |
| Like anything, again, anything that DuckDB can read and write from in using our file system | |
| abstraction, which is an abstraction that DuckDB has. | |
| But anything that is a DuckDB file system abstraction and can write can be used to store DuckLake | |
| files. | |
| So it's up to you. | |
| Again, the simplest possible setup, you just point it to a local folder. | |
| Yeah. | |
| I want to say the canonical cloud-based setup would be like you pull up an Amazon RDS hosted | |
| Postgres and you use S3 as a storage backend and that will work as well. | |
| And I've actually tried this extensively. | |
| But if you wanted to run this on some Google service, that could work if you wanted to | |
| run this on anything else, really. | |
| You can also mix and match, right? | |
| You can have, as I mentioned earlier, you can have a local DuckDB file for the metadata. | |
| Maybe you don't trust S3 with your encryption keys or trust Amazon with your encryption keys, | |
| but then you can point that to, still point that to an untrusted S3 bucket. | |
| Exactly. | |
| And our point about this, you probably have this already in your organization. | |
| Yes. | |
| You probably have a blob store, you know, most likely the Microsoft people sold you Azure | |
| already. | |
| Yes. | |
| You probably already have that and you probably have a Postgres or something else like that | |
| somewhere. | |
| Probably the Microsoft people sold you on it as well because they bought Citus, right? | |
| So, you most likely have these things, but we have not talked about the third aspect. | |
| We talked about how you can scale storage, we can store metadata. | |
| But how do you scale compute? | |
| Well, that's kind of trivial because you can have an arbitrary amount of DuckDB instances | |
| doing this. | |
| Yes. | |
| Yeah. | |
| It's fully parallel. | |
| It's fully, you know, transactionally safe. | |
| It can deal with conflicts. | |
| It's, you can essentially have an arbitrary amount of DuckDB instances that talk to these, | |
| to your metadata server and storage. | |
| Exactly. | |
| So, you can have DuckDB running on your laptop. | |
| Your coworkers can have it run on their laptop. | |
| You can have services like Easy2 instances that run DuckDB there as well or Lambda functions. | |
| Wherever you can run DuckDB, you can instantiate compute for DuckLink, essentially. | |
| Soon in the browser. | |
| In the browser. | |
| Yeah. | |
| Yeah. | |
| And I think we, yeah, so that's what the Duck Lake extension does. | |
| But I think we should probably at this point zoom back out a bit. | |
| Yeah. | |
| So, this extension is available. | |
| You can use it. | |
| It's actually free and open source, we should mention. | |
| It is free and open source. | |
| It's MIT licensed. | |
| Absolutely. | |
| Like DuckDB, there is no like, this is not like a commercial expansion that we're doing | |
| or anything like that. | |
| We're still working under the same model. | |
| But I think what I wanted to maybe to wrap up a little bit is wanted to talk about what | |
| this means for DuckDB. | |
| Yeah. | |
| Because we've talked about this earlier is how DuckDB has so far been limited a bit in | |
| its multiplayer capability. | |
| Yeah. | |
| And we, being the international expert on protocols, didn't want to implement a client protocol. | |
| Yes. | |
| To fix this. | |
| But we have now, with DuckLake, have a actually quite revolutionary solution for this multiplayer | |
| DuckDB problem. | |
| Exactly. | |
| Because instead of saying, we'll basically spread out in the client server sense where | |
| we have, you know, a single database instance, like an old school data warehouse. | |
| Like at this point, to me, this is funny because the architecture seems so ancient, right? | |
| You put a server, maybe that's running in a clustered way somewhere in a data center. | |
| And then you have all these clients connecting to it. | |
| And there's some load balancing going on in order to, you know, auto scaling that whole | |
| nonsense. | |
| To talk to it. | |
| We have turned this completely around with DuckLake, where everyone runs a DuckDB instance. | |
| The compute you run, like, on your, wherever you want, really. | |
| Yeah, yeah. | |
| Like, that's what you scale. | |
| You scale, first and foremost, like, you don't have a client, you have the compute node in | |
| your local, close to you, close in your application, on your phone, in your app, whatever. | |
| You are the compute node. | |
| You are the compute node. | |
| Well, you don't, I'm not going to make you run SQL queries by hand. | |
| The only, we can do this at this point. | |
| You can stare at a SQL query and execute in our brain, yes. | |
| But you are the compute node. | |
| Like, you are, wherever your application runs, is it your app server? | |
| Is it your application? | |
| Is it your local, you know, laptop, whatever? | |
| You are the compute node. | |
| And this is already how people use DuckDB, right? | |
| Yeah. | |
| Like, this is already how organizations deploy DuckDB. | |
| They run DuckDB locally. | |
| What changes now, I think, is that now this is no longer a read-only sort of thing. | |
| Now it is a full-on data warehouse sort of thing. | |
| It is absolutely a data warehouse solution, but it is a 2025 data warehouse solution. | |
| It is the modern data stack now with updates. | |
| Now with updates. | |
| Yeah. | |
| Like. | |
| And in a sane way, in a transactional way, without, you know, having to do crazy things about | |
| small updates, without having to, you know, worry about, you know, the amount, the degree | |
| of parallelism. | |
| Yes. | |
| It doesn't matter, right? | |
| You have, so this is our, so DuckLake is really, if you think about both the standard convention | |
| and the implementation is our answer to multiplayer DuckDB. | |
| Yes. | |
| You can basically take an existing traditional data warehouse setup where you have like a | |
| server, again, you know, maybe self-hosted server of some sort that you talk to with clients, | |
| huge problems of fairness and, you know, resource scheduling, all of that, and replace that | |
| with DuckLake where you have a single, much more lightweight, much, much, orders of magnitude, | |
| more lightweight metadata server. | |
| Yes. | |
| A cheap, you know, storage backend that stores files. | |
| Yeah. | |
| Great. | |
| And you have your local DuckDB instances that are basically running these transactions in | |
| a globally synchronized way. | |
| Yes. | |
| Globally synchronized, sane, asset compliant, much faster than what you can do with existing | |
| solutions. | |
| Yeah. | |
| That's, that's where we're, that's what we're on about. | |
| Exactly. | |
| Yeah. | |
| Yeah. | |
| It's, it's, it's pretty crazy for us. | |
| I think so. | |
| I think it's the mix. | |
| It is, it is the next big step for DuckDB, right? | |
| I think it is the next big step. | |
| Yes. | |
| Right. | |
| And I mean, you, you pointed this out because I came to you with this crazy idea at some | |
| point, right? | |
| And then you were sitting on it for a week and you came back, no, no, Hannes, we need to | |
| do this. | |
| This is gonna be the next big step for DuckDB. | |
| I, yeah. | |
| And I, I absolutely believe so. | |
| Yeah. | |
| I think this is kind of what people have been asking for and what a lot of people have been | |
| building themselves. | |
| Right. | |
| Right. | |
| Because you can do this with DuckDB already. | |
| Yeah. | |
| As I said, like the DuckLake extension is actually not complex. | |
| No. | |
| It uses all the functionality DuckDB already has. | |
| DuckDB can already talk to Pulsegres. | |
| It can already talk to SQLite. | |
| It can already write Parquet files, right? | |
| These, this has existed. | |
| We, we know we can do this. | |
| What changes now is that there's a standard way of doing this. | |
| That's easy to set up, easy to use, portable, and you don't need to invent it. | |
| Like not every single organization needs to invent this again if they want to do multiplayer | |
| DuckDB. | |
| I just go install DuckLake. | |
| Exactly. | |
| Single line. | |
| Install DuckLake. | |
| I mean, how, how more, how much more easy can we make it? | |
| Yeah. | |
| I don't know. | |
| All right. | |
| So let's maybe talk before wrapping up. | |
| Let's talk about the next steps for DuckLake. | |
| Yeah. | |
| Yeah. | |
| Yeah. | |
| So first we have to launch it. | |
| First we have to launch. | |
| Yeah. | |
| That's going to be on, that's not happening yet. | |
| That's going to be next week. | |
| Yeah. | |
| So you don't, you know, you know, this is still secret. | |
| We are walking around with this big secret. | |
| We've been thinking about this for maybe a year at this point. | |
| I think it's been a while that it's been floating around. | |
| Yeah. | |
| It's like, I think it materialized like, | |
| more than half a year ago. | |
| Yeah. | |
| And then. | |
| It's exciting. | |
| And then I think three months ago or four months ago, | |
| we were like, okay, we should, we should do it. | |
| Yeah. | |
| Right. | |
| Yeah. | |
| So next steps for DuckLake. | |
| First, we have to release it. | |
| Yes. | |
| We're very, very curious what you think of this. | |
| Yes. | |
| The reaction. | |
| I mean. | |
| Yeah. | |
| It's like, it's like, like, it's like the party anxiety. | |
| You organize a party and then you hope people come. | |
| We hope you come. | |
| We think it's revolutionary. | |
| Yeah. | |
| Of course. | |
| Yeah. | |
| We're not the ones making that decision. | |
| But we are experts. | |
| Yes. | |
| Yes. | |
| So we have a chance. | |
| We have a chance. | |
| We have a chance. | |
| So we want to expand a bit. | |
| Mark has already mentioned this. | |
| The implementation is going to gain the ability to import and export from existing iceberg. | |
| Yes. | |
| Files because we use the same files. | |
| So we, it can be a metadata only import. | |
| It can be a very lightweight import and export. | |
| Exactly. | |
| Yeah. | |
| We also wanted to expand the implementation. | |
| Again, that's not the concept. | |
| That's the implementation. | |
| The implementation to be able to talk to more databases. | |
| Yeah. | |
| Right. | |
| So we're working on a generic ODBC integration for DuckDB that would allow DuckLake then to | |
| talk to arbitrary databases. | |
| Like if you wanted to glue your DB2 to DuckLake, you know. | |
| Why not? | |
| Why not? | |
| Why not? | |
| Knock yourself out. | |
| We wanted to also work a bit on rollbacks and restoring of old versions. | |
| Not rollback. | |
| The undo. | |
| The undo of transactions where we basically pretend like a sequence of snapshots hasn't happened | |
| and we just go back to an existing one. | |
| This can be done with the existing implementation. | |
| Yes. | |
| But it's not super elegant yet. | |
| So we want to work on that. | |
| And then, again, as it happened with DuckDB, right? | |
| We released DuckDB in 2018? | |
| 19. | |
| In 2019. | |
| In 2018 we started on. | |
| Sigma 2019. | |
| Yeah. | |
| Exactly. | |
| And we have, you know, we have, of course, been following our agenda, but we also have | |
| been very much influenced by what people wanted. | |
| Exactly. | |
| And obviously we're going to do the exact same thing with DuckLake where we're going to, | |
| you know, listen to you. | |
| Yes. | |
| You. | |
| And build the things that people care about. | |
| But we think it's a really compelling sort of project solution for art problem of architecture. | |
| Architecture. | |
| Yeah. | |
| Yeah. | |
| It's almost as if we did database architectures. | |
| Yes. | |
| Database architectures. | |
| Yes. | |
| We did come from the database architectures. | |
| By the way, that's the research group we originally came from called database architectures. | |
| Shockingly, we're still doing that. | |
| Yes. | |
| Okay. | |
| But that's, I think that's it. | |
| Did you want to add anything? | |
| No. | |
| I think this is super exciting. | |
| I'm very excited to see what people think and what people will use this for. | |
| Because I think when we released DuckDB, we did not anticipate half of the use cases that | |
| people have come up with. | |
| I have a suspicion this might be the same. | |
| Yeah. | |
| Like I can see a lot of very interesting architectures come forth out of this. | |
| Yeah. | |
| Yeah, me too. | |
| Cool. | |
| So that's it. | |
| Thanks, everyone. | |
| Thank you. |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment