Choose DuckDB rather than SQLite

(tracewayapp.com)

82 points | by rubenvanwyk 4 hours ago ago

28 comments

  • brightball 4 hours ago ago

    > DuckDB's columnar engine

    That is workload specific. Title should be "Choose DuckDB rather than SQLite for Analytics" IMHO

  • shubhamjain 4 hours ago ago

    Despite its obvious advantages, the biggest drawback of DuckDB is its concurrency model [1]. If a process opens a database in read-write mode, it acquires an exclusive lock on the file. This prevents even simple read operations from other processes as long as the writer remains open. Maybe there's a simple workaround I haven't come across, but I found it to be quite a productivity killer.

    So yes, all these benchmarks are great, but it wasn't so fun working with DuckDB when I had to close duckdb cli, just so a query in another script could run.

    [1]: https://duckdb.org/docs/current/connect/concurrency

  • ethin 4 hours ago ago

    How are these two DB engines even comparable other than at the edges? They handle two completely separate workload types: one is more a general-purpose DB engine and the other is specifically for columnar datasets, analytics and the like -- of course a hand-tuned DB engine is going to destroy SQLite on any reasonable benchmark: SQLite wouldn't be optimized for that hand-tuned use-case whereas something like DuckDB is.

  • tptacek 4 hours ago ago

    Title, which is already synthetic, should be "Choose DuckDB rather than SQLite for Analytics Workloads".

  • k3liutZu 3 hours ago ago

    I couldn't read the article as it read like AI.

    And I am fatigued by the AI style in all code comments, reviews, PRs etc :(

  • oathvz 4 hours ago ago

    This is a bait and switch article. Compare apples and oranges, "oh btw look at our product".

  • lanstin 3 hours ago ago

    DuckDB is modern but written in C++ and crashes in production more than SQLite, which is old and it doesn’t really crash. you have to build in resilience to use DuckDb.

  • biophysboy 3 hours ago ago

    As many others have said here, you should just think about transactional vs analytical as well as single user in-process vs multi-user client-server when making choices.

    That said, I do think duckdb has a wider range of use cases than people here might think. It can whip through fairly large datasets (I use it for ~1B row tables all the time)

  • adsharma 3 hours ago ago

    SQLite can be replicated. rqlite, dqlite, litestream etc

    DuckDB can't be. PR was sent a year ago. Blocked on the same concurrency model issue in the other sub thread.

    Specifically on windows, the database can't read its own WAL file from a different thread in the same process.

    Love DuckDB for being permissively open source, great tech and performance!

  • egeozcan 4 hours ago ago

    TL;DR: SQLite was doing OLAP work it was never built for (and it was okay at it, I must add), and the perf. ceiling moves two orders of magnitude when you use something that's more fit for the purpose (DuckDB in this case).

    TLDRTL;DR: If everything you do is column-store territory, use a column-store.

  • crustycoder 3 hours ago ago

    Heartening to see so many "This is an apple, that is an orange" comments. Spot on folks!

  • raro11 3 hours ago ago

    > $16.49/month server [...] hetzner CCX13

    That server is now $51.09 for those wondering See https://news.ycombinator.com/item?id=48540844

  • pixelesque 4 hours ago ago

    Even for row-based data?

  • 4 hours ago ago
    [deleted]
  • 4 hours ago ago
    [deleted]
  • 3 hours ago ago
    [deleted]
  • cynicalsecurity 4 hours ago ago

    Great job on comparing apples to oranges. DuckDB is a columnar OLAP engine, SQLite is row-oriented OLTP. DuckDB should stomp SQLite in that particular use case.

  • otterley 4 hours ago ago

    AI slop. The content might be valuable but the framing makes it too painful to read.

    Please, folks, write with your own voice -- especially if it's for your business blog. It's good for you as an author (practice makes perfect) and it's good for your readers (whom you want to influence).

  • d1l 3 hours ago ago

    The AI slop is tiring man wtf. It’s so fucking lazy. The benchmark is comparing apples to oranges and doesn’t seem to be aware of it, and the way it’s written just reeks of LLM.

  • an hour ago ago
    [deleted]
  • datadrivenangel 3 hours ago ago

    This is AI slop, but I really want to know more about their write patterns and how they were doing batches.

  • esafak 3 hours ago ago

    How about writing with sqlite and querying with duckdb-sqlite, using a replica or WAL, to avoid locks?

  • 4 hours ago ago
    [deleted]
  • clumsysmurf 4 hours ago ago

    Too bad Android doesn't have JDBC APIs. Getting the native binaries compiled on Android is the easy part, but there is no straightforward way to access it from Java. The Room APIs are tied to SQLite as well.

  • sharpvik 4 hours ago ago

    great research there

  • cronping 4 hours ago ago

    [flagged]

  • bbqbbqbbq 4 hours ago ago

    [dead]

  • knuckleheads 4 hours ago ago

    I ran this through an AI checker and it flagged half of it immediately. @dang, I know Substack just enabled Pangram integration, is there anyway you could get Y Combinator to spring for a Pangram subscription for the front page or something ?