The Spectrum Dispatch News

technology

DuckDB 2.0 Alpha Shows Speed Gains in Async I/O, Recursive CTEs, and VARIANT Type

The upcoming DuckDB 2.0 release brings asynchronous I/O for faster S3 queries, a rewritten recursive CTE engine for deep hierarchy traversals, and a VARIANT data type that shreds J

DuckDB 2.0 Alpha Shows Speed Gains in Async I/O, Recursive CTEs, and VARIANT Type

According to the source, DuckDB 2.0 is expected later this year with an alpha already available for testing. The author benchmarked the new version on an M5 laptop using a home internet connection that equally affected both releases, noting that users should run their own tests before quoting the numbers.

DuckDB 2.0 Alpha Shows Speed Gains in Async I/O, Recursive CTEs, and VARIANT Type

One of the highlighted improvements is asynchronous I/O for reading data from Amazon S3. The source explains that in DuckDB 1.5.5 each worker alternated between downloading a row group and decoding it, leaving the CPU idle while waiting for the network. In 2.0 a separate pool of threads prefetches row groups, keeping the decode workers busy. As a result, a query that reads a 2.2‑GB Parquet file (one column, 228 million rows) dropped from 18.8 seconds in 1.5.5 to 7.7 seconds in the 2.0 alpha. Similar speedups appear for larger workloads: 23 Parquet files totaling 13.6 GB went from 11.8 seconds to 3.9 seconds, and a 1.7‑GB CSV file improved from 116 seconds to 55 seconds. The source notes that for very small files (around 1 MB each) the gain is minimal because the overhead is dominated by per‑file round trips, and it advises against storing a data lake as thousands of tiny Parquet files.

The second area of focus is the recursive common table expression (CTE) engine. The source says the team rewrote this component and claims up to a 40× speed increase for graph‑reachability workloads. In a test walking the ancestry of a 20 000‑commit Git history, DuckDB 1.5.5 required between 1.8 and 16 seconds across runs, while the 2.0 alpha consistently finished in about 0.10 second. The improvement comes from reading the source table once, building a lookup on the parent column, and then only probing the few rows newly discovered in each iteration, rather than re‑scanning the whole table every round. The source recommends keeping hierarchy tables as simple parent/child integer columns and using USING KEY when the recursion carries extra values such as depth or cost.

Finally, the source introduces VARIANT as a first‑class data type that replaces the need to store JSON as opaque text. When a VARIANT column is written to disk, DuckDB “shreds” fields that appear with the same type in most rows into real columns underneath, leaving only the irregular or infrequent fields in a binary remainder. This layout stores the consistent part of the data like a normal table, which reduces storage overhead and speeds up queries that access the shredded fields. The source illustrates the concept with a typical event JSON containing purchase details, user information, properties, and tags, noting that fields such as event type, user.id, and props.amount would be shredded because they are consistently typed across millions of rows.

Overall, the source characterizes DuckDB 2.0 as delivering measurable performance gains for common data‑engineering tasks without requiring changes to existing queries, provided the data are shaped to take advantage of the new features.

Key facts

Sources

← All posts