DuckDB 2.0 alpha版本已推出,相比1.5.5版本在多个关键指标上实现了显著性能提升。1
通过引入异步I/O技术,该版本对云存储查询的优化最为突出。S3数据查询速度提升2-3倍,通过read_ahead_depth参数实现,该参数默认设置为-1以自动启用。1递归CTE引擎经历了重写,使得深层级父子数据查询速度提升40倍,特别适用于git历史和组织结构等场景。1新增的VARIANT数据类型采用字段分解技术,存储空间相比JSON字符串减少2.7倍,字段查询速度提升6倍,较1.5.5版本快78倍。1
此外,DuckDB 2.0还引入了多项新功能特性。外部文件缓存溢出功能在内存限制为300MB的场景下表现明显,854MB Parquet文件的第二次读取时间从23.9秒降至0.35秒。1该版本新增Triggers、嵌套schemas、CTE中的DML操作以及Spark SQL兼容模式等功能。1
DuckDB has released its 2.0 alpha version, introducing multiple optimizations that substantially accelerate query performance compared to version 1.5.5.1 The update leverages asynchronous I/O technology to boost S3 data query speeds by 2–3 times, with the feature automatically enabled through the read_ahead_depth parameter set to its default value of -1.1
A complete rewrite of the recursive CTE engine delivers a 40-fold performance improvement, making it particularly effective for querying hierarchical data structures such as git histories and organizational charts.1 Additionally, DuckDB 2.0 introduces a new VARIANT data type that achieves storage efficiency 2.7 times better than JSON strings while enabling field queries to execute 6 times faster than JSON equivalents—representing a 78-fold speedup over version 1.5.5.1 The update also implements an external file spillover cache feature that dramatically reduces read times for large files under memory constraints; when handling an 854 MB Parquet file with only 300 MB of available memory, the second read time decreased from 23.9 seconds to 0.35 seconds.1
Beyond performance enhancements, DuckDB 2.0 introduces several new capabilities including Triggers, nested schemas, DML operations within CTEs, and a Spark SQL compatibility mode.1
评论
还没有评论,欢迎留下第一条。