llama.cpp项目的提示词查询解码(prompt lookup decoding)功能近日获得显著性能优化。1开发者通过一系列技术改进,包括消除不必要的map拷贝、替换hash map实现、采用排序向量和constmap等措施,将解码速度提升最高达42倍,同时内存使用量减少了2.6倍。1在Apple M4 Pro(14核,48GB内存)设备上的测试显示,使用constmap优化后,加载时间从3.76秒大幅降低至0.23秒。1
随后,开发者Daniel Lemire提交的优化提案在此基础上再次实现了4.2倍的性能提升,使总体加速倍数达到140倍。1这些优化应用于WikiText-103语料库对应的467 MB静态缓存文件,在模型上下文大小为4096 tokens的配置下验证了显著效果。1
Developers have achieved significant performance improvements in llama.cpp's prompt lookup decoding through a series of targeted optimizations 1. The initial improvements, led by Hayder Tirmazi, delivered up to a 42-fold speed increase and reduced memory consumption by up to 2.6 times 1. These enhancements were accomplished through multiple technical refinements, including eliminating unnecessary map copies, replacing hash map implementations with more efficient alternatives, and adopting sorted vectors and const maps 1.
A subsequent pull request contributed by Daniel Lemire further accelerated the system by an additional 4.2 times on top of the existing optimizations, bringing the cumulative performance gain to 140 times 1. Testing conducted on an Apple M4 Pro with 14 cores and 48GB of memory, using a model context size of 4,096 tokens, demonstrated these improvements in practice 1. The optimization work reduced the loading time for static cache from 3.76 seconds to 0.23 seconds through the implementation of const map optimization 1. The static cache for the WikiText-103 corpus required 541 MB of storage, with the corresponding cache file consuming 467 MB 1.
评论
还没有评论,欢迎留下第一条。