开发者 Niko1221 开源推出了 Strata 引擎,使消费级显卡能够运行量化的大模型。1 这款引擎支持在 12GB 及以上显存的消费级显卡上运行 Qwen3.8-Flash-Next 模型。1 该模型拥有 125B 参数,包含 51B 参数的 n-gram 嵌入表,原生支持 262K token 上下文长度。1
通过将 MoE 模型载入 RAM、仅在 VRAM 中加载高频专家以及使用投机解码等技术,Strata 引擎在消费级硬件上实现了显著的推理性能。1 在 RTX 5070(12GB VRAM)、Ryzen 5 7600 处理器和 64GB 内存的配置下,Q2_0 量化版达到 94 词元/秒的推理速度。1 IQ3_S 量化版本的推理速度为 53 词元/秒,而 AMD RX 9070 XT(16GB VRAM)的 Q2_0 版本可达 60 词元/秒。1 该引擎的最低硬件要求为:12GB 以上 VRAM 显卡、32GB 以上内存和 80GB 以上存储空间。1
Developer Niko1221 has open-sourced the Strata inference engine, enabling consumers to run the quantized Qwen3.8-Flash-Next large language model on affordable graphics cards 1. The model contains 125 billion parameters, including a 51-billion-parameter n-gram embedding table, and natively supports a context length of 262K tokens 1.
On an RTX 5070 with 12GB of video memory paired with a Ryzen 5 7600 processor and 64GB of system RAM, the Q2_0 quantized version of the model achieves an inference speed of 94 tokens per second 1. The same GPU configuration with IQ3_S quantization produces 53 tokens per second, while an AMD RX 9070 XT with 16GB of VRAM reaches 60 tokens per second using Q2_0 quantization 1. Strata accomplishes this performance by loading the MoE model into system RAM, keeping only high-frequency experts in VRAM, and employing speculative decoding techniques 1.
The engine requires a minimum of 12GB of VRAM on the graphics card, 32GB of system RAM, and 80GB of storage space 1. This approach makes running large-scale language models accessible to users with conventional consumer hardware.
评论
还没有评论,欢迎留下第一条。