Strata是一个开源工具项目,可在配备RTX 4090等消费级显卡的个人电脑上运行125亿参数的Qwen 3.8-Flash-Next大模型,推理速度达到100 token/秒1。该项目通过在GPU、内存和固态硬盘间动态分配工作负载,使原本需要服务器级硬件才能支撑的大模型在家用计算机上可行1。
该模型在RTX 3090(24GB显存)上的预期推理速度为100-140 token/秒1。模型总大小约70GB,需要35-55GB内存才能完整加载1。同时支持32K token的上下文窗口和4K token的答案生成能力1。
Strata采用MIT许可证以完全开源形式发布,并支持NVIDIA和AMD两类显卡,兼容Windows和Linux操作系统1。用户可在本地完整运行该工具进行聊天、代码编写和图像处理任务,全程无数据外传1。
Strata, an open-source project, enables users to run the Qwen 3.8 Flash Next large language model with 12.5 billion parameters on standard personal computers equipped with consumer-grade GPUs like the RTX 4090 1. The tool achieves inference speeds of 100 tokens per second on RTX 4090 hardware, with RTX 3090 devices (24GB) expected to deliver between 100 and 140 tokens per second 1.
The project addresses the challenge of deploying enterprise-scale models on consumer hardware through intelligent workload distribution across GPU memory, RAM, and SSD storage 1. The Qwen model requires approximately 70GB of storage and between 35 to 55GB of memory for loading 1. The system supports 32,000-token context windows and can generate responses up to 4,000 tokens 1. All processing occurs locally without any data transmission to external servers 1.
Strata is released as free open-source software under the MIT license and supports both NVIDIA and AMD graphics cards on Windows and Linux operating systems 1. The tool enables various applications including chatbot interactions, code generation, and image processing directly on consumer machines 1.
评论
还没有评论,欢迎留下第一条。