Reddit 用户通过自制软件实现了 MacBook Pro 与 iPhone 17 Pro Max 的协同运行,借助 USB-C 接口将两台设备连接,共同执行 Qwen3.8-27B AI 模型推理任务1。该方案通过拆分计算任务、充分利用 iPhone 的 A19 Pro GPU 和神经网络引擎,使得模型的预填充性能获得显著提升1。在不同上下文长度下,性能改进幅度存在差异:8K 上下文时性能提升 35%,16K 上下文时最高提升 44%,32K 上下文时提升 29%1。
GPU 运算的引入带来了显著的推理速度改进1。iPhone 端在启用 GPU 加速后的运行速度相比未使用 GPU 时提升约 2.4 倍,单 token 写入耗时从 279 毫秒缩短至 176 毫秒1。这一自制方案已由开发者在 GitHub 开源发布,项目名称为"backburner"1。
A Reddit user has demonstrated a novel approach to accelerating artificial intelligence workloads by connecting an iPhone 17 Pro Max to a MacBook Pro via USB-C, enabling the two devices to jointly run the Qwen 3.8-27B AI model 1. The custom software splits computational tasks between the laptop and smartphone, leveraging the iPhone's A19 Pro GPU and neural engine to enhance performance 1.
The results show significant gains in prefill performance depending on context length: a 35% improvement at 8K context, 44% at 16K context, and 29% at 32K context 1. On the iPhone side, GPU acceleration increased processing speed by approximately 2.4 times compared to CPU-only execution 1. For a 140K context length, the time to write a single token was reduced from 279 milliseconds to 176 milliseconds 1.
The solution addresses a practical constraint of the MacBook Pro M4 Pro, which features only 24GB of unified memory 1. The open-source software, named "backburner," has been released on GitHub for developers to experiment with similar multi-device inference setups 1.
评论
还没有评论,欢迎留下第一条。