Banseok Lee 和 Youngmin Kim 发布了 LittleBit,一项通过潜在因子分解技术将大语言模型压缩至 sub-1-bit 体系的方法1。该技术支持 0.1 至 1.0 bits-per-weight 的压缩范围1,已被纳入 NeurIPS 2025 会议1。
改进版本 LittleBit-2 通过潜在几何对齐进一步优化了压缩效果1,采用 Internal Latent Rotation with Joint Iterative Quantization(Joint-ITQ)进行初始化优化1,并已入选 ICML 20261。该方法将密集权重矩阵分解为低秩潜在因子,对因子进行二值化,并通过轻量级学习尺度恢复幅度信息1。在部署层面,LittleBit-2 仅修改初始化阶段,推理时无额外开销1,同时在推理时保持原始模型架构1。
该压缩方法已支持包括 OPT、Llama 系列、Phi-4、Qwen 系列、Gemma 系列在内的多个主流模型1。复现论文结果推荐使用 Python 3.12 和 transformers 4.51.x1。
Researchers Banseok Lee and Youngmin Kim have released open-source implementations of LittleBit, a technique for compressing large language models to sub-1-bit precision through latent factorization 1. The method decomposes dense weight matrices into low-rank latent factors, binarizes these factors, and recovers magnitude information through lightweight learned scales while maintaining the original model architecture during inference 1.
LittleBit achieves compression rates spanning 0.1 to 1.0 bits-per-weight and has been accepted to NeurIPS 2025 1. An improved version, LittleBit-2, introduces latent geometric alignment and employs Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ) for initialization optimization, with acceptance to ICML 2026 1. Notably, LittleBit-2 modifies only the initialization phase, incurring no additional inference overhead during deployment 1.
The technique supports compression of multiple model families including OPT, Llama series, Phi-4, Qwen series, and Gemma series 1. For reproducing the paper's results, the developers recommend using Python 3.12 with transformers version 4.51.x 1.
评论
还没有评论,欢迎留下第一条。