小米 MiMo 大模型负责人罗福莉宣布,MiMo-V3 将采用全新架构,其核心组件 HySparse 2 已于今日发布1。新架构通过 KV 桥接和 KV 重用等技术创新,在处理超长文本时性能显著提升1。
在 100 万 Token 长度的长上下文处理中,新架构将预填充计算量降低至原来的 1/5.02,KV 缓存缩小至原来的 1/4.51。同时,HySparse 2 采用了 Token 级别的选择机制替代原有的块级别选择,并用近期 Token 的强制窗口替代了独立的 SWA 分支1。此前小米发布并开源了 Xiaomi MiMo-V2.6 系列全模态模型,包含 Pro 与 Flash 两个原生全模态模型1。
Luo Fulie, head of Xiaomi's MiMo large language model division, has announced that MiMo-V3 will adopt an entirely new architecture, with the core component HySparse 2 releasing today.1 The redesigned framework introduces technological innovations including KV bridging and KV reuse mechanisms that substantially enhance efficiency at extended context lengths.1
At a context window of 1 million tokens, the new architecture reduces prefill computation (FLOPs) to 1/5.02 of its previous level.1 Similarly, KV cache requirements shrink to 1/4.5 of the original size.1 The architecture incorporates token-level selection in place of block-level selection and replaces a separate sliding window attention branch with a forced window using recent tokens.1 These refinements deliver improved retrieval performance in long-context scenarios.1
The announcement follows Xiaomi's release and open-sourcing of the MiMo-V2.6 series, which includes Pro and Flash variants as native multimodal models.1
评论
还没有评论,欢迎留下第一条。