小米MiMo大模型负责人罗福莉宣布,团队在4月开源MiMo-v2.5后沉寂了近半年,专注于研究一个核心课题——强化学习究竟能扩展到多远1。MiMo-V2.6目前处于强化学习中间阶段,扩展了计算、环境工具和评分三项能力1。
在训练规模方面,每步处理约20亿个Token,采用1,568个Prompt乘以16条Rollout的配置1。MiMo-V2.6的训练成本已超过125万美元1。小米表示将在接下来几周内逐步开源相关训练细节1。
Xiaomi's large language model division head Luo Fuli announced that following the open-source release of MiMo-v2.5 in April, the team has spent the past six months intensively researching the scalability potential of reinforcement learning 1. The company is now showcasing MiMo-V2.6, which remains in an intermediate stage of reinforcement learning training and has incorporated expanded capabilities in computation, environmental tools, and scoring mechanisms 1.
The development of MiMo-V2.6 has already incurred training costs exceeding $1.25 million 1. The model processes approximately 2 billion tokens per step across 1,568 prompts with 16 rollouts each 1. Luo Fuli described the team's extended focus, stating that they had "remained quiet for nearly half a year to research one critical question—how far can reinforcement learning actually scale" 1. Xiaomi plans to progressively open-source details of the training process in the coming weeks 1.
评论
还没有评论,欢迎留下第一条。