Black Forest Labs 与 mimic robotics 宣布推出 FLUX-mimic,这是一款基于 FLUX 3 多模态基础模型的视频-动作预测系统 1。FLUX 3 在统一训练框架中可同时生成图像、视频和音频,并能预测机器人动作 1。
在模型开发中,视频预测在联合训练的总计算成本中占比超过 95% 1。研究团队通过 Self-Flow 实验发现,动作预测相比无此模块的视频模型能节省约一半的训练步数 1。添加动作预测后,模型在 3500 步内恢复了视频生成质量,性能下降随后完全消失 1。
在推理性能上,FLUX-mimic 的 backbone 在单个 NVIDIA RTX 5090 GPU 上可在 80 毫秒内完成从输入到世界表示的处理 1。mimic 优化的完整部署系统实现了 101 毫秒的机器人反应时间 1。相比视觉-语言-动作模型,mimic-video 论文报告的样本效率提升达到 10 倍 1。
该系统已在奥迪生产线上开展测试部署,用于处理零件装配、电子控制单元插装、部件组装和柔性材料处理等任务 1。奥迪的 Christoph Schneider 表示:"这些机器人解决了传统机器人不可能处理的复杂软体操作任务" 1。
Black Forest Labs and mimic robotics have unveiled FLUX-mimic, a next-generation video-action model built on the FLUX 3 multimodal foundation model 1. FLUX 3, trained jointly, generates images, video, and audio while simultaneously predicting robot actions 1.
The system has been deployed for testing on Audi production lines, where it handles complex tasks traditionally difficult for conventional robotics, including soft material manipulation and precision assembly operations 1. According to Christoph Schneider of Audi, "these robots solve complex soft-body manipulation tasks that traditional robots cannot possibly handle" 1. Deployment tasks encompass parts assembly, electronic control unit insertion, component assembly, and flexible material handling 1.
Performance and Efficiency
During joint training, video prediction accounts for over 95% of total computational cost 1. When action prediction was added, the model recovered full video generation quality within 3,500 training steps, with performance degradation subsequently disappearing 1. The FLUX-mimic backbone processes input to world representation in 80 milliseconds on a single NVIDIA RTX 5090 GPU 1, while mimic's optimized full deployment system achieves a robot reaction time of 101 milliseconds 1.
Training efficiency gains are notable: the Self-Flow approach reduced training steps by half compared to video models without Self-Flow 1, and the mimic-video approach demonstrated a tenfold improvement in sample efficiency relative to vision-language-action models 1.
评论
还没有评论,欢迎留下第一条。