研究人员发布了OpenWAM框架,这是一个支持视频预测与机器人控制灵活组合的开放基础设施1。该框架在32块NVIDIA B200 GPU上进行了14天的预训练,使用约334万轨迹和14.64k小时源视频1。通过采用混合专家架构并应用反事实数据集增强,OpenWAM在LIBERO基准测试中实现了98.6%的平均成功率,在LIBERO-Long上超过95%1。
框架的核心创新在于支持独立训练的组件在推理时的灵活组合1。相比原始Wan2.2初始化,机器人视频预训练在LIBERO-Long上分别为VTA和Joint带来29.4个百分点和34.4个百分点的性能提升1。研究团队创建了包含32,000个反事实片段跨越10个任务的LIBERO-Long-CF数据集1,其中冻结的反事实逆动力学模型在新任务上达到84.0%的平均成功率,相比仅依赖演示的47.0%实现了显著提升1。在执行精度方面,端效应器位置在平均3.34厘米内(中位数1.90厘米),反事实监督将RGB均方误差降低34.5%,并将16个替代方案中正确未来的识别率从21.1%提升至71.3%1。
Researchers have released OpenWAM, an open infrastructure framework designed for studying world-action models that enables flexible combinations of video prediction and robot control capabilities 1. The framework underwent pretraining over 14 days using 32 NVIDIA B200 GPUs on approximately 3.34 million trajectories and 14,640 hours of source video 1.
The VTA model achieved a 98.6% average success rate on the LIBERO benchmark, with performance exceeding 95% on LIBERO-Long tasks 1. When compared to the original Wan2.2 initialization, robot video pretraining demonstrated substantial improvements on LIBERO-Long, boosting performance by 29.4 percentage points for VTA and 34.4 percentage points for Joint training 1. To enhance model robustness, the researchers created LIBERO-Long-CF, a counterfactual dataset containing 32,000 counterfactual segments spanning 10 tasks 1. A frozen counterfactual inverse dynamics model applied to new tasks achieved an 84.0% average success rate, representing a significant improvement over the 47.0% baseline with demonstrations alone 1.
The framework's compositional approach demonstrates practical precision in robotic execution, with end-effector positioning averaging within 3.34 centimeters (median error of 1.90 centimeters) 1. Counterfactual supervision improved visual prediction quality, reducing RGB mean squared error by 34.5% and increasing the identification rate of correct future frames from 21.1% to 71.3% across 16 alternative predictions 1. In 61.6% of perturbation cases, objects moved beyond their demonstrated configurations, underscoring the model's ability to generalize beyond initial training conditions 1.
评论
还没有评论,欢迎留下第一条。