Google DeepMind推出了Gemini Robotics 2,一个主打"全身智能"的机器人通用模型[1]。该模型包含三个子模型:ER 2负责环境理解与任务规划,Robotics 2 VLA模型将视觉与语言转换为动作,On-Device 2则将动作模型本地化并可用少于200个样例在几小时内适配新机器人[1]。相比前代只控制上半身的系统,新模型实现了机器人双腿、躯干、双臂和手指的协调控制,使机器人能在移动中执行复杂操作[1]。
配备五指灵巧手的Apollo 2从桌面、地面和货架取物的平均成功率分别为68.4%、45.7%和76.3%[1]。在更精细的操作上,五指灵巧手完成拧下灯泡、拧上灯泡、扎垃圾袋、使用簸箕和封自封袋的成功率分别为92%、36%、44%、32%和40%[1]。搭载两指夹爪的Franka Duo机器人在通用抓放、工具配套和精密插入任务中的成功率分别达到74.2%、78.9%和89.6%[1]。谷歌表示同一模型已用于搭载不同灵巧手的Apollo 2和使用两指夹爪的Franka Duo[1]。
Google DeepMind has introduced Gemini Robotics 2, a unified machine learning model designed to enable comprehensive body coordination in robotic systems [1]. The system represents a significant advancement from previous generations, which focused primarily on upper-body control, by now enabling coordinated movement and manipulation across a robot's legs, torso, arms, and fingers [1].
The Gemini Robotics 2 framework comprises three specialized components: the ER 2 module handles environmental perception and task planning, the Robotics 2 Vision-Language-Action (VLA) model translates visual and linguistic inputs into motor commands, and the On-Device 2 system localizes the action model onto individual robot hardware, requiring fewer than 200 examples and hours of data to adapt to new robot platforms [1]. Google reports that the same foundational model has been successfully deployed across different robotic configurations, including the Apollo 2 equipped with five-fingered dexterous hands and the Franka Duo with dual-finger grippers [1].
Performance testing reveals varying success rates across manipulation tasks. The Apollo 2 achieves average success rates of 68.4% for tabletop object retrieval, 45.7% for ground-level pickup, and 76.3% for shelf retrieval [1]. The five-fingered dexterous hand demonstrates task-specific performance, including 92% success in unscrewing a light bulb, 36% in screwing on a light bulb, 44% in puncturing a garbage bag, 32% in dustpan use, and 40% in sealing a resealable bag [1]. The dual-finger gripper configuration achieves 74.2% success in general grasping and placement, 78.9% in tool-specific manipulation, and 89.6% in precision insertion tasks [1].