Google DeepMind推出了Gemini Robotics 2机器人通用模型,采用"全身智能"架构,通过统一的视觉-语言-动作(VLA)模型来协调人形机器人的双腿、躯干、双臂和手指完成复杂动作[1]。这一设计突破了以往将机器人各部位作为独立模块控制的方式,转而用单一模型统一全身动作控制[1]。
该系统由三个子模块组成:ER 2负责环境理解和任务规划,Robotics 2 VLA模型将视觉与语言信息转化为具体动作指令,On-Device 2则将动作模型部署至本地设备并支持快速适配[1]。在官方测试中,Apollo机器人从桌面、地面和货架取物的平均成功率分别达到68.4%、45.7%和76.3%[1]。多指灵巧手的任务表现方面,拧下灯泡的成功率为92%,拧上灯泡、扎垃圾袋、使用簸箕、封自封袋的成功率分别为36%、44%、32%和40%[1]。
On-Device 2模型的适配能力突出,仅用少于200个样例和几小时的数据就能适配新的机器人本体[1]。该模型已被应用于搭载不同灵巧手的Apollo 2和使用两指夹爪的Franka Duo等多种机器人平台[1]。官方发布的演示视频中所有机器人均为自主运行,动作按真实速度播放[1]。
Google DeepMind has introduced Gemini Robotics 2, a universal model designed to coordinate humanoid robots' legs, torso, arms, and fingers for executing complex tasks through integrated full-body control [1]. Rather than treating individual limbs as separate modules, the system employs a single vision-language-action (VLA) model to unify motion control across the entire body [1].
The architecture comprises three specialized components: ER 2 handles environmental understanding and task planning, the Robotics 2 VLA model converts visual and language inputs into motor actions, and On-Device 2 compresses the action model onto local hardware while enabling rapid adaptation to new robot platforms [1]. Testing with the Apollo robot demonstrates varying success rates across different retrieval scenarios—68.4% for objects on tables, 45.7% for items on floors, and 76.3% for shelf pickups [1]. Fine-grained manipulation tasks show more variable results, with the system achieving 92% success in unscrewing lightbulbs, but lower rates for installation (36%), tying garbage bags (44%), using dustpans (32%), and sealing zip-lock bags (40%) [1].
The On-Device 2 model requires fewer than 200 sample examples and only hours of training data to adapt to new robot platforms [1]. The same model has already been deployed across different hardware configurations, including Apollo 2 equipped with various dexterous hands and Franka Duo with two-finger grippers [1]. All demonstrations presented feature robots operating autonomously with actions displayed at actual execution speed [1].