研究人员提出了Transformer Transformer模型,这是一个统一的AI系统,能够根据目标运动演示自动生成优化的机器人设计,包括每个连接、关节、马达和惯性属性[1]。该模型在cloth flinging任务中实现了显著性能提升,相比ALOHA2原始设计,追踪误差降低73%,最大关节速度降低30%[1]。
该模型基于RoboTokens统一标记化框架,采用扩散变换器架构,并通过Dynamics Self-Guidance机制实现零样本奖励优化[1]。RoboTokens框架涵盖11个MuJoCo机器人,质量范围从0.65 kg至67.5 kg,拥有6到35个活跃关节[1]。每个机器人被转换为28-101个RoboTokens序列[1]。同一网络通过改变掩码方式可扮演三个角色:生成器、评论器和跨实体控制器[1]。
研究表明,性能平台期约在一分钟推理时间后出现[1]。在强化学习专家数据生成方面,每个离散设计选择需要16小时的A100计算成本[1]。
Researchers have introduced Transformer Transformer, a unified artificial intelligence system capable of automatically generating optimized robot designs based on target motion demonstrations [1]. The model specifies every connection, joint, motor, and inertial property to match the demonstrated motion, representing a significant advance in robot co-design automation [1].
Testing on a cloth flinging task demonstrated substantial performance improvements compared to the original ALOHA2 design, with tracking error reduced by 73% and maximum joint velocity decreased by 30% [1]. The system operates on the RoboTokens unified tokenization framework, which encompasses 11 MuJoCo robots with masses ranging from 0.65 kilograms to 67.5 kilograms and between 6 and 35 active joints [1]. Each robot is converted into a sequence of 28 to 101 RoboTokens [1].
The architecture employs a diffusion transformer design enhanced by a Dynamics Self-Guidance mechanism that enables zero-shot reward optimization [1]. A single network performs three distinct roles by varying masking strategies: generator, critic, and cross-entity controller [1]. Performance plateaus approximately after one minute of inference time [1]. Creating reinforcement learning expert data for the system carries significant computational costs, requiring 16 hours of A100 GPU computing per discrete design choice [1].