李飞飞团队近日发布名为T-Rex的机器人研究项目,通过为AI系统集成高频触觉反馈能力,使机器人能够完成翻书页、拧灯泡、挤牙膏等复杂精细任务[1]。该研究基于100小时包含触觉信号的机器人操作数据[1],在12项任务中的平均成功率达到65%,比最强对照模型高出30个百分点[1]。
李飞飞指出,AI的下一步发展需要理解物体的位置、相互关系,以及人类行动对世界的影响[1]。在多感官融合架构中,视觉负责识别"我要做什么",而触觉通过更高频率的信号不断修正"我现在做得对不对"[1]。这一研究表明,要使AI有效进入物理世界执行复杂操作,必须突破单纯视觉识别的局限,整合视觉、触觉和模拟器等多感官系统[1]。
Researchers led by Li Fei-Fei have unveiled T-Rex, a machine learning framework that integrates high-frequency tactile feedback to enable robots to perform delicate manipulation tasks such as turning book pages, screwing light bulbs, and squeezing toothpaste tubes [1]. The research demonstrates that artificial intelligence must move beyond visual recognition alone and instead combine multiple sensory modalities—vision, touch, and simulation—to effectively operate in the physical world and execute complex operations [1].
The T-Rex system was trained on 100 hours of robot operation data that included tactile signals [1]. In evaluation across 12 different tasks, the model achieved an average success rate of 65 percent, outperforming the strongest baseline model by 30 percentage points [1]. According to Li Fei-Fei's team, the framework divides sensory responsibilities such that vision answers the question "what should I do," while tactile feedback operates at higher frequency to continuously correct "am I doing this correctly" [1]. The researchers emphasize that AI's next step requires understanding where objects are located, how they relate to each other, and how the world changes when humans take action within it [1].
Figure has introduced the Helix 02 robot, equipped with fingertip sensors capable of detecting forces as low as 3 grams [1].