具身智能行业正面临一场数据危机。与大语言模型可从互联网海量获取文本数据不同,具身智能需要千万至亿小时级别的真实物理交互数据才能获得通用泛化能力,但全球可用的高质量物理交互数据总量仅约50万小时,缺口超过99%[1]。这一瓶颈已成为制约行业发展的核心障碍,促使企业探索多条技术路线来解决数据积累难题。
为了突破数据瓶颈,业界正在尝试多条采集途径。北京人形机器人创新中心在数采场地投入数千万元,部署了120台机器人[1];智元机器人则部署约200台机器人,配备数百名数据采集员[1]。另一方面,乐聚OpenLET已开源超6万分钟的真机数据集[1]。然而,现有路线面临"规模、质量、成本"难以兼顾的困局。仿真环境下任务完成率可达89.4%,但在真实家庭场景中仅约12.4%[1],说明虚实间存在显著差异。
针对数据获取的瓶颈,行业正在探索通过闭环自我迭代实现"数据奇点"的可能性。LWD系统训练的策略在长程复杂任务中平均成功率达0.95,超越传统行为克隆的0.76[1],展现了优化训练方法的潜力。这些探索表明,具身智能企业正在寻求突破数据困局的新路径。
The embodied artificial intelligence industry faces a critical bottleneck that may prove more damaging than computational constraints: an acute shortage of high-quality physical interaction data.[1] While large language models can readily access text data from the internet, embodied AI systems require millions to billions of hours of real-world physical interaction data to achieve general-purpose capabilities.[1] Currently, the global supply of usable high-quality data for such applications totals only approximately 500,000 hours, leaving a shortfall exceeding 99 percent.[1]
The industry is pursuing three primary technical approaches to address this challenge, each presenting distinct tradeoffs. Remote teleoperation, UMI data collection systems, and simulated synthesis represent the main pathways forward, yet they collectively struggle with an intractable triangle of constraints: scale, quality, and cost cannot be simultaneously optimized.[1] Major players have invested significantly in data collection infrastructure—Beijing's humanoid robot innovation center has deployed 120 robots across facilities requiring tens of millions of yuan in investment, while another robotics company operates approximately 200 robots supported by hundreds of data collection personnel.[1] The severity of the quality gap is evident in performance metrics: systems trained in simulated environments achieve task completion rates of 89.4 percent, but this drops sharply to approximately 12.4 percent when deployed in actual household settings.[1]
Some organizations are beginning to explore alternative pathways through self-iterative learning cycles that could potentially unlock exponential data generation. Leju Robotics has open-sourced over 60,000 minutes of real-world robot operation data, and systems trained using certain advanced methodologies have demonstrated average success rates of 0.95 in complex long-horizon tasks, substantially outperforming traditional behavior cloning approaches that achieve 0.76.[1]