比亚迪汽车新技术研究院公开发表AI论文HyWorldVLA,这是一个融合Vision-Language-Action与混合世界模型的自动驾驶基础模型[1]。该模型在NAVSIM v1公开基准测试中取得PDMS 90.59的业界最佳成绩[1],体现了比亚迪在物理AI基础模型研究方面的进展。
从消融实验数据来看,混合世界模型的架构设计起到了关键作用[1]。去掉画面预测功能后,PDMS从90.59下降至87.50;去掉隐空间则降至89.91[1]。在雨雾场景测试中,HyWorldVLA达到86.87的性能表现,而纯像素方法仅为60.65[1],显示出该模型的鲁棒性优势。
论文由比亚迪汽车新技术研究院独立完成,没有外部高校或研究机构合作署名[1]。团队核心成员包括来自哈工大机器人技术与系统国家重点实验室的Liulong Ma以及具有哈工大和CMU背景的Hongbiao Zhu[1],反映出比亚迪在聚集机器人领域人才上的投入。这一论文的发表标志着比亚迪从传统汽车电子向机器人与自动驾驶一体化AI研发的转变。
BYD Automobile's New Technology Research Institute has publicly released an artificial intelligence research paper introducing HyWorldVLA, a foundational model for autonomous driving that combines Vision-Language-Action capabilities with a hybrid world model architecture [1]. The model achieved a state-of-the-art PDMS score of 90.59 on the NAVSIM v1 public benchmark, demonstrating BYD's advancement in physical AI foundation model research [1].
Ablation studies conducted on the hybrid world model revealed the significance of its components: removing frame prediction caused the PDMS score to decline to 87.50, while removing the latent space reduced it to 89.91 [1]. In challenging weather conditions, HyWorldVLA attained a score of 86.87 in rain and fog scenarios, substantially outperforming the pure pixel-based WoTE method, which achieved only 60.65 in the same tests [1].
The research was completed independently by BYD Automobile's New Technology Research Institute without external collaboration from universities or research institutions [1]. The core team members include Liulong Ma, who previously worked at Harbin Institute of Technology's State Key Laboratory of Robotics and Systems, and Hongbiao Zhu, who has background experience from both Harbin Institute of Technology and Carnegie Mellon University [1]. This research marks BYD's strategic transition from traditional automotive electronics toward integrated AI development combining robotics and autonomous driving [1].