JuliaHub 对 OpenAI 的 GPT 5.6 系列模型(Terra、Sol、Luna)和 Anthropic 的 Claude Fable 5 进行了系统的物理 AI 能力评测[1]。研究采用 Dyad AI 框架,在五个密封物理建模与仿真问题上运行了 52 次评分试验[1]。
在物理建模准确性方面,Claude Fable 5 表现最佳,难度加权得分达 0.889,是唯一完全通过四个核心问题的 12 次试验的模型[1]。而 GPT 5.6 Sol 的得分为 0.814,虽然略低于 Fable 5,但在性价比上优势明显——其成本仅为 Fable 5 的五分之一,分别为每次试验 1.74 美元和 9.60 美元[1]。值得注意的是,所有模型都能生成可编译运行的代码,但物理准确性存在显著差异[1]。
研究还揭示了框架工具对模型性能的影响远超模型本身的差异[1]。使用 Dyad 框架的模型得分达 0.899,而通用编码代理的得分仅为 0.533,框架差异高达 0.366 点,超过最佳与最差模型之间的差异 0.162 点[1]。五个评测问题包括以 NASA HL-20 飞行器为代表的难度最高的问题,目前没有模型能完全解决[1]。
JuliaHub conducted a comprehensive evaluation of physical AI capabilities across OpenAI's GPT-5.6 series (Terra, Sol, and Luna) and Anthropic's Claude Fable 5, using the Dyad AI framework to run 52 scoring trials across five sealed problems.[1] The benchmark tested the models' ability to solve physics modeling and simulation challenges, with results highlighting distinct strengths in accuracy versus cost efficiency.
Claude Fable 5 achieved the highest accuracy score of 0.889 in physical modeling precision and was the only model to completely solve all 12 trials across four core problems.[1] However, it came at a significantly higher cost, requiring $9.60 per trial.[1] GPT-5.6 Sol, by contrast, delivered a competitive score of 0.814 while costing only $1.74 per trial—one-fifth of Fable's expense—making it the most cost-effective option.[1] The evaluation encompassed five distinct problems, including modeling of the NASA HL-20 aircraft, which posed the greatest difficulty and was not fully solved by any model.[1]
A notable finding emerged regarding the influence of framework choice: performance variance between the Dyad framework (0.899) and general coding agents (0.533) reached 0.366 points, exceeding the performance gap between the best and worst models at 0.162 points.[1] While all models demonstrated the ability to generate compilable and executable code, significant divergence appeared in physical accuracy across implementations.[1]