JuliaHub 对 OpenAI 的 GPT 5.6 系列模型(Terra、Sol、Luna)和 Anthropic 的 Claude Fable 5 进行了系统的物理 AI 能力评测1。研究采用 Dyad AI 框架,在五个密封物理建模与仿真问题上运行了 52 次评分试验1。
在物理建模准确性方面,Claude Fable 5 表现最佳,难度加权得分达 0.889,是唯一完全通过四个核心问题的 12 次试验的模型1。而 GPT 5.6 Sol 的得分为 0.814,虽然略低于 Fable 5,但在性价比上优势明显——其成本仅为 Fable 5 的五分之一,分别为每次试验 1.74 美元和 9.60 美元1。值得注意的是,所有模型都能生成可编译运行的代码,但物理准确性存在显著差异1。
研究还揭示了框架工具对模型性能的影响远超模型本身的差异1。使用 Dyad 框架的模型得分达 0.899,而通用编码代理的得分仅为 0.533,框架差异高达 0.366 点,超过最佳与最差模型之间的差异 0.162 点1。五个评测问题包括以 NASA HL-20 飞行器为代表的难度最高的问题,目前没有模型能完全解决1。
JuliaHub conducted a comprehensive evaluation of physical AI capabilities across OpenAI's GPT-5.6 series (Terra, Sol, and Luna) and Anthropic's Claude Fable 5, using the Dyad AI framework to run 52 scoring trials across five sealed problems.1 The benchmark tested the models' ability to solve physics modeling and simulation challenges, with results highlighting distinct strengths in accuracy versus cost efficiency.
Claude Fable 5 achieved the highest accuracy score of 0.889 in physical modeling precision and was the only model to completely solve all 12 trials across four core problems.1 However, it came at a significantly higher cost, requiring $9.60 per trial.1 GPT-5.6 Sol, by contrast, delivered a competitive score of 0.814 while costing only $1.74 per trial—one-fifth of Fable's expense—making it the most cost-effective option.1 The evaluation encompassed five distinct problems, including modeling of the NASA HL-20 aircraft, which posed the greatest difficulty and was not fully solved by any model.1
A notable finding emerged regarding the influence of framework choice: performance variance between the Dyad framework (0.899) and general coding agents (0.533) reached 0.366 points, exceeding the performance gap between the best and worst models at 0.162 points.1 While all models demonstrated the ability to generate compilable and executable code, significant divergence appeared in physical accuracy across implementations.1
评论
还没有评论,欢迎留下第一条。