研究人员通过系统实验检验了一个假设:顶级AI实验室是否针对"生成pelican骑自行车"这一著名基准任务对模型进行了优化。1
该实验测试了7个前沿AI模型——GPT-5.6 Terra、Claude Sonnet 5、Gemini 3.5 Flash、Grok 4.5、Qwen3.7-Max、GLM-5.2和DeepSeek V4 Pro。1研究人员让这些模型各生成了1,008张SVG图像,涵盖8种动物与6种交通工具的组合,每种组合包含3个样本。1
分析结果表明,pelican和bicycle都不是各模型表现最优的对象。在8种动物的绘制质量排名中,pelican位列第6;在6种交通工具中,bicycle排名倒数第二。1pelican与bicycle的组合在全部48种组合中则排到第42位。1
统计检验未能找到针对性优化的证据。固定效应回归分析显示,pelican效应的p值最小为0.25;在bicycle效应中,仅有Gemini达到p<0.05的显著性水平,但这一结果无法通过多重比较修正。1
即使在看似异常的细节上,也未观察到有针对性的调整。所有21张pelican-bicycle组合的图像均朝向右边,但在整体数据集中,60%的图像同样朝向右边。1
这项实验的成本相对低廉,仅需约80美元的API费用。1
Researcher Dylan Castillo conducted an experiment to investigate whether leading artificial intelligence laboratories have optimized their models to excel at generating images of pelicans riding bicycles, a phenomenon he termed "pelicanmaxxing." 1
The study tested seven frontier AI models—GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro—by generating a total of 1,008 SVG images across eight animal types and six transportation methods, with three samples per combination. 1
The results provided little support for the pelicanmaxxing hypothesis. Pelican ranked sixth among the eight animals tested, while bicycle ranked fifth out of six transportation methods. 1 Across all 48 possible animal-vehicle combinations, the pelican-bicycle pairing ranked 42nd. 1
Statistical analysis using fixed-effects regression revealed no meaningful evidence of targeted optimization. All pelican effect p-values were at minimum 0.25, and among bicycle effects, only Gemini achieved p<0.05—a result that could not withstand multiple comparison correction. 1
One observation stood out: all 21 pelican-bicycle images were oriented rightward, yet 60 percent of images across the entire dataset faced the same direction, suggesting no anomalous bias specific to the combination. 1
The experiment was conducted on a modest budget of approximately 80 dollars in API costs. 1
评论
还没有评论,欢迎留下第一条。