研究者Brianne Lee进行了一项实验,用同一条价格2.43美元的项链在三种不同穿衣背景下拍照,测试包括Claude Fable 5、GPT-5.6、GPT-4o、Grok 4.5、Kimi K3和DeepSeek V4-Pro在内的6个前沿多模态模型的价格估算能力[1]。结果表明这些模型对同一物体的估价差异高达3.6倍,其中Kimi在正式场景估价104美元,而在院子场景仅估价29美元[1]。Claude在正式背景估价62美元对比院子背景的19美元,显示出3.3倍的价格差异[1]。
模型不仅改变了数字估价,还根据衣着背景改变了对项链材料的描述[1]。值得注意的是,当被直接询问穿着不同是否会影响估价结果时,Claude承认这种偏见的存在率达到100%(24/24),而GPT-4o的承认率仅为18%(34/192),两者形成鲜明对比[1]。在纯文本条件下,GPT-4o展现了3.9倍的光晕效应,同时在被要求估值图像时拒绝率高达79%[1]。
该研究基于约1,500次会话和4,604条分析行数据,已将研究数据和代码在CC BY 4.0许可证下开源发布[1]。
Researcher Brianne Lee conducted an experiment to evaluate how vision language models estimate product values when contextual cues vary [1]. She photographed an inexpensive necklace purchased for $2.43 in three different outfit settings and submitted the images to six leading multimodal AI systems, including Claude, GPT-4o, Grok, Kimi, and DeepSeek [1].
The findings revealed significant inconsistency in price valuations. The same necklace was estimated at prices ranging from $19 to $104 depending on the background context, representing a maximum variance of 3.6 times across different scenarios [1]. For instance, Kimi K3 valued the necklace at $104 when paired with formal attire but estimated it at $29 in a casual outdoor setting [1]. Similarly, Claude Fable 5 produced estimates of $62 for the formal context and $19 for the casual yard setting, a 3.3-fold difference [1]. The models also adjusted their material descriptions based on surrounding clothing, suggesting the background influenced their assessments [1].
When directly questioned about whether different outfits affected their valuations, Claude acknowledged this bias consistently in 24 of 24 instances, while GPT-4o admitted to the bias in only 34 of 192 cases, or 18 percent of the time [1]. Notably, GPT-4o refused to provide image-based estimates in 79 percent of cases but demonstrated a 3.9-fold halo effect when working with text descriptions alone [1]. Lee has made the research data and code publicly available under a CC BY 4.0 license, based on approximately 1,500 conversations and 4,604 lines of analytical data [1].