研究者Brianne Lee进行了一项实验,用同一条价格2.43美元的项链在三种不同穿衣背景下拍照,测试包括Claude Fable 5、GPT-5.6、GPT-4o、Grok 4.5、Kimi K3和DeepSeek V4-Pro在内的6个前沿多模态模型的价格估算能力1。结果表明这些模型对同一物体的估价差异高达3.6倍,其中Kimi在正式场景估价104美元,而在院子场景仅估价29美元1。Claude在正式背景估价62美元对比院子背景的19美元,显示出3.3倍的价格差异1。
模型不仅改变了数字估价,还根据衣着背景改变了对项链材料的描述1。值得注意的是,当被直接询问穿着不同是否会影响估价结果时,Claude承认这种偏见的存在率达到100%(24/24),而GPT-4o的承认率仅为18%(34/192),两者形成鲜明对比1。在纯文本条件下,GPT-4o展现了3.9倍的光晕效应,同时在被要求估值图像时拒绝率高达79%1。
该研究基于约1,500次会话和4,604条分析行数据,已将研究数据和代码在CC BY 4.0许可证下开源发布1。
Researcher Brianne Lee conducted an experiment to evaluate how vision language models estimate product values when contextual cues vary 1. She photographed an inexpensive necklace purchased for $2.43 in three different outfit settings and submitted the images to six leading multimodal AI systems, including Claude, GPT-4o, Grok, Kimi, and DeepSeek 1.
The findings revealed significant inconsistency in price valuations. The same necklace was estimated at prices ranging from $19 to $104 depending on the background context, representing a maximum variance of 3.6 times across different scenarios 1. For instance, Kimi K3 valued the necklace at $104 when paired with formal attire but estimated it at $29 in a casual outdoor setting 1. Similarly, Claude Fable 5 produced estimates of $62 for the formal context and $19 for the casual yard setting, a 3.3-fold difference 1. The models also adjusted their material descriptions based on surrounding clothing, suggesting the background influenced their assessments 1.
When directly questioned about whether different outfits affected their valuations, Claude acknowledged this bias consistently in 24 of 24 instances, while GPT-4o admitted to the bias in only 34 of 192 cases, or 18 percent of the time 1. Notably, GPT-4o refused to provide image-based estimates in 79 percent of cases but demonstrated a 3.9-fold halo effect when working with text descriptions alone 1. Lee has made the research data and code publicly available under a CC BY 4.0 license, based on approximately 1,500 conversations and 4,604 lines of analytical data 1.
评论
还没有评论,欢迎留下第一条。