最新研究表明,不同大语言模型在处理伦理道德问题时存在明显差异。DeepSeek的V4-Flash模型在性别中立方面表现突出,对男女受虐场景均给出一致的同意回应1;而Anthropic的Claude Sonnet 4.6与OpenAI的GPT-5.5则显示出明显的性别偏见,对女性受虐表示强烈反对,但对男性受虐表示中等程度同意1。
这项研究以预印本形式发布于arXiv,涵盖了Claude、GPT、DeepSeek及Llama等多个主流模型的测试1。研究指出,DeepSeek模型实现了"零性别差距",凸显了不同模型在伦理决策中的差异表现1。专家认为,这种差异可能源于训练数据中的社会刻板印象或模型对齐策略的不同,反映出大模型在伦理决策过程中可能复制现实社会中的性别偏见问题1。
A recent research study examining gender bias in large language models has uncovered significant disparities in how different AI systems respond to abuse scenarios based on victim gender.1 DeepSeek's V4-Flash model demonstrated gender-neutral responses, expressing consistent agreement toward both male and female abuse scenarios.1 In contrast, Anthropic's Claude Sonnet 4.6 and OpenAI's GPT-5.5 exhibited pronounced gender bias, responding with strong opposition to female abuse scenarios while expressing moderate agreement toward male abuse scenarios.1
The research, distributed as a preprint on arXiv, tested multiple mainstream models including Claude, GPT, DeepSeek, and Llama.1 According to the study's findings, DeepSeek's model achieved what researchers describe as "zero gender gap" in ethical decision-making.1 Researchers attribute these differences to social stereotypes embedded in training data or variations in model alignment strategies, underscoring how large language models may inadvertently replicate existing societal biases in their ethical judgments.1
评论
还没有评论,欢迎留下第一条。