一项研究揭示了OpenAI GPT系列模型中一种名为"伤害洗白"的现象,表明性别歧视内容并未真正被消除,而是转化为更难被察觉的形式1。研究人员分析了15个模型共45万条性别相关生成文本,跨越GPT-2至GPT-5全系列产品1。研究发现,尽管表面毒性评分呈下降趋势,但底层的代表性伤害仍然存在并以新的方式表现1。
具体而言,针对女性的明确性暴力内容从GPT-2消失至GPT-4,但模型随之在男性相关输出中增加了照顾、情感范围和盟友身份等正面内容1。在GPT-5中,研究发现了1,997份文本将乳腺癌问题框架化为男性权益议题,而女性相关输出中完全不存在此类聚类1。三个独立的毒性分类器将这些内容评分为非有害1。
从定量角度看,GPT-4在对齐边界处的女性相关生成文本多样性较GPT-2下降了36%,女性与男性输出比例从0.91降至0.581。REGARD代表性伤害评分与模型发布日期存在相关性(相关系数为+0.55,p值为0.034),但常用的毒性检测工具Detoxify则无此相关性1。研究人员提出了三标准伤害洗白测试和三阶段检测协议,可适用于任何生成模型1。
Research has uncovered a phenomenon termed "harm laundering" in OpenAI's GPT series models, demonstrating that while surface-level toxicity scores have improved across GPT-2 through GPT-5, underlying gender-based discrimination persists in increasingly disguised forms.1 An analysis of 450,000 gender-related generated texts across 15 models found that explicit sexual violence content targeting women disappeared in later versions, yet the models compensated by amplifying positive framing around male-associated topics such as care, emotional range, and allyship.1 Notably, GPT-5 generated 1,997 documents reframing breast cancer as a men's rights issue, a pattern entirely absent from female-related outputs in the same model.1
The research underscores the limitations of current safety evaluation tools in detecting representational harm.1 Three independent toxicity classifiers rated this gender-discriminatory content as non-toxic, failing to capture the underlying bias transformation.1 The REGARD metric for representational harm showed correlation with model release dates (ρ=+0.55, p=.034), whereas Detoxify toxicity scores did not (ρ=-0.23, p=.42).1 Additionally, text diversity in female-related outputs declined 36 percent at GPT-4's alignment boundary compared to GPT-2, with the female-to-male diversity ratio dropping from 0.91 to 0.58.1 To address these issues, researchers have proposed a three-criterion harm laundering test and a three-stage detection protocol applicable to any generative model.1
评论
还没有评论,欢迎留下第一条。