一项研究发现,用于验证AI生成临床记录的大语言模型评判器存在严重缺陷,无法有效检测记录中遗漏的信息。1研究团队在包含500对单一错误记录的基准测试中发现,这些评判器对遗漏内容的检测率仅为0.50-0.63,而对添加或篡改内容的检测率则达到0.79-0.94。1其中基准测试包含298个确定遗漏的事实案例和202个添加或更改对照。1
为解决这一问题,研究人员开发了一种管道方法,通过将验证任务重构为逐项事实列表核对来改进检测能力。1这种改进方案的误报率为2.7%,而单次调用方法的误报率为6.2%但成本仅为其十分之一。1在同一批案例中,管道方法的检测率为24.6%,而单次调用方法为36.9%(p=0.002)。1医学作者对70项进行了验证,在两种方法意见不一致的10个案例中全部支持管道方法(p=0.002)。1研究团队已公开发布了基准数据集、提示词和判断结果。1
Researchers have identified a critical limitation in large language models used to evaluate AI-generated clinical records: they struggle to detect missing information while performing well at catching additions or alterations.1 A benchmark test of 500 single-error record pairs—comprising 298 cases with omitted facts and 202 with added or modified content—revealed that eight different judge designs achieved detection rates of only 0.50–0.63 for omissions, compared to 0.79–0.94 for additions and changes.1
To address this asymmetry, the research team developed a structured verification approach that reformulates the evaluation task as itemized fact-list validation.1 This pipeline method achieved a false positive rate of 2.7 percent, compared to 6.2 percent for a single-call approach that required only one-tenth the computational cost.1 In head-to-head testing, the pipeline method detected omissions in 24.6 percent of cases versus 36.9 percent for the single-call method (p=0.002).1 Medical authors validated 70 cases and unanimously supported the pipeline method across all 10 instances where the two approaches disagreed (p=0.002).1 The researchers have publicly released their benchmark dataset, prompts, and evaluation results to support further development in this area.1
评论
还没有评论,欢迎留下第一条。