研究人员提出了一种依赖感知的标签聚合方法,用于改进多个大语言模型(LLM)评判系统中的判断聚合1。该方法使用Ising模型对评判者之间的相关性进行建模1,相比传统的加权多数投票法,在准确率上实现了9%-14%的提升1。
在三项具体任务的性能对比中,新方法表现显著1。相关性分类任务中,准确率从0.820/0.804提升至0.9121;毒性分类任务中从0.694/0.695提升至0.7921;总结评估任务中从0.737/0.561提升至0.8061。所有评判模型均在温度为零的设置下运行,确保评判过程无随机性1。
这项研究成果由包括Shiva Kasiviswanathan在内的研究团队完成1,已发表于今年的国际机器学习大会(ICML)1。
Researchers have developed a dependency-aware label aggregation method designed to improve how multiple language model judges reach consensus in evaluation systems.1 The approach employs an Ising model to capture correlations between different judges, demonstrating significant improvements over traditional weighted majority voting across three distinct tasks.1
In comparative testing, the new method achieved substantially higher accuracy rates: 0.912 versus 0.820 and 0.804 on relevance classification, 0.792 versus 0.694 and 0.695 on toxicity classification, and 0.806 versus 0.737 and 0.561 on summarization evaluation.1 Overall, the technique yielded improvements of 9 to 14 percentage points relative to the best baseline approaches.1 The research was conducted in collaboration with Shiva Kasiviswanathan and presented at this year's International Conference on Machine Learning (ICML).1 All evaluation models operated under zero-temperature settings, eliminating randomness from the judgment process.1
评论
还没有评论,欢迎留下第一条。