作者介绍了一套通过约束输出令牌来优化语言模型多选题回答能力的方法1。该方案采用约束解码技术,将模型输出限制在固定选项集合(A、B、C、D、E)内1。
在CommonsenseQA数据集上的实验中,使用Qwen3-1.7B模型的基础版本取得了59.38%的精度(725/1221题正确),宏F1指标为0.58441。经过微调后,模型性能显著提升至62.41%的精度(762/1221题正确),宏F1提高到0.62341。
针对模型过度自信的问题,作者采用了温度缩放方法进行校准1。实验发现,校准前模型在0.9-1.0置信度区间内的准确率仅为70%1。通过应用温度值3.797的温度缩放,校准后该置信度区间的准确率提升至95.41%,使模型的置信度与准确率更好地对应1。
A developer has demonstrated how to construct a decision model that improves language model performance on multiple-choice questions through constrained output decoding 1. Using the Qwen3-1.7B model tested on the CommonsenseQA dataset, the approach restricts model outputs to a fixed set of options (A, B, C, D, E), achieving an initial accuracy of 59.38% (725 out of 1,221 test cases) with a macro F1 score of 0.5844 1.
Fine-tuning the model further boosted performance to 62.41% (762 out of 1,221 cases) with a macro F1 score of 0.6234 1. However, the model exhibited overconfidence calibration issues, correctly answering only 70% of questions in the 0.9–1.0 confidence interval 1. To address this misalignment between confidence levels and actual accuracy, the developer applied temperature scaling with a temperature value of 3.797, which recalibrated the model's confidence estimates 1. Following this calibration adjustment, accuracy for predictions in the highest confidence bracket improved significantly to 95.41% 1.
评论
还没有评论,欢迎留下第一条。