研究人员以GPT-5.5 Pro为教师模型,对Qwen 3.8等开源模型进行了推理前缀填充对比实验1。实验通过在目标模型的推理通道中插入GPT-5.5 Pro推理过程的前1%内容,衡量目标模型答案中与教师模型的重叠程度1。
针对45个问题的评估涵盖三大类别:15个STEM问题、15个非STEM问题和15个合成谜题1。实验结果显示,Qwen 3.8相比之前的实验向GPT-5.5 Pro移动了+18.18个百分点,其中包括在私有合成谜题上的大幅效应,这暗示该模型可能学习自GPT-5.5 Pro或相关GPT模型1。
与此同时,Kimi K3与GPT-5.5 Pro的重叠率最高,无前缀情况下为31.11%,添加前缀后达到35.65%,前缀添加仅产生+4.54个百分点的增长1。研究采用的测量方法是计算目标模型答案前100个token中教师模型可见答案的单元组、二元组和三元组源回想率的平均值1。
Researchers have conducted reasoning prefix-filling experiments using GPT-5.5 Pro as a teacher model to evaluate open-source models including Qwen 3.8 1. The study involved inserting the first one percent of GPT-5.5 Pro's reasoning process into the inference channel of target models and measuring the degree of overlap between the target model's answers and those of the teacher model 1.
Qwen 3.8 demonstrated a notable shift toward GPT-5.5 Pro, moving +18.18 percentage points compared to previous experiments, suggesting the model may have learned from GPT-5.5 Pro or related GPT models 1. The assessment covered 45 questions spanning three categories: 15 STEM problems, 15 non-STEM problems, and 15 synthetic puzzles 1. The +18.18 percentage point movement includes substantial effects on private synthetic puzzles 1.
In contrast, Kimi K3 exhibited the highest overlap rate with GPT-5.5 Pro at 31.11 percent without prefix injection and 35.65 percent with prefix injection, representing an increase of only +4.54 percentage points 1. The measurement methodology calculated the average recall of unigrams, bigrams, and trigrams from the teacher model's visible answers within the first 100 tokens of the target model's responses 1.
评论
还没有评论,欢迎留下第一条。