研究人员推出ProVer框架,以解决GRPO(群体相对策略优化)算法在大语言模型代理训练中存在的信用分配问题1。该框架的核心创新在于使用智能评判员对比成功和失败的轨迹来定位关键决策点,随后通过验证这些决策段的优势来优化策略,从而避免了对每个中间状态进行逐步评估的高成本1。验证过程通过估计当前策略在决策段前后的终止成功率差异来衡量优势1。
在ALFWorld、WebShop和SearchQA三个基准任务的实验中,ProVer相比原始GRPO算法取得了显著成效1。在Qwen3.5-2B模型上实现了9.91%的相对改进,在Qwen3.5-4B模型上实现了7.12%的相对改进1。该研究于2026年9月28日提交1。
Researchers have introduced ProVer, a framework designed to tackle credit assignment problems in GRPO (Group Relative Policy Optimization) when training large language model agents.1 Rather than evaluating every intermediate state, the approach leverages an intelligent judge to compare successful and failed trajectories, identifying critical decision points and then verifying the advantage of policy decisions at those specific segments to optimize training.1
The ProVer framework demonstrates measurable improvements across multiple benchmarks.1 On ALFWorld, WebShop, and SearchQA tasks, ProVer achieves performance gains of 9.91% on Qwen3.5-2B and 7.12% on Qwen3.5-4B relative to standard GRPO.1 The verification mechanism estimates advantage by comparing the terminal success rates of the current policy before and after decision segments, avoiding the computational expense of step-by-step evaluation.1 Notably, the model's judgment is used solely to select verification locations rather than being directly trusted for training decisions, representing a key innovation in the approach.1
评论
还没有评论,欢迎留下第一条。