聚闻统一资讯入口
← 返回

ProVer:为智能体 RL 精准分配信用

AIAIHOT采集时间:2026-10-02 08:00原文 ↗
elvis· @omarsar0 · X·2026-10-02 16:00· 53 分钟前AI 评分47
AI 导读

ProVer 提出用 LLM judge 定位轨迹中的关键片段、再由 rollout 决定该步信用大小的信用分配方法,解决 GRPO 给轨迹中每个 token 相同 advantage、无法区分决定性步骤的问题。

正文

Good paper on credit assignment for agent RL.

The main finding is that you want an LLM judge to choose where to check a trajectory, and the rollouts to decide how much credit that step gets.

GRPO gives every token in a trajectory the same advantage, so the training signal cannot tell the decisive step from the rest.

ProVer has a judge compare successful and failed rollouts and name the segment it thinks caused the difference. It then samples continuations from just before and just after that segment and uses the change in success rate as the segment's advantage.

Across ALFWorld, WebShop and SearchQA, this gives relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. It still helps when the judge is a smaller model.

来源:elvis · x.com

Agent 智能体推理能力论文研究#论文/研究#Agent#推理
聚闻 · 统一资讯入口 · 聚合新闻与公众号阅读 · V1.1.0