When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO 文章

ArXiv CS.AI2026-08-05PAPERen作者: Zhe Cao, Miaowen Wen, Fangjiong Chen

详细信息

来源站点
ArXiv CS.AI
作者
Zhe Cao, Miaowen Wen, Fangjiong Chen
文章类型
PAPER
语言
en
发布日期
2026-08-05

摘要

arXiv:2608.03467v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this completion-level uniformity creates structure-level skew: recurring correct solution forms accumulate positive coefficient mass in proportion to how often they are sampled, while rare forms receive limited credit. We formalize this behavior as multiplicity-induced structure-level credit concentration and introduce a partition- conditioned rule that redistributes positive advantages accord- ing to cluster rarity. Cue-GRPO instantiates this rule with- out auxiliary-model inference by using deterministic Strategy Cues to construct rollout-local partitions of verified-correct traces. Across Qwen2.5-Math-7B and Llama-3.1-8B-Instruct, Cue-GRPO improves AIME repeated-sampling performance, with the largest gains at high sampling budgets.