详细信息
- 来源站点
- ArXiv CS.CL
- 作者
- Ming Shen, Zhikun Xu, Jacob Dineen, Xiao Ye, Ben Zhou
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-05
摘要
arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word reasoning methods make this tension explicit: they reward a model for producing a rationale that supports one context-conditioned continuation, which can turn a pre-existing preference into a confident, self-justifying trajectory. We introduce BOW, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space. BOW's core reward is mediated by the generated trajectory. The policy conditions on the full context, but a frozen scorer assigns the core reward from the trajectory alone, without receiving the original context as a separate input. The trajectory may restate relevant context; the bottleneck is the missing direct context-to-scorer path in the core reward.
相关事件
暂无数据
相关人物
暂无数据