BOW: Training Language Models to Reason Over Plausible Next Words 文章

ArXiv CS.CL2026-08-05PAPERen作者: Ming Shen, Zhikun Xu, Jacob Dineen, Xiao Ye, Ben Zhou

详细信息

来源站点
ArXiv CS.CL
作者
Ming Shen, Zhikun Xu, Jacob Dineen, Xiao Ye, Ben Zhou
文章类型
PAPER
语言
en
发布日期
2026-08-05

摘要

arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word reasoning methods make this tension explicit: they reward a model for producing a rationale that supports one context-conditioned continuation, which can turn a pre-existing preference into a confident, self-justifying trajectory. We introduce BOW, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space. BOW's core reward is mediated by the generated trajectory. The policy conditions on the full context, but a frozen scorer assigns the core reward from the trajectory alone, without receiving the original context as a separate input. The trajectory may restate relevant context; the bottleneck is the missing direct context-to-scorer path in the core reward.