RL Post-Training Builds Compositional Reasoning Strategies 文章

ArXiv CS.CL2026-07-09PAPERen作者: Azwar Abdulsalam, Nishil Patel, Andrew Saxe

详细信息

来源站点
ArXiv CS.CL
作者
Azwar Abdulsalam, Nishil Patel, Andrew Saxe
文章类型
PAPER
语言
en
发布日期
2026-07-09

摘要

arXiv:2607.07646v1 Announce Type: cross Abstract: Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite can be audited. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained on a Trace-based reasoning task with only a binary final-answer reward. RL solves held-out problems that remain rarely solved by the pretrained model even under much larger sampling budgets, while rejection fine-tuning improves early but plateaus. Trace analysis shows that RL reorganizes primitive competence through a phased compositional mechanism: it first strengthens primitive reductions, then discovers valid composed procedures.

相关事件

暂无数据

相关人物

暂无数据