详细信息
- 来源站点
- ArXiv CS.AI
- 作者
- Haoran Liu, Yuwei Zhang, Xiyao Li, Bohan Lyu, Jingbo Shang
- 文章类型
- NEWS
- 语言
- en
- 发布日期
- 2026-06-11
摘要
arXiv:2606.11559v1 Announce Type: new Abstract: Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns. Recent on-policy self-distillation methods offer a promising alternative by converting privileged feedback into dense token-level supervision through a self-teacher. Our study is motivated by the unexpected performance degradation observed when naively extending this paradigm to multi-turn settings, which we attribute to a lack of alignment between privileged feedback, such as successful trajectories or terminal outcomes, and the student's current decision context. We introduce HERO, a hindsight-enhanced self-distillation framework that uses next environment observations as locally aligned feedback.