HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation 文章

ArXiv CS.AI2026-06-11NEWSen作者: Haoran Liu, Yuwei Zhang, Xiyao Li, Bohan Lyu, Jingbo Shang

详细信息

来源站点
ArXiv CS.AI
作者
Haoran Liu, Yuwei Zhang, Xiyao Li, Bohan Lyu, Jingbo Shang
文章类型
NEWS
语言
en
发布日期
2026-06-11

摘要

arXiv:2606.11559v1 Announce Type: new Abstract: Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns. Recent on-policy self-distillation methods offer a promising alternative by converting privileged feedback into dense token-level supervision through a self-teacher. Our study is motivated by the unexpected performance degradation observed when naively extending this paradigm to multi-turn settings, which we attribute to a lack of alignment between privileged feedback, such as successful trajectories or terminal outcomes, and the student's current decision context. We introduce HERO, a hindsight-enhanced self-distillation framework that uses next environment observations as locally aligned feedback.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据