Not All Transitions Matter: Evidence from PPO 文章

ArXiv CS.AI2026-05-26NEWSen作者: Ajhesh Basnet

详细信息

来源站点: ArXiv CS.AI
作者: Ajhesh Basnet
文章类型: NEWS
语言: en
发布日期: 2026-05-26

摘要

arXiv:2605.24071v1 Announce Type: cross Abstract: Training a reinforcement learning agent on-policy means collecting fresh experience at every update, and that experience comes with a hidden problem. Each state in a rollout is the direct output of the previous one, causally chained together by the agent's own actions. Because of this, consecutive transitions are never truly independent. They carry overlapping information, and the gradient signal the network receives ends up far more repetitive than the batch size suggests. The same directions get reinforced over and over, the value network struggles to keep up as the policy shifts, and training becomes quietly unstable in ways that reward curves alone rarely reveal. This paper asks whether that redundancy can simply be removed.

Not All Transitions Matter: Evidence from PPO 文章

详细信息

摘要

相关事件

相关公司查看全部 (4)

相关人物

相关产品查看全部 (8)

相关技术查看全部 (16)