A Closer Look at Invalid Action Masking in Policy Gradient Algorithms 论文

2022Proceedings of the ... International Florida Artificial Intelligence Research Society Conference引用 343顶会

Reinforcement Learning in RoboticsAdversarial Robustness in Machine LearningDomain Adaptation and Few-Shot Learning

机器人 Domain Adaptation and Few-Shot Learning Reinforcement Learning in Robotics Adversarial Robustness in Machine Learning

关系图谱

作者

摘要

In recent years, Deep Reinforcement Learning (DRL) algorithms have achieved state-of-the-art performance in many challenging strategy games. Because these games have complicated rules, an action sampled from the full discrete action distribution predicted by the learned policy is likely to be invalid according to the game rules (e.g., walking into a wall). The usual approach to deal with this problem in policy gradient algorithms is to “mask out” invalid actions and just sample from the set of valid actions. The implications of this process, however, remain under-investigated. In this paper, we 1) show theoretical justification for such a practice, 2) empirically demonstrate its importance as the space of invalid actions grows, and 3) provide further insights by evaluating different action masking regimes, such as removing masking after an agent has been trained using masking.

作者查看全部 (2)

Santiago Ontañón

Shengyi Huang

A Closer Look at Invalid Action Masking in Policy Gradient Algorithms 论文

摘要

作者查看全部 (2)

相关技术查看全部 (3)

相关事件

相关文章