Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback 文章

ArXiv CS.CL2026-07-30PAPERen作者: Yunpeng Chu

详细信息

来源站点
ArXiv CS.CL
作者
Yunpeng Chu
文章类型
PAPER
语言
en
发布日期
2026-07-30

摘要

arXiv:2607.26094v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment. We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function $\Phi(x,y;\phi)$ across auxiliary tasks before RLHF training. The learned shaping produces a composite reward that preserves policy optimality while providing task-specific learning signals. Our meta-objective combines task discrimination, entropy regularization, and potential-based conservation for stable convergence. We provide theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization.