Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards 事件

PRODUCT_LAUNCH2026-05-28影响: MEDIUM

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be checked automatically. However, many important tasks are only partially verifiable: prompts contain multiple requirements, responses may satisfy some but not all of them, or no single reference answer might exist. We introduce Soft-RLV

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards · 相关报道