Self-Trained Verification for Training- and Test-Time Self-Improvement 事件

PRODUCT_LAUNCH2026-05-29影响: MEDIUM

Self-Trained Verification for Training- and Test-Time Self-Improvement arXiv:2605.30290v1 Announce Type: cross Abstract: Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate while accuracy stagnates, and when feedback is t