CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation 文章

ArXiv CS.AI2026-07-21PAPERen作者: Satyam Kumar, Saurabh Jha

详细信息

来源站点
ArXiv CS.AI
作者
Satyam Kumar, Saurabh Jha
文章类型
PAPER
语言
en
发布日期
2026-07-21

摘要

arXiv:2607.16955v1 Announce Type: cross Abstract: On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling, where time-only forward/reverse-KL interpolation ignores the student's coverage state; and (iii) binary reward sparsity, where pass/fail signals discard information from partially correct traces. We present CADENCE, a unified framework with a targeted fix for each. Its DRIFT mechanism schedules a per-token convex mixture of forward-KL and reverse-KL surrogate objectives on student-sampled trajectories (per-token surrogates, not sequence-level KL gradient estimators). Six components extend it: (A) COVA, a coverage-adaptive $\beta$ schedule accelerating the forward-to-reverse transition;

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据