The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale 文章

ArXiv CS.CL2026-08-06PAPERen作者: Mingguang Chen, Bo Qu, Licheng Wang

详细信息

来源站点
ArXiv CS.CL
作者
Mingguang Chen, Bo Qu, Licheng Wang
文章类型
PAPER
语言
en
发布日期
2026-08-06

摘要

arXiv:2608.04355v1 Announce Type: new Abstract: Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boundary, and test the failure causally rather than only observationally. Across Qwen3.5 (0.8B-9B), Gemma-4-12B, and two frontier models via API (Tencent Hy3, Nvidia Nemotron-3-Ultra-550B) in 29 primary cells plus a frontier arm, we decompose the always-revise accuracy shift into a content margin (both answers parseable) and format-recovery/loss margins (parseability changes). On 12 cells with meaningful unparseable-answer rates, format effects exceed content effects (Wilcoxon p=1.7e-3). To test this causally, we force already-generated reasoning through grammar-constrained decoding so every answer is parseable by construction: across 14 cells this closes a median 71% of the gap between the naive total effect and the content-margin estimate, with two cells converging exactly and a residual…

摘要可能不完整,可查看原文