When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction 文章

ArXiv CS.CL2026-08-10PAPERen作者: Yuqing Yang, Robin Jia

详细信息

来源站点
ArXiv CS.CL
作者
Yuqing Yang, Robin Jia
文章类型
PAPER
语言
en
发布日期
2026-08-10

摘要

arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in their previously generated false assertions. Using model-specific testbeds, we find that while LLMs are capable of retraction, they do so only rarely, even when they can recognize their mistakes when asked in a separate interaction. We identify a reliable predictor of retraction: the model's momentary belief, as measured by a linear probe on its internal representation. The probe is trained to predict the correctness of answers on external datasets unrelated to retraction, then applied to settings where models should retract. A model retracts only when it "believes" its answers to be incorrect during generation; these beliefs frequently diverge from models' parametric knowledge as measured by factoid questions.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据