The Knowing-Saying Gap: When Probes See Errors that Confidence Misses 文章

ArXiv CS.CL2026-08-11PAPERen作者: Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk

详细信息

来源站点
ArXiv CS.CL
作者
Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk
文章类型
PAPER
语言
en
发布日期
2026-08-11

摘要

arXiv:2608.07528v1 Announce Type: cross Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered "persistence beats peak" hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据