Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States 文章

ArXiv CS.CL2026-07-30PAPERen作者: Weiyi Kong, Zhuoran Li

详细信息

来源站点
ArXiv CS.CL
作者
Weiyi Kong, Zhuoran Li
文章类型
PAPER
语言
en
发布日期
2026-07-30

摘要

arXiv:2607.26929v1 Announce Type: new Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than matching the evidence to the causal question. We introduce paired prompts that repeat the same diagnostic evidence verbatim while changing the causal target. Each prompt is labeled Favors, Challenges, Unresolved, or Wrong Target according to how the evidence bears on the causal question. A pair is recovered only when both prompts are classified correctly. Using linear readouts trained on a separate development set, we analyze the final-token hidden state from the penultimate transformer block of Qwen2.5-7B-Instruct, Qwen3-8B, and Llama-3.1-8B-Instruct.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据