Caliper: Probing Lexical Anchors versus Causal Structure in LLMs 事件
PRODUCT_LAUNCH2026-06-04影响: MEDIUM
Caliper: Probing Lexical Anchors versus Causal Structure in LLMs arXiv:2606.04915v1 Announce Type: new Abstract: Large language models reach 50 to 70% accuracy on causal reasoning benchmarks such as CLadder, but it is unclear whether this reflects structural reasoning or lexical pattern matching. We introduce Caliper, a controlled perturbation that replaces semantic variable names with placeholder tokens while preserving the causal graph and probabilistic specification of each question. Across