Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets 文章

ArXiv CS.CL2026-07-28PAPERen作者: Zongyou Yang, Yinghan Hou

详细信息

来源站点
ArXiv CS.CL
作者
Zongyou Yang, Yinghan Hou
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.24268v1 Announce Type: new Abstract: Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framework that separates scorer-independent execution evidence, including termination, answer exposure, parseability, and completion length, from scorer-dependent correctness. Across 2,550 outputs from five fixed Qwen and DeepSeek configurations on MATH and ARC-Challenge, matched 2,048-token limits produce sharply different execution mixtures: 49 of 450 Qwen MATH outputs terminate without a final answer, compared with 5 of 300 DeepSeek MATH outputs and none of the 750 ARC outputs. Among the same 300 DeepSeek MATH question-model pairs, no missing-final length termination is observed at 8,192 tokens.

相关事件

暂无数据

相关公司查看全部 (1)

A
arXivNONPROFIT

相关人物

暂无数据