Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap 文章

ArXiv CS.CL2026-08-06PAPERen作者: Ankit Goyal, Jaideep Ray

详细信息

来源站点
ArXiv CS.CL
作者
Ankit Goyal, Jaideep Ray
文章类型
PAPER
语言
en
发布日期
2026-08-06

摘要

arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the native-vs-translate gap on MGSM (German, Thai, Swahili) is a token-budget artifact for Qwen3-8B and Llama-3.1-8B-Instruct under four prompting strategies. The measured gap swings by up to 57 points across budgets, length normalization moves it by up to 38.9 points where the cap binds, and at tight caps normalization can reverse which strategy scores higher. We prospectively froze the sweep's three Qwen peaks and its near-zero value at 1024 and evaluated them on 540,000 independently hard-capped decodes: a second frozen family of six Holm-corrected tests rejects every null. The frozen test at $B^*=1024$ still fails to reject because native accuracy has already saturated there;

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关技术

暂无数据