Understanding Tone-Dependent Inference Cost in Large Language Models 文章

ArXiv CS.CL2026-07-28PAPERen作者: Akhil Kumar, Om Dobariya

详细信息

来源站点
ArXiv CS.CL
作者
Akhil Kumar, Om Dobariya
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.23915v1 Announce Type: new Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据