Benchmarking Large Language Models for News Summarization 论文

2024Transactions of the Association for Computational Linguistics引用 302顶会

Topic ModelingAdvanced Text Analysis TechniquesNatural Language Processing Techniques

Natural Language Processing Techniques Topic Modeling Advanced Text Analysis Techniques

作者

摘要

Abstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, not model size, is the key to the LLM’s zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LLM summaries are judged to be on par with human written summaries.

作者查看全部 (5)

Tatsunori Hashimoto

Kathleen McKeown

Percy Liang

Esin Durmus

Benchmarking Large Language Models for News Summarization 论文

摘要

作者查看全部 (5)

相关技术查看全部 (2)

相关事件

相关文章