Benchmarking Large Language Models for News Summarization

Zhang, Tianyi; Ladhak, Faisal; Durmus, Esin; Liang, Percy; McKeown, Kathleen; Hashimoto, Tatsunori B.

Computer Science > Computation and Language

arXiv:2301.13848 (cs)

[Submitted on 31 Jan 2023]

Title:Benchmarking Large Language Models for News Summarization

Authors:Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, Tatsunori B. Hashimoto

View PDF

Abstract:Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, and not model size, is the key to the LLM's zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LMM summaries are judged to be on par with human written summaries.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2301.13848 [cs.CL]
	(or arXiv:2301.13848v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2301.13848

Submission history

From: Tianyi Zhang [view email]
[v1] Tue, 31 Jan 2023 18:46:19 UTC (715 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2023-01

Change to browse by:

cs
cs.AI
cs.LG

References & Citations

1 blog link

(what is this?)

export BibTeX citation

Computer Science > Computation and Language

Title:Benchmarking Large Language Models for News Summarization

Submission history

Access Paper:

References & Citations

1 blog link

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Benchmarking Large Language Models for News Summarization

Submission history

Access Paper:

References & Citations

1 blog link

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators