Abstractive Summarization of Reddit Posts with Multi-level Memory Networks

Byeongchang Kim; Hyunwoo Kim; Gunhee Kim

doi:10.18653/v1/N19-1260

Abstractive Summarization of Reddit Posts with Multi-level Memory Networks

Byeongchang Kim, Hyunwoo Kim, Gunhee Kim

Abstract

We address the problem of abstractive summarization in two directions: proposing a novel dataset and a new model. First, we collect Reddit TIFU dataset, consisting of 120K posts from the online discussion forum Reddit. We use such informal crowd-generated posts as text source, in contrast with existing datasets that mostly use formal documents as source such as news articles. Thus, our dataset could less suffer from some biases that key sentences usually located at the beginning of the text and favorable summary candidates are already inside the text in similar forms. Second, we propose a novel abstractive summarization model named multi-level memory networks (MMN), equipped with multi-level memory to store the information of text from different levels of abstraction. With quantitative evaluation and user studies via Amazon Mechanical Turk, we show the Reddit TIFU dataset is highly abstractive and the MMN outperforms the state-of-the-art summarization models.

Anthology ID:: N19-1260
Volume:: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)
Month:: June
Year:: 2019
Address:: Minneapolis, Minnesota
Editors:: Jill Burstein, Christy Doran, Thamar Solorio
Venue:: NAACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 2519–2531
Language:
URL:: https://aclanthology.org/N19-1260
DOI:: 10.18653/v1/N19-1260
Bibkey:
Cite (ACL):: Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim. 2019. Abstractive Summarization of Reddit Posts with Multi-level Memory Networks. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2519–2531, Minneapolis, Minnesota. Association for Computational Linguistics.
Cite (Informal):: Abstractive Summarization of Reddit Posts with Multi-level Memory Networks (Kim et al., NAACL 2019)
Copy Citation:
PDF:: https://aclanthology.org/N19-1260.pdf
Data: Reddit TIFU, NEWSROOM

PDF Cite Search