Simple Entity-Centric Questions Challenge Dense Retrievers

Sciavolino, Christopher; Zhong, Zexuan; Lee, Jinhyuk; Chen, Danqi

Computer Science > Computation and Language

arXiv:2109.08535 (cs)

[Submitted on 17 Sep 2021 (v1), last revised 22 Feb 2022 (this version, v3)]

Title:Simple Entity-Centric Questions Challenge Dense Retrievers

Authors:Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, Danqi Chen

View PDF

Abstract:Open-domain question answering has exploded in popularity recently due to the success of dense retrieval models, which have surpassed sparse models using only a few supervised training examples. However, in this paper, we demonstrate current dense models are not yet the holy grail of retrieval. We first construct EntityQuestions, a set of simple, entity-rich questions based on facts from Wikidata (e.g., "Where was Arve Furset born?"), and observe that dense retrievers drastically underperform sparse methods. We investigate this issue and uncover that dense retrievers can only generalize to common entities unless the question pattern is explicitly observed during training. We discuss two simple solutions towards addressing this critical problem. First, we demonstrate that data augmentation is unable to fix the generalization problem. Second, we argue a more robust passage encoder helps facilitate better question adaptation using specialized question encoders. We hope our work can shed light on the challenges in creating a robust, universal dense retriever that works well across different input distributions.

Comments:	EMNLP 2021. The code and data is publicly available at this https URL
Subjects:	Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:	arXiv:2109.08535 [cs.CL]
	(or arXiv:2109.08535v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2109.08535

Submission history

From: Zexuan Zhong [view email]
[v1] Fri, 17 Sep 2021 13:19:03 UTC (1,702 KB)
[v2] Sun, 26 Sep 2021 20:50:23 UTC (1,702 KB)
[v3] Tue, 22 Feb 2022 03:03:43 UTC (3,410 KB)

Computer Science > Computation and Language

Title:Simple Entity-Centric Questions Challenge Dense Retrievers

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Simple Entity-Centric Questions Challenge Dense Retrievers

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators