research-article

Crowdsourcing for book search evaluation: impact of hit design on comparative system ranking

Authors:

Gabriella Kazai,

Jaap Kamps,

Marijn Koolen,

Natasa Milic-FraylingAuthors Info & Claims

SIGIR '11: Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval

Pages 205 - 214

https://doi.org/10.1145/2009916.2009947

Published: 24 July 2011 Publication History

Get Access

Abstract

The evaluation of information retrieval (IR) systems over special collections, such as large book repositories, is out of reach of traditional methods that rely upon editorial relevance judgments. Increasingly, the use of crowdsourcing to collect relevance labels has been regarded as a viable alternative that scales with modest costs. However, crowdsourcing suffers from undesirable worker practices and low quality contributions. In this paper we investigate the design and implementation of effective crowdsourcing tasks in the context of book search evaluation. We observe the impact of aspects of the Human Intelligence Task (HIT) design on the quality of relevance labels provided by the crowd. We assess the output in terms of label agreement with a gold standard data set and observe the effect of the crowdsourced relevance judgments on the resulting system rankings. This enables us to observe the effect of crowdsourcing on the entire IR evaluation process. Using the test set and experimental runs from the INEX 2010 Book Track, we find that varying the HIT design, and the pooling and document ordering strategies leads to considerable differences in agreement with the gold set labels. We then observe the impact of the crowdsourced relevance label sets on the relative system rankings using four IR performance metrics. System rankings based on MAP and Bpref remain less affected by different label sets while the Precision@10 and nDCG@10 lead to dramatically different system rankings, especially for labels acquired from HITs with weaker quality controls. Overall, we find that crowdsourcing can be an effective tool for the evaluation of IR systems, provided that care is taken when designing the HITs.

References

[1]

O. Alonso and R. A. Baeza-Yates. Design and implementation of relevance assessments using crowdsourcing. In Advances in Information Retrieval -- 33rd European Conference on IR Research (ECIR 2011), volume 6611 of LNCS, pages 153--164. Springer, 2011.

Abstract

References

Cited By

Index Terms

Recommendations

Social book search: comparing topical relevance judgements and book suggestions for evaluation

Understanding book search behavior on the web

Book search experiments: investigating IR methods for the indexing and retrieval of books

Comments

Information

Published In

Sponsors

Publisher

Publication History

Permissions

Check for updates

Author Tags

Qualifiers

Conference

Acceptance Rates

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Cited By

Get Access

Login options

Full Access

View options

PDF

eReader

Figures

Other

Share

Share this Publication link

Share on social media

Affiliations