research-article

Constructing and Evaluating a Novel Crowdsourcing-based Paraphrased Opinion Spam Dataset

Authors:

Donghyeon Park,

Jaewoo KangAuthors Info & Claims

WWW '17: Proceedings of the 26th International Conference on World Wide Web

Pages 827 - 836

https://doi.org/10.1145/3038912.3052607

Published: 03 April 2017 Publication History

Abstract

Opinion spam, intentionally written by spammers who do not have actual experience with services or products, has recently become a factor that undermines the credibility of information online. In recent years, studies have attempted to detect opinion spam using machine learning algorithms. However, limitations of gold-standard spam datasets still prove to be a major obstacle in opinion spam research. In this paper, we introduce a novel dataset called Paraphrased OPinion Spam (POPS), which contains a new type of review spam that imitates real human opinions using crowdsourcing. To create such a seemingly truthful review spam dataset, we asked task participants to paraphrase truthful reviews, and include factual information and domain knowledge in their reviews. The classification experiments and semantic analysis results show that our POPS dataset most linguistically and semantically resembles truthful reviews. We believe that our new deceptive opinion spam dataset will help advance opinion spam research.

References

[1]

S. Banerjee and A. Y. Chua. Applauses in hotel reviews: Genuine or deceptive? In Science and Information Conference (SAI), pages 938--942. IEEE, 2014.

[2]

D. Das, N. Schneider, D. Chen, and N. A. Smith. Probabilistic frame-semantic parsing. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, HLT '10, pages 948--956, Stroudsburg, PA, USA, 2010. Association for Computational Linguistics.

Digital Library

[3]

J. C. Duchi, E. Hazan, and Y. Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12:2121--2159, 2011.

Digital Library

[4]

S. Feng, R. Banerjee, and Y. Choi. Syntactic stylometry for deception detection. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2, pages 171--175. Association for Computational Linguistics, 2012.

Digital Library

[5]

V. W. Feng and G. Hirst. Detecting deceptive opinions with profile compatibility. In The 6th International Joint Conference on Natural Language Processing (IJCNLP), pages 338--346. Association for Computational Linguistics, 2013.

[6]

C. J. Fillmore. Frame semantics and the nature of language. Annals of the New York Academy of Sciences, 280(1):20--32, 1976.

[7]

S. Gokhman, J. Hancock, P. Prabhu, M. Ott, and C. Cardie. In search of a gold standard in studies of deception. In Proceedings of the Workshop on Computational Approaches to Deception Detection, pages 23--30. Association for Computational Linguistics, 2012.

Digital Library

[8]

A. Graves, N. Jaitly, and A. Mohamed. Hybrid speech recognition with deep bidirectional LS™. In 2013 IEEE Workshop on Automatic Speech Recognition and Understanding, Olomouc, Czech Republic, December 8-12, 2013, pages 273--278, 2013.

[9]

A. Heydari, M. ali Tavakoli, N. Salim, and Z. Heydari. Detection of review spam: A survey. Expert Systems with Applications, 42(7):3634--3642, 2015.

Digital Library

[10]

M. Hu and B. Liu. Mining and summarizing customer reviews. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 168--177. ACM, 2004.

Digital Library

[11]

N. Jindal and B. Liu. Analyzing and detecting review spam. In Data Mining, 2007. ICDM 2007. Seventh IEEE International Conference on, pages 547--552. IEEE, 2007.

Digital Library

[12]

N. Jindal and B. Liu. Review spam detection. In Proceedings of the 16th international conference on World Wide Web, pages 1189--1190. ACM, 2007.

Digital Library

[13]

N. Jindal and B. Liu. Opinion spam and analysis. In Proceedings of the 2008 International Conference on Web Search and Data Mining, pages 219--230. ACM, 2008.

Digital Library

[14]

S. Kim, H. Chang, S. Lee, M. Yu, and J. Kang. Deep semantic frame-based deceptive opinion spam analysis. In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management, CIKM '15, pages 1131--1140. ACM, 2015.

Digital Library

[15]

Y. Kim. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1746--1751, Doha, Qatar, October 2014. Association for Computational Linguistics.

[16]

S. Lai, L. Xu, K. Liu, and J. Zhao. Recurrent convolutional neural networks for text classification. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, AAAI'15, pages 2267--2273. AAAI Press, 2015.

Digital Library

[17]

J. Li, M.-T. Luong, and D. Jurafsky. A hierarchical neural autoencoder for paragraphs and documents. arXiv preprint arXiv:1506.01057, 2015.

[18]

J. Li, M. Ott, and C. Cardie. Identifying manipulated offerings on review portals. In EMNLP, pages 1933--1942. ACM, 2013.

[19]

J. Li, M. Ott, C. Cardie, and E. Hovy. Towards a general rule for identifying deceptive opinion spam. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1566--1576. Association for Computational Linguistics, 2014.

[20]

N. Lipka and B. Stein. Identifying featured articles in wikipedia: writing style matters. In Proceedings of the 19th international conference on World wide web, pages 1147--1148. ACM, 2010.

Digital Library

[21]

A. Mukherjee, B. Liu, and N. Glance. Spotting fake reviewer groups in consumer reviews. In Proceedings of the 21st international conference on World Wide Web, pages 191--200. ACM, 2012.

Digital Library

[22]

A. Mukherjee, V. Venkataraman, B. Liu, and N. S. Glance. What yelp fake review filter might be doing. In The 7th International AAAI Conference on Weblogs and Social Media. AAAI, 2013.

[23]

M. Ott, C. Cardie, and J. Hancock. Estimating the prevalence of deception in online review communities. In Proceedings of the 21st international conference on World Wide Web, pages 201--210. ACM, 2012.

Digital Library

[24]

M. Ott, C. Cardie, and J. T. Hancock. Negative deceptive opinion spam. In HLT-NAACL, pages 497--501. Association for Computational Linguistics, 2013.

[25]

M. Ott, Y. Choi, C. Cardie, and J. T. Hancock. Finding deceptive opinion spam by any stretch of the imagination. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, pages 309--319. Association for Computational Linguistics, 2011.

Digital Library

[26]

H. Sun, A. Morales, and X. Yan. Synthetic review spamming and defense. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 1088--1096. ACM, 2013.

Digital Library

Cited By

Rani MSumathy S(2022)A Study on Diverse Methods and Performance Measures in Sentiment AnalysisRecent Patents on Engineering10.2174/187221211499920101915495416:3Online publication date: May-2022
https://doi.org/10.2174/1872212114999201019154954
Yin BWei X(2022)Efficient Crowdsourced Pareto-Optimal Queries Over Partial Orders With Quality GuaranteeIEEE Transactions on Emerging Topics in Computing10.1109/TETC.2020.301719810:1(297-311)Online publication date: 1-Jan-2022
https://doi.org/10.1109/TETC.2020.3017198
Alallaq NAl-khiza’ay MHan X(2019)Group topic-author model for efficient discovery of latent social astroturfing groups in tourism domainCybersecurity10.1186/s42400-019-0029-82:1Online publication date: 25-Mar-2019
https://doi.org/10.1186/s42400-019-0029-8
Show More Cited By

Index Terms

Constructing and Evaluating a Novel Crowdsourcing-based Paraphrased Opinion Spam Dataset

Recommendations

Deep Semantic Frame-Based Deceptive Opinion Spam Analysis
CIKM '15: Proceedings of the 24th ACM International on Conference on Information and Knowledge Management

User-generated content is becoming increasingly valuable to both individuals and businesses due to its usefulness and influence in e-commerce markets. As consumers rely more on such information, posting deceptive opinions, which can be deliberately used ...
Opinion spam and analysis
WSDM '08: Proceedings of the 2008 International Conference on Web Search and Data Mining

Evaluative texts on the Web have become a valuable source of opinions on products, services, events, individuals, etc. Recently, many researchers have studied such opinion sources as product reviews, forum posts, and blogs. However, existing research ...
Neural networks for deceptive opinion spam detection

The products reviews are increasingly used by individuals and organizations for purchase and business decisions. Driven by the desire of profit, spammers produce synthesized reviews to promote some products or demote competitors products. So deceptive ...

Comments

Information & Contributors

Information

Published In

cover image ACM Other conferences

WWW '17: Proceedings of the 26th International Conference on World Wide Web

April 2017

1678 pages

ISBN:9781450349130

General Chairs:
Rick Barrett
W3Events
,
Rick Cummings
Murdoch University
,
Program Chairs:
Eugene Agichtein
Emory University
,
Evgeniy Gabrilovich
Google Research

Copyright © 2017 Copyright is held by the International World Wide Web Conference Committee (IW3C2).

Sponsors

IW3C2: International World Wide Web Conference Committee

In-Cooperation

SIGWEB: ACM Special Interest Group on Hypertext, Hypermedia, and Web

Publisher

International World Wide Web Conferences Steering Committee

Republic and Canton of Geneva, Switzerland

Publication History

Published: 03 April 2017

Permissions

Request permissions for this article.

Request Permissions

Check for updates

Author Tags

Qualifiers

Research-article

Funding Sources

National Research Foundation of Korea (NRF)

Conference

WWW '17

Sponsor:

IW3C2

WWW '17: 26th International World Wide Web Conference

April 3 - 7, 2017

Perth, Australia

Acceptance Rates

WWW '17 Paper Acceptance Rate 164 of 966 submissions, 17%;

Overall Acceptance Rate 1,899 of 8,196 submissions, 23%

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

5
Total Citations
View Citations
260
Total Downloads

Downloads (Last 12 months)9
Downloads (Last 6 weeks)0

Reflects downloads up to 25 Dec 2024

Other Metrics

View Author Metrics

Citations

Cited By

Rani MSumathy S(2022)A Study on Diverse Methods and Performance Measures in Sentiment AnalysisRecent Patents on Engineering10.2174/187221211499920101915495416:3Online publication date: May-2022
https://doi.org/10.2174/1872212114999201019154954
Yin BWei X(2022)Efficient Crowdsourced Pareto-Optimal Queries Over Partial Orders With Quality GuaranteeIEEE Transactions on Emerging Topics in Computing10.1109/TETC.2020.301719810:1(297-311)Online publication date: 1-Jan-2022
https://doi.org/10.1109/TETC.2020.3017198
Alallaq NAl-khiza’ay MHan X(2019)Group topic-author model for efficient discovery of latent social astroturfing groups in tourism domainCybersecurity10.1186/s42400-019-0029-82:1Online publication date: 25-Mar-2019
https://doi.org/10.1186/s42400-019-0029-8
Vogler NPearl L(2019)Using linguistically defined specific details to detect deception across domainsNatural Language Engineering10.1017/S1351324919000408(1-25)Online publication date: 1-Aug-2019
https://doi.org/10.1017/S1351324919000408
Saumya SSingh J(2018)Detection of spam reviews: a sentiment analysis approachCSI Transactions on ICT10.1007/s40012-018-0193-06:2(137-148)Online publication date: 15-May-2018
https://doi.org/10.1007/s40012-018-0193-0

View Options

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Publication

View options

PDF

View or Download as a PDF file.

eReader

View online with eReader.

Media

Figures

Other

Tables

View Table of Contents