Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts

Carton, Samuel; Mei, Qiaozhu; Resnick, Paul

Computer Science > Computation and Language

arXiv:1809.01499v1 (cs)

[Submitted on 1 Sep 2018 (this version), latest version 19 Oct 2018 (v2)]

Title:Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts

Authors:Samuel Carton, Qiaozhu Mei, Paul Resnick

View PDF

Abstract:We introduce an adversarial method for producing high-recall explanations of neural text classifier decisions. Building on an existing architecture for extractive explanations via hard attention, we add an adversarial layer which scans the residual of the attention for remaining predictive signal. Motivated by the important domain of detecting personal attacks in social media comments, we additionally demonstrate the importance of manually setting a semantically appropriate `default' behavior for the model by explicitly manipulating its bias term. We develop a validation set of human-annotated personal attacks to evaluate the impact of these changes.

Comments:	Accepted to EMNLP 2018; code and data available soon
Subjects:	Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1809.01499 [cs.CL]
	(or arXiv:1809.01499v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1809.01499

Submission history

From: Samuel Carton [view email]
[v1] Sat, 1 Sep 2018 00:15:30 UTC (989 KB)
[v2] Fri, 19 Oct 2018 20:59:09 UTC (989 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2018-09

Change to browse by:

cs
cs.IR
cs.LG
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Samuel Carton
Qiaozhu Mei
Paul Resnick

export BibTeX citation

Computer Science > Computation and Language

Title:Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators