Discovering and Categorising Language Biases in Reddit

Ferrer, Xavier; van Nuenen, Tom; Such, Jose M.; Criado, Natalia

Computer Science > Computation and Language

arXiv:2008.02754 (cs)

[Submitted on 6 Aug 2020 (v1), last revised 13 Aug 2020 (this version, v2)]

Title:Discovering and Categorising Language Biases in Reddit

Authors:Xavier Ferrer, Tom van Nuenen, Jose M. Such, Natalia Criado

View PDF

Abstract:We present a data-driven approach using word embeddings to discover and categorise language biases on the discussion platform Reddit. As spaces for isolated user communities, platforms such as Reddit are increasingly connected to issues of racism, sexism and other forms of discrimination. Hence, there is a need to monitor the language of these groups. One of the most promising AI approaches to trace linguistic biases in large textual datasets involves word embeddings, which transform text into high-dimensional dense vectors and capture semantic relations between words. Yet, previous studies require predefined sets of potential biases to study, e.g., whether gender is more or less associated with particular types of jobs. This makes these approaches unfit to deal with smaller and community-centric datasets such as those on Reddit, which contain smaller vocabularies and slang, as well as biases that may be particular to that community. This paper proposes a data-driven approach to automatically discover language biases encoded in the vocabulary of online discourse communities on Reddit. In our approach, protected attributes are connected to evaluative words found in the data, which are then categorised through a semantic analysis system. We verify the effectiveness of our method by comparing the biases we discover in the Google News dataset with those found in previous literature. We then successfully discover gender bias, religion bias, and ethnic bias in different Reddit communities. We conclude by discussing potential application scenarios and limitations of this data-driven bias discovery method.

Comments:	Author's copy of the paper accepted at the International AAAI Conference on Web and Social Media (ICWSM 2021)
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG); Social and Information Networks (cs.SI)
MSC classes:	68T50, 68T09, 91D30
Cite as:	arXiv:2008.02754 [cs.CL]
	(or arXiv:2008.02754v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2008.02754
Journal reference:	International AAAI Conference on Web and Social Media (ICWSM 2021)

Submission history

From: Xavier Ferrer Aran [view email]
[v1] Thu, 6 Aug 2020 16:42:10 UTC (275 KB)
[v2] Thu, 13 Aug 2020 18:38:21 UTC (324 KB)

Computer Science > Computation and Language

Title:Discovering and Categorising Language Biases in Reddit

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Discovering and Categorising Language Biases in Reddit

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators