Dialectograms: Machine Learning Differences between Discursive Communities

Enggaard, Thyge; Lohse, August; Pedersen, Morten Axel; Lehmann, Sune

Computer Science > Computation and Language

arXiv:2302.05657 (cs)

[Submitted on 11 Feb 2023]

Title:Dialectograms: Machine Learning Differences between Discursive Communities

Authors:Thyge Enggaard (1), August Lohse (1), Morten Axel Pedersen (1 and 2), Sune Lehmann (1 and 3) ((1) Copenhagen Center for Social Data Science, University of Copenhagen, Denmark, (2) Department of Anthropology, University of Copenhagen, Denmark, (3) DTU Compute, Technical University of Denmark, Denmark)

View PDF

Abstract:Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word embeddings are complex, high-dimensional spaces and a focus on identifying differences only captures a fraction of their richness. Here, we take a step towards leveraging the richness of the full embedding space, by using word embeddings to map out how words are used differently. Specifically, we describe the construction of dialectograms, an unsupervised way to visually explore the characteristic ways in which each community use a focal word. Based on these dialectograms, we provide a new measure of the degree to which words are used differently that overcomes the tendency for existing measures to pick out low frequent or polysemous words. We apply our methods to explore the discourses of two US political subreddits and show how our methods identify stark affective polarisation of politicians and political entities, differences in the assessment of proper political action as well as disagreement about whether certain issues require political intervention at all.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2302.05657 [cs.CL]
	(or arXiv:2302.05657v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2302.05657

Submission history

From: Thyge Enggaard [view email]
[v1] Sat, 11 Feb 2023 11:32:08 UTC (13,784 KB)

Computer Science > Computation and Language

Title:Dialectograms: Machine Learning Differences between Discursive Communities

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Dialectograms: Machine Learning Differences between Discursive Communities

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators