Attentive Representation Learning with Adversarial Training for Short Text Clustering

Zhang, Wei; Dong, Chao; Yin, Jianhua; Wang, Jianyong

Computer Science > Machine Learning

arXiv:1912.03720 (cs)

[Submitted on 8 Dec 2019 (v1), last revised 20 Jan 2021 (this version, v2)]

Title:Attentive Representation Learning with Adversarial Training for Short Text Clustering

Authors:Wei Zhang, Chao Dong, Jianhua Yin, Jianyong Wang

View PDF

Abstract:Short text clustering has far-reaching effects on semantic analysis, showing its importance for multiple applications such as corpus summarization and information retrieval. However, it inevitably encounters the severe sparsity of short text representations, making the previous clustering approaches still far from satisfactory. In this paper, we present a novel attentive representation learning model for shot text clustering, wherein cluster-level attention is proposed to capture the correlations between text representations and cluster representations. Relying on this, the representation learning and clustering for short texts are seamlessly integrated into a unified model. To further ensure robust model training for short texts, we apply adversarial training to the unsupervised clustering setting, by injecting perturbations into the cluster representations. The model parameters and perturbations are optimized alternately through a minimax game. Extensive experiments on four real-world short text datasets demonstrate the superiority of the proposed model over several strong competitors, verifying that robust adversarial training yields substantial performance gains.

Comments:	14pages, to appear in IEEE TKDE
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:	arXiv:1912.03720 [cs.LG]
	(or arXiv:1912.03720v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1912.03720

Submission history

From: Wei Zhang [view email]
[v1] Sun, 8 Dec 2019 17:22:07 UTC (231 KB)
[v2] Wed, 20 Jan 2021 01:18:30 UTC (15,833 KB)

Computer Science > Machine Learning

Title:Attentive Representation Learning with Adversarial Training for Short Text Clustering

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Attentive Representation Learning with Adversarial Training for Short Text Clustering

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators