A Novel Document Generation Process for Topic Detection based on Hierarchical Latent Tree Models

Chen, Peixian; Chen, Zhourong; Zhang, Nevin L.

Computer Science > Computation and Language

arXiv:1712.04116 (cs)

[Submitted on 12 Dec 2017 (v1), last revised 28 Jun 2019 (this version, v3)]

Title:A Novel Document Generation Process for Topic Detection based on Hierarchical Latent Tree Models

Authors:Peixian Chen, Zhourong Chen, Nevin L. Zhang

View PDF

Abstract:We propose a novel document generation process based on hierarchical latent tree models (HLTMs) learned from data. An HLTM has a layer of observed word variables at the bottom and multiple layers of latent variables on top. For each document, we first sample values for the latent variables layer by layer via logic sampling, then draw relative frequencies for the words conditioned on the values of the latent variables, and finally generate words for the document using the relative word frequencies. The motivation for the work is to take word counts into consideration with HLTMs. In comparison with LDA-based hierarchical document generation processes, the new process achieves drastically better model fit with much fewer parameters. It also yields more meaningful topics and topic hierarchies. It is the new state-of-the-art for the hierarchical topic detection.

Subjects:	Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
Cite as:	arXiv:1712.04116 [cs.CL]
	(or arXiv:1712.04116v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1712.04116

Submission history

From: Peixian Chen [view email]
[v1] Tue, 12 Dec 2017 04:07:10 UTC (881 KB)
[v2] Wed, 13 Dec 2017 02:46:02 UTC (881 KB)
[v3] Fri, 28 Jun 2019 03:15:45 UTC (1,413 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2017-12

Change to browse by:

cs
cs.IR
cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

Peixian Chen
Zhourong Chen
Nevin L. Zhang

export BibTeX citation

Computer Science > Computation and Language

Title:A Novel Document Generation Process for Topic Detection based on Hierarchical Latent Tree Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:A Novel Document Generation Process for Topic Detection based on Hierarchical Latent Tree Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators