Learning Topics using Semantic Locality

Ziyi Zhao; Krittaphat Pugdeethosapol; Sheng Lin; Zhe Li; Caiwen Ding,; Yanzhi Wang; Qinru Qiu

arXiv:1804.04205·cs.LG·April 13, 2018

Learning Topics using Semantic Locality

Ziyi Zhao, Krittaphat Pugdeethosapol, Sheng Lin, Zhe Li, Caiwen Ding,, Yanzhi Wang, Qinru Qiu

PDF

Open Access

TL;DR

This paper introduces a novel feature extraction method for topic modeling that enhances the semantic quality of topics by filtering and merging word pairs, leading to improved accuracy over existing models.

Contribution

A new three-step feature extraction technique for topic modeling that improves topic quality by semantic filtering and merging of word pairs.

Findings

01

Improves topic accuracy by up to 12.99%.

02

Effective on datasets like OMDb, Reuters, and 20NewsGroup.

03

Outperforms LDA and RBM in experiments.

Abstract

The topic modeling discovers the latent topic probability of the given text documents. To generate the more meaningful topic that better represents the given document, we proposed a new feature extraction technique which can be used in the data preprocessing stage. The method consists of three steps. First, it generates the word/word-pair from every single document. Second, it applies a two-way TF-IDF algorithm to word/word-pair for semantic filtering. Third, it uses the K-means algorithm to merge the word pairs that have the similar semantic meaning. Experiments are carried out on the Open Movie Database (OMDb), Reuters Dataset and 20NewsGroup Dataset. The mean Average Precision score is used as the evaluation metric. Comparing our results with other state-of-the-art topic models, such as Latent Dirichlet allocation and traditional Restricted Boltzmann Machines. Our proposed data…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Text and Document Classification Technologies · Advanced Text Analysis Techniques